python爬虫（requests）

最新推荐文章于 2024-04-09 15:07:35 发布

Nothenhe

最新推荐文章于 2024-04-09 15:07:35 发布

阅读量121

点赞数

本文链接：https://blog.csdn.net/Nothenhe/article/details/104402008

版权

python爬虫（requests）

网址：http://www.biquge.cm/2/2042/
ps：该网址以暂停解析

准备工作

导入第三方库：requests bs4

爬虫笔记：

导入requests库也可以使用urllib库这里使用requests库比urllib库简单，内部是由requests库写的，以及数据筛选库如：bs4 或者lxml
编写header字典，用于反反爬虫不然容易被检测。
实例化res对象，通过requests.get(url=“网址”,headers=header),获得网站内容
将获取的内容使用xpath或者beautifulsoup进行解析
存储解析到的有效数据

代码部分;

# 导入requests库
import requests
# 导入文件操作库
import codecs
from bs4 import BeautifulSoup
import sys
import importlib
importlib.reload(sys)

# 给请求指定一个请求头来模拟chrome浏览器
global headers
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/54.0.2840.99 Safari/537.36'}
server = 'http://www.biquge.cm'
# 星辰变地址
book = 'http://www.biquge.cm/2/2042/'
# 定义存储位置
global save_path
save_path = 'G:/星辰变'


# 获取章节内容
def get_contents(chapter):
    req = requests.get(url=chapter)
    html = req.content
    html_doc = str(html, 'gbk')
    bf = BeautifulSoup(html_doc, 'html.parser')
    texts = bf.find_all('div', id="content")
    # 获取div标签id属性content的内容 \xa0 是不间断空白符 &nbsp;
    content = texts[0].text.replace('\xa0' * 4, '\n')
    return content


# 写入文件
def write_txt(chapter, content, code):
    with codecs.open(chapter, 'a', encoding=code)as f:
        f.write(content)


# 主方法
def main():
    res = requests.get(book, headers=headers)
    html = res.content
    html_doc = str(html, 'gbk')
    # 使用自带的html.parser解析
    soup = BeautifulSoup(html_doc, 'html.parser')
    # 获取所有的章节
    a = soup.find('div', id='list').find_all('a')
    print('总章节数: %d ' % len(a))
    for each in a:
        try:
            chapter = server + each.get('href')
            content = get_contents(chapter)
            chapter = save_path + "/" + each.string + ".txt"
            write_txt(chapter, content, 'utf8')
        except Exception as e:
            print(e)


if __name__ == '__main__':
    main()