爬虫基础模板

最新推荐文章于 2024-01-18 21:58:19 发布

小鸡快跑哒咯哒

最新推荐文章于 2024-01-18 21:58:19 发布

阅读量199

点赞数

分类专栏：爬虫学习

本文链接：https://blog.csdn.net/weixin_42140690/article/details/102621445

版权

爬虫学习专栏收录该内容

4 篇文章 0 订阅

订阅专栏

import requests # 调用requests库
from bs4 import BeautifulSoup # 调用BeautifulSoup库
res =requests.get('https://localprod.pandateacher.com/python-manuscript/crawler-html/spider-men5.0.html')
# 返回一个response对象，赋值给res
print('响应状态码:',res.status_code) #检查请求是否正确响应
res.encoding='gbk'
#定义Response对象的编码为gbk
html=res.text
# 把res解析为字符串
soup = BeautifulSoup( html,'html.parser')
# 把网页解析为BeautifulSoup对象
items = soup.find_all(class_='books')   # 通过匹配属性class='books'提取出我们想要的元素
for item in items:                      # 遍历列表items
    kind = item.find('h2')               # 在列表中的每个元素里，匹配标签<h2>提取出数据
    title = item.find(class_='title')     #  在列表中的每个元素里，匹配属性class_='title'提取出数据
    brief = item.find(class_='info')      # 在列表中的每个元素里，匹配属性class_='info'提取出数据
    print(kind.text,'\n',title.text,'\n',title['href'],'\n',brief.text) # 打印书籍的类型、名字、链接和简介的文字

小鸡快跑哒咯哒

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
爬虫基础模板

import requests # 调用requests库from bs4 import BeautifulSoup # 调用BeautifulSoup库res =requests.get('https://localprod.pandateacher.com/python-manuscript/crawler-html/spider-men5.0.html')# 返回一个response对...
复制链接

扫一扫

专栏目录