python爬虫基础--04 BeautifulSoup库基础（1）

最新推荐文章于 2024-10-05 08:53:18 发布

weixin_44679200

最新推荐文章于 2024-10-05 08:53:18 发布

阅读量157

点赞数

分类专栏：爬虫文章标签：爬虫 BeautifulSoup库

爬虫专栏收录该内容

5 篇文章 0 订阅

订阅专栏

1.安装BeautifulSoup库

pip install bs4

bs4在使用时候需要一个第三方库

pip install lxml

2.基础使用

（参考 https://www.cnblogs.com/bobo-zhang/p/9682516.html）
使用流程：
- 导包：from bs4 import BeautifulSoup
- 使用方式：可以将一个html文档，转化为BeautifulSoup对象，然后通过对象的方法或者属性去查找指定的节点内容
（1）转化本地文件：

 soup = BeautifulSoup(open('本地文件'), 'lxml')

（2）转化网络文件：

  soup = BeautifulSoup('字符串类型或者字节类型', 'lxml')

（3）打印soup对象显示内容为html文件中的内容

基础巩固：
（1）根据标签名查找
- soup.a 只能找到第一个符合要求的标签
（2）获取属性
- soup.a.attrs 获取a所有的属性和属性值，返回一个字典
- soup.a.attrs['href'] 获取href属性
- soup.a['href'] 也可简写为这种形式
（3）获取内容
- soup.a.string
- soup.a.text
- soup.a.get_text()
【注意】如果标签还有标签，那么string获取到的结果为None，而其它两个，可以获取文本内容
（4）find：找到第一个符合要求的标签
- soup.find('a') 找到第一个符合要求的
- soup.find('a', title="xxx")
- soup.find('a', alt="xxx")
- soup.find('a', class_="xxx")
- soup.find('a', id="xxx")
（5）find_all：找到所有符合要求的标签
- soup.find_all('a')
- soup.find_all(['a','b']) 找到所有的a和b标签
- soup.find_all('a', limit=2) 限制前两个
（6）根据选择器选择指定的内容