BeautifulSoup库的基本元素

最新推荐文章于 2022-10-22 10:37:53 发布

NY_YN

最新推荐文章于 2022-10-22 10:37:53 发布

阅读量415

点赞数

文章标签： html python xml

本文链接：https://blog.csdn.net/ny_yn/article/details/111576704

版权

BeautifulSoup库的基本元素

BeautifulSoup库的理解
html是由<>储存信息的，不同标签之间存在上下游关系形成标签树(一组类）
BeautifulSoup库是解析、遍历、维护“标签树”的功能库

<html>
<body>
<p class='title">...</p>
</body>
</html>

BeautifulSoup的基本元素
BeautifulSoup库
BeautifulSoup库，也叫beautifulsou4或bs4，引用方法：

from bs4 import BeautifulSoup

import bs4

BeautifulSoup对应一个HTML/XML文档的全部内容
在这里插入图片描述

from bs4 import BeautifulSoup
soup = BeautifulSoup("<html>data</html>","html.parser")#html.parser:指定解析器
soup2 = BeautifulSoup(open("D://demo.html"),"html.parser")

beautifulsoup的四种解析器

解析器	使用方法	条件
bs4的HTML解析器	BeautifulSoup(mk,‘html.parser’)	安装bs4
Ixml的HTML解析器	BeautifulSoup(mk,‘Ixml’)	pip installer Ixml
Ixml的XML解析器	BeautifulSoup(mk,‘xml’)	pip installer Ixml
html5lib的解析器	BeautifulSoup(mk,‘html5lib’	pip installer html5lib

以上的解析器都能有效解析html

BeautifulSoup类的基本元素

基本元素	说明
Tag	标签，最基本的信息组织单元，分别用<>和</>标明开头和结尾
Name	标签的名字，< p>…</ p >的名字是’p’,格式：< tag>.name
Attributes	标签的属性，字典形式组织，格式：< tag >.attrs
NavigableString	标签非属性字符串，<>…</>中字符串，格式:< tag >.string
Comment	标签内字符串的注释部分，一种特殊的Comment类型

在这里插入图片描述