Python爬虫 BeautifulSoup库详解及实践全部相关代码

最新推荐文章于 2023-05-14 09:00:00 发布

胡乱huluan

最新推荐文章于 2023-05-14 09:00:00 发布

阅读量572

点赞数

分类专栏： # 网安—Python爬虫文章标签： python 爬虫代码

本文链接：https://blog.csdn.net/qq_44867435/article/details/104412042

版权

Python爬虫（六）

学习Python爬虫过程中的心得体会以及知识点的整理，方便我自己查找，也希望可以和大家一起交流。

—— BeautifulSoup库详解及实践相关代码 ——

1.BeautifulSoup库详解

from bs4 import BeautifulSoup
import bs4
import re

# 待分析字符串
html_doc = """ 
<html> 
    <head> 
        <title>The Dormouse's story</title> 
    </head> 
    <body> 
    <p class="title aq"> 
        <b> 
            The Dormouse's story 
        </b> 
    </p> 

    <p class="story">Once upon a time there were three little sisters; and their names were 
        <a href="http://example.com/elsie" class="sister" id="link1">Elsie</a>, 
        <a href="http://example.com/lacie" class="sister" id="link2">Lacie</a>  
        and 
        <a href="http://example.com/tillie" class="sister" id="link3">Tillie</a>; 
        and they lived at the bottom of a well. 
    </p> 

    <p class="story">...</p> 
    </body>
</html>
"""
# 每一段代码中注释部分即为运行结果
# html字符串创建BeautifulSoup对象
soup = BeautifulSoup(html_doc, 'html.parser', from_encoding='utf-8')

# 输出第一个 title 标签
print(soup.title)
# <title>The Dormouse's story</title>

# 输出第一个 title 标签的标签名称
print(soup.title.name)
# title

# 输出第一个 title 标签的包含内容
print(soup.title.string)
# The Dormouse's story

# 输出第一个