使用Python爬虫库BeautifulSoup遍历文档树并对标签进行操作详解

最新推荐文章于 2021-12-19 21:23:40 发布

python进步学习者

最新推荐文章于 2021-12-19 21:23:40 发布

阅读量6.5k

点赞数 6

分类专栏： python教程文章标签：编程语言 python

本文链接：https://blog.csdn.net/haoxun05/article/details/104506265

版权

本文详细介绍了使用Python爬虫库BeautifulSoup如何遍历文档树，并对标签进行操作。内容包括子节点的查找（如通过名字、contents属性、children、descendants、string、strings和stripped_strings）、父节点的获取（parent和parents）、兄弟节点（next_sibling、previous_sibling、next_siblings和previous_siblings），以及回退与前进的操作（next_element、previous_element、next_elements和previous_elements）。

摘要由CSDN通过智能技术生成

今天为大家介绍下Python爬虫库BeautifulSoup遍历文档树并对标签进行操作的详细方法与函数
下面就是使用Python爬虫库BeautifulSoup对文档树进行遍历并对标签进行操作的实例，都是最基础的内容

html_doc = """
<html><head><title>The Dormouse's story</title></head>
 
<p class="title"><b>The Dormouse's story</b></p>
 
<p class="story">Once upon a time there were three little sisters; and their names were
<a href="http://example.com/elsie" rel="external nofollow" rel="external nofollow" rel="external nofollow" rel="external nofollow" class="sister" id="link1">Elsie</a>,
<a href="http://example.com/lacie" rel="external nofollow" rel="external nofollow" rel="external nofollow" rel="external nofollow" rel="external nofollow" rel="external nofollow" class="sister" id="link2">Lacie</a> and
<a href="http://example.com/tillie" rel="external nofollow" rel="external nofollow" rel="external nofollow" rel="external nofollow" class="sister" id="link3">Tillie</a>;
and they lived at the bottom of a well.</p>
 
<p class="story">...</p>
"""
 
from bs4 import BeautifulSoup
soup = BeautifulSoup(html_doc,'lxml')

一、子节点

一个Tag可能包含多个字符串或者其他Tag，这些都是这个Tag的子节点.BeautifulSoup提供了许多操作和遍历子结点的属性。

1.通过Tag的名字来获得Tag

print(soup.head)
print(soup.title)

<head><title>The Dormouse's story</title></head>
<title>The Dormouse's story</title>

通过名字的方法只能获得第一个Tag，如果要获得所有的某种Tag可以使用find_all方法

soup.find_all('a')

[<a class="sister" href="http://example.com/elsie" rel="external nofollow" rel="external nofollow" rel="external nofollow" rel="external nofollow" id="link1">Elsie</a>,
 <a class="sister" href="http://example.com/lacie" rel="external nofollow" rel="external nofollow" rel="external nofollow" rel="external nofollow" rel=