python爬虫：获取标签内部全部文本

最新推荐文章于 2024-08-08 16:43:27 发布

李汶峰

最新推荐文章于 2024-08-08 16:43:27 发布

阅读量2.2w

点赞数

分类专栏：学习

本文链接：https://blog.csdn.net/qq_37245397/article/details/81408304

版权

学习专栏收录该内容

18 篇文章 1 订阅

订阅专栏

取出以下字符串：亲测链接

我要取出text内容，怎么取呢，很多方法，bs4也可以，正则也可以，动态selenium也可以，这次我们先实现xpath，xpath的确很强大，不多说，上程序。

通过text获取文本

import reqiests
from lxml import etree
url = 'https://tieba.baidu.com/p/5815118868?pn=&red_tag=1075036600'
headers = {'User-Agent':'Mozilla/5.0 (Windows NT 6.1; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/67.0.3396.99 Safari/537.36'}
response = requests.get(url,headers=headers).tetx
code = etree.HTML(response)
info = code.xpath('//div['//div[@class="d_post_content_main  d_post_content_firstfloor"]/div/cc/div/text()')
#/text()获取标签的文本   //text()获取标签以及子标签的文本   
print(info)#获取的文本还要进行美化修改

使用xpath('string(.)')获取文本

import reqiests
from lxml import etree
url = 'https://tieba.baidu.com/p/5815118868?pn=&red_tag=1075036600'
headers = {'User-Agent':'Mozilla/5.0 (Windows NT 6.1; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/67.0.3396.99 Safari/537.36'}
response = requests.get(url,headers=headers).tetx
code = etree.HTML(response)
code.xpath('//div[@class="d_post_content_main  d_post_content_firstfloor"]')