python中beautifulsoup怎么找出网页链接,Python：如何使用BeautifulSoup从HTML页面中提取URL？...

最新推荐文章于 2024-04-16 15:21:32 发布

魏金华

最新推荐文章于 2024-04-16 15:21:32 发布

阅读量1.1k

点赞数

文章标签： python中beautifulsoup怎么找出网页链接

我需要得到< a href =>具有类article-additional-info的所有div的值

我是BeautifulSoup的新手

所以我需要网址

"http://www.thehindu.com/news/national/gangrape-case-two-lawyers-claim-to-be-engaged-by-accused/article4332680.ece"

"http://www.thehindu.com/news/cities/Delhi/power-discoms-demand-yet-another-hike-in-charges/article4331482.ece"

实现这一目标的最佳方法是什么？

最佳答案

根据您的标准,它返回三个URL(而不是两个) – 您想要过滤掉第三个吗？

基本思想是迭代HTML,只抽取你的类中的那些元素,然后迭代该类中的所有链接,拉出实际的链接：

In [1]: from bs4 import BeautifulSoup

In [2]: html = # your HTML

In [3]: soup = BeautifulSoup(html)

In [4]: for item in soup.find_all(attrs={'class': 'article-additional-info'}):

...: for link in item.find_all('a'):

...: print link.get('href')

...:

http://www.thehindu.com/news/national/gangrape-case-two-lawyers-claim-to-be-engaged-by-accused/article4332680.ece

http://www.thehindu.com/news/cities/Delhi/power-discoms-demand-yet-another-hike-in-charges/article4331482.ece

http://www.thehindu.com/news/cities/Delhi/power-discoms-demand-yet-another-hike-in-charges/article4331482.ece#comments

这会将您的搜索范围限制为仅包含article-additional-info类标记的元素,并在其中查找所有锚点(a)标记并获取其相应的href链接.

总结

如果觉得编程之家网站内容还不错，欢迎将编程之家网站推荐给程序员好友。

本图文内容来源于网友网络收集整理提供，作为学习参考使用，版权属于原作者。

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
python中beautifulsoup怎么找出网页链接,Python：如何使用BeautifulSoup从HTML页面中提取URL？...

我需要得到< a href =>具有类article-additional-info的所有div的值我是BeautifulSoup的新手所以我需要网址"http://www.thehindu.com/news/national/gangrape-case-two-lawyers-claim-to-be-engaged-by-accused/article4332680.ece""htt...
复制链接

扫一扫

评论

被折叠的条评论为什么被折叠?

到【灌水乐园】发言

查看更多评论

添加红包

成就一亿技术人!

hope_wisdom

发出的红包

实付元

使用余额支付

点击重新获取

扫码支付

钱包余额 0

抵扣说明：

1.余额是钱包充值的虚拟货币，按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载，可以购买VIP、付费专栏及课程。