python beautifulsoup获取特定html源码

最新推荐文章于 2022-11-14 12:00:00 发布

weixin_30278311

最新推荐文章于 2022-11-14 12:00:00 发布

阅读量509

点赞数

文章标签： python

原文链接：http://www.cnblogs.com/vickey-wu/p/6843411.html

版权

beautifulsoup 获取特定html源码（无需登录页面）

import re
from bs4 import BeautifulSoup
import urllib2

url = 'http://www.cnblogs.com/vickey-wu/'
# connect to a URL
web = urllib2.urlopen(url)
# read html code
html = web.read()
# print html
soup = BeautifulSoup(html,'html.parser')
prety = soup.prettify()
# print prety
pointed_div = soup.findAll(name="div", attrs={"class":re.compile("forFlow")})　　　　# 筛选标签为div且属性class为forFlow的源码
print pointed_div

转载于:https://www.cnblogs.com/vickey-wu/p/6843411.html

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

立即使用

weixin_30278311

关注关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
python beautifulsoup获取特定html源码

beautifulsoup 获取特定html源码（无需登录页面）import refrom bs4 import BeautifulSoupimport urllib2url = 'http://www.cnblogs.com/vickey-wu/'# connect to a URLweb = urllib2.urlopen(url)# read html codehtml = web.read...
复制链接

扫一扫