python抓取网页内容_如何用Python爬虫抓取网页内容?

最新推荐文章于 2023-11-04 14:31:29 发布

weixin_39916379

最新推荐文章于 2023-11-04 14:31:29 发布

阅读量92

点赞数

文章标签： python抓取网页内容

展开全部

首先,你要安装requests和5261BeautifulSoup4,然后执行如下代码.import requests

from bs4 import BeautifulSoup

iurl = 'http://news.sina.com.cn/c/nd/2017-08-03/doc-ifyitapp0128744.shtml'

res = requests.get(iurl)

res.encoding = 'utf-8'

#print(len(res.text))

soup = BeautifulSoup(res.text,'html.parser')

#标题4102

H1 = soup.select('#artibodyTitle')[0].text

#来源

time_source = soup.select('.time-source')[0].text

#来源

origin = soup.select('#artibody p')[0].text.strip()

#原标题

oriTitle = soup.select('#artibody p')[1].text.strip()

#内容

raw_content = soup.select('#artibody p')[2:19]

content = []

for paragraph in raw_content:

content.append(paragraph.text.strip())

'@'.join(content)

#责任编辑

ae = soup.select('.article-editor')[0].text

这样1653就可以了

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

关注关注