python爬虫爬取新闻标题_Python正则抓取新闻标题和链接的方法示例

最新推荐文章于 2022-10-21 09:59:08 发布

weixin_39620334

最新推荐文章于 2022-10-21 09:59:08 发布

阅读量817

点赞数

文章标签： python爬虫爬取新闻标题

本文实例讲述了Python正则抓取新闻标题和链接的方法。分享给大家供大家参考，具体如下：

#-*-coding:utf-8-*-

import re

from urllib import urlretrieve

from urllib import urlopen

#获取网页信息

doc = urlopen("http://www.itongji.cn/news/").read() #自己找的一个大数据的新闻网站

#抓取新闻标题和链接

def extract_title(info):

pat = '

'

title = re.findall(pat, info)

titles='\n'.join(title)

#print titles

#修改指定字符串

titles1=titles.replace('class="title"','title')

titles2=titles1.replace('>',':')

titles3=titles2.replace('href','url:')

titles4=titles3.replace('="/','"http://www.itongji.cn/')

#写入文件

save=open('xinwen.txt','w')

save.write(titles4)

save.close()

titles = extract_title(doc)

PS：这里再为大家提供2款非常方便的正则表达式工具供大家参考使用：

希望本文所述对大家Python程序设计有所帮助。

weixin_39620334

关注

0
点赞
踩
1

收藏

觉得还不错? 一键收藏
0
评论
python爬虫爬取新闻标题_Python正则抓取新闻标题和链接的方法示例

本文实例讲述了Python正则抓取新闻标题和链接的方法。分享给大家供大家参考，具体如下：#-*-coding:utf-8-*-import refrom urllib import urlretrievefrom urllib import urlopen#获取网页信息doc = urlopen("http://www.itongji.cn/news/").read() #自己找的一个大数据的新闻...
复制链接

扫一扫

评论

被折叠的条评论为什么被折叠?

到【灌水乐园】发言

查看更多评论

添加红包

成就一亿技术人!

hope_wisdom

发出的红包

实付元

使用余额支付

点击重新获取

扫码支付

钱包余额 0

抵扣说明：

1.余额是钱包充值的虚拟货币，按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载，可以购买VIP、付费专栏及课程。