金融大数据Python爬虫——(按时间爬取、一次性批量爬取多页、一次性批量爬取多家公司多页)爬取百度新闻标题、网址、日期和新闻来源(数据爬取、清洗)

我是X大魔王

已于 2023-01-03 13:31:40 修改

阅读量2.8k

点赞数 2

分类专栏： Python👻 文章标签： python 金融爬虫

于 2023-01-01 14:36:53 首次发布

转载必须声明

本文链接：https://blog.csdn.net/Xmumu_/article/details/128511656

版权

好几个月没写博文了，有空来玩玩爬虫，之前接触了一个爬虫的项目，感触挺深的，当时有个爬取巨潮网的操作，网上的代码天花乱坠，最后还是要靠自己，今天这篇算是入门级别，欢迎收藏评论。🐳🐳🐳🐳🐳

按默认顺序

效果

在这里插入图片描述

代码

import requests
import re

headers = {
   'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/108.0.0.0 Safari/537.36'}

url = 'http://www.baidu.com/s?tn=news&rtt=1&bsst=1&cl=2&wd=阿里巴巴'  # 把链接中rtt参数换成4即是按时间排序，默认为1按焦点排序，3.4.1小节也有讲到
res = requests.get(url, headers=headers).text  # 加上headers用来告诉网站这是通过一个浏览器进行的访问
# print(res)

p_href = '<h3 class="news-title_1YtI1 "><a href="(.*?)"'   #提取新闻网址
href = re.findall(p_href, res, re.S)
p_title = '<h3 class="news-title_1YtI1 ">.*?>(.*?)</a>'
title = re.findall(p_title, res, re.S)
p_date = '<span class="c-color-gray2 c-font-normal c-gap-right-xsmall" .*?>(.*?)</span>'
date = re.findall(p_date, res)
p_source = '<span class="c-color-gray" .*?>(.*?)</span>'
source = re.findall(p_source, res)

# print(title)
# print(href)
# print(date)
# print(source)
#
for i in range(len(title)):  # range(len(title)),这里因为知道len(title) = 10，所以也可以写成for i in range(10)
    title[i] = title[i].strip()  # strip()函数用来取消字符串两端的换行或者空格，不过目前（2020-10）并没有换行或空格，所以其实不写这一行也没事
    title[i] = re.sub('<.*?>', '', title[i])  # 核心，用re.sub()函数来替换不重要的内容
    print(str(i + 1) + '.' + title

最低0.47元/天解锁文章

我是X大魔王

关注

2
点赞
踩
11

收藏

觉得还不错? 一键收藏
打赏
2
评论
金融大数据Python爬虫——(按时间爬取、一次性批量爬取多页、一次性批量爬取多家公司多页)爬取百度新闻标题、网址、日期和新闻来源(数据爬取、清洗)

好几个月没写博文了，有空来玩玩爬虫，之前接触了一个爬虫的项目，感触挺深的，当时有个爬取巨潮网的操作，网上的代码天花乱坠，最后还是要靠自己，今天这篇算是入门级别，欢迎收藏评论。🐳🐳🐳🐳🐳
复制链接

扫一扫