爬取知乎发现首页 requests+pyquery

最新推荐文章于 2021-05-04 22:14:07 发布

mooe1011

最新推荐文章于 2021-05-04 22:14:07 发布

阅读量416

点赞数

分类专栏： Python

本文链接：https://blog.csdn.net/mooe1011/article/details/91593598

版权

本人的第一次爬虫，爬取后的文章保存在txt文件里。

参考：python3 网络爬虫开发实战

import requests
from requests import RequestException
from pyquery import PyQuery as pq

url = 'https://www.zhihu.com/explore'
headers = {
    'User-Agent':
        'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit'
        '/537.36 (KHTML, like Gecko) Chrome/74.0.3729.169 Safari/537.36'
}

def get_one_page(url):
    try:
        response = requests.get(url, headers=headers)
        if response.status_code == 200:
            return response.text
        return None
    except RequestException:
        return None


def write_result(content):
    doc = pq(content)
    items  = doc('.explore-tab .feed-item').items()
    with open('explore.txt','a',encoding='utf-8') as file:
        for item in items:
            question = item.find('h2').text()

最低0.47元/天解锁文章

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

立即使用

mooe1011

关注关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
爬取知乎发现首页 requests+pyquery

本人的第一次爬虫，爬取后的文章保存在txt文件里。参考：python3 网络爬虫开发实战import requestsfrom requests import RequestExceptionfrom pyquery import PyQuery as pqurl = 'https://www.zhihu.com/explore'headers = { 'User-Ag...
复制链接

扫一扫