爬取糗事百科[文字]栏前十页

最新推荐文章于 2020-11-05 14:27:17 发布

aspiring123

最新推荐文章于 2020-11-05 14:27:17 发布

阅读量172

点赞数

分类专栏： Python 爬虫文章标签： crawer 爬虫

本文链接：https://blog.csdn.net/qq_39198486/article/details/81366049

版权

Python 爬虫专栏收录该内容

10 篇文章 0 订阅

订阅专栏

import urllib.request
import re



def jokeCrawer(url):
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Maxthon/4.4.3.4000 Chrome/30.0.1599.101 Safari/537.36"
    }
    req = urllib.request.Request(url, headers=headers)
    response = urllib.request.urlopen(req)

    HTML = response.read().decode('utf-8')

    pat = r'<div class="author clearfix">(.*?)<span class="stats-vote"><i class="number">'
    re_joke = re.compile(pat, re.S)
    divsList = re_joke.findall(HTML)
    # print(divsList)
    # print(len(divsList))
    dic = {}
    for div in divsList:
        # 用户名
        re_u = re.compile(r'<h2>(.*?)</h2>', re.S)
        username = re_u.findall(div)
        username = username[0]

        # 段子
        re_d = re.compile(r'<div class="content">\n<span>(.*?)</span>', re.S)
        duanzi = re_d.findall(div)
        duanzi = duanzi[0]
        # print(duanzi)

        dic[username] = duanzi


    return dic



    # with open(r'E:\all-workspace\qianfeng\0802-爬虫简介与json\file\file3.html', 'w') as f:
    #     f.write(HTML)

for i in range(1,10):
    url = "https://www.qiushibaike.com/text/page/str(i)/"
    info = jokeCrawer(url)
    for k, v in info.items():
        print(k+"说"+ v)

aspiring123

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
爬取糗事百科[文字]栏前十页

import urllib.requestimport redef jokeCrawer(url): headers = { "User-Agent": "Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Maxthon/4.4.3.4000 Chrome/30.0...
复制链接

扫一扫

专栏目录