scrapy入门教程

最新推荐文章于 2020-09-01 00:46:14 发布

一个程序员的自我积累

最新推荐文章于 2020-09-01 00:46:14 发布

阅读量128

点赞数

分类专栏：爬虫

本文链接：https://blog.csdn.net/weixin_41834904/article/details/82121022

版权

爬虫专栏收录该内容

4 篇文章 0 订阅

订阅专栏

import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"
    start_urls = [
        'http://quotes.toscrape.com/tag/humor/',
    ]

    def parse(self, response):
        for quote in response.css('div.quote'):
            yield {
                'text': quote.css('span.text::text').extract_first(),
                'author': quote.xpath('span/small/text()').extract_first(),
            }

        next_page = response.css('li.next a::attr("href")').extract_first()
        if next_page is not None:
            next_page = response.urljoin(next_page)
            yield scrapy.Request(next_page, callback=self.parse)

# 运行命令：scrapy runspider quotes_spider.py -o quotes.json

一个程序员的自我积累

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
scrapy入门教程

import scrapyclass QuotesSpider(scrapy.Spider): name = "quotes" start_urls = [ 'http://quotes.toscrape.com/tag/humor/', ] def parse(self, response): for quote in r...
复制链接

扫一扫

专栏目录