使用scrapy爬取豆瓣top250并导出

最新推荐文章于 2023-10-23 00:00:00 发布

OTF893

最新推荐文章于 2023-10-23 00:00:00 发布

阅读量897

点赞数 1

本文链接：https://blog.csdn.net/weixin_62484294/article/details/121362217

版权

import scrapy

class DoubanSpider(scrapy.Spider):
    name = 'douban'
    allowed_domains = ['douban.com']
    start_urls = [f'https://movie.douban.com/top250?start=0']
    def parse(self, response):
        title=response.xpath('//a/span[@class="title"][1]/text()').getall()
        start=response.xpath('//span[@class="rating_num"]/text()').getall()
        #使用原生的Python写法 写入txt文档
        # with open('moive.txt','w',encoding='utf-8') as f:
        #     for t,s in zip(title,start):
        #         #print(f'{t}{s}')
        #         f.write(f'{t},{s}\n')
        #
        #导出的另一种格式写法
# FEED_EXPORT_ENCODING = 'utf-8' 导出格式json使用
        for t,s in zip(title,start):
            yield {
                'title':t,
                'star':s
            }
#scrapy crawl douban -o wen.csv  保存csv格式的文件 也就是excel
#scrapy crawl douban -o wen.json 保存json格式的文件 
#scrapy crawl douban -o wen.txt  以文本方式保存
#打开设置
#打开pipelines 添加open close
# ITEM_PIPELINES = {
#    'scrapy01.pipelines.Scrapy01Pipeline': 300,
# }

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

立即使用

OTF893

关注关注

1
点赞
踩
1

收藏

觉得还不错? 一键收藏
0
评论
使用scrapy爬取豆瓣top250并导出

import scrapyclass DoubanSpider(scrapy.Spider): name = 'douban' allowed_domains = ['douban.com'] start_urls = [f'https://movie.douban.com/top250?start=0'] def parse(self, response): title=response.xpath('//a/span[@class="title"][.
复制链接

扫一扫