行百里者半九十 —— scrapy 框架关于下载中间件的补充

本文介绍了如何在Scrapy框架中实现下载中间件,以拦截请求并应用UA池和代理池。通过随机选择UA伪装用户代理,并在请求遇到异常时动态切换HTTP或HTTPS代理,确保爬虫的匿名性和稳定性。代码中展示了process_request方法用于设置UA,process_exception方法用于处理请求异常并切换代理。
摘要由CSDN通过智能技术生成

下载中间件拦截请求

需求

《行百里者半九十 —— scrapy 框架(6)》一文中我们介绍了下载中间件的作用,并演示了其中拦截响应的代码实现。

现在我们来试着实现拦截请求的代码实现,也就是UA池和代理池的实现。因为免费 IP 总是失效,所以在这里只提供中间件部分的代码实现,不提供运行结果。

正因为此,代码可能有所疏漏,还望各位看官海涵。

中间件部分代码实现

# Define here the models for your spider middleware
#
# See documentation in:
# https://docs.scrapy.org/en/latest/topics/spider-middleware.html

from scrapy import signals

# useful for handling different item types with a single interface
from itemadapter import is_item, ItemAdapter


# class MidproSpiderMiddleware:
#     # Not all methods need to be defined. If a method is not defined,
#     # scrapy acts as if the spider middleware does not modify the
#     # passed objects.
#
#     @classmethod
#     def from_crawler(cls, crawler):
#         # This method is used by Scrapy to create your spiders.
#         s = cls()
#         crawler.signals.connect(s.spider_opened, signal=signals.spider_opened)
#         return s
#
#     def process_spider_input(self, response, spider):
#         # Called for each response that goes through the spider
#         # middleware and into the spider.
#
#         # Should return None or raise an exception.
#         return None
#
#     def process_spider_output(self, response, result, spider):
#         # Called with the results returned from the Spider, after
#         # it has processed the response.
#
#         # Must return an iterable of Request, or item objects.
#         for i in result:
#             yield i
#
#     def process_spider_exception(self, response, exception, spider):
#         # Called when a spider or process_spider_input() method
#         # (from other spider middleware) raises an exception.
#
#         # Should return either None or an iterable of Request or item objects.
#         pass
#
#     def process_start_requests(self, start_requests, spider):
#         # Called with the start requests of the spider, and works
#         # similarly to the process_spider_output() method, except
#         # that it doesn’t have a response associated.
#
#         # Must return only requests (not items).
#         for r in start_requests:
#             yield r
#
#     def spider_opened(self, spider):
#         spider.logger.info('Spider opened: %s' % spider.name)

import random

class MidproDownloaderMiddleware:
    # 代理池
    proxy_http = []

    proxy_https = []

    # 拦截请求
    def process_request(self, request, spider):
        # UA伪装
        user_agent_list = [] # UA池

        request.headers["User-Agent"] = random.choice(user_agent_list)

        return None

    # 拦截所有响应
    def process_response(self, request, response, spider):
        # Called with the response returned from the downloader.

        # Must either;
        # - return a Response object
        # - return a Request object
        # - or raise IgnoreRequest
        return response

    # 拦截所有异常
    def process_exception(self, request, exception, spider):
        if request.url.split(":")[0] == "http":
            request.meta["proxy"] = "http://" + random.choice(self.proxy_http)
        else:
            request.meta["proxy"] = "https://" + random.choice(self.proxy_https)

        return request # 将修正之后的请求对象重新进行请求发送
评论
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包
实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值