关于反反爬虫技术：对限制连续请求时间的处理

python2021_

已于 2022-04-03 09:08:42 修改

阅读量420

点赞数 1

文章标签： python 爬虫反爬虫网络爬虫

于 2022-04-03 09:07:47 首次发布

本文链接：https://blog.csdn.net/python2021_/article/details/123932582

版权

一般的反爬措施是在多次请求之间增加随机的间隔时间，即设置一定的延时。但如果请求后存在缓存，就可以省略设置延迟，这样一定程度地缩短了爬虫程序的耗时。

下面利用requests_cache实现模拟浏览器缓存行为来访问网站，具体逻辑如下：存在缓存，就直接走，不存在缓存，就停一下再走

示例代码

用勾子函数根据缓存行为设置访问时间

import requests_cache
import time
requests_cache.install_cache()  #默认按照浏览器的缓存进行
requests_cache.clear()
def make_throttle_hook(timeout=0.1):
    def hook(response, *args, **kwargs):
        print(response.text)
          # 判断没有缓存时就添加延时
        if not getattr(response, 'from_cache', False):
                print(f'Wait {timeout} s!')
                time.sleep(timeout)
        else:
               print(f'exists cache: {response.from_cache}')
        return response
    return hook
if __name__ == '__main__':
    requests_cache.install_cache()
    requests_cache.clear()
    session = requests_cache.CachedSession() # 创建缓存会话
    session.hooks = {'response': make_throttle_hook(2)} # 配置钩子函数
    print('first requests'.center(50,'*'))
    session.get('http://httpbin.org/get')
    print('second requests'.center(50,'*'))
    session.get('http://httpbin.org/get')

有关requests_cache的更多用法，参考下面requests_cache说明

爬虫相关库

1. 爬虫常用的测试网站：httpbin.org

httpbin.org 这个网站能测试 HTTP 请求和响应的各种信息，比如 cookie、ip、headers 和登录验证等，且支持 GET、POST 等多种方法，对 web 开发和测试很有帮助。它用 Python + Flask 编写，是一个开源项目。

官方网站：http://httpbin.org/

开源地址：https://github.com/Runscope/httpbin

2. requests-cache

requests-cache，是 requests 库的一个扩展包，利用它可以非常方便地实现请求的缓存，直接得到对应的爬取结果。