python使用代理IP访问网络爬取数据

最新推荐文章于 2024-04-29 14:48:17 发布

D_ry

最新推荐文章于 2024-04-29 14:48:17 发布

阅读量1.7k

点赞数

分类专栏：爬虫

本文链接：https://blog.csdn.net/qq_36936510/article/details/104662537

版权

示例1：Python 3.X HTTP代理调用·爬虫（动态）代理IP

'''
Python 3.x
描述：本DEMO演示了使用爬虫（动态）代理IP请求网页的过程，代码使用了多线程
逻辑：每隔5秒从API接口获取IP，对于每一个IP开启一个线程去抓取网页源码
'''
import requests
import time
import threading
from requests.packages import urllib3

ips = []

# 爬数据的线程类
class CrawlThread(threading.Thread):
    def __init__(self,proxyip):
        super(CrawlThread, self).__init__()
        self.proxyip=proxyip
    def run(self):
        # 开始计时
        start = time.time()
        #消除关闭证书验证的警告
        urllib3.disable_warnings()
        #使用代理IP请求网址，注意第三个参数verify=False意思是跳过SSL验证（可以防止报SSL错误）
        html=requests.get(url=targetUrl, proxies={
   "http" : 'http://' + self.proxyip, "https" : 'https://' + self.proxyip}, verify=False, timeout=15).content.decode()
        # 结束计时
        end = time.time()
        # 输出内容
        print(threading.current_thread().getName() +  "使用代理IP, 耗时 " + str(end - start) + "毫秒 " + self.proxyip + " 获取到如下HTML内容：\n" + html + "\n*************")

# 获取代理IP的线程类
class GetIpThread(threading.Thread):
    def __init__(self,fetchSecond):
        super(GetIpThread, self).__init__()
        self.fetchSecond=fetchSecond
    def run(self):
        global ips
        while True:
            # 获取IP列表
            res = requests.get(apiUrl).content.decode()
            # 按照\n分割获取到的IP
            ips = res.split('\n'

最低0.47元/天解锁文章

D_ry

关注

0
点赞
踩
3

收藏

觉得还不错? 一键收藏
打赏
0
评论
python使用代理IP访问网络爬取数据

示例1：Python 3.X HTTP代理调用·爬虫（动态）代理IP'''Python 3.x描述：本DEMO演示了使用爬虫（动态）代理IP请求网页的过程，代码使用了多线程逻辑：每隔5秒从API接口获取IP，对于每一个IP开启一个线程去抓取网页源码'''import requestsimport timeimport threadingfrom requests.packages...
复制链接

扫一扫