爬虫神器xpath的用法(三)

最新推荐文章于 2024-08-26 18:04:13 发布

weixin_30797027

最新推荐文章于 2024-08-26 18:04:13 发布

阅读量69

点赞数

文章标签：爬虫

原文链接：http://www.cnblogs.com/gide/p/5246809.html

版权

xpath的多线程爬虫

#encoding=utf-8
'''
pool = Pool(4) cpu的核数为4核
results = pool.map(爬取函数，网址列表)
'''
from multiprocessing.dummy import Pool as ThreadPool
import requests
import time

def getsource(url):
    html = requests.get(url)

urls = []

for i in range(1,21):
    newpage = 'http://tieba.baidu.com/p/3522395718?pn=' + str(i)
    urls.append(newpage)

time1 = time.time()
for i in urls:
    print i
    getsource(i)
time2 = time.time()
print u'单线程耗时：' + str(time2-time1)

pool = ThreadPool(4)
time3 = time.time()
results = pool.map(getsource, urls)
pool.close()
pool.join()
time4 = time.time()
print u'并行耗时：' + str(time4-time3)

输出：

单线程耗时：12.0818030834
并行耗时：3.58480286598

转载于:https://www.cnblogs.com/gide/p/5246809.html

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

立即使用

weixin_30797027

关注关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
爬虫神器xpath的用法(三)

xpath的多线程爬虫#encoding=utf-8'''pool = Pool(4) cpu的核数为4核results = pool.map(爬取函数，网址列表)'''from multiprocessing.dummy import Pool as ThreadPoolimport requestsimport timedef getsource(u...
复制链接

扫一扫