python爬虫面向对象_python面向对象多线程爬虫爬取搜狐页面的实例代码

最新推荐文章于 2021-04-11 17:00:00 发布

weixin_39525118

最新推荐文章于 2021-04-11 17:00:00 发布

阅读量163

点赞数

文章标签： python爬虫面向对象

首先我们需要几个包：requests, lxml, bs4, pymongo, redis

1. 创建爬虫对象，具有的几个行为：抓取页面，解析页面，抽取页面，储存页面

class Spider(object):

def __init__(self):

# 状态(是否工作)

self.status = SpiderStatus.IDLE

# 抓取页面

def fetch(self, current_url):

pass

# 解析页面

def parse(self, html_page):

pass

# 抽取页面

def extract(self, html_page):

pass

# 储存页面

def store(self, data_dict):

pass

2. 设置爬虫属性，没有在爬取和在爬取中，我们用一个类封装， @unique使里面元素独一无二，Enum和unique需要从 enum里面导入：

@unique

class SpiderStatus(Enum):

IDLE = 0

WORKING = 1

3. 重写多线程的类：

class SpiderThread(Thread):

def __init__(self, spider, tasks):

super().__init__(daemon=True)

self.spider = spider

self.tasks = tasks

def run(self):

while True:

pass

4. 现在爬虫的基本结构已经做完了，在main函数创建tasks， Queue需要从queue里面导入：

def main():

# list没有锁，所以使用Queue比较安全, task_queue=[]也可以使用,Queue 是先进先出结构, 即 FIFO

task_queue = Queue()

# 往队列放种子url, 即搜狐手机端的url

task_queue.put('http://m.sohu,com/')

# 指定起多少个线程

spider_threads = [SpiderThread(Spider(), task_queue) for _ in range(10)] for spider_thread in spider_threads:

spider_thread.start()

# 控制主线程不能停下,如果队列里有东西，任务不能停, 或者spider处于工作状态，也不能停

while task_queue.empty() or is_any_alive(spider_threads):

pass

print('Over')

4-1. 而 is_any_threads则是判断线程里是否有spider还活着，所以我们再写一个函数来封装一下:

def is_any_alive(spider_threads):

return any([spider_thread.spider.status == SpiderStatus.WORKING

for spider_thread in spider_threads])

5. 所有的结构已经全部写完，接下来就是可以填补爬虫部分的代码，在SpiderThread(Thread)里面，开始写爬虫运行 run 的方法，即线程起来后，要做的事情：

def run(self):

while True:

# 获取url

current_url = self.tasks_queue.get()

visited_urls.add(current_url)

# 把爬虫的status改成working

self.spider.status = SpiderStatus.WORKING

# 获取页面

html_page = self.spider.fetch(current_url)

# 判断页面是否为空

if html_page not in [None, '']:

weixin_39525118

关注

0
点赞
踩
1

收藏

觉得还不错? 一键收藏
0
评论
python爬虫面向对象_python面向对象多线程爬虫爬取搜狐页面的实例代码

首先我们需要几个包：requests, lxml, bs4, pymongo, redis1. 创建爬虫对象，具有的几个行为：抓取页面，解析页面，抽取页面，储存页面class Spider(object):def __init__(self):# 状态(是否工作)self.status = SpiderStatus.IDLE# 抓取页面def fetch(self, current_url):pa...
复制链接

扫一扫

评论

被折叠的条评论为什么被折叠?

到【灌水乐园】发言

查看更多评论

添加红包

成就一亿技术人!

hope_wisdom

发出的红包

实付元

使用余额支付

点击重新获取

扫码支付

钱包余额 0

抵扣说明：

1.余额是钱包充值的虚拟货币，按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载，可以购买VIP、付费专栏及课程。