我的第一只爬虫
数据源
打开糗百主页,查看html源文件
代码
抓取作者名字
#coding=utf-8
import urllib
import urllib2
import re
page = 2
url = 'http://www.qiushibaike.com/hot/page/' + str(page)
user_agent = 'Mozilla/4.0 (compatible; MSIE 5.5; Windows NT)'
headers = { 'User-Agent' : user_agent }
try:
request = urllib2.Request(url,headers = headers)
response = urllib2.urlopen(request)
content = response.read().decode('utf-8')
pattern = re.compile('<div.*?author.*?<a.*?</a>.*?<a.*?title="(.*?)">.*?<h2>(.*?)</h2>.*?</a>.*?</div>',re.S)
items = re.findall(pattern,content)
for item in items:
print item[0]
except urllib2.URLError, e:
if hasattr(e,"code"):
print e.code
if hasattr(e,"reason"):
print e.reason
结果
Python 2.7.2 |EPD_free 7.2-2 (32-bit)| (default, Sep 14 2011, 11:02:05) [MSC v.1500 32 bit (Intel)] on win32
Type "copyright", "credits" or "license()" for more information.
>>> ================================ RESTART ================================
>>>
挖鼻孔的老虎
loser...........
Dan喵
@胖妞
向阳河
单名一个饭字
欲湖冰心
王爷有人
陌路莫回。媚娘
向阳河
♂ART
⌒oOㄣ季向晚哥哥
这个冬天冻成狗
哈哈大好时光
王冰痕
壞壊_
智商都用来卖萌啦
二女子、
(驹迷)超越
_后来!