正确代码如下:
#coding=utf-8
import urllib
#在python3.3里面,用urllib.request代替urllib2
import urllib.request
import re
url = "http://tieba.baidu.com/p/2460150866"
page = urllib.request.urlopen(url)
html = page.read()
print(html) #python3中只能用print(html) python2中能写print html
#正则匹配
reg = r'src="(.+?\.jpg)" pic_ext'
imgre = re.compile(reg)
imglist = re.findall(imgre, html.decode('utf-8'))
x = 0
print("start dowload pic")
for imgurl in imglist:
print(imgurl)
resp = urllib.request.urlopen(imgurl)
respHtml = resp.read()
picFile = open('%s.jpg' % x, "wb")
picFile.write(respHtml)
picFile.close()
x = x+1
print("done")
#!/usr/python3
import re
import urllib.request
def gethtml(url):
page=urllib.request.urlopen(url)
html=page.read()
return html
def getimg(html):
reg = r'src="(.*?\.jpg)"'
img=re.compile(reg)
html=html.decode('utf-8') # python3
imglist=re.findall(img,html)
x = 0
for imgurl in imglist:
urllib.request.urlretrieve(imgurl,'%s.jpg'%x)
x = x+1
html=gethtml("http://news.ifeng.com/a/20161115/50243265.html")
print(getimg(html))
代码中红色字体部分均为Python3.0及以上版本在学到爬虫是需要注意的,如果没有这些红色的代码的话可能会出现以下情况:
1.TypeError: cannot use a string pattern on a bytes-like object 这种情况解决方法就是加上html=html.decode(‘utf-8’)#python3这句代码;
2.AttributeError: module ‘urllib’ has no attribute 'urlopen’这种情况的解决办法就是将urllib改成urllib.request就行了。
参考:https://blog.csdn.net/tzs_1041218129/article/details/52228905
https://www.cnblogs.com/areyouready/p/9032251.html
674

被折叠的 条评论
为什么被折叠?



