手写python爬虫

最新推荐文章于 2022-07-09 20:53:50 发布

三名狂客

最新推荐文章于 2022-07-09 20:53:50 发布

阅读量1.1k

点赞数

分类专栏： python爬虫文章标签： python 网络爬虫实现的原理 python爬虫

本文链接：https://blog.csdn.net/zuochao_2013/article/details/75271155

版权

一、图片爬虫

(1)京东手机图片的抓取

import re
import urllib.request
def craw(url,page):
    html1=urllib.request.urlopen(url).read()
    html1=str(html1)
    pat1='<div id="plist".+? <div class="page clearfix">'
    result1=re.compile(pat1).findall(html1)
    result1=result1[0]
    pat2='<img width="220" height="220" data-img="1" data-lazy-img="//(.+?\.jpg)">'
    imagelist=re.compile(pat2).findall(result1)
    x=1
    for imageurl in imagelist:
        imagename="E:/picture/"+str(page)+str(x)+".jpg"
        imageurl="http://"+imageurl
        try:
            urllib.request.urlretrieve(imageurl,filename=imagename)
        except urllib.error.URLError as e:
            if hasattr(e,"code"):
                x+=1
            if hasattr(e,"reason"):
                x+=1
        x+=1
        print(x)
for i in range(1,79):
    url="http://list.jd.com/list.html?cat=9987,653,655&page="+str(i)
    craw(url,i)

最低0.47元/天解锁文章

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

立即使用

三名狂客

关注关注

0
点赞
踩
2

收藏

觉得还不错? 一键收藏
0
评论
手写python爬虫

一、图片爬虫 (1)京东手机图片的抓取import reimport urllib.requestdef craw(url,page): html1=urllib.request.urlopen(url).read() html1=str(html1) pat1='' result1=re.compile(pat1).findall(html1)
复制链接

扫一扫