Python进阶实战爬虫：对于房天下租房信息进行爬取

最新推荐文章于 2023-10-24 10:52:07 发布

学习-永无止境

最新推荐文章于 2023-10-24 10:52:07 发布

阅读量358

点赞数

分类专栏： Python零基础学习教程文章标签： python 数据挖掘

本文链接：https://blog.csdn.net/weixin_45974628/article/details/103647431

版权

本文介绍了使用Python进行房天下租房信息的爬取，详细讲解了爬取过程，包括代码实现和后续分区信息的抓取。

摘要由CSDN通过智能技术生成

对于房天下租房信息进行爬取

代码

import re

import requests
from lxml.html import etree

url_xpath = '//dd/p[1]/a[1]/@href'
title_xpath = '//dd/p[1]/a[1]/@title'
data_xpaht = '//dd/p[2]/text()'
headers = {
    'rpferpr': 'https://sh.zu.fang.com/',
    'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/75.0.3770.90 Safari/537.36'
}
rp = requests.get('https://sh.zu.fang.com/', headers=headers)
rp.encoding = rp.apparent_encoding
html = etree.HTML(rp.text)
url = html.xpath(url_xpath)
title = html.xpath(title_xpath)
data = re.findall('<p class="font15 mt12 bold">(.*?)</p>', rp.text, re.S)
mold_lis = []
house_type_lis = []
area_lis = []
for a in data:
    a = re.sub('�O', '平方米', a)
    mold = re.findall('\