创新项目实训(四)

最新推荐文章于 2024-08-28 16:14:15 发布

Reika_xiang

最新推荐文章于 2024-08-28 16:14:15 发布

阅读量109

点赞数

分类专栏：创新项目实训文章标签： python 爬虫

本文链接：https://blog.csdn.net/a0939763286/article/details/115982561

版权

Python 爬虫 Booking.com 数据获取 BeautifulSoup

关键词由CSDN通过智能技术生成

创新项目实训专栏收录该内容

12 篇文章 0 订阅

订阅专栏

创新项目实训(四)

前言

我们组打算搭建一个国内旅游比价网站，
而我负责的部份是各大订酒店网站的数据获取及整理

主要参考版上的经验分享+自己的修改理解
小白0经验入门记录、边爬边学习ing
有错误或更好的建议都可以指教讨论

缤客Booking.com

采用python+request+beautifulsoup
目标获取酒店名称、星级、用户评分、评论数、最低价格

上个结果图
阳春的展示结果

在这里插入图片描述

正片开始

搜索上海的酒店，不用登入就能看到价格

url删减后可以剩以下参数
https://www.booking.com/searchresults.zh-cn.html?checkin_month={入住月份}&checkin_monthday={入住日期}&checkin_year={入住年份}&checkout_month={退房月份}&checkout_monthday={退房日期}&checkout_year={退房年份}&group_adults={大人数}&group_children={小孩数}&no_rooms={房间数}&ss={城市名}

在这里插入图片描述

翻到第二页，可以看到后面多了row&offset 在这里插入图片描述
再翻到第三页，可以看到offset变成50

所以可以靠更改offset获取每页的数据

ok开始写代码

先获取每页的资料

def getdata(city,checkin,checkout,start):
    url = 'https://www.booking.com/searchresults.zh-cn.html?'
    header ={
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36",
        }
    param ={
        'ss': city,
        'checkin_monthday': checkin[8:10],
        'checkin_year_month': checkin[0:7],
        'checkout_monthday': checkout[8:10],
        'checkout_year_month': checkout[0:7],
        'group_adults': 2,
        'group_children': 0,
        'no_rooms': 1,
    }
	#第一页
    if start==0:
        res = requests.get(url=url, headers=header, params=param)
        total_item = "".join(re.findall(r'b\_available\_hotels\:(.*?)\,', res.text))
        print('酒店总数',total_item)
        page = math.ceil(int(total_item) / 25)
        return res.text,page
    else:
        param['offset'] =start
        res = requests.get(url=url, headers=header, params=param)
        return res.text

再来用bs4来提取我们要的数据，现确认下元素在哪里

在这里插入图片描述

接下来直接提取就可以了

def printList(list,day):
    soup = BeautifulSoup(list, 'html.parser')
    Hotels = soup.select('.sr_property_block')

    for name in Hotels:
        text = name.select('.sr-hotel__name')[0].get_text().strip() + name.select('.bui-review-score__badge')[
            0].get_text().strip()
        str = 'https://www.booking.com/' + name.select('.hotel_name_link')[0]['href'].strip()
        print('----------------------------------------')
        print(text)
        print(name.select('.hotel_image')[0]['data-highres'])
        print(str.rstrip(';highlight_room=#hotelTmpl').strip())
        print(name.select('.bui-review-score__text')[0].get_text().strip())
        price =int(int(name.select('.bui-price-display__value')[0].get_text().strip().strip('元').replace(',',''))/day)
        print('CNY',price)

但星级的部份，因为不是每间酒店都有星级，用beautifulsoup获取的话匹配到，还在想有没有其他办法能解决

Reika_xiang

关注

0
点赞
踩
1

收藏

觉得还不错? 一键收藏
0
评论
创新项目实训(四)

创新项目实训(四)前言我们组打算搭建一个国内旅游比价网站，而我负责的部份是各大订酒店网站的数据获取及整理主要参考版上的经验分享+自己的修改理解小白0经验入门记录、边爬边学习ing有错误或更好的建议都可以指教讨论缤客Booking.com采用python+request目标获取酒店名称、星级、用户评分、评论数、最低价格...
复制链接

扫一扫

专栏目录