写爬虫时,需要的html和用requests.get返回的html不一样导致后面用bs老出错
requests.get()获取不到正确的源代码HTML
# 1. 获取网页数据
url = 'https://movie.douban.com/top250'
headers = {
'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.51 Safari/537.36'
}
response = requests.get(url, headers=headers)
# 2. 解析数据
soup = BeautifulSoup(response.text, 'lxml')
这个不行
# 指定要爬取的网站
url = 'http://www.360doc.com/index.html?type=36&classid=19'
soup = getsoup(url)
print(soup)
# 错了这么多,soup中竟没有
imgList =soup.select('.c5_ul3>li') # 上下两标签内容 .class名>下级 标签
试了下下面的网址不行,更换headers一样不同:
#获取网页数据
url = 'https://movie.douban.com/tv'
headers = {
'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.51 Safari/537.36'
}
response = requests.get(url, headers=headers)
# 2. 解析数据
soup = BeautifulSoup(response.text, 'lxml')
这个库,没看出来为什么,有的网页可以,有的却是错的