简言:学习了爬虫基础后对爬虫代码理解后进行编程。
收获:对于数据类型的了解更加深入,学习了txt文件的存储以及读取
摘要:
python爬取豆瓣网内容然后进行数据分析
编程
导入模块
import requests
import re
爬虫搭建
start = 0
result = []
headers = {
'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/89.0.4389.90 Safari/537.36 '
}
for i in range(0,10):
#250部电影,一页25部
html = requests.get('https://movie.douban.com/top250?start='+str(start)+'&filter=',headers=headers)
result.append(html.text)
start+=25
按F12,不难发现第2页start =25,第3页start=50,这里我没有看到第一页内容我推测第一页start=0.根据这个写一个for循环进行分页爬取
将result写入txt文件
with open('data/doubantop250_result.txt','w',encoding= 'utf-8') as f:
f.write(str(result))#将result化为字符串写入f文件
f = open('data/doubantop250_result.txt',encoding= 'utf-8')
data = f.readlines()
f.close()
print(data)
空dataframe
import pandas as pd
df = pd.DataFrame([])
name = []
year = []
n = []
score = []
obj = re.compile(r'<li>.*?<div class="item">.*?<span class="title">(?P<name>.*?)'
r'</span>.*?<p class="">.*?<br>(?P<year>.*?) .*?<span '
r'class="rating_num" property="v:average">(?P<score>.*?)</span>.*?'
r'<span>(?P<n>.*?)人评价</span>',re.S)
r = obj.finditer(str(result))
for i in r:
name.append(i.group('name'))
year.append(i.group('year').strip())
n.append(i.group('n'))
score.append(i.group('score'))
df['name'] = name
df['year'] = year
df['score'] = score
df['n'] = n
print(df)