代码内容 By ywz2008008
上一篇内容读出了页面的内容,虽然乱乱的一团,但还是成功了。
今天继续读出精准的内容,重点用到了:BeautifulSoup、lxml模块。
#!/usr/bin/python
# -*- coding: UTF-8 -*-
import requests
from bs4 import BeautifulSoup
movie_url = 'https://movie.douban.com/subject/1292052/'
def download_page(url):
headers = {
'User-Agent':'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_12)AppleWebKit/537.36 (KHTML, like Gecko) Chrome/47.0.2526.80 Safari/537.36'
}
data = requests.get(url, headers = headers).content
return data
def paser_html(html):
soup = BeautifulSoup(html, 'lxml')
title = soup.find(property =