Python爬虫实战--爬取天气数据保存

最新推荐文章于 2024-07-24 14:36:42 发布

ChrisKyrie

最新推荐文章于 2024-07-24 14:36:42 发布

阅读量1.6k

点赞数 1

分类专栏： Python

本文链接：https://blog.csdn.net/sinat_28826891/article/details/104502510

版权

Python 专栏收录该内容

4 篇文章 0 订阅

订阅专栏

个人总结的爬虫（爬取数据）的简单步骤：

1、获取待爬取网页的html信息

2、解析爬取的html信息，得到相关的数据

3、保存数据

# coding : UTF-8


import requests
import csv
import random
import time
import socket
import http.client
from bs4 import BeautifulSoup


def get_content(url, data=None):
    header = {
        'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8,',
        'Accept-Encoding': 'gzip, deflate',
        'Accept-Language': 'zh-CN,zh;q=0.9',
        'Connection': 'keep-alive',
        'User-Agent': 'Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/78.0.3904.70'
    }
    timeout = random.choice(range(80, 180))
    while True:
        try:
            rep = requests.get(url, headers=header, timeout=timeout)
            rep.encoding = 'utf-8'
            break
        # 异常处理
        except socket.timeout as e:
            print('3:', e)
            time.sleep(random.choice(range(8, 15)))

        except socket.error as e:
            print('4:', e)
            time.sleep(random.choice(range(20, 60)))

        except http.client.BadStatusLine as e:
            print('5:', e)
            time.sleep(random.choice(range(30, 80)))

        except http.client.IncompleteRead as e:
            print('6:', e)
            time.sleep(random.choice(range(5, 15)))

    return rep.text


def get_data(html_text):
    final = []
    bs = BeautifulSoup(html_text, "html.parser")
    body = bs.body
    data = body.find('div', {'class': 'c7d'})
    ul = data.find('ul')
    li = ul.find_all('li')

    for day in li:
        temp = []
        date = day.find('h1').string
        temp.append(date)
        inf = day.find_all('p')
        temp.append(inf[0].string, )
        if inf[1].find('span') is None:
            temperature_highest = None
        else:
            temperature_highest = inf[1].find('span').string
            temperature_highest = temperature_highest.replace('℃', '')
        temperature_lowest = inf[1].find('i').string
        temperature_lowest = temperature_lowest.replace('℃', '')
        temp.append(temperature_highest)
        temp.append(temperature_lowest)
        wind = inf[2].find('i').string #获取风的级数
        temp.append(wind)
        final.append(temp)

    return final


def write_data(data, name):
    file_name = name
    with open(file_name, 'a', errors='ignore', newline='') as f:
        f_csv = csv.writer(f)
        f_csv.writerows(data)


if __name__ == '__main__':
    url = 'http://www.weather.com.cn/weather/101220802.shtml'
    html = get_content(url)
    result = get_data(html)
    write_data(result, 'weather.csv')

参考博客：https://www.cnblogs.com/chengxuyuanaa/p/11955775.html

ChrisKyrie

关注

1
点赞
踩
12

收藏

觉得还不错? 一键收藏
0
评论
Python爬虫实战--爬取天气数据保存

个人总结的爬虫（爬取数据）的简单步骤：1、获取待爬取网页的html信息2、解析爬取的html信息，得到相关的数据3、保存数据# coding : UTF-8import requestsimport csvimport randomimport timeimport socketimport http.clientfrom bs4 import Beautifu...
复制链接

扫一扫

专栏目录