爬虫（3）爬取数据再处理

最新推荐文章于 2023-08-26 08:45:00 发布

听说不挂科

最新推荐文章于 2023-08-26 08:45:00 发布

阅读量366

点赞数 1

分类专栏： python

本文链接：https://blog.csdn.net/qq_53029299/article/details/114850455

版权

python 专栏收录该内容

63 篇文章 9 订阅

订阅专栏

上次我们爬取了1960年世界的GDP
但是还是有一些数据需要去除的，比如空，还有有空格的地方，还有广告位等等，这里我们去除这些东西

from selenium import webdriver
from bs4 import BeautifulSoup

driver=webdriver.Chrome()
url="https://www.kylc.com/stats/global/yearly/g_gdp/1960.html"
xpath="/html/body/div[2]/div[1]/div[5]/div[1]/div/div/div/table"
driver.get(url)
tablel=driver.find_element_by_xpath(xpath).get_attribute('innerHTML')
soup=BeautifulSoup(tablel,"html.parser")
table=soup.find_all('tr')
for row in table:
    cols=[col.text for col in row.find_all('td')]
    if len(cols)==0 or not cols[0].isdigit():
        continue
    print(cols)

这里加了
if len(cols)==0 or not cols[0].isdigit():
continue
目的是去除空，列表第一个元素是空格的行，还有广告位
结果如下
在这里插入图片描述

听说不挂科

关注

1
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
爬虫（3）爬取数据再处理

上次我们爬取了1960年世界的GDP但是还是有一些数据需要去除的，比如空，还有有空格的地方，还有广告位等等，这里我们去除这些东西from selenium import webdriverfrom bs4 import BeautifulSoupdriver=webdriver.Chrome()url="https://www.kylc.com/stats/global/yearly/g_gdp/1960.html"xpath="/html/body/div[2]/div[1]/div[5]/
复制链接

扫一扫