初学python爬虫学习笔记——爬取网页中小说标题

最新推荐文章于 2025-02-25 15:00:00 发布

白芷加茯苓

最新推荐文章于 2025-02-25 15:00:00 发布

阅读量1.3k

点赞数 1

分类专栏： Python学习记录文章标签： python 爬虫学习

本文链接：https://blog.csdn.net/qq_43627631/article/details/132765448

版权

Python学习记录专栏收录该内容

7 篇文章

订阅专栏

初学python爬虫学习笔记——爬取网页中小说标题

一、要爬取的网站小说如下图

在这里插入图片描述

二、打开网页的“检查”，查看html页面

发现每个标题是列表下的一个个超链接，从183.html到869.html
可以使用for循环依次得到：

x = range(183,600)
for i in x:
    print(soup.find('a', href="http://www.kanxshuo.com/11/182/"+str(i)+".html").get_text())

在这里插入图片描述

三、具体代码如下：

import requests
import random
from bs4 import BeautifulSoup
# 要爬取的网站
url = "http://www.kanxshuo.com/11/182/"
# 发出访问请求，获得对应网页
response = requests.get(url)
print(response)

# 将获得的页面解析内容写入soup备用
soup = BeautifulSoup(response.content, 'lxml')

# 解析网站数据
# print(soup)

# 根据目标，首先要获得小说的标题和章节标题
# <a href="http://www.kanxshuo.com/11/182/211.html" title="第一卷 第二十九章 神祗遗闻">第一卷 第二十九章 神祗遗闻</a>
t1 = soup.find('a', href="http://www.kanxshuo.com/11/182/").get_text()
t2 = soup.find(id='booklistBox')
print(soup.find('a', href="http://www.kanxshuo.com/11/182/"+"183"+".html").get_text())
x = range(183,600)
for i in x:
    print(soup.find('a', href="http://www.kanxshuo.com/11/182/"+str(i)+".html").get_text())