爬虫学的好，牢饭吃得好（爬虫实例）

大模型官方资料

已于 2024-04-18 15:41:07 修改

阅读量1k

点赞数 22

文章标签：爬虫开发语言 python

于 2024-04-18 15:40:57 首次发布

本文链接：https://blog.csdn.net/xzp740813/article/details/137924037

版权

鉴于本人喜欢爬虫，最近看了一些爬虫的基础，几个爬虫入门实例。下面给你们看，大佬勿喷

主要知识点:

1.标题web是如何交互的
2.requests库的get、post函数的应用
3.response对象的相关函数，属性
4.python文件的打开，保存

好，接下来先安装requests库
在pycharm命令行输入

pip install requests

安装好了以后咱先爬个baidu首页

# 爬虫示例,爬取百度页面

import requests #导入爬虫的库，不然调用不了爬虫的函数

response = requests.get("http://www.baidu.com")  #生成一个response对象

response.encoding = response.apparent_encoding #设置编码格式

print("状态码:"+ str( response.status_code ) ) #打印状态码

print(response.text)#输出爬取的信息

get方法实例

# get方法实例

import requests #先导入爬虫的库，不然调用不了爬虫的函数

response = requests.get("http://httpbin.org/get")  #get方法

print( response.status_code ) #状态码

print( response.text )

post方法实例

# post方法实例

import requests #先导入爬虫的库，不然调用不了爬虫的函数

response = requests.post("http://httpbin.org/post")  #post方法访问

print( response.status_code ) #状态码

print( response.text )

get传参方法实例

#  get传参方法实例

import requests #先导入爬虫的库，不然调用不了爬虫的函数

response = requests.get("http://httpbin.org/get?name=hezhi&age=20")  # get传参

print( response.status_code ) #状态码

print( response.text )

post传参方法实例

#  post传参方法实例

import requests #先导入爬虫的库，不然调用不了爬虫的函数

data = {
	"name":"hezhi",
	"age":20
}
response = requests.post( "http://httpbin.org/post" , params=data )  # post传参

print( response.status_code ) #状态码

print( response.text )

绕过反爬机制，以zhihu为例


import requests #先导入爬虫的库，不然调用不了爬虫的函数

response = requests.get( "http://www.zhihu.com")  #第一次访问知乎，不设置头部信息

print( "第一次,不设头部信息,状态码:"+response.status_code )# 没写headers，不能正常爬取，状态码不是 200

#下面是可以正常爬取的区别，更改了User-Agent字段

headers = {

		"User-Agent":"Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/80.0.3987.122 Safari/537.36"

}#设置头部信息,伪装浏览器

response = requests.get( "http://www.zhihu.com" , headers=headers )  #get方法访问,传入headers参数，

print( response.status_code ) # 200！访问成功的状态码

print( response.text )

保存百度图片到本地

#保存百度图片到本地

import requests #先导入爬虫的库，不然调用不了爬虫的函数

response = requests.get("https://www.baidu.com/img/baidu_jgylogo3.gif")  #get方法的到图片响应

file = open("D:\\爬虫\\baidu_logo.gif","wb") #打开一个文件,wb表示以二进制格式打开一个文件只用于写入

file.write(response.content) #写入文件

file.close()#关闭操作，运行完毕后去你的目录看一眼有没有保存成功

愿你早日成为爬虫大佬

如果大家对Python感兴趣，这套python学习资料一定对你有用

对于0基础小白入门：