php爬冲ajax,python - 关于scrapy爬虫AJAX页面

weixin_39569543

于 2021-08-06 11:00:41 发布

阅读量126

点赞数

文章标签： php爬冲ajax

问题：

爬取信息页面为：知乎话题广场

当点击加载的时候，用Chrome 开发者工具，可以看到Network中，实际请求的链接为：

FormData为：

urlencode：

然后我的代码为：

...

data = response.css('.zh-general-list::attr(data-init)').extract()

param = json.loads(data[0])

topic_id = param['params']['topic_id']

# hash_id = param['params']['hash_id']

hash_id = ""

for i in range(32):

params = json.dumps({"topic_id": topic_id,"hash_id": hash_id, "offset":i*20})

payload = {"method": "next", "params": params, "_xsrf":_xsrf}

print payload

yield scrapy.Request(

url="https://m.zhihu.com/node/TopicsPlazzaListV2?" + urlencode(payload),

headers=headers,

meta={

"proxy": proxy,

"cookiejar": response.meta["cookiejar"],

callback=self.get_topic_url,

)

执行爬虫之后，返回的是：

{'_xsrf': u'161c70f5f7e324b92c2d1a6fd2e80198', 'params': '{"hash_id": "", "offset": 140, "topic_id": 253}', 'method': 'next'}

^C^C^C{'_xsrf': u'161c70f5f7e324b92c2d1a6fd2e80198', 'params': '{"hash_id": "", "offset": 160, "topic_id": 253}', 'method': 'next'}

2016-05-09 11:09:36 [scrapy] DEBUG: Retrying (failed 1 times): []

是不是url哪里错了？请指点一二。

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

关注关注