urllib.robotparser — Parser for robots.txt

最新推荐文章于 2023-12-08 16:06:51 发布

weixin_33974433

最新推荐文章于 2023-12-08 16:06:51 发布

阅读量90

点赞数

原文链接：http://www.cnblogs.com/gswang/p/7476009.html

版权

import urllib.robotparser
>>> rp = urllib.robotparser.RobotFileParser() >>> rp.set_url("http://www.musi-cal.com/robots.txt") >>> rp.read() >>> rrate = rp.request_rate("*") >>> rrate.requests 3 >>> rrate.seconds 20 >>> rp.crawl_delay("*") 6 >>> rp.can_fetch("*", "http://www.musi-cal.com/cgi-bin/search?city=San+Francisco") False >>> rp.can_fetch("*", "http://www.musi-cal.com/") True