python selenium 加载完毕_python+selenium获取页面加载的所有静态资源文件链接操作...

最新推荐文章于 2023-05-16 21:08:38 发布

李子坝的风

最新推荐文章于 2023-05-16 21:08:38 发布

阅读量594

点赞数

文章标签： Python3 Selenium 静态资源链接获取 Chrome无头浏览器

本文链接：https://blog.csdn.net/weixin_35252979/article/details/113312691

版权

这篇文章主要介绍了python3+selenium获取页面加载的所有静态资源文件链接操作，具有很好的参考价值，希望对大家有所帮助。一起跟随小编过来看看吧

直接上代码：

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.desired_capabilities import DesiredCapabilities

d = DesiredCapabilities.CHROME
chrome_options = Options()
#使用无头浏览器
chrome_options.add_argument('--headless')
chrome_options.add_argument('--user-agent=Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/71.0.3578.98 Safari/537.36')
#浏览器启动默认最大化
chrome_options.add_argument("--start-maximized");
#该处替换自己的chrome驱动地址
browser = webdriver.Chrome("D://googleDever//chromedriver.exe",chrome_options=chrome_options,desired_capabilities=d)
browser.set_page_load_timeout(150)
browser.get("https://www.xxx.com")
#静态资源链接存储集合
urls = []
#获取静态资源有效链接
for log in browser.get_log('performance'):
	 if 'message' not in log:
			continue
	 log_entry = json.loads(log['message'])
	 try:
		#该处过滤了data:开头的base64编码引用和document页面链接
			if "data:" not in log_entry['message']['params']['request']['url'] and 'Document' not in log_entry['message']['params']['type']:
				urls.append(log_entry['message']['params']['request']['url'])
	 except Exception as e:
			pass
 print(urls)

打印结果为页面渲染时加载的静态资源文件链接：