WebMagic 爬虫框架使用教程

最新推荐文章于 2024-08-23 09:09:48 发布

孔朦煦

最新推荐文章于 2024-08-23 09:09:48 发布

阅读量306

点赞数 4

本文链接：https://blog.csdn.net/gitblog_00849/article/details/141015750

版权

WebMagic 爬虫框架使用教程

项目地址:https://gitcode.com/gh_mirrors/we/webmagic

项目介绍

WebMagic 是一个简单灵活的 Java 爬虫框架，它基于模块化设计，易于扩展和维护。WebMagic 提供了简洁的 API，使得开发者可以快速上手，并且支持多线程和分布式爬取。

项目快速启动

环境准备

Java 开发环境
Maven 依赖管理工具

添加 Maven 依赖

在你的 pom.xml 文件中添加以下依赖：

<dependency>
    <groupId>us.codecraft</groupId>
    <artifactId>webmagic-core</artifactId>
    <version>0.7.3</version>
</dependency>
<dependency>
    <groupId>us.codecraft</groupId>
    <artifactId>webmagic-extension</artifactId>
    <version>0.7.3</version>
</dependency>

编写爬虫代码

以下是一个简单的爬虫示例：

import us.codecraft.webmagic.Page;
import us.codecraft.webmagic.Site;
import us.codecraft.webmagic.Spider;
import us.codecraft.webmagic.processor.PageProcessor;

public class GithubRepoPageProcessor implements PageProcessor {
    private Site site = Site.me().setRetryTimes(3).setSleepTime(1000).setTimeOut(10000);

    @Override
    public void process(Page page) {
        page.addTargetRequests(page.getHtml().links().regex("(https://github\\.com/[\\w\\-]+/[\\w\\-]+)").all());
        page.addTargetRequests(page.getHtml().links().regex("(https://github\\.com/[\\w\\-])").all());
        page.putField("author", page.getUrl().regex("https://github\\.com/(\\w+)/ *").toString());
        page.putField("name", page.getHtml().xpath("//h1[@class='entry-title public']/strong/a/text()").toString());
        if (page.getResultItems().get("name") == null) {
            // 跳过这个页面
            page.setSkip(true);
        }
    }

    @Override
    public Site getSite() {
        return site;
    }

    public static void main(String[] args) {
        Spider.create(new GithubRepoPageProcessor())
                .addUrl("https://github.com/code4craft")
                .thread(5)
                .run();
    }
}