使用轻量级JAVA 爬虫Gecco工具抓取新闻DEMO

最新推荐文章于 2022-09-20 13:26:38 发布

赵侠客

最新推荐文章于 2022-09-20 13:26:38 发布

阅读量5.5k

点赞数 3

分类专栏：搜索引擎 Java 文章标签： java 爬虫 gecco 轻量

本文链接：https://blog.csdn.net/whzhaochao/article/details/51096367

版权

Java 同时被 2 个专栏收录

68 篇文章 1 订阅

订阅专栏

搜索引擎

12 篇文章 0 订阅

订阅专栏

写在前面

最近看到Gecoo爬虫工具，感觉比较简单好用，所有写个DEMO测试一下，抓取网站
http://zj.zjol.com.cn/home.html，主要抓取新闻的标题和发布时间做为抓取测试对象。抓取HTML节点通过像Jquery选择器一样选择节点，非常方便，Gecco代码主要利用注解实现来实现URL匹配，看起来比较简洁美观。

Gecoo GitHub地址
https://github.com/xtuhcy/gecco
Gecoo 作者博客
http://my.oschina.net/u/2336761/blog?fromerr=ZuKKo3fH

添加Maven依赖

        <dependency>
            <groupId>com.geccocrawler</groupId>
            <artifactId>gecco</artifactId>
            <version>1.0.8</version>
        </dependency>

编写抓取列表页面

@Gecco(matchUrl = "http://zj.zjol.com.cn/home.html?pageIndex={pageIndex}&pageSize={pageSize}",pipelines = "zJNewsListPipelines")
public class ZJNewsGeccoList implements HtmlBean {
    @Request
    private HttpRequest request;
    @RequestParameter
    private int pageIndex;
    @RequestParameter
    private int pageSize;
    @HtmlField(cssPath = "#content > div > div > div.con_index > div.r.main_mod > div > ul > li  > dl > dt > a")
    private List<HrefBean> newList;
}


@PipelineName("zJNewsListPipelines")
public class ZJNewsListPipelines implements Pipeline<ZJNewsGeccoList> {
    public void process(ZJNewsGeccoList zjNewsGeccoList) {
        HttpRequest request=zjNewsGeccoList.getRequest();
        for (HrefBean bean:zjNewsGeccoList.getNewList()){
            //进入祥情页面抓取
       SchedulerContext.into(request.subRequest("http://zj.zjol.com.cn"+bean.getUrl()));
        }
        int page=zjNewsGeccoList.getPageIndex()+1;
        String nextUrl = "http://zj.zjol.com.cn/home.html?pageIndex="+page+"&pageSize=100";
        //抓取下一页
        SchedulerContext.into(request.subRequest(nextUrl));
    }
}

编写抓取祥情页面


@Gecco(matchUrl = "http://zj.zjol.com.cn/news/{code}.html" ,pipelines = "zjNewsDetailPipeline")
public class ZJNewsDetail implements HtmlBean {

    @Text
    @HtmlField(cssPath = "#headline")
    private String title ;

    @Text
    @HtmlField(cssPath = "#content > div > div.news_con > div.news-content > div:nth-child(1) > div > p.go-left.post-time.c-gray")
    private String createTime;
}

@PipelineName("zjNewsDetailPipeline")
public class ZJNewsDetailPipeline implements Pipeline<ZJNewsDetail> {
    public void process(ZJNewsDetail zjNewsDetail) {
        System.out.println(zjNewsDetail.getTitle()+"  "+zjNewsDetail.getCreateTime());
    }
}

启动主函数


public class Main {
    public static void main(String [] rags){
        GeccoEngine.create()
                //工程的包路径
                .classpath("com.zhaochao.gecco.zj")
                //开始抓取的页面地址
                .start("http://zj.zjol.com.cn/home.html?pageIndex=1&pageSize=100")
                //开启几个爬虫线程
                .thread(10)
                //单个爬虫每次抓取完一个请求后的间隔时间
                .interval(10)
                //使用pc端userAgent
                .mobile(false)
                //开始运行
                .run();
    }
}

抓取结果

这里写图片描述

项目完成代码

http://git.oschina.net/whzhaochao/geccoDemo

赵侠客

关注

3
点赞
踩
2

收藏

觉得还不错? 一键收藏
1
评论
使用轻量级JAVA 爬虫Gecco工具抓取新闻DEMO

写在前面最近看到Gecoo爬虫工具，感觉比较简单好像，所有写个DEMO测试一下，抓取网站 http://zj.zjol.com.cn/home.html，主要抓取新闻的标题和发布时间做为抓取测试对象。Gecoo GitHub地址 https://github.com/xtuhcy/gecco Gecoo 作者博客 http://my.oschina.net/u/2336761/blog?fr
复制链接

扫一扫