MapReduce使用PathFilter进行文件过滤

ctp6666999

于 2020-11-15 14:39:03 发布

阅读量640

点赞数

分类专栏：大数据文章标签： mapreduce 过滤器 hadoop

本文链接：https://blog.csdn.net/qq_44776813/article/details/109703695

版权

前言

在使用MapReduce对数据进行处理的过程中，难免会遇到在一个文件夹中避开某类文件的问题，在本篇博客中我们使用PathFilter路径过滤器过滤掉*.txt文件。这里使用词频统计来做一个简单的小例子.

一、定义Mapper类

import org.apache.hadoop.io.IntWritable;
import org.apache.hadoop.io.LongWritable;
import org.apache.hadoop.io.Text;
import org.apache.hadoop.mapreduce.Mapper;

import java.io.IOException;

public class WcMapper extends Mapper<LongWritable, Text,Text, IntWritable> {
   
    Text k = new Text();
    IntWritable v = new IntWritable(1);
    @Override
    protected void map(LongWritable key, Text value, Context context) throws IOException, InterruptedException {
   
        String line = value.toString();
        String[] words = line.split(" ");
        for (String word : words) {
   
            k