Hadoop统计、排序、去重代码

最新推荐文章于 2023-11-23 14:23:51 发布

轮胎技术Tyretek

最新推荐文章于 2023-11-23 14:23:51 发布

阅读量106

点赞数

文章标签： hadoop 大数据 java 分布式开发语言

本文链接：https://blog.csdn.net/weixin_42581846/article/details/129577560

版权

好的，在 Hadoop 中统计、排序、去重常常使用 MapReduce 来实现。下面是一个简单的例子：

统计单词出现的次数：

public static class TokenizerMapper
    extends Mapper<Object, Text, Text, IntWritable>{

  private final static IntWritable one = new IntWritable(1);
  private Text word = new Text();

  public void map(Object key, Text value, Context context
                  ) throws IOException, InterruptedException {
    StringTokenizer itr = new StringTokenizer(value.toString());
    while (itr.hasMoreTokens()) {
      word.set(itr.nextToken());
      context.write(word, one);
    }
  }
}

public static class IntSumReducer
    extends Reducer<Text,IntWritable,Text,IntWritable> {
  private IntWritable result = new IntWritable();

  public void reduce(Text key, Iterable<IntWritable> values,
                     Context context
                     ) throws IOException, InterruptedException {
    int sum = 0;
    for (IntWritable val : values) {
      sum += val.get();
    }
    result.set(sum);
    context.write(key, result);
  }
}

轮胎技术Tyretek

关注

0
点赞
踩
1

收藏

觉得还不错? 一键收藏
0
评论
Hadoop统计、排序、去重代码

好的，在 Hadoop 中统计、排序、去重常常使用 MapReduce 来实现。下面是一个简单的例子：统计单词出现的次数：public static class TokenizerMapper extends Mapper<Object, Text, Text, IntWritable>{ private final static IntWritable one = ...
复制链接

扫一扫