spark算子中reduceByKey和groupByKey两者的区别

最新推荐文章于 2023-10-28 08:30:00 发布

一过人_

最新推荐文章于 2023-10-28 08:30:00 发布

阅读量1.7k

点赞数

分类专栏： spark 源码分析文章标签： spark 大数据

本文链接：https://blog.csdn.net/newhandnew413/article/details/107732775

版权

spark 同时被 2 个专栏收录

10 篇文章 1 订阅

订阅专栏

源码分析

10 篇文章 1 订阅

订阅专栏

spark中算子应该是重点中的重点了，今天我们来分析一下两个算子reduceByKey和groupByKey

这两个算子都属于k-v类型的算子

我们先来看看这两个算子的作用是什么？

reduceByKey是通过key对数据进行聚合

groupByKey是通过key对数据进行分组

这两个都需要对数据进行打乱重组，所以都会有shuffle

两者的区别：

reduceByKey：在shuffle之前有combine（预聚合）操作，返回结果是RDD[k,v]。
groupByKey：直接进行shuffle。

我们来看看reduceByKey的源码，为什么会在shuffle之前有预聚合的操作

def reduceByKey(func: (V, V) => V): RDD[(K, V)] = self.withScope {
    reduceByKey(defaultPartitioner(self), func)
  }

我们看到方法里又进行了一步操作，再进去看看

  def reduceByKey(partitioner: Partitioner, func: (V, V) => V): RDD[(K, V)] = self.withScope {
    combineByKeyWithClassTag[V]((v: V) => v, func, func, partitioner)
  }

我们看groupByKey的源码

  def groupByKey(partitioner: Partitioner): RDD[(K, Iterable[V])] = self.withScope {
    // groupByKey shouldn't use map side combine because map side combine does not
    // reduce the amount of data shuffled and requires all map side data be inserted
    // into a hash table, leading to more objects in the old gen.
    val createCombiner = (v: V) => CompactBuffer(v)
    val mergeValue = (buf: CompactBuffer[V], v: V) => buf += v
    val mergeCombiners = (c1: CompactBuffer[V], c2: CompactBuffer[V]) => c1 ++= c2
    val bufs = combineByKeyWithClassTag[CompactBuffer[V]](
      createCombiner, mergeValue, mergeCombiners, partitioner, mapSideCombine = false)
    bufs.asInstanceOf[RDD[(K, Iterable[V])]]
  }