spark2.3源码分析之RDD算子combineByKey

最新推荐文章于 2021-04-08 22:52:14 发布

zhifeng687

最新推荐文章于 2021-04-08 22:52:14 发布

阅读量313

点赞数

分类专栏： spark

本文链接：https://blog.csdn.net/qq_26222859/article/details/54099062

版权

spark 专栏收录该内容

30 篇文章 4 订阅

订阅专栏

概述

combineByKey通过使用聚合函数，把键值对中每一个key对应的value进行聚合，将JavaPairRDD[(K, V)]转换成JavaPairRDD[(K, C)]的结果，C是聚合后的类型。

用户需要提供以下三个函数：

createCombiner：在一个partition中创建每个key对应的累加器的初始值。
mergeValue：在一个partition中，将同一个key对应的value合并该key对应的累加器中。
mergeCombiners：合并不同partition中，同一个key对应的累加器。

除此之外，用户可以定义RDD的partitioner，在shuffle过程中使用的serializer，一级是否使用map端的聚合。

combineByKey函数

def combineByKey[C](createCombiner: JFunction[V, C],
      mergeValue: JFunction2[C, V, C],
      mergeCombiners: JFunction2[C, C, C],
      partitioner: Partitioner,
      mapSideCombine: Boolean,
      serializer: Serializer): JavaPairRDD[K, C] = {
      implicit val ctag: ClassTag[C] = fakeClassTag
    fromRDD(rdd.combineByKeyWithClassTag(
      createCombiner,
      mergeValue,
      mergeCombiners,
      partitioner,
      mapSideCombine,
      serializer
    ))
  }

combineByKeyWithClassTag创建Aggregator和ShuffledRDD

def combineByKeyWithClassTag[C](
      createCombiner: V => C,
      mergeValue: (C, V) => C,
      mergeCombiners: (C, C) => C,
      partitioner: Partitioner,
      mapSideCombine: Boolean = true,
      serializer: Serializer = null)(implicit ct: ClassTag[C]): RDD[(K, C)] = self.withScope {
    require(mergeCombiners != null, "mergeCombiners must be defined") // required as of Spark 0.9.0
    if (keyClass.isArray) {
      if (mapSideCombine) {
        throw new SparkException("Cannot use map-side combining with array keys.")
      }
      if (partitioner.isInstanceOf[HashPartitioner]) {
        throw new SparkException("HashPartitioner cannot partition array keys.")
      }
    }
//创建Aggregator
    val aggregator = new Aggregator[K, V, C](
      self.context.clean(createCombiner),
      self.context.clean(mergeValue),
      self.context.clean(mergeCombiners))
    if (self.partitioner == Some(partitioner)) {
      self.mapPartitions(iter => {
        val context = TaskContext.get()
        new InterruptibleIterator(context, aggregator.combineValuesByKey(iter, context))
      }, preservesPartitioning = true)
    } else {
//创建ShuffledRDD
      new ShuffledRDD[K, V, C](self, partitioner)
        .setSerializer(serializer)
        .setAggregator(aggregator)
        .setMapSideCombine(mapSideCombine)
    }
  }

Aggregator

Aggregator基于createCombiner、mergeValue、mergeCombiners函数和ExternalAppendOnlyMap数据结构实现数据聚合。

/**
 * :: DeveloperApi ::
 * A set of functions used to aggregate data.
 *
 * @param createCombiner function to create the initial value of the aggregation.
 * @param mergeValue function to merge a new value into the aggregation result.
 * @param mergeCombiners function to merge outputs from multiple mergeValue function.
 */
@DeveloperApi
case class Aggregator[K, V, C] (
    createCombiner: V => C,
    mergeValue: (C, V) => C,
    mergeCombiners: (C, C) => C) {

//根据相同的key合并它们的value
  def combineValuesByKey(
      iter: Iterator[_ <: Product2[K, V]],
      context: TaskContext): Iterator[(K, C)] = {
    val combiners = new ExternalAppendOnlyMap[K, V, C](createCombiner, mergeValue, mergeCombiners)
    combiners.insertAll(iter)
    updateMetrics(context, combiners)
    combiners.iterator
  }

//根据相同的key合并对应的combiner
  def combineCombinersByKey(
      iter: Iterator[_ <: Product2[K, C]],
      context: TaskContext): Iterator[(K, C)] = {
    val combiners = new ExternalAppendOnlyMap[K, C, C](identity, mergeCombiners, mergeCombiners)
    combiners.insertAll(iter)
    updateMetrics(context, combiners)
    combiners.iterator
  }

}

zhifeng687

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
spark2.3源码分析之RDD算子combineByKey

概述combineByKey通过使用聚合函数，把键值对中每一个key对应的value进行聚合，将JavaPairRDD[(K, V)]转换成JavaPairRDD[(K, C)]的结果，C是聚合后的类型。用户需要提供以下三个函数：createCombiner：在一个partition中创建每个key对应的累加器的初始值。 mergeValue：在一个partition中，将同一个ke...
复制链接

扫一扫

专栏目录