Spark2.3 RDD之 distinct 源码浅谈

最新推荐文章于 2022-04-22 22:09:40 发布

DPnice

最新推荐文章于 2022-04-22 22:09:40 发布

阅读量2.1k

点赞数

分类专栏： spark 文章标签： distinct spark scala

本文链接：https://blog.csdn.net/DPnice/article/details/80097520

版权

spark 专栏收录该内容

19 篇文章 0 订阅

订阅专栏

distinct 源码：

/**

 * Return a new RDD containing the distinct elements in this RDD.
 */
def distinct(numPartitions: Int)(implicit ord: Ordering[T] = null): RDD[T] = withScope {
  map(x => (x, null)).reduceByKey((x, y) => x, numPartitions).map(_._1)
}

/**
 * Return a new RDD containing the distinct elements in this RDD.
 */
def distinct(): RDD[T] = withScope {
  distinct(partitions.length)
}

这个去重算子比较直观了，其实就是把map 和 reduceByKey 封装了一下。distinct有一个可选参数numPartitions，这个参数是你期望的分区数。

例子：

object DistinctTest extends App {

  val sparkConf = new SparkConf().
    setAppName("TreeAggregateTest")
    .setMaster("local[6]")

  val spark = SparkSession
    .builder()
    .config(sparkConf)
    .getOrCreate()

  val value: RDD[Int] = spark.sparkContext.parallelize(List(1, 2, 3, 5, 8, 9), 3)
  println(value.distinct(1).getNumPartitions)
}

最后的结果分区被重置为1。

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

立即使用

DPnice

关注关注

0
点赞
踩
1

收藏

觉得还不错? 一键收藏
0
评论
Spark2.3 RDD之 distinct 源码浅谈

distinct 源码：/** * Return a new RDD containing the distinct elements in this RDD. */def distinct(numPartitions: Int)(implicit ord: Ordering[T] = null): RDD[T] = withScope { map(x =&gt; (x, null))...
复制链接

扫一扫