Spark处理数据如何获得行号

最新推荐文章于 2022-06-17 09:17:58 发布

白杨树

最新推荐文章于 2022-06-17 09:17:58 发布

阅读量3.1k

点赞数

分类专栏：大数据(Big Data) 文章标签： spark

本文链接：https://blog.csdn.net/hongchangfirst/article/details/80175839

版权

因为Spark并行的处理数据，所以你不能在自己的driver program中计数到底是处理到第几个。Spark提供了zipWithIndex可以给你提供索引号。这个索引号是全局有序和唯一的。

public RDD<scala.Tuple2<T,Object>> zipWithIndex()
Zips this RDD with its element indices. The ordering is first based on the partition index and then the ordering of items within each partition. So the first item in the first partition gets index 0, and the last item in the last partition receives the largest index.
This is similar to Scala's zipWithIndex but it uses Long instead of Int as the index type. This method needs to trigger a spark job when this RDD contains more than one partitions.

Note that some RDDs, such as those returned by groupBy(), do not guarantee order of elements in a partition. The index assigned to each element is therefore not guaranteed, and may even change if the RDD is reevaluated. If a fixed ordering is required to guarantee the same index assignments, you should sort the RDD with sortByKey() or save i

最低0.47元/天解锁文章

白杨树

关注

0
点赞
踩
1

收藏

觉得还不错? 一键收藏
0
评论
Spark处理数据如何获得行号

因为Spark并行的处理数据，所以你不能在自己的driver program中计数到底是处理到第几个。Spark提供了zipWithIndex可以给你提供索引号。这个索引号是全局有序和唯一的。public RDD&lt;scala.Tuple2&lt;T,Object&gt;&gt; zipWithIndex()Zips this RDD with its element indices. The...
复制链接

扫一扫

专栏目录