实验内容:
对于两个输入文件A和B,编写Spark独立程序,对两个文件进行合并,并剔除其中重复的内容,得到一个新文件C。下面是输入文件和输出文件的样例:
输入文件A的样例如下:
20170101 x
20170102 y
20170103 x
20170104 y
20170105 z
20170106 z
输入文件B的样例如下:
20170101 y
20170102 y
20170103 x
20170104 z
20170105 y
根据输入的文件A和B合并得到的输出文件C的样例如下:
20170101 x
20170101 y
20170102 y
20170103 x
20170104 y
20170104 z
20170105 y
20170105 z
20170106 z
通过使用sbt工具将整个应用程序打包成jar包,并将jar包通过spark-submit提交到spark中运行。
代码:
import org.apache.spark.SparkContext
import org.apache.spark.SparkContext._
import org.apache.spark.SparkConf
import org.apache.spark.HashPartitioner
object FileMerge{
def main(args: Array[String]) {
val logFile = "file:///home/hadoop/A.txt,file:///home/hadoop/B.txt"
val conf = new SparkConf().setAppName("FileMerge")
val sc = new SparkContext(conf)
val logData = sc.textFile(logFile,2)
val res = logData.distinct()
res.saveAsTextFile("/home/hadoop/C.txt")
}
}
simple.sbt
name :="Simple Project"
version := "1.0"
scalaVersion := "2.12.10"
libraryDependencies += "org.apache.spark" %% "spark-core" % "3.1.2"
打包运行:
# /usr/local/sbt/sbt package
# spark-submit --class "FileMerge" ./target/scala-2.12/simple-project_2.12-1.0.jar