给定数据如下:
班级ID 姓名 年龄 性别 科目 成绩 12 张三 25 男 chinese 50 12 张三 25 男 math 60 12 张三 25 男 english 70 12 李四 20 男 chinese 50 12 李四 20 男 math 50 12 李四 20 男 english 50 12 王芳 19 女 chinese 70 12 王芳 19 女 math 70 12 王芳 19 女 english 70 13 张大三 25 男 chinese 60 13 张大三 25 男 math 60 13 张大三 25 男 english 70 13 李大四 20 男 chinese 50 13 李大四 20 男 math 60 13 李大四 20 男 english 50 13 王小芳 19 女 chinese 70 13 王小芳 19 女 math 80 13 王小芳 19 女 english 70 |
需求:
1. 一共有多少人参加考试? 2. 一共有多个男生参加考试? 3. 12班有多少人参加考试? 4. 语文科目的平均成绩是多少? 5. 单个人平均成绩是多少? 6. 12班平均成绩是多少? 7. 全校语文成绩最高分是多少? 8. 总成绩大于150分的12班的女生有几个? 9. 总成绩大于150分,且数学大于等于70,且年龄大于等于20岁的学生的平均成绩是多少? |
这里收集了2个人做的方式,其中有的会重复,但是主要是看方法
方式一:
需求如下:
1. 一共有多少人参加考试?
val file = sc.textFile("file:///jar/score")
val name = file.map(x => {val line = x.split(" ");line(0) + "," + line(1)})
val numPeo = name.distinct.count()
1.1 一共有多少个小于20岁的人参加考试?
val file = sc.textFile("file:///jar/score")
val age = file.map(x => {val line = x.split(" ");line(0) + "," + line(1) + "," + line(2)})
val numPeo = age.distinct.filter(_.split(",")(2).toInt<20).count()
1.2 一共有多少个等于20岁的人参加考试?
val file = sc.textFile("file:///jar/score")
val age = file.map(x => {val line = x.split(" ");line(0) + "," + line(1) + "," + line(2)})
val numPeo = age.distinct.filter(_.split(",")(2).toInt == 20).count()
1.3 一共有多少个大于20岁的人参加考试?
val file = sc.textFile("file:///jar/score")
val age = file.map(x => {val line = x.split(" ");line(0) + "," + line(1) + "," + line(2)})
val numPeo = age.distinct.filter(_.split(",")(2).toInt == 20).count()
2. 一共有多个男生参加考试?
val file = sc.textFile("file:///jar/score")
val sex = file.map(x => {val line = x.split(" ");line(0) + "," + line(1) + "," + line(3)})
val numPeo = sex.distinct.filter(_.split(",")(2) == "男").count()
2.1 一共有多少个女生参加考试?
val file = sc.textFile("file:///jar/score")
val sex = file.map(x => {val line = x.split(" ");line(0) + "," + line(1) + "," + line(3)})
val numPeo = sex.distinct.filter(_.split(",")(2) == "女").count()
3. 12班有多少人参加考试?
val file = sc.textFile("file:///jar/score")
val classNum = file.map(x => {val line = x.split(" ");line(0) + "," + line(1) })
val numPeo = classNum.distinct.filter(_.split(",")(0).toInt == 12).count()
sc.makeRDD(Array(numPeo)).saveAsTextFile("file:///jar/result/class12numPeo")
3.1 13班有多少人参加考试?
val file = sc.textFile("file:///jar/score")
val classNum = file.map(x => {val line = x.split(" ");line(0) + "," + line(1) })
val numPeo = classNum.distinct.filter(_.split(",")(0).toInt == 13).count()
sc.makeRDD(Array(numPeo)).saveAsTextFile("file:///jar/result/class13numPeo")
4. 语文科目的平均成绩是多少?
val chineseLine = file.map(x => {val line = x.split(" "); line(4)+ "," + line(5)})
val chineseGennal = chineseLine.filter(_.split(",")(0) == "chinese")
val chineseLength = chineseGennal.count.toInt//6
val chineseSum = chineseGennal.map(_.split(",")(1).toInt).reduce(_ + _)//350
val chineseAvg = chineseSum/chineseLength//58
sc.makeRDD(Array(chineseGennal.map(_.split(",")(1).toInt)
.reduce(_ + _)/chineseGennal.count.toInt))
.saveAsTextFile("file:///jar/result/chineseAvg")
4.1 数学科目的平均成绩是多少?
val mathLine = file.map(x => {val line = x.split(" "); line(4)+ "," + line(5)})
val mathGennal = mathLine.filter(_.split(",")(0) == "math")
val mathLength = mathGennal.count.toInt
val mathSum = mathGennal.map(_.split(",")(1).toInt).reduce(_ + _)
val mathAvg = mathSum/mathLength
sc.makeRDD(Array(mathGennal.map(_.split(",")(1).toInt)
.reduce(_ + _)/mathGennal.count.toInt))
.saveAsTextFile("file:///jar/result/mathAvg")
4.2 英语科目的平均成绩是多少?
val englishLine = file.map(x => {val line = x.split(" "); line(4)+ "," + line(5)})
val englishGennal = englishLine.filter(_.split(",")(0) == "english")
val englishLength = englishGennal.count.toInt
val englishSum = englishGennal.map(_.split(",")(1).toInt).reduce(_ + _)
val englishAvg = englishSum/englishLength
sc.makeRDD(Array(englishGennal.map(_.split(",")(1).toInt)
.reduce(_ + _)/englishG