Spark standalone 与 GlusterFS 配合使用

最新推荐文章于 2022-09-29 19:55:26 发布

FeelBreak

最新推荐文章于 2022-09-29 19:55:26 发布

阅读量327

点赞数

分类专栏：大数据 Spark

本文链接：https://blog.csdn.net/DANTE54/article/details/100702055

版权

大数据同时被 2 个专栏收录

4 篇文章 0 订阅

订阅专栏

Spark

4 篇文章 0 订阅

订阅专栏

Spark with glusterfs

测试设备架构

在这里插入图片描述

测试环境搭建过程

搭建GlusterFS，测试环境中用的是两个节点做GlusterFS，备份数是两份
搭建Spark Standalone环境，三台机器做Spark Standalone集群，其中每台GlusterFS上都要配置为Spark的Worker机
三台Spark机器上都要挂载glusterfs的文件到同一个目录
```
mount -t glusterfs slave1:/test /data/mnt/glusterfs/
```

Spark WordCount 代码

代码：

	package spark.wordcount
	import org.apache.spark.{SparkConf, SparkContext}
	object SparkWordCount {
		def main(args: Array[String]):Unit = {
			val conf = new SparkConf().setAppName("Spark Word Count")
			val sc = new SparkContext(conf)
			val startTime:Long = sc.startTime
			println(startTime)
			val words = sc.textFile("/data/mnt/glusterfs/test/")
			words.map(c => (c, 1)).reduceByKey(_ + _).collect().foreach(println(_))
		}
	}

提交jar包命令脚本：

	$SPARK_HOME/bin/spark-submit \
		--master spark://Master:7077 \
		--num-executors 2 \
		--class spark.wordcount.SparkWordCount \
		--conf spark.dynamicAllocation.enabled=false \
		ScalaWordCount.jar

Tips:

注意挂载glusterfs目录到每一个Spark节点，Master也必须挂载
注意提交任务时，添加参数–conf spark.dynamicAllocation.enabled=false，目的是为了如下问题解决问题

FeelBreak

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
Spark standalone 与 GlusterFS 配合使用

Spark with glusterfs测试设备架构测试环境搭建过程搭建GlusterFS，测试环境中用的是两个节点做GlusterFS，备份数是两份搭建Spark Standalone环境，三台机器做Spark Standalone集群，其中每台GlusterFS上都要配置为Spark的Worker机三台Spark机器上都要挂载glusterfs的文件到同一个目录mount -t ...
复制链接

扫一扫

专栏目录