Steps in Setting hadoop in GHC

最新推荐文章于 2022-07-25 19:44:07 发布

niniumc

最新推荐文章于 2022-07-25 19:44:07 发布

阅读量649

点赞数

分类专栏： machine learning

本文链接：https://blog.csdn.net/nnumc/article/details/19958595

版权

machine learning 专栏收录该内容

6 篇文章 0 订阅

订阅专栏

After spending some time on it, I finally could run the hadoop stuff in GHC. Therefore, I would like to share it with those who are still struggling in setting.

1. longin: ssh andrew_id@ghc09.ghc.andrew.cmu.edu

2. set the .bashrc:

$ ls -a

$ vim .bashrc (then copy the setting from the website http://curtis.ml.cmu.edu/w/courses/index.php/Hadoop_cluster_information)

3. enter bash:

$ bash

$ hadoop fs -copyFromLocal nb.jar /user/andrew_id (This is to copy the file from your local disk to the hadoop, make sure you upload the file to cluster first, could use scp)

$ export HADOOP_CLASSPATH=$HADOOP_CLASSPATH:./nb.jar

$ hadoop jar /usr/local/hadoop/contrib/streaming/hadoop-streaming-1.0.1.jar -input RCV1.small_test.txt -file nb.jar -output output -mapper "/usr/bin/java -cp ./lib/nb.jar NBTrainMapper" -reducer "/usr/bin/java -cp ./lib/nb.jar NBTrainReducer" (Here the input file is either your small test file or full dataset, the output folder name should be a folder does not exist, like in AWS.)

Then you could see the running process. You could also view all the material in console. http://ghc03.ghc.andrew.cmu.edu:50075/browseDirectory.jsp?dir=/user&namenodeInfoPort=50070, you could find your user name here, and after you copy the file to it, you could also see it.

Hope it helps.