配置hadoop环境变量
涉及到几个xml配置文件
hadoop-env.sh:配置hadoop依赖的环境,如:jkd
core-site-xml:core的配置项,例如hdfs和mapreduce常用的i/o设置等
hdfs-site.xml:hadoop守护进程的配置项,包括namenode、辅助namenode和datanode等
yarn-site.xml:mapreduce配置项
master 记录运行辅助namenode的机器列表
slave 记录运行datanode和tasktracker的机器列表
1.配置hadoop-env.sh中JAVA_HOME
2.配置core-site-xml
<property>
<name>fs.defaultFS</name>
<value>hdfs://localhost:9000</value>
</property>
3.配置hdfs-site.xml
<property>
<name>dfs.namenode.name.dir</name>
<value>file:/usr/local/hadoop/dfs/name</value>
</property>
<property>
<name>dfs.datanode.data.dir</name>
<value>file:/usr/local/hadoop/dfs/data</value>
</property>
<property>
<name>dfs.replication</name>
<value>1</value>
</property>
<property>
<name>dfs.permissions</name>
<value>false</value>
</property>
4.配置yarn-site.xml
<property>
<name>yarn.resourcemanager.resource-tracker.address</name>
<value>localhost:8031</value>
</property>
<property>
<name>yarn.resourcemanager.address</name>
<value>localhost:8032</value>
</property>
<property>
<name>yarn.resourcemanager.scheduler.address</name>
<value>localhost:8030</value>
</property>
<property>
<name>yarn.resourcemanager.admin.address</name>
<value>localhost:8033</value>
</property>
<property>
<name>yarn.resourcemanager.webapp.address</name>
<value>localhost:8088</value>
</property>
<property>
<name>yarn.nodemanager.aux-services</name>
<value>mapreduce.shuffle</value>
</property>
<property>
<name>yarn.nodemanager.aux-services.mapreduce.shuffle.class</name>
<value>org.apache.hadoop.mapred.ShuffleHandler</value>
</property>
5.将mapred-site.xml.temporary mv 为mapred-site.xml
<property>
<name>mapreduce.framework.name</name>
<value>yarn</value>
</property>
<property>
<name>mapreduce.jobhistory.address</name>
<value>localhost:10020</value>
</property>
<property>
<name>mapreduce.jobhistory.webapp.address</name>
<value>localhost:19888</value>
</property>
6.开启Hadoop /sbin/start-all.sh
7.开启hadoop jobhistory的方式为 /sbin/mr-jobhistory-daemon.sh start historyserver