这不同的地方也是两者在底层架构区别的体现。
hive的参数hive.mapred.mode是控制hive执行mapred的方式的,有两个选项:strict和nonstrict,默认值是nonstrict。
这个两个值对order by的执行有着很大的影响。
测试用例
hive> select * from test09;
OK
100 tom
200 mary
300 kate
400 tim
Time taken: 0.061 seconds
我们先来看看nonstrict的情况。
hive> set hive.mapred.mode=nonstrict;
hive> select * from test09 order by id;
Total MapReduce jobs = 1
Launching Job 1 out of 1
Number of reduce tasks determined at compile time: 1
In order to change the average load for a reducer (in bytes):
set hive.exec.reducers.bytes.per.reducer=
In order to limit the maximum number of reducers:
set hive.exec.reducers.max=
In order to set a constant number of reducers:
set mapred.reduce.tasks=
Starting Job = job_201105020924_0065, Tracking URL = http://hadoop00:50030/jobdetails.jsp?jobid=job_201105020924_0065
Kill Command = /home/hjl/hadoop/bin/../bin/hadoop job -Dmapred.job.tracker=hadoop00:9001 -kill job_201105020924_0065
2011-05-03 03:37:41,270 Stage-1 map = 0%, reduce = 0%
2011-05-03 03:37:43,292 Stage-1 map = 50%, reduce = 0%
2011-05-03 03:37:45,314 Stage-1 map = 100%, reduce = 0%
2011-05-03 03:37:50,360 Stage-1 map = 100%, reduce = 100%
Ended Job = job_201105020924_0065
OK
100 tom
200 mary
300 kate
400 tim
Time taken: 15.049 seconds
这个时候order by可以正常的执行,hive启动了一个reduce进行处理,事实上也只能启动一个reduce。
在来看看strict的情况
hive> set hive.mapred.mode=strict;
hive> select * from test09 order by id;
FAILED: Error in semantic analysis: line 1:30 In strict mode, limit must be specified if ORDER BY is present id
这个时候提示你,在strict模式下如果执行order by的操作必须要指定limit。
因为执行order by的时候只能启动单个reduce执行,如果排序的结果集过大,那么执行时间会非常漫长。
hive> select * from test09 order by id limit 4;
Total MapReduce jobs = 1
Launching Job 1 out of 1
Number of reduce tasks determined at compile time: 1
In order to change the average load for a reducer (in bytes):
set hive.exec.reducers.bytes.per.reducer=
In order to limit the maximum number of reducers:
set hive.exec.reducers.max=
In order to set a constant number of reducers:
set mapred.reduce.tasks=
Starting Job = job_201105020924_0067, Tracking URL = http://hadoop00:50030/jobdetails.jsp?jobid=job_201105020924_0067
Kill Command = /home/hjl/hadoop/bin/../bin/hadoop job -Dmapred.job.tracker=hadoop00:9001 -kill job_201105020924_0067
2011-05-03 04:18:26,828 Stage-1 map = 0%, reduce = 0%
2011-05-03 04:18:27,842 Stage-1 map = 50%, reduce = 0%
2011-05-03 04:18:29,864 Stage-1 map = 100%, reduce = 0%
2011-05-03 04:18:35,916 Stage-1 map = 100%, reduce = 100%
Ended Job = job_201105020924_0067
OK
100 tom
200 mary
300 kate
400 tim
Time taken: 15.706 seconds
加上limit后,SQL成功执行。
本文转自http://www.oratea.net/?p=622
这不同的地方也是两者在底层架构区别的体现。
hive的参数hive.mapred.mode是控制hive执行mapred的方式的,有两个选项:strict和nonstrict,默认值是nonstrict。
这个两个值对order by的执行有着很大的影响。
测试用例
hive> select * from test09;
OK
100 tom
200 mary
300 kate
400 tim
Time taken: 0.061 seconds
我们先来看看nonstrict的情况。
hive> set hive.mapred.mode=nonstrict;
hive> select * from test09 order by id;
Total MapReduce jobs = 1
Launching Job 1 out of 1
Number of reduce tasks determined at compile time: 1
In order to change the average load for a reducer (in bytes):
set hive.exec.reducers.bytes.per.reducer=
In order to limit the maximum number of reducers:
set hive.exec.reducers.max=
In order to set a constant number of reducers:
set mapred.reduce.tasks=
Starting Job = job_201105020924_0065, Tracking URL = http://hadoop00:50030/jobdetails.jsp?jobid=job_201105020924_0065
Kill Command = /home/hjl/hadoop/bin/../bin/hadoop job -Dmapred.job.tracker=hadoop00:9001 -kill job_201105020924_0065
2011-05-03 03:37:41,270 Stage-1 map = 0%, reduce = 0%
2011-05-03 03:37:43,292 Stage-1 map = 50%, reduce = 0%
2011-05-03 03:37:45,314 Stage-1 map = 100%, reduce = 0%
2011-05-03 03:37:50,360 Stage-1 map = 100%, reduce = 100%
Ended Job = job_201105020924_0065
OK
100 tom
200 mary
300 kate
400 tim
Time taken: 15.049 seconds
这个时候order by可以正常的执行,hive启动了一个reduce进行处理,事实上也只能启动一个reduce。
在来看看strict的情况
hive> set hive.mapred.mode=strict;
hive> select * from test09 order by id;
FAILED: Error in semantic analysis: line 1:30 In strict mode, limit must be specified if ORDER BY is present id
这个时候提示你,在strict模式下如果执行order by的操作必须要指定limit。
因为执行order by的时候只能启动单个reduce执行,如果排序的结果集过大,那么执行时间会非常漫长。
hive> select * from test09 order by id limit 4;
Total MapReduce jobs = 1
Launching Job 1 out of 1
Number of reduce tasks determined at compile time: 1
In order to change the average load for a reducer (in bytes):
set hive.exec.reducers.bytes.per.reducer=
In order to limit the maximum number of reducers:
set hive.exec.reducers.max=
In order to set a constant number of reducers:
set mapred.reduce.tasks=
Starting Job = job_201105020924_0067, Tracking URL = http://hadoop00:50030/jobdetails.jsp?jobid=job_201105020924_0067
Kill Command = /home/hjl/hadoop/bin/../bin/hadoop job -Dmapred.job.tracker=hadoop00:9001 -kill job_201105020924_0067
2011-05-03 04:18:26,828 Stage-1 map = 0%, reduce = 0%
2011-05-03 04:18:27,842 Stage-1 map = 50%, reduce = 0%
2011-05-03 04:18:29,864 Stage-1 map = 100%, reduce = 0%
2011-05-03 04:18:35,916 Stage-1 map = 100%, reduce = 100%
Ended Job = job_201105020924_0067
OK
100 tom
200 mary
300 kate
400 tim
Time taken: 15.706 seconds
加上limit后,SQL成功执行。
本文转自http://www.oratea.net/?p=622