hive分桶

最新推荐文章于 2024-07-15 10:59:50 发布

anke5156

最新推荐文章于 2024-07-15 10:59:50 发布

阅读量197

点赞数

文章标签： hive 大数据 sql

本文链接：https://blog.csdn.net/u012762281/article/details/106099665

版权

hive分桶

1.创建分桶表

drop table stu_buck;
create table stu_buck(Sno int,Sname string,Sex string,Sage int,Sdept string)
clustered by(Sno)
sorted by(Sno DESC)
into 4 buckets
row format delimited
fields terminated by ',';

2.设置变量

设置分桶为true, 设置reduce数量是分桶的数量个数
set hive.enforce.bucketing = true;
set mapreduce.job.reduces=4;

insert overwrite table student_buck
select * from student cluster by(Sno) sort by(Sage);  报错,cluster 和 sort 不能共存

3.开始往创建的分通表插入数据

插入数据需要是已分桶, 且排序的
可以使用distribute by(sno) sort by(sno asc) 或是排序和分桶的字段相同的时候使用Cluster by(字段)
注意使用cluster by 就等同于分桶+排序(sort)

insert into table stu_buck
select Sno,Sname,Sex,Sage,Sdept from student distribute by(Sno) sort by(Sno asc);

insert overwrite table stu_buck
select * from student distribute by(Sno) sort by(Sno asc);

insert overwrite table stu_buck
select * from student cluster by(Sno);

anke5156

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
hive分桶

hive分桶1.创建分桶表drop table stu_buck;create table stu_buck(Sno int,Sname string,Sex string,Sage int,Sdept string)clustered by(Sno)sorted by(Sno DESC)into 4 bucketsrow format delimitedfields terminated by ',';2.设置变量设置分桶为true, 设置reduce数量是分桶的数量个数set
复制链接

扫一扫