flume 实时读取本地文件到hdfs

最新推荐文章于 2023-02-03 00:00:13 发布

小哇666

最新推荐文章于 2023-02-03 00:00:13 发布

阅读量719

点赞数

分类专栏： # flume

本文链接：https://blog.csdn.net/qq_41712271/article/details/103939224

版权

flume 专栏收录该内容

10 篇文章 0 订阅

订阅专栏

1．Flume要想将数据输出到HDFS，必须持有Hadoop相关jar包
将commons-configuration-1.6.jar、
hadoop-auth-2.7.2.jar、
hadoop-common-2.7.2.jar、
hadoop-hdfs-2.7.2.jar、
commons-io-2.4.jar、
htrace-core-3.1.0-incubating.jar
拷贝到/flume安装目录下/lib文件夹下。

2．创建 flume-hdfs.conf文件
source的类型选择注意
注：要想读取Linux系统中的文件，就得按照Linux命令的规则执行命令。
由于log日志在Linux系统中所以读取文件的类型选择：exec即execute执行的意思。
表示执行Linux命令来读取文件。

生成到hdfs文件的说明：
一个小时生成一个文件夹，一分钟或者134217700字节（128M）生成一个文件
例如/flume/20191105/10/logs-1592369,但是在文件滚动的一分钟内，文件是以.temp结尾的。

#Name the components on this agent
a2.sources = r2
a2.sinks = k2
a2.channels = c2
 
# Describe/configure the source
a2.sources.r2.type = exec
a2.sources.r2.command = tail -F /opt/test.log
a2.sources.r2.shell = /bin/bash -c
 
# Describe the sink
a2.sinks.k2.type = hdfs
a2.sinks.k2.hdfs.path = hdfs://hadoop01:8020/flume/%Y%m%d/%H
#上传文件的前缀
a2.sinks.k2.hdfs.filePrefix = logs-
#是否按照时间滚动文件夹
a2.sinks.k2.hdfs.round = true
#多少时间单位创建一个新的文件夹
a2.sinks.k2.hdfs.roundValue = 1
#重新定义时间单位
a2.sinks.k2.hdfs.roundUnit = hour
#是否使用本地时间戳
a2.sinks.k2.hdfs.useLocalTimeStamp = true
#积攒多少个Event才flush到HDFS一次
a2.sinks.k2.hdfs.batchSize = 1000
#设置文件类型，可支持压缩
a2.sinks.k2.hdfs.fileType = DataStream
#多久生成一个新的文件
a2.sinks.k2.hdfs.rollInterval = 600
#设置每个文件的滚动大小
a2.sinks.k2.hdfs.rollSize = 134217700
#文件的滚动与Event数量无关
a2.sinks.k2.hdfs.rollCount = 0
#最小冗余数
a2.sinks.k2.hdfs.minBlockReplicas = 1
 
# Use a channel which buffers events in memory
a2.channels.c2.type = memory
a2.channels.c2.capacity = 1000
a2.channels.c2.transactionCapacity = 100
 
# Bind the source and sink to the channel
a2.sources.r2.channels = c2
a2.sinks.k2.channel = c2

注意：
对于所有与时间相关的转义序列，Event Header中必须存在以 “timestamp”的key（除非hdfs.useLocalTimeStamp设置为true，此方法会使用TimestampInterceptor自动添加timestamp）。
a2.sinks.k2.hdfs.useLocalTimeStamp = true

3．运行flume配置文件
bin/flume-ng agent -c conf/ -f job/flume-hdfs.conf -n a2

小哇666

关注

0
点赞
踩
2

收藏

觉得还不错? 一键收藏
0
评论
flume 实时读取本地文件到hdfs

1．Flume要想将数据输出到HDFS，必须持有Hadoop相关jar包将commons-configuration-1.6.jar、hadoop-auth-2.7.2.jar、hadoop-common-2.7.2.jar、hadoop-hdfs-2.7.2.jar、commons-io-2.4.jar、htrace-core-3.1.0-incubating.jar拷贝到/flu...
复制链接

扫一扫