DataX配置MySQL到同步全量表到HDFS的一键脚本

法外狂徒张百万

已于 2022-08-01 11:52:00 修改

阅读量1.2k

点赞数 2

文章标签： mysql hdfs 数据库

于 2022-06-17 21:33:26 首次发布

本文链接：https://blog.csdn.net/qq_42417269/article/details/125340224

版权

DataX配置MySQL到同步全量表到HDFS的一键脚本

一般该类的配置文件，我一般放在用户的bin目录下，当然工作中是没有root用户的，就是个人家目录的bin下：

[wyp@hadoop102 bin]$ vim ~/bin/gen_import_config.py

插入如下python脚本：


# coding=utf-8
import os
import sys
import MySQLdb

#MySQL相关配置，需根据实际情况作出修改
mysql_host = "hadoop102"
mysql_port = "3306"
mysql_user = "root"
mysql_passwd = "123456"

#HDFS NameNode相关配置，需根据实际情况作出修改
hdfs_nn_host = "hadoop102"
hdfs_nn_port = "8020"

#生成配置文件的目标路径，可根据实际情况作出修改
output_path = "/opt/module/datax/job/import"


def get_connection():
    return MySQLdb.connect(host=mysql_host, port=int(mysql_port), user=mysql_user, passwd=mysql_passwd)


def get_mysql_meta(database, table):
    connection = get_connection()
    cursor = connection.cursor()
    cursor.execute(sql, [database, table])
    fetchall = cursor.fetchall()
    cursor.close()
    connection.close()
    return fetchall


def get_mysql_columns(database, table):
    return map(lambda x: x[0], get_mysql_meta(database, table))


def get_hive_columns(database, table):
    def type_mapping(mysql_type):
        mappings = {
            "bigint": "bigint",
            "int": "bigint",
            "smallint": "bigint",
            "tinyint": "bigint",
            "decimal": "string",
            "double": "double",
            "float": "float",
            "binary": "string",
            "char": "string",
            "varchar": "string",
            "datetime": "string",
            "time": "string",
            "timestamp": "string",
            "date": "string",
            "text": "string"
        }
        return mappings[mysql_type]

    meta = get_mysql_meta(database, table)
    return map(lambda x: {"name": x[0], "type": type_mapping(x[1].lower())}, meta)


def generate_json(source_database, source_table):
    job = {
        "job": {
            "setting": {
                "speed": {
                    "channel": 3
                },
                "errorLimit": {
                    "record": 0,
                    "percentage": 0.02
                }
            },
            "content": [{
                "reader": {
                    "name": "mysqlreader",
                    "parameter": {
                        "username": mysql_user,
                        "password": mysql_passwd,
                        "column": get_mysql_columns(source_database, source_table),
                        "splitPk": "",
                        "connection": [{
                            "table": [source_table],
                            "jdbcUrl": ["jdbc:mysql://" + mysql_host + ":" + mysql_port + "/" + source_database]
                        }]
                    }
                },
                "writer": {
                    "name": "hdfswriter",
                    "parameter": {
                        "defaultFS": "hdfs://" + hdfs_nn_host + ":" + hdfs_nn_port,
                        "fileType": "text",
                        "path": "${targetdir}",
                        "fileName": source_table,
                        "column": get_hive_columns(source_database, source_table),
                        "writeMode": "append",
                        "fieldDelimiter": "\t",
                        "compress": "gzip"
                    }
                }
            }]
        }
    }
    if not os.path.exists(output_path):
        os.makedirs(output_path)
    with open(os.path.join(output_path, ".".join([source_database, source_table, "json"])), "w") as f:
        json.dump(job, f)


def main(args):
    source_database = ""
    source_table = ""

    options, arguments = getopt.getopt(args, '-d:-t:', ['sourcedb=', 'sourcetbl='])
    for opt_name, opt_value in options:
        if opt_name in ('-d', '--sourcedb'):
            source_database = opt_value
        if opt_name in ('-t', '--sourcetbl'):
            source_table = opt_value

    generate_json(source_database, source_table)


if __name__ == '__main__':
    main(sys.argv[1:])

脚本使用说明：

# 通过-d传入数据库名，-t传入表名，执行上述命令即可生成该表的DataX同步配置文件。
python gen_import_config.py -d database -t table

由于需要使用Python访问Mysql数据库，故需安装驱动，命令如下：

# 使用root权限去yum上下载python连接mysql的驱动
[wyp@hadoop102 bin]$ sudo yum install -y MySQL-python

在~/bin目录下创建gen_import_config.sh脚本：

[wyp@hadoop102 bin]$ vim ~/bin/gen_import_config.sh

# 配置你需要全量上传的数据库名和数据表名 使用说明在上方
#!/bin/bash

python ~/bin/gen_import_config.py -d wyp -t activity_info
python ~/bin/gen_import_config.py -d wyp -t activity_rule
python ~/bin/gen_import_config.py -d wyp -t base_category1
python ~/bin/gen_import_config.py -d wyp -t base_category2
python ~/bin/gen_import_config.py -d wyp -t base_category3
python ~/bin/gen_import_config.py -d wyp -t base_dic
python ~/bin/gen_import_config.py -d wyp -t base_province
python ~/bin/gen_import_config.py -d wyp -t base_region
python ~/bin/gen_import_config.py -d wyp -t base_trademark
python ~/bin/gen_import_config.py -d wyp -t cart_info
python ~/bin/gen_import_config.py -d wyp -t coupon_info
python ~/bin/gen_import_config.py -d wyp -t sku_attr_value
python ~/bin/gen_import_config.py -d wyp -t sku_info
python ~/bin/gen_import_config.py -d wyp -t sku_sale_attr_value
python ~/bin/gen_import_config.py -d wyp -t spu_info

为gen_import_config.sh脚本增加执行权限

[wyp@hadoop102 bin]$ chmod 777 ~/bin/gen_import_config.sh

执行gen_import_config.sh脚本，生成配置文件

[wyp@hadoop102 bin]$ gen_import_config.sh

查看配置文件生成的情况：

# 默认是生成在/opt/module/datax/job/import/ ，如果需要修改
ll /opt/module/datax/job/import/ 
结果如下：
-rw-rw-r-- 1 wyp wyp  957 10月 15 22:17 wyp.activity_info.json
-rw-rw-r-- 1 wyp wyp 1049 10月 15 22:17 wyp.activity_rule.json
-rw-rw-r-- 1 wyp wyp  651 10月 15 22:17 wyp.base_category1.json
-rw-rw-r-- 1 wyp wyp  711 10月 15 22:17 wyp.base_category2.json
-rw-rw-r-- 1 wyp wyp  711 10月 15 22:17 wyp.base_category3.json
-rw-rw-r-- 1 wyp wyp  835 10月 15 22:17 wyp.base_dic.json
-rw-rw-r-- 1 wyp wyp  865 10月 15 22:17 wyp.base_province.json
-rw-rw-r-- 1 wyp wyp  659 10月 15 22:17 wyp.base_region.json
-rw-rw-r-- 1 wyp wyp  709 10月 15 22:17 wyp.base_trademark.json
-rw-rw-r-- 1 wyp wyp 1301 10月 15 22:17 wyp.cart_info.json
-rw-rw-r-- 1 wyp wyp 1545 10月 15 22:17 wyp.coupon_info.json
-rw-rw-r-- 1 wyp wyp  867 10月 15 22:17 wyp.sku_attr_value.json
-rw-rw-r-- 1 wyp wyp 1121 10月 15 22:17 wyp.sku_info.json
-rw-rw-r-- 1 wyp wyp  985 10月 15 22:17 wyp.sku_sale_attr_value.json
-rw-rw-r-- 1 wyp wyp  811 10月 15 22:17 wyp.spu_info.json

但是，datax提交任务的hdfs路径必须存在，否则上传就要失败。而且datax没有自动创建的功能，让人手工去一个一个目录创建，这个工作量是很恐怖的，那就写一个能自动生成目录的脚本：

全量表数据同步脚本

# 依旧在家目录的bin下创建
vim ~/bin/mysql_to_hdfs_full.sh 
#!/bin/bash

DATAX_HOME=/opt/module/datax

# 如果传入日期则do_date等于传入的日期，否则等于前一天日期
if [ -n "$2" ] ;then
    do_date=$2
else
    do_date=`date -d "-1 day" +%F`
fi

#处理目标路径，此处的处理逻辑是，如果目标路径不存在，则创建；若存在，则清空，目的是保证同步任务可重复执行
handle_targetdir() {
  hadoop fs -test -e $1
  if [[ $? -eq 1 ]]; then
    echo "路径$1不存在，正在创建......"
    hadoop fs -mkdir -p $1
  else
    echo "路径$1已经存在"
    fs_count=$(hadoop fs -count $1)
    content_size=$(echo $fs_count | awk '{print $3}')
    if [[ $content_size -eq 0 ]]; then
      echo "路径$1为空"
    else
      echo "路径$1不为空，正在清空......"
      hadoop fs -rm -r -f $1/*
    fi
  fi
}

#数据同步
import_data() {
  datax_config=$1
  target_dir=$2

  handle_targetdir $target_dir
  python $DATAX_HOME/bin/datax.py -p"-Dtargetdir=$target_dir" $datax_config
}

case $1 in
"activity_info")
  import_data /opt/module/datax/job/import/wyp.activity_info.json /origin_data/wyp/db/activity_info_full/$do_date
  ;;
"activity_rule")
  import_data /opt/module/datax/job/import/wyp.activity_rule.json /origin_data/wyp/db/activity_rule_full/$do_date
  ;;
"base_category1")
  import_data /opt/module/datax/job/import/wyp.base_category1.json /origin_data/wyp/db/base_category1_full/$do_date
  ;;
"base_category2")
  import_data /opt/module/datax/job/import/wyp.base_category2.json /origin_data/wyp/db/base_category2_full/$do_date
  ;;
"base_category3")
  import_data /opt/module/datax/job/import/wyp.base_category3.json /origin_data/wyp/db/base_category3_full/$do_date
  ;;
"base_dic")
  import_data /opt/module/datax/job/import/wyp.base_dic.json /origin_data/wyp/db/base_dic_full/$do_date
  ;;
"base_province")
  import_data /opt/module/datax/job/import/wyp.base_province.json /origin_data/wyp/db/base_province_full/$do_date
  ;;
"base_region")
  import_data /opt/module/datax/job/import/wyp.base_region.json /origin_data/wyp/db/base_region_full/$do_date
  ;;
"base_trademark")
  import_data /opt/module/datax/job/import/wyp.base_trademark.json /origin_data/wyp/db/base_trademark_full/$do_date
  ;;
"cart_info")
  import_data /opt/module/datax/job/import/wyp.cart_info.json /origin_data/wyp/db/cart_info_full/$do_date
  ;;
"coupon_info")
  import_data /opt/module/datax/job/import/wyp.coupon_info.json /origin_data/wyp/db/coupon_info_full/$do_date
  ;;
"sku_attr_value")
  import_data /opt/module/datax/job/import/wyp.sku_attr_value.json /origin_data/wyp/db/sku_attr_value_full/$do_date
  ;;
"sku_info")
  import_data /opt/module/datax/job/import/wyp.sku_info.json /origin_data/wyp/db/sku_info_full/$do_date
  ;;
"sku_sale_attr_value")
  import_data /opt/module/datax/job/import/wyp.sku_sale_attr_value.json /origin_data/wyp/db/sku_sale_attr_value_full/$do_date
  ;;
"spu_info")
  import_data /opt/module/datax/job/import/wyp.spu_info.json /origin_data/wyp/db/spu_info_full/$do_date
  ;;
"all")
  import_data /opt/module/datax/job/import/wyp.activity_info.json /origin_data/wyp/db/activity_info_full/$do_date
  import_data /opt/module/datax/job/import/wyp.activity_rule.json /origin_data/wyp/db/activity_rule_full/$do_date
  import_data /opt/module/datax/job/import/wyp.base_category1.json /origin_data/wyp/db/base_category1_full/$do_date
  import_data /opt/module/datax/job/import/wyp.base_category2.json /origin_data/wyp/db/base_category2_full/$do_date
  import_data /opt/module/datax/job/import/wyp.base_category3.json /origin_data/wyp/db/base_category3_full/$do_date
  import_data /opt/module/datax/job/import/wyp.base_dic.json /origin_data/wyp/db/base_dic_full/$do_date
  import_data /opt/module/datax/job/import/wyp.base_province.json /origin_data/wyp/db/base_province_full/$do_date
  import_data /opt/module/datax/job/import/wyp.base_region.json /origin_data/wyp/db/base_region_full/$do_date
  import_data /opt/module/datax/job/import/wyp.base_trademark.json /origin_data/wyp/db/base_trademark_full/$do_date
  import_data /opt/module/datax/job/import/wyp.cart_info.json /origin_data/wyp/db/cart_info_full/$do_date
  import_data /opt/module/datax/job/import/wyp.coupon_info.json /origin_data/wyp/db/coupon_info_full/$do_date
  import_data /opt/module/datax/job/import/wyp.sku_attr_value.json /origin_data/wyp/db/sku_attr_value_full/$do_date
  import_data /opt/module/datax/job/import/wyp.sku_info.json /origin_data/wyp/db/sku_info_full/$do_date
  import_data /opt/module/datax/job/import/wyp.sku_sale_attr_value.json /origin_data/wyp/db/sku_sale_attr_value_full/$do_date
  import_data /opt/module/datax/job/import/wyp.spu_info.json /origin_data/wyp/db/spu_info_full/$do_date
  ;;
esac

赋予权限：

chmod 777 ~/bin/mysql_to_hdfs_full.sh

测试一下数据，传参数时：

# 也可以对单表进行操作，这里只是稍作演示，对单表进行操作的时候，填写表明参数即可，同样表名或者脚本名不同时，自己手动改一下。
mysql_to_hdfs_full.sh all 2022-06-17

这个脚本的目的在于，执行上方的提交任务的命令，并且检查分区表是否存在，不存在就创建。实际上命令的最后一个参数，不指定时是默认在hdfs上创建前一天的目录的，一般离线数仓全量同步的任务是在凌晨同步前一天的数据，所以这个设定也是比较合理的。
在个人测试的时候，可以指定一下时间参数，看看效果，项目中用调度器的时候可以不指定，让他自己自动生成。