异常信息:
Job aborted due to stage failure: Total size of serialized results of 17509 tasks (2.0 GiB) is bigger than spark.driver.maxResultSize (2.0 GiB)
解决方案:
spark.driver.maxResultSize默认大小为1G,指的是每个Spark action(如collect)所有分区的序列化结果的总大小限制,就是说,executor给driver返回的结果过大。解决办法有两个
(1)避免使用类似的方法,比如countByValue,countByKey等
(2)调大参数:spark.driver.maxResultSize 4g
HiveTask添加参数的方法
在脚本最后的ht.exec_sql里增加参数,例如:
ht.exec_sql(schema_name = 'mydb', table_name = 'mytable', sql = sql, merge_flag = True, merge_part_dir =['dt='+ data_day_str], merge_type='mr', exec_engine='spark', spark_args=['--conf spark.driver.maxResultSize=4g'])