mysql游标分段遍历统计数量

汉堡儿小可爱

已于 2022-02-08 14:39:15 修改

阅读量966

点赞数

分类专栏： MYSQL python 文章标签： mysql 数据库 sql

于 2022-01-15 23:34:55 首次发布

本文链接：https://blog.csdn.net/m0_56646269/article/details/122517988

版权

python 同时被 2 个专栏收录

10 篇文章 0 订阅

订阅专栏

MYSQL

3 篇文章 0 订阅

订阅专栏

小白写个小代码。

因为数据库有的表数据量很大，例如几千万，每条语句的查询时长有限制，不然会影响别人使用此数据库的速度，所以建议使用mysql游标分段遍历统计数量，设置合适的分段范围，尽量减少查询时间，不影响别人。

# -*- coding:utf-8 -*-
"""
作者：孙敏
日期：2022年01月15日
"""
import pymysql
from tqdm import tqdm
import re

#填写对应的mysql连接信息
host = 'localhost'
port = 3306
user = 'root'
password = '1234'
database = 'sunmin'
#连接mysql数据库
db =pymysql.connect(host=host,port=port,user=user,password=password,
                         database=database,charset='utf8')
#使用游标
cursor = db.cursor()


#使用sql语句对id进行分段查询database对应数据库中的表，将查询结果直接打印
def total_count(start,end,step):
    allnumber = [] #为求和函数做准备：先定义一个空列表，后续往此列表内赋值，然后使用sum函数求和
    for i in tqdm(range(start,end,step),'进度'):#代表id范围和步长
        section = list((i,i+step-1))
        cursor.execute(sql,section)#执行单个sql语句
        first_record = cursor.fetchone()#取出查询语句的第一个元素
        allnumber.append(int(first_record[0]))
    totalcount = sum(allnumber)
    print('sql查询所得总数量：'+str(totalcount))



#使用sql语句对id进行分段查询database对应数据库中的表，将查询结果生成excel表格
def output_as_a_excel_table(path,start,end,step):
    xlsx_list = []
    for i in tqdm(range(start,end,step),'进度'):#代表id范围和步长
        section = list((i,i+step-1))
        cursor.execute(sql,section)#执行单个sql语句
        first_record = cursor.fetchone()#取出查询语句的第一个元素
        y = []#定义一个列表，主要使输出结果的xlsx_list列表的元素还是列表
        b=str(section)
        y.append(b)#为xlsx_list列表元素的子列表添加第一个元素
        c=str(first_record[0])
        y.append(c)#为xlsx_list列表元素的子列表添加第二个元素
        xlsx_list.append(y)#为xlsx_list列表添加子列表元素
    result = open(path, 'w', encoding='gbk')
    # 参数'w'代表往指定表格写入数据，会先将表格中原本的内容清空
    # 若把参数’w'修改为‘a+',即可实现在原本内容的基础上，增加新写入的内容
    result.write('id区间\t对应数量\n')
    for m in range(len(xlsx_list)):
        for n in range(len(xlsx_list[m])):
            result.write(str(xlsx_list[m][n]))
            result.write('\t')  # '\t'表示每写入一个元素后，会移动到同行的下一个单元格
        result.write('\n')# 换行操作
    result.close()

if __name__ == '__main__':
    #设置多个sql语句
    sqllist = ["SELECT count(id) FROM t_list where (id between %s and %s)","SELECT count(id) FROM t_content where (id between %s and %s)"]
    for sql in sqllist:
        startid = 1#设置起始id
        step = 100000#设置步长
        #查找对应表的最大id
        tablename = re.findall("FROM(.*?)where", sql)[0]
        idsql = 'select id from' + tablename + 'order by id desc limit 1'
        cursor.execute(idsql)  # 执行单个sql语句
        endidtuple = cursor.fetchone()  # 取出查询语句的第一个元素
        endid = endidtuple[0]
        #调用函数
        total_count(startid,endid,step)
        # 设置存储路径
        # path = 'D:\python\测试1.xls'
        # output_as_a_excel_table(path,startid,endid,step)

cursor.close()#关闭游标
db.close()#关闭数据库连接

汉堡儿小可爱

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
mysql游标分段遍历统计数量

小白写个小代码。因为数据库有的表数据量很大，例如几千万，每条语句的查询时长有限制，不然会影响别人使用此数据库的速度，所以建议使用mysql游标分段遍历统计数量，设置合适的分段范围，尽量减少查询时间，不影响别人。# -*- coding:utf-8 -*-"""作者：孙敏日期：2022年01月15日"""import pymysqlfrom tqdm import tqdm#填写对应的mysql连接信息host = 'localhost'port = 3306user = 'r
复制链接

扫一扫