数据分析之pandas-第3章分组

最新推荐文章于 2021-05-24 10:35:46 发布

wield_jjz

最新推荐文章于 2021-05-24 10:35:46 发布

阅读量333

点赞数

分类专栏： pandas 数据处理文章标签：数据分析

本文链接：https://blog.csdn.net/wield_jjz/article/details/105770919

版权

pandas学习-第3章分组

一、SAC过程

内涵
SAC指的是分组操作中的split-apply-combine过程

split指基于某一些规则，将数据拆成若干组
apply是指对每一组独立地使用函数
combine指将每一组的结果组合成某一类数据结构

apply过程
在该过程中，我们实际往往会遇到四类问题：

整合（Aggregation）——即分组计算统计量（如求均值、求每组元素个数）
变换（Transformation）——即分组对每个单元的数据进行操作（如元素标准化）
过滤（Filtration）——即按照某些规则筛选出一些组（如选出组内某一指标小于50的组）
综合问题——即前面提及的三种问题的混合

二、groupby函数

分组函数的基本内容：

# （a）根据某一列分组
grouped_single = df.groupby('School')
#经过groupby后会生成一个groupby对象，该对象本身不会返回任何东西，只有当相应的方法被调用才会起作用
# 例如取出某一个组：
grouped_single.get_group('S_1').head()

# （b）根据某几列分组
grouped_mul = df.groupby(['School','Class'])
grouped_mul.get_group(('S_2','C_4'))

# （c）组容量与组数
# 组容量
grouped_single.size()
grouped_mul.size()
# 组数
grouped_single.ngroups
grouped_mul.ngroups

# （d）组的遍历
for name,group in grouped_single:
    print(name) #组名
    display(group.head())
 
# （e）level参数（用于多级索引）和axis参数
df.set_index(['Gender','School']).groupby(level=1,axis=0).get_group('S_1').head()
#level=0对应Gender，level=1对应School
？？？

groupby对象的特点

#（a）查看所有可调用的方法
#由此可见，groupby对象可以使用相当多的函数，灵活程度很高
print([attr for attr in dir(grouped_single) if not attr.startswith('_')])
'''
['Address', 'Class', 'Gender', 'Height', 'Math', 'Physics', 'School', 'Weight', 'agg', 'aggregate', 'all', 'any', 'apply', 'backfill', 'bfill', 'boxplot', 'corr', 'corrwith', 'count', 'cov', 'cumcount', 'cummax', 'cummin', 'cumprod', 'cumsum', 'describe', 'diff', 'dtypes', 'expanding', 'ffill', 'fillna', 'filter', 'first', 'get_group', 'groups', 'head', 'hist', 'idxmax', 'idxmin', 'indices', 'last', 'mad', 'max', 'mean', 'median', 'min', 'ndim', 'ngroup', 'ngroups', 'nth', 'nunique', 'ohlc', 'pad', 'pct_change', 'pipe', 'plot', 'prod', 'quantile', 'rank', 'resample', 'rolling', 'sem', 'shift', 'size', 'skew', 'std', 'sum', 'tail', 'take', 'transform', 'tshift', 'var']
'''

#（b）分组对象的head和first
#对分组对象使用head函数，返回的是每个组的前几行，而不是数据集前几行
grouped_single.head(2)
#first显示的是以分组为索引的每组的第一个分组信息
grouped_single.first()

#（c）分组依据
#对于groupby函数而言，分组的依据是非常自由的，只要是与数据框长度相同的列表即可，同时支持函数型分组
df.groupby(np.random.choice(['a','b','c'],df.shape[0])).get_group('a').head()
#相当于将np.random.choice(['a','b','c'],df.shape[0])当做新的一列进行分组
???

#从原理上说，我们可以看到利用函数时，传入的对象就是索引，因此根据这一特性可以做一些复杂的操作
df[:5].groupby(lambda x:print(x)).head(0)

# 根据奇偶行分组
df.groupby(lambda x:'奇数行' if not df.index.get_loc(x)%2==1 else '偶数行').groups
'''
{'偶数行': Int64Index([1102, 1104, 1201, 1203, 1205, 1302, 1304, 2101, 2103, 2105, 2202,
             2204, 2301, 2303, 2305, 2402, 2404],
            dtype='int64', name='ID'),
 '奇数行': Int64Index([1101, 1103, 1105, 1202, 1204, 1301, 1303, 1305, 2102, 2104, 2201,
             2203, 2205, 2302, 2304, 2401, 2403, 2405],
            dtype='int64', name='ID')}
'''

#如果是多层索引，那么lambda表达式中的输入就是元组，下面实现的功能为查看两所学校中男女生分别均分是否及格
#注意：此处只是演示groupby的用法，实际操作不会这样写
math_score = df.set_index(['Gender','School'])['Math'].sort_index()
grouped_score = df.set_index(['Gender','School']).sort_index().\
            groupby(lambda x:(x,'均分及格' if math_score[x].mean()>=60 else '均分不及格'))
for name,_ in grouped_score:print(name)
'''
(('F', 'S_1'), '均分及格')
(('F', 'S_2'), '均分及格')
(('M', 'S_1'), '均分及格')
(('M', 'S_2'), '均分不及格')
'''

# （d）groupby的[]操作
#可以用[]选出groupby对象的某个或者某几个列，上面的均分比较可以如下简洁地写出：
df.groupby(['Gender','School'])['Math'].mean()>=60
'''
Gender  Scho

最低0.47元/天解锁文章

wield_jjz

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
数据分析之pandas-第3章分组

pandas学习-第3章分组一、SAC过程内涵SAC指的是分组操作中的split-apply-combine过程split指基于某一些规则，将数据拆成若干组apply是指对每一组独立地使用函数combine指将每一组的结果组合成某一类数据结构apply过程在该过程中，我们实际往往会遇到四类问题：整合（Aggregation）——即分组计算统计量（如求均值、求每组元...
复制链接

扫一扫