Pandas Dataframe 分割字符串

最新推荐文章于 2023-06-02 11:24:33 发布

huihui12a

最新推荐文章于 2023-06-02 11:24:33 发布

阅读量710

点赞数

文章标签： python

本文链接：https://blog.csdn.net/zhangxiaohuiNO1/article/details/125487132

版权

遇到分割后列表中每个子列表中元素个数不同的情况的处理

import pandas as pd
import matplotlib.pyplot as plt
import numpy as np

path = "D:\深度学习\python数据分析\pandas\IMDB-Movie-Data.csv"
d1 = pd.read_csv(path)

# 演员人数
print(d1["Actors"])
temp_actors = d1['Actors'].str.split(", ").tolist()
print(temp_actors)
"""actors = []
for temp_actor in temp_actors:
    actors.extend(temp_actor)"""
actors = [i for j in temp_actors for i in j]
print(actors)
print(len(set(actors)))

原Series数据形式

执行分割后结果，观察到列表中每个子列表中元素个数不同

temp_actors =np.array(d1['Actors'].str.split(", ").tolist())

[['Chris Pratt', 'Vin Diesel', 'Bradley Cooper', 'Zoe Saldana'], ['Noomi Rapace', 'Logan Marshall-Green', 'Michael Fassbender', 'Charlize Theron'], ['James McAvoy', 'Anya Taylor-Joy', 'Haley Lu Richardson', 'Jessica Sula'], ['Matthew McConaughey,Reese Witherspoon', 'Seth MacFarlane', 'Scarlett Johansson'], ['Will Smith', 'Jared Leto', 'Margot Robbie', 'Viola Davis'], ['Matt Damon', 'Tian Jing', 'Willem Dafoe', 'Andy Lau'], ['Ryan Gosling', 'Emma Stone', 'Rosemarie DeWitt', 'J.K. Simmons'], ['Essie Davis', 'Andrea Riseborough', 'Julian Barratt,Kenneth Branagh'], ['Charlie Hunnam', 'Robert Pattinson', 'Sienna Miller', 'Tom Holland'], ['Jennifer Lawrence', 'Chris Pratt', 'Michael Sheen,Laurence Fishburne'], ['Eddie Redmayne', 'Katherine Waterston', 'Alison Sudol,Dan Fogler']]

将列表展开，结果示例为

# 方式1
actors = []
for temp_actor in temp_actors:
    actors.extend(temp_actor)"""
# 方式2
actors = [i for j in temp_actors for i in j]

['Chris Pratt', 'Vin Diesel', 'Bradley Cooper', 'Zoe Saldana', 'Noomi Rapace']

huihui12a

关注

0
点赞
踩
3

收藏

觉得还不错? 一键收藏
0
评论
Pandas Dataframe 分割字符串

Pandas Dataframe 分割字符串，遇到分割后列表中每个子列表中元素个数不同的情况的处理
复制链接

扫一扫