简单的问题:我读过的所有教程都向您展示了如何使用其中一种方法将并行计算的结果输出到一个列表(或者最多是一个字典)ipython.parallel或多重处理。在
你能给我举一个简单的例子,用这两个库将计算结果输出到共享的pandas数据帧吗?在import pandas as pd
import multiprocessing as mp
LARGE_FILE = "D:\\my_large_file.txt"
CHUNKSIZE = 100000 # processing 100,000 rows at a time
def process_frame(df):
# process data frame
return len(df)
if __name__ == '__main__':
reader = pd.read_table(LARGE_FILE, chunksize=CHUNKSIZE)
pool = mp.Pool(4) # use 4 processes
funclist = []
for df in reader:
# process each data frame
f = pool.apply_async(process_frame,[df])
funclist.append(f)
result = 0
for f in funclist:
result += f.get(timeout=10) # timeout in 10 seconds
print "There are %d rows of data"%(result)