perl与python性能,文本处理-Python vs Perl性能

Here is my Perl and Python script to do some simple text processing from about 21 log files, each about 300 KB to 1 MB (maximum) x 5 times repeated (total of 125 files, due to the log repeated 5 times).

Python Code (code modified to use compiled re and using re.I)

#!/usr/bin/python

import re

import fileinput

exists_re = re.compile(r'^(.*?) INFO.*Such a record already exists', re.I)

location_re = re.compile(r'^AwbLocation (.*?) insert into', re.I)

for line in fileinput.input():

fn = fileinput.filename()

currline = line.rstrip()

mprev = exists_re.search(currline)

if(mprev):

xlogtime = mprev.group(1)

mcurr = location_re.search(currline)

if(mcurr):

print fn, xlogtime, mcurr.group(1)

Perl Code

#!/usr/bin/perl

while (<>) {

chomp;

if (m/^(.*?) INFO.*Such a record already exists/i) {

$xlogtime = $1;

}

if (m/^AwbLocation (.*?) insert into/i) {

print "$ARGV $xlogtime $1\n";

}

}

And, on my PC both code generates exactly the same result file of 10,790 lines. And, here is the timing done on Cygwin's Perl and Python implementations.

User@UserHP /cygdrive/d/tmp/Clipboard

# time /tmp/scripts/python/afs/process_file.py *log* *log* *log* *log* *log* >

summarypy.log

real 0m8.185s

user 0m8.018s

sys 0m0.092s

User@UserHP /cygdrive/d/tmp/Clipboard

# time /tmp/scripts/python/afs/process_file.pl *log* *log* *log* *log* *log* >

summarypl.log

real 0m1.481s

user 0m1.294s

sys 0m0.124s

Originally, it took 10.2 seconds using Python and only 1.9 secs using Perl for this simple text processing.

(UPDATE) but, after the compiled re version of Python, it now takes 8.2 seconds in Python and 1.5 seconds in Perl. Still Perl is much faster.

Is there a way to improve the speed of Python at all OR it is obvious that Perl will be the speedy one for simple text processing.

By the way this was not the only test I did for simple text processing... And, each different way I make the source code, always always Perl wins by a large margin. And, not once did Python performed better for simple m/regex/ match and print stuff.

Please do not suggest to use C, C++, Assembly, other flavours of

Python, etc.

I am looking for a solution using Standard Python with its built-in

modules compared against Standard Perl (not even using the modules).

Boy, I wish to use Python for all my tasks due to its readability, but

to give up speed, I don't think so.

So, please suggest how can the code be improved to have comparable

results with Perl.

UPDATE: 2012-10-18

As other users suggested, Perl has its place and Python has its.

So, for this question, one can safely conclude that for simple regex match on each line for hundreds or thousands of text files and writing the results to a file (or printing to screen), Perl will always, always WIN in performance for this job. It as simple as that.

Please note that when I say Perl wins in performance... only standard Perl and Python is compared... not resorting to some obscure modules (obscure for a normal user like me) and also not calling C, C++, assembly libraries from Python or Perl. We don't have time to learn all these extra steps and installation for a simple text matching job.

So, Perl rocks for text processing and regex.

Python has its place to rock in other places.

Update 2013-05-29: An excellent article that does similar comparison is here. Perl again wins for simple text matching... And for more details, read the article.

解决方案

This is exactly the sort of stuff that Perl was designed to do, so it doesn't surprise me that it's faster.

One easy optimization in your Python code would be to precompile those regexes, so they aren't getting recompiled each time.

exists_re = re.compile(r'^(.*?) INFO.*Such a record already exists')

location_re = re.compile(r'^AwbLocation (.*?) insert into')

And then in your loop:

mprev = exists_re.search(currline)

and

mcurr = location_re.search(currline)

That by itself won't magically bring your Python script in line with your Perl script, but repeatedly calling re in a loop without compiling first is bad practice in Python.

  • 0
    点赞
  • 0
    收藏
    觉得还不错? 一键收藏
  • 0
    评论
评论
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包
实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值