ubuntu 12.04 下安装 PyTesser 进行OCR识别

最新推荐文章于 2018-10-21 20:28:00 发布

鹧鸪菜

最新推荐文章于 2018-10-21 20:28:00 发布

阅读量1.8k

点赞数

分类专栏： Linux/os/Developer

Linux/os/Developer 专栏收录该内容

14 篇文章 0 订阅

订阅专栏

ubuntu 12.04 下安装 PyTesser 进行OCR识别

安装所需的库

sudo apt-get install libpng12-dev
sudo apt-get install libjpeg62-dev
sudo apt-get install libtiff4-dev

sudo apt-get install gcc
sudo apt-get install g++
sudo apt-get install automake

pytesser 调用了 tesseract，因此需要安装 tesseract，安装 tesseract 需要安装 leptonica，否则编译tesseract 的时候出现 "configure: error: leptonica not found"。

以下都是解压编译安装的老步骤：

./configure
make -j4
sudo make install

解压tessdata目录下的文件（9个）到 "/usr/local/share/tessdata"目录下

注意：这个网址下载到的只有一个，不能用，使用中会报错，http://tesseract-ocr.googlecode.com/files/eng.traineddata.gz

下载安装 pytesser

http://code.google.com/p/pytesser/

最新的是 pytesser_v0.0.1.zip

测试pytesser

到pytesser的安装目录，创建一个test.py，python test.py 查看结果。

from pytesser import *
#im = Image.open('fnord.tif')
#im = Image.open('phototest.tif')
#im = Image.open('eurotext.tif')
im = Image.open('fonts_test.png')
text = image_to_string(im)
print text

tesseract 目录还有其他tif文件，也可以复制过来测试，上面测试的tif，png文件正确识别出文字。

pytesser的验证码识别能力较低，只能对规规矩矩不歪不斜数字和字母验证码进行识别。测试了几个网站的验证码，显示 Empty page，看来用它来识别验证码是无望了。

测试发现提高对比度后再识别有助于提高识别准确率。

enhancer = ImageEnhance.Contrast(im)
im = enhancer.enhance(4)

http://www.cnblogs.com/congbo/archive/2012/10/31/2746943.html

参考：

http://www.oschina.net/question/54100_59400

http://ubuntuforums.org/showthread.php?p=10248384

鹧鸪菜

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
ubuntu 12.04 下安装 PyTesser 进行OCR识别

ubuntu 12.04 下安装 PyTesser 进行OCR识别安装所需的库sudo apt-get install libpng12-devsudo apt-get install libjpeg62-devsudo apt-get install libtiff4-devsudo apt-get install gccsudo apt-get ins
复制链接

扫一扫