opencv-ocr识别

最新推荐文章于 2024-08-29 14:45:23 发布

ic之芯

最新推荐文章于 2024-08-29 14:45:23 发布

阅读量83

点赞数 1

分类专栏： opencv 文章标签： opencv ocr python

本文链接：https://blog.csdn.net/qq_65838372/article/details/132457724

版权

opencv 专栏收录该内容

21 篇文章 1 订阅 ¥19.90 ¥99.00

订阅专栏

超级会员免费看

本文主要介绍了使用OpenCV库结合OCR技术进行文字识别的过程。首先，我们预处理图像以提高文字识别的准确性，然后利用Tesseract OCR引擎进行文字检测与提取。通过Python实现，详细步骤包括图像二值化、噪声去除等，最终成功识别并输出图像中的文本。

摘要由CSDN通过智能技术生成

# https://digi.bib.uni-mannheim.de/tesseract/
# 配置环境变量如E:\Program Files (x86)\Tesseract-OCR
# tesseract -v进行测试
# tesseract XXX.png 得到结果 
# pip install pytesseract
# anaconda lib site-packges pytesseract pytesseract.py
# tesseract_cmd 修改为绝对路径即可
from PIL import Image
import pytesseract
import cv2
import os

preprocess = 'blur' #thresh

image = cv2.imread('images\page.jpg')
gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)

if preprocess == "thresh":
    gray = cv2.threshold(gray, 0, 255,cv2.THRESH_BINARY | cv2.THRESH_OTSU)[1]

if preprocess == "blur":
    gray = cv2.medianBlur(gray, 3)
    
filename = "{}.png".format(os.getpid())
cv2.imwrite(filename, gray)
    
text = pytesseract.image_to_string(Image.open(filename))
print(text)
os.remove(filename)

cv2.imshow("Image", image)
cv2.imshow("Output", gray)
cv2.waitKey(0)