SpringBoot整合图片文字识别

最新推荐文章于 2024-08-01 10:09:35 发布

㏒灵韵№

最新推荐文章于 2024-08-01 10:09:35 发布

阅读量1.1k

点赞数

本文链接：https://blog.csdn.net/daai5201314/article/details/128724699

版权

项目模块专栏收录该内容

13 篇文章 1 订阅

订阅专栏

OCR是一种光学字符识别技术，用于将纸质文档转换为可编辑的电子文本。Tesseract是一个开源OCR引擎，支持多种编程语言调用，如Java和Python。Tess4J是它的Java封装库，本文介绍了如何在Java项目中使用Tess4J进行OCR识别，包括导入依赖、设置字体库和执行识别的步骤，并提供了代码示例和配置YML文件的方法。

摘要由CSDN通过智能技术生成

什么是OCR?

OCR （Optical Character Recognition，光学字符识别）是指电子设备（例如扫描仪或数码相机）检查纸上打印的字符，通过检测暗、亮的模式确定其形状，然后用字符识别方法将形状翻译成计算机文字的过程

方案	说明
百度OCR	收费
Tesseract-OCR	Google维护的开源OCR引擎，支持Java，Python等语言调用
Tess4J	封装了Tesseract-OCR ，支持Java调用

Tess4j案例

①：创建项目导入tess4j对应的依赖

<dependency>
    <groupId>net.sourceforge.tess4j</groupId>
    <artifactId>tess4j</artifactId>
    <version>4.1.1</version>
</dependency>

②：导入中文字体库，把资料中的tessdata文件夹拷贝到自己的工作空间下
链接：https://pan.baidu.com/s/1nMU-LVNp6Tta71AaNGj_sg
提取码：1111
–来自百度网盘超级会员V2的分享

③：编写测试类进行测试

import net.sourceforge.tess4j.ITesseract;
import net.sourceforge.tess4j.Tesseract;

import java.io.File;

public class Application {

    public static void main(String[] args) {
        try {
            //获取本地图片
            File file = new File("D:\\26.png");
            //创建Tesseract对象
            ITesseract tesseract = new Tesseract();
            //设置字体库路径
            tesseract.setDatapath("D:\\workspace\\tessdata");
            //中文识别
            tesseract.setLanguage("chi_sim");
            //执行ocr识别
            String result = tesseract.doOCR(file);
            //替换回车和tal键  使结果为一行
            result = result.replaceAll("\\r|\\n","-").replaceAll(" ","");
            System.out.println("识别的结果为："+result);
        } catch (Exception e) {
            e.printStackTrace();
        }
    }
}

抽取成工具类：


import lombok.Getter;
import lombok.Setter;
import net.sourceforge.tess4j.ITesseract;
import net.sourceforge.tess4j.Tesseract;
import net.sourceforge.tess4j.TesseractException;
import org.springframework.boot.context.properties.ConfigurationProperties;
import org.springframework.stereotype.Component;

import java.awt.image.BufferedImage;

@Getter
@Setter
@Component
@ConfigurationProperties(prefix = "tess4j")
public class Tess4jClient {

    private String dataPath;
    private String language;

    public String doOCR(BufferedImage image) throws TesseractException {
        //创建Tesseract对象
        ITesseract tesseract = new Tesseract();
        //设置字体库路径
        tesseract.setDatapath(dataPath);
        //中文识别
        tesseract.setLanguage(language);
        //执行ocr识别
        String result = tesseract.doOCR(image);
        //替换回车和tal键  使结果为一行
        result = result.replaceAll("\\r|\\n", "-").replaceAll(" ", "");
        return result;
    }

}

yml配置：

tess4j:
  data-path: D:\workspace\tessdata
  language: chi_sim

㏒灵韵№

关注

0
点赞
踩
5

收藏

觉得还不错? 一键收藏
0
评论
复制链接

分享到 QQ

分享到新浪微博

扫一扫

专栏目录