Python实现提取表格（1）

weixin_49320263

已于 2024-03-11 20:50:56 修改

阅读量1.2k

点赞数 9

分类专栏： python笔记文章标签： python

于 2024-03-10 15:22:03 首次发布

本文链接：https://blog.csdn.net/weixin_49320263/article/details/136592731

版权

python笔记专栏收录该内容

5 篇文章

订阅专栏

现实中常常会遇到从图片或pdf中提取表格，下面我们解释如何使用python实现。

#从图片中提取表格
from img2table.document import Image
# Instantiation of the image
img = Image(src="1.jpg")
from img2table.ocr import TesseractOCR
# Instantiation of the OCR, Tesseract, which requires prior installation
languages = "eng+chi_sim"
ocr = TesseractOCR(lang=languages)
# Table identification
img_tables = img.extract_tables(ocr=ocr,borderless_tables=True)
# Result of table identification
img_tables

#从PDF中提取表格
from img2table.document import PDF
from img2table.ocr import TesseractOCR
# Instantiation of the pdf
pdf = PDF(src="11.pdf")
# Instantiation of the OCR, Tesseract, which requires prior installation
ocr = TesseractOCR(n_threads=3,lang="chi_sim")
# Table identification and extraction
pdf_tables = pdf.extract_tables(ocr=ocr,borderless_tables=True,min_confidence=99)
# We can also create an excel file with the tables
pdf.to_xlsx(dest='tables.xlsx',ocr=ocr,borderless_tables=True,min_confidence=99)

使用paddleocr精确度更高，需要安装pip install img2table[paddle]

#调用PaddleOCR---准确度更高
from img2table.ocr import PaddleOCR
ocr = PaddleOCR(lang="ch")
img = Image(src="3.png")
# Table identification
img_tables = img.extract_tables(ocr=ocr,borderless_tables=True)
# Result of table identification
img_tables

高级玩法---训练自己的识别字库：可参考下面的链接中的文章，非常详细，尝试可以。

jTessBoxEditor工具安装和使用操作-CSDN博客