pdfplumber 提取pdf中文 Python

最新推荐文章于 2025-07-11 14:31:37 发布

李同学va

最新推荐文章于 2025-07-11 14:31:37 发布

阅读量1.2k

点赞数

CC 4.0 BY-SA版权

分类专栏：工作 Python 文章标签：大数据 pdf pdfplumber

本文链接：https://blog.csdn.net/weixin_43629889/article/details/125674053

工作同时被 2 个专栏收录

3 篇文章

订阅专栏

Python

2 篇文章

订阅专栏

该博客介绍了如何利用python的pdfplumber库来提取PDF文件中的文本内容。代码示例展示了如何打开PDF，遍历每一页并提取文本，但请注意此库无法解析图片中的文字，对于图片内容的识别需要借助OCR技术。

基于pdfplumber库来识别pdf中文字内容

无法识别pdf中图片的内容如果需要解析图片内容需要使用OCR技术

1. 详细代码以及注释

import pdfplumber


def extract_content(pdf_path):
    # 内容提取，使用 pdfplumber 打开 PDF，用于提取文本
    with pdfplumber.open(pdf_path) as pdf_file:

        content = ''
        print(len(pdf_file.pages))
        # len(pdf.pages)为PDF文档页数，一页页解析
        for i in range(len(pdf_file.pages)):
            print("当前第 %s 页" % i)
            # pdf.pages[i] 是读取PDF文档第i+1页
            page_text = pdf_file.pages[i]
            # page.extract_text()函数即读取文本内容
            page_content = page_text.extract_text()
            if page_content:
                content = content + page_content + "\n"
    

if __name__ == '__main__':
    pdf_file = '1.pdf'
    extract_content(pdf_file)