python文本分析pdf_使用Python进行PDF解析-提取格式化和纯文本

最新推荐文章于 2024-07-10 11:49:26 发布

weixin_39840606

最新推荐文章于 2024-07-10 11:49:26 发布

阅读量138

点赞数

文章标签： python文本分析pdf

I'm looking for a PDF library which will allow me to extract the text from a PDF document. I've looked at PyPDF, and this can extract the text from a PDF document very nicely. The problem with this is that if there are tables in the document, the text in the tables is extracted in-line with the rest of the document text. This can be problematic because it produces sections of text that aren't useful and look garbled (for instance, lots of numbers mashed together).

I'd like to extract the text from a PDF document, excluding any tables and special formatting. Is there a library out there that does this?

解决方案

You can also take a look at PDFMiner (or for older versions of Python see PDFMiner).

A particular feature of interest in PDFMiner is that you can control how it regroups text parts when extracting them. You do this by specifying the space between lines, words, characters, etc. So, maybe by tweaking this you can achieve what you want (that depends of the variability of your documents). PDFMiner can also give you the location of the text in the page, it can extract data by Object ID and other stuff. So dig in PDFMiner and be creative!

But your problem is really not an easy one to solve because, in a PDF, the text is not continuous, but made from a lot of small groups of characters positioned absolutely in the page. The focus of PDF is to keep the layout intact. It's not content oriented but presentation oriented.