日语转文字python,如何读取亚洲语言（中文，日文，泰文等）的PDF文件并以python字符串形式存储...

最新推荐文章于 2023-12-18 10:17:13 发布

渚熏

最新推荐文章于 2023-12-18 10:17:13 发布

阅读量398

点赞数

文章标签：日语转文字python

I am using PyPDF2 to read PDF files in python. While it works well for languages in English and European languages (with alphabets in english), the library fails to read Asian languages like Japanese and Chinese. I tried encode('utf-8'), decode('utf-8') but nothing seems to work. It just prints a blank string on extraction of the text.

I have tried other libraries like textract and PDFMiner but no success yet.

When I copy the text from PDF and paste it on a notebook, the characters turn into some random format text (probably in a different encoding).

def convert_pdf_to_text(filename):

text = ''

pdf = PyPDF2.PdfFileReader(open(filename, "rb"))

if pdf.isEncrypted:

pdf.decrypt('')

for page in pdf.pages:

text = text + page.extractText()

return text

Can anyone point me in the right direction?

解决方案

I too faced similar issue. I could resolve it by using 'tika-python' library.

import tika

tika.initVM()

from tika import parser

parsed = parser.from_file('fileName.pdf')

print(parsed["metadata"])

print(parsed["content"])

You can find more information about the library over here

关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
日语转文字python,如何读取亚洲语言（中文，日文，泰文等）的PDF文件并以python字符串形式存储...

I am using PyPDF2 to read PDF files in python. While it works well for languages in English and European languages (with alphabets in english), the library fails to read Asian languages like Japanese ...
复制链接

扫一扫

评论

被折叠的条评论为什么被折叠?

到【灌水乐园】发言

查看更多评论

添加红包

成就一亿技术人!

hope_wisdom

发出的红包

实付元

使用余额支付

点击重新获取

扫码支付

钱包余额 0

抵扣说明：

1.余额是钱包充值的虚拟货币，按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载，可以购买VIP、付费专栏及课程。