python 读取pdf cid,如何处理PDFMiner提取的文本中的CID？

最新推荐文章于 2022-10-25 09:58:38 发布

VIP文章好食捷

最新推荐文章于 2022-10-25 09:58:38 发布

阅读量700

点赞数

文章标签： python 读取pdf cid

I've some PDFs which are in Hindi, and have extractable text. I used pdfminer.six for python 3.6, to do the extraction. The output looks like:

As one can see, there are a number of characters that are converted into the form "(cid :number)".

On further analysis, I found out that a PDF contains CMAPs which map character codes to glyph indices. So, a CID is a character identity for the glyph it maps to, inside the CMAP table.

But how are these character codes related to Unicode values? Basically, how is a PDF viewer able to show the glyph using this mapping?

Moreover, according to a comment to this similar question, this process may not be legal. But I'm not trying to steal someone's font. I want t

最低0.47元/天解锁文章

优惠劵

好食捷

关注关注

0
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
python 读取pdf cid,如何处理PDFMiner提取的文本中的CID？

I've some PDFs which are in Hindi, and have extractable text. I used pdfminer.six for python 3.6, to do the extraction. The output looks like:As one can see, there are a number of characters that are ...
复制链接

扫一扫