java如何调用diffpdf,使用Java比较两个pdf文件（方法）

最新推荐文章于 2024-05-22 10:02:39 发布

发狂小傻喵

最新推荐文章于 2024-05-22 10:02:39 发布

阅读量1.2k

点赞数

文章标签： java如何调用diffpdf

i need to write a java class that compares two pdf files and points out the differences(differences in text/position/font)

using some sort of highlighting.

my initial approach was use pdfbox to parse the file using pdfbox and store the extracted text using in some data structure that would help me with comparing.

Is there any java library that can extract the text,preserve the formatting,help me with indexing and comparing.Can i use tika/ google's diff-match for this.

tika extracts text in the form of xhtml but how can i compare two xhtml files?

解决方案

I had to compare tons of pdf files in my project. my requirement was to compare the pdf files by pixel by pixel. After a lot of googling and as i could not find anything good, I ended up creating my own pdf utility for this purpose.

Please check this blog for more details & jar download.

确定要放弃本次机会？

福利倒计时

: :

立减 ¥

普通VIP年卡可用

立即使用

发狂小傻喵

关注关注

0
点赞
踩
1

收藏

觉得还不错? 一键收藏
0
评论
java如何调用diffpdf,使用Java比较两个pdf文件（方法）

i need to write a java class that compares two pdf files and points out the differences(differences in text/position/font)using some sort of highlighting.my initial approach was use pdfbox to parse th...
复制链接

扫一扫