xml读取异常Invalid byte 1 of 1-byte UTF-8 sequence

xml读取异常Invalid byte 1 of 1-byte UTF-8 sequence


说简单点当你解析别人的xml格式出现这个错误可能就是别人在生成xml时没有保存为utf-8的字符编码格式。

在中文版的window下java的默认的编码为GBK,也就是所虽然我们标识了要将xml保存为utf-8格式但实际上文件是以GBK格式来保存的,所以这也就是为什么能够我们使用GBK、GB2312编码来生成xml文件能正确的被解析,而以UTF-8格式生成的文件不能被xml解析器所解析的原因。


xml解析时遇到的编码异常:

  1. org.dom4j.DocumentException:Invalidbyte1of1-byteUTF-8sequence.Nestedexception:Invalidbyte1of1-byteUTF-8sequence.
  2. atorg.dom4j.io.SAXReader.read(SAXReader.java:484)
  3. atorg.dom4j.io.SAXReader.read(SAXReader.java:321)
  4. atcom.dataoperate.PaseXml.pXml(PaseXml.java:28)
  5. atcom.dataoperate.JdbcOp.insertDb(JdbcOp.java:30)
  6. atcom.dataoperate.JdbcOp.main(JdbcOp.java:89)
  7. Nestedexception:
  8. com.sun.org.apache.xerces.internal.impl.io.MalformedByteSequenceException:Invalidbyte1of1-byteUTF-8sequence.
  9. atcom.sun.org.apache.xerces.internal.impl.io.UTF8Reader.invalidByte(UTF8Reader.java:684)
  10. atcom.sun.org.apache.xerces.internal.impl.io.UTF8Reader.read(UTF8Reader.java:554)
  11. atcom.sun.org.apache.xerces.internal.impl.XMLEntityScanner.load(XMLEntityScanner.java:1742)
  12. atcom.sun.org.apache.xerces.internal.impl.XMLEntityScanner.peekChar(XMLEntityScanner.java:487)
  13. atcom.sun.org.apache.xerces.internal.impl.XMLDocumentFragmentScannerImpl$FragmentContentDriver.next(XMLDocumentFragmentScannerImpl.java:2687)
  14. atcom.sun.org.apache.xerces.internal.impl.XMLDocumentScannerImpl.next(XMLDocumentScannerImpl.java:648)
  15. atcom.sun.org.apache.xerces.internal.impl.XMLNSDocumentScannerImpl.next(XMLNSDocumentScannerImpl.java:140)
  16. atcom.sun.org.apache.xerces.internal.impl.XMLDocumentFragmentScannerImpl.scanDocument(XMLDocumentFragmentScannerImpl.java:511)
  17. atcom.sun.org.apache.xerces.internal.parsers.XML11Configuration.parse(XML11Configuration.java:808)
  18. atcom.sun.org.apache.xerces.internal.parsers.XML11Configuration.parse(XML11Configuration.java:737)
  19. atcom.sun.org.apache.xerces.internal.parsers.XMLParser.parse(XMLParser.java:119)
  20. atcom.sun.org.apache.xerces.internal.parsers.AbstractSAXParser.parse(AbstractSAXParser.java:1205)
  21. atcom.sun.org.apache.xerces.internal.jaxp.SAXParserImpl$JAXPSAXParser.parse(SAXParserImpl.java:522)
  22. atorg.dom4j.io.SAXReader.read(SAXReader.java:465)
  23. atorg.dom4j.io.SAXReader.read(SAXReader.java:321)
  24. atcom.dataoperate.PaseXml.pXml(PaseXml.java:28)
  25. atcom.dataoperate.JdbcOp.insertDb(JdbcOp.java:30)
  26. atcom.dataoperate.JdbcOp.main(JdbcOp.java:89)
  27. Nestedexception:com.sun.org.apache.xerces.internal.impl.io.MalformedByteSequenceException:Invalidbyte1of1-byteUTF-8sequence.
  28. atcom.sun.org.apache.xerces.internal.impl.io.UTF8Reader.invalidByte(UTF8Reader.java:684)
  29. atcom.sun.org.apache.xerces.internal.impl.io.UTF8Reader.read(UTF8Reader.java:554)
  30. atcom.sun.org.apache.xerces.internal.impl.XMLEntityScanner.load(XMLEntityScanner.java:1742)
  31. atcom.sun.org.apache.xerces.internal.impl.XMLEntityScanner.peekChar(XMLEntityScanner.java:487)
  32. atcom.sun.org.apache.xerces.internal.impl.XMLDocumentFragmentScannerImpl$FragmentContentDriver.next(XMLDocumentFragmentScannerImpl.java:2687)
  33. atcom.sun.org.apache.xerces.internal.impl.XMLDocumentScannerImpl.next(XMLDocumentScannerImpl.java:648)
  34. atcom.sun.org.apache.xerces.internal.impl.XMLNSDocumentScannerImpl.next(XMLNSDocumentScannerImpl.java:140)
  35. atcom.sun.org.apache.xerces.internal.impl.XMLDocumentFragmentScannerImpl.scanDocument(XMLDocumentFragmentScannerImpl.java:511)
  36. atcom.sun.org.apache.xerces.internal.parsers.XML11Configuration.parse(XML11Configuration.java:808)
  37. atcom.sun.org.apache.xerces.internal.parsers.XML11Configuration.parse(XML11Configuration.java:737)
  38. atcom.sun.org.apache.xerces.internal.parsers.XMLParser.parse(XMLParser.java:119)
  39. atcom.sun.org.apache.xerces.internal.parsers.AbstractSAXParser.parse(AbstractSAXParser.java:1205)
  40. atcom.sun.org.apache.xerces.internal.jaxp.SAXParserImpl$JAXPSAXParser.parse(SAXParserImpl.java:522)
  41. atorg.dom4j.io.SAXReader.read(SAXReader.java:465)
  42. atorg.dom4j.io.SAXReader.read(SAXReader.java:321)
  43. atcom.dataoperate.PaseXml.pXml(PaseXml.java:28)
  44. atcom.dataoperate.JdbcOp.insertDb(JdbcOp.java:30)
  45. atcom.dataoperate.JdbcOp.main(JdbcOp.java:89)
解决:

1、最简单就是把<?xml version="1.0" encoding="UTF-8"?>改成<?xml version="1.0" encoding="gbk"?>

2、或者把xml打开另存的时候把字符集改为UTF-8后保存

3、在代码解析的时候先把xml重新写一遍

[javascript] view plain copy
  1. SAXReaderreader=newSAXReader();
  2. org.dom4j.Documentdocument=reader.read("D:\\ha.xml");
  3. OutputFormatof=newOutputFormat();
  4. of.setEncoding("UTF-8");//改变编码方式
  5. XMLWriterwriter=newXMLWriter(newFileWriter"d:\\dom4j.xml"),of);

4、直接dom4j读取的时候用io来读,修改字符编码

  1. FileInputStreamin=newFileInputStream(newFile(fileName));
  2. Readerread=newInputStreamReader(in,"gbk");
  3. Documentdocument=reader.read(read);
评论
添加红包

请填写红包祝福语或标题

红包个数最小为10个

红包金额最低5元

当前余额3.43前往充值 >
需支付:10.00
成就一亿技术人!
领取后你会自动成为博主和红包主的粉丝 规则
hope_wisdom
发出的红包
实付
使用余额支付
点击重新获取
扫码支付
钱包余额 0

抵扣说明:

1.余额是钱包充值的虚拟货币,按照1:1的比例进行支付金额的抵扣。
2.余额无法直接购买下载,可以购买VIP、付费专栏及课程。

余额充值