判断文件字符集的简单方法

最新推荐文章于 2024-05-30 06:40:03 发布

haha0515

最新推荐文章于 2024-05-30 06:40:03 发布

阅读量305

点赞数 1

分类专栏： java 文章标签： HTML

本文链接：https://blog.csdn.net/haha0515/article/details/83710702

版权

java 专栏收录该内容

16 篇文章 0 订阅

订阅专栏

/**
	 * 
	 *	  ANSI：　　　　　　　　无格式定义；
	 *	  Unicode： 　　　　　　前两个字节为FFFE
	 *	  Unicode big endian：　前两字节为FEFF　 
	 *	  UTF-8：　 　　　　　　前两字节为EFBB
	 * @param file
	 * @return
	 */
	public static String get_charset(File file) {
		String charset = "GBK";
		byte[] first3Bytes = new byte[3];
		try {
			boolean checked = false;
			BufferedInputStream bis = new BufferedInputStream(
					new FileInputStream(file));
			bis.mark(0);
			int read = bis.read(first3Bytes, 0, 3);
			if (read == -1)
				return charset;
			if (first3Bytes[0] == (byte) 0xFF && first3Bytes[1] == (byte) 0xFE) {
				charset = "UTF-16LE";
				checked = true;
			} else if (first3Bytes[0] == (byte) 0xFE
					&& first3Bytes[1] == (byte) 0xFF) {
				charset = "UTF-16BE";
				checked = true;
			} else if (first3Bytes[0] == (byte) 0xEF
					&& first3Bytes[1] == (byte) 0xBB
					&& first3Bytes[2] == (byte) 0xBF) {
				charset = "UTF-8";
				checked = true;
			}
			bis.reset();
			if (!checked) {
				// int len = 0;
				int loc = 0;

				while ((read = bis.read()) != -1) {
					loc++;
					if (read >= 0xF0)
						break;
					if (0x80 <= read && read <= 0xBF) // 单独出现BF以下的，也算是GBK
						break;
					if (0xC0 <= read && read <= 0xDF) {
						read = bis.read();
						if (0x80 <= read && read <= 0xBF) // 双字节 (0xC0 - 0xDF)
															// (0x80
							// - 0xBF),也可能在GB编码内
							continue;
						else
							break;
					} else if (0xE0 <= read && read <= 0xEF) {// 也有可能出错，但是几率较小
						read = bis.read();
						if (0x80 <= read && read <= 0xBF) {
							read = bis.read();
							if (0x80 <= read && read <= 0xBF) {
								charset = "UTF-8";
								break;
							} else
								break;
						} else
							break;
					}
				}
				// System.out.println( loc + " " + Integer.toHexString( read )
				// );
			}

			bis.close();
		} catch (Exception e) {
			e.printStackTrace();
		}

		return charset;
	}

转至：http://ajava.org/code/I18N/14816.html

haha0515

关注

1
点赞
踩
0

收藏

觉得还不错? 一键收藏
0
评论
判断文件字符集的简单方法

/** * * ANSI：　　　　　　　　无格式定义； * Unicode：　　　　　　前两个字节为FFFE * Unicode big endian：　前两字节为FEFF　 * UTF-8：　　　　　　　前两字节为EFBB * @param file * @return */ public static String g...
复制链接

扫一扫