运用前面几篇文章中的分词算法,可以把中文中的词语分出来,但是不同算法可能得到的分词结果不一样,到底如何确定那种分词效果最好呢。我们在这篇文章中探讨一种判断分词效果好坏的方法。
在分词的时候,有些单字是不成词的。可以搜索所有分词的可能性,然后对每一种分词结果进行统计,规则如下:每分出一个词就给分词统计结果加一,如果遇到不成词的单字就再给该分词结果加一。得到的结果分别计算出来。选出一个得分最低的就是所要分词的字符串的结果。
下面就用代码实现一下上面的思想。分别用正向最大匹配和逆向最大匹配,然后在对分词结果进行统计,选出一个最好的结果
package com;
import java.util.ArrayList;
import java.util.List;
public class Segmentation4 {
private List<String> dictionary = new ArrayList<String>();
private List<String> notWords = new ArrayList<String>();
private static String request = "他说的确实在理";
public void setDictionary() {
dictionary.add("的确");
dictionary.add("确实");
dictionary.add("实在");
dictionary.add("在理");
}
public void setNotWords() {
notWords.add("确");
notWords.add("实");
notWords.add("理");
}
public String leftMax() {
String response = "";
String s = "";
for(int i=0; i<request.length(); i++) {
s += request.charAt(i);
if(isIn(s, dictionary) && aheadCount(s, dictionary)==1) {
response += (s + "/");
s = "";
} else if(aheadCount(s, dictionary) > 0) {
} else {
response += (s + "/");
s = "";
}
}
return response;
}
private boolean isIn(String s, List<String> list) {
for(int i=0; i<list.size(); i++) {
if(s.equals(list.get(i))) return true;
}
return false;
}
private int aheadCount(String s, List<String> list) {
int count = 0;
for(int i=0; i<list.size(); i++) {
if((s.length()<=list.get(i).length()) && (s.equals(list.get(i).substring(0, s.length())))) count ++;
}
return count;
}
public String rightMax() {
String response = "";
String s = "";
for(int i=request.length()-1; i>=0; i--) {
s = request.charAt(i) + s;
if(isIn(s, dictionary) && tailCount(s, dictionary)==1) {
response = (s + "/") + response;
s = "";
} else if(tailCount(s, dictionary) > 0) {
} else {
response = (s + "/") + response;
s = "";
}
}
return response;
}
private int tailCount(String s, List<String> list) {
int count = 0;
for(int i=0; i<list.size(); i++) {
if((s.length()<=list.get(i).length()) && (s.equals(list.get(i).substring(list.get(i).length()-s.length(), list.get(i).length())))) count ++;
}
return count;
}
public int getCount(String s) {
String[] words = s.split("/");
int count = words.length;
for(String word : words) {
if(isIn(word, notWords)) count++;
}
return count;
}
public static void main(String[] args) {
System.out.println(request);
String response;
Segmentation4 seg = new Segmentation4();
seg.setDictionary();
seg.setNotWords();
String response1 = seg.leftMax();
System.out.println(response1);
String response2 = seg.rightMax();
System.out.println(response2);
if(seg.getCount(response1)<=seg.getCount(response2))
response = response1;
else response = response2;
System.out.println(response);
}
}
上面程序运行的结果是:
他说的确实在理
他/说/的确/实在/理/
他/说/的/确实/在理/
他/说/的/确实/在理/