命名实体识别NER

最新推荐文章于 2023-06-22 16:14:03 发布

sole_cc

最新推荐文章于 2023-06-22 16:14:03 发布

阅读量1.8k

点赞数

文章标签：斯坦福大学 entity

java 专栏收录该内容

13 篇文章 0 订阅

订阅专栏

自然语言分析之命名实体识别_Stanford Named Entity Recognizer (NER)简单实例

Stanford Named Entity Recognizer (NER)是斯坦福大学自然语言研究小组发布的成果之一，主页是：http://nlp.stanford.edu/software/CRF-NER.shtml

Stanford NER 是一个Java实现的命名实体识别（以下简称NER）)程序。NER将文本中的实体按类标记出来，例如人名，公司名，地区，基因和蛋白质的名字等。

NER基于一个训练而得的Model工作，用于训练的数据即大量人工标记好的文本，理论上用于训练的数据量越大，NER的识别效果就越好。

斯坦福小组给出了三个训练好的Model：

Location, Person, Organization
Location, Person, Organization, Misc
Time, Location, Organization, Person, Money, Percent, Date

但不幸的是，这三个Model都不能被扩展，用于训练的数据也不公开，所以想要一个适应自己需求的Model我们只能从头训练。

下面我们就用一个简单的例子来摸索一下如何训练一个新的Model并用它尝试识别几个简单的句子。

首先从主页下载Stanford Named Entity Recognizer：http://nlp.stanford.edu/software/stanford-ner-2013-11-12.zip 解压即可。

文中提到的训练数据、配置文件、Model和完整工程可从网盘下载：http://pan.baidu.com/s/1xNAqD

准备训练数据

这是本例所使用的训练数据，包含标记好的两句话：Today is Friday. Tomorrow is 11/30.
如图，句子中的每个单词独立成行，Tab后跟该单词的类别，默认为O。例中，我们标记了Friday为WEEK，11/30为DATA，其余均为默认。

训练获得Model

官方说虽然所有的参数都可以通过命令行的方式指定，但更推荐用配置文件的方式。

 
        #location of the training file 
       
        trainFile=testdata.tsv 
       
        #location where you would like to save (serialize to) your 
       
        #classifier; adding .gz at the end automatically gzips the file, 
       
        #making it faster and smaller 
       
        serializeTo=ner-model.ser.gz 
       
        #structure of your training file; this tells the classifier 
       
        #that the word is in column 0 and the correct answer is in 
       
        #column 1 
       
        map=word=0,answer=1 
       
        #these are the features we'd like to train with 
       
        #some are discussed below, the rest can be 
       
        #understood by looking at NERFeatureFactory 
       
        useClassFeature=true 
       
        useWord=true 
       
        useNGrams=true 
       
        #no ngrams will be included that do not contain either the 
       
        #beginning or end of the word 
       
        noMidNGrams=true 
       
        useDisjunctive=true 
       
        maxNGramLeng=6 
       
        usePrev=true 
       
        useNext=true 
       
        useSequences=true 
       
        usePrevSequences=true 
       
        maxLeft=1 
       
        #the next 4 deal with word shape features 
       
        useTypeSeqs=true 
       
        useTypeSeqs2=true 
       
        useTypeySequences=true 
       
        wordShape=chris2useLC

其中trainFile =testdata.tsv指定了我们用于训练的数据，serializeTo = ner-model.ser.gz指定了输出model的名字，其余参数也均有注释说明。

我们将该配置文件austen.prop和训练数据testdata.tsv都放到stanford-ner-2013-11-12文件夹中。

cmd中将当前目录切换到stanford-ner-2013-11-12文件夹，并执行命令：
java -cp stanford-ner.jar edu.stanford.nlp.ie.crf.CRFClassifier -prop austen.prop

执行成功后，我们就可以看到目录下多了ner-model.ser.gz，这就是通过我们的训练数据得到的model。

尝试识别

 
        importjava.io.IOException; 
       
        importedu.stanford.nlp.ie.AbstractSequenceClassifier; 
       
        importedu.stanford.nlp.ie.crf.CRFClassifier; 
       
        importedu.stanford.nlp.ling.CoreLabel; 
       
        publicclassNERDemo{ 
       
            publicstaticvoidmain(String[]args)throwsIOException{         
       
                StringserializedClassifier="ner-model.ser.gz"; 
       
                AbstractSequenceClassifierclassifier=CRFClassifier.getClassifierNoExceptions(serializedClassifier); 
       
                Strings1="The next day is Friday"; 
       
                System.out.println(classifier.classifyToString(s1)); 
       
                Strings2="Today is 11/29"; 
       
                System.out.println(classifier.classifyToString(s2)); 
       
            } 
       
        }