Neural disambiguation of lemma and part of speech in morphologically rich languages.

José María Hoya Quecedo,Maximilian W. Koppatz,Giacomo Furlan,Roman Yangarber

Neural disambiguation of lemma and part of speech in morphologically rich languages.

2020

José María Hoya Quecedo
Maximilian W. Koppatz
Giacomo Furlan
Roman Yangarber

We consider the problem of disambiguating the lemma and part of speech of ambiguous words in morphologically rich languages. We propose a method for disambiguating ambiguous words in context, using a large un-annotated corpus of text, and a morphological analyser -- with no manual disambiguation or data annotation. We assume that the morphological analyser produces multiple analyses for ambiguous words. The idea is to train recurrent neural networks on the output that the morphological analyser produces for unambiguous words. We present performance on POS and lemma disambiguation that reaches or surpasses the state of the art -- including supervised models -- using no manually annotated data. We evaluate the method on several morphologically rich languages.

Keywords:

Data Annotation
Analyser
Natural language processing
Lemma (mathematics)
Artificial intelligence
Part of speech
Computer science
Recurrent neural network

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations