Subwords-Only Alternatives to fastText for Morphologically Rich Languages

Tsolak Ghukasyan,Yeva Yeshilbashyan,Karen Avetisyan

Subwords-Only Alternatives to fastText for Morphologically Rich Languages

2021

Tsolak Ghukasyan
Yeva Yeshilbashyan
Karen Avetisyan

In this work, we present purely subword-based alternatives to fastText word embedding algorithm The alternatives are modifications of the original fastText model, but rely on subword information only, eliminating the reliance on word-level vectors and at the same time helping to dramatically reduce the size of embeddings. Proposed models differ in their subword information extraction method: character n-grams, suffixes, and the byte-pair encoding units. We test the models in the task of morphological analysis and lemmatization for 3 morphologically rich languages: Finnish, Russian, and German. The results are compared with other recent subword-based models, demonstrating consistently higher results.

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations