RobeCzech: Czech RoBERTa, a Monolingual Contextualized Language Representation Model.

Milan Straka,Jakub Náplava,Jana Straková,David Samuel

RobeCzech: Czech RoBERTa, a Monolingual Contextualized Language Representation Model.

2021

Milan Straka
Jakub Náplava
Jana Straková
David Samuel

We present RobeCzech, a monolingual RoBERTa language representation model trained on Czech data. RoBERTa is a robustly optimized Transformer-based pretraining approach. We show that RobeCzech considerably outperforms equally-sized multilingual and Czech-trained contextualized language representation models, surpasses current state of the art in all five evaluated NLP tasks and reaches state-of-the-art results in four of them. The RobeCzech model is released publicly at https://hdl.handle.net/11234/1-3691 and https://huggingface.co/ufal/robeczech-base.

Keywords:

Czech
transformer
Natural language processing
Artificial intelligence
language representation
State (computer science)
Computer science

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations