Wiki-40B: Multilingual Language Model Dataset

Mandy Guo,Zihang Dai,Denny Vrandecic,Rami Al-Rfou

Wiki-40B: Multilingual Language Model Dataset

2020

Mandy Guo
Zihang Dai
Denny Vrandecic
Rami Al-Rfou

We propose a new multilingual language model benchmark that is composed of 40+ languages spanning several scripts and linguistic families. With around 40 billion characters, we hope this new resource will accelerate the research of multilingual modeling. We train monolingual causal language models using a state-of-the-art model (Transformer-XL) establishing baselines for many languages. We also introduce the task of multilingual causal language modeling where we train our model on the combined text of 40+ languages from Wikipedia with different vocabulary sizes and evaluate on the languages individually. We released the cleaned-up text of 40+ Wikipedia language editions, the corresponding trained monolingual language models, and several multilingual language models with different fixed vocabulary sizes.

Keywords:

Language model
Artificial intelligence
Natural language processing
Computer science

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations