XGLM 564M XGLM 564M is a multilingual autoregressive language model (with 564 million parameters) trained on a balanced corpus of a diverse set of 30 languages totaling 500 billion sub tokens. It was introduced in the paper Few shot Learning with Multilingual Language Models by Xi Victoria Lin\ , Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, Xian Li\ (\ Equal Contribution). The original implementation was released in this repository. Training Data Statistics The training data statistics of XGLM 564M is shown in the table below. ISO 639 1 family name tokens ratio ratio w/ lowRes upsampling : : : : : : en Indo European English 803526736124 0.489906 0.3259 ru Indo European Russian 147791898098 0.0901079 0.0602 zh Sino Tibetan Chinese 132770494630 0.0809494 0.0483 de Indo European German 89223707856 0.0543992 0.0363 es Indo European Spanish 87303083105 0.0532282 0.0353 fr Indo European French 77419639775 0.0472023 0.0313 ja Japonic Japanese 660543645…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy