FlauBERT: Unsupervised Language Model Pre training for French FlauBERT is a French BERT trained on a very large and heterogeneous French corpus. Models of different sizes are trained using the new CNRS (French National Centre for Scientific Research) Jean Zay supercomputer. Along with FlauBERT comes FLUE : an evaluation setup for French NLP systems similar to the popular GLUE benchmark. The goal is to enable further reproducible experiments in the future and to share models and progress on the French language.For more details please refer to the official website. FlauBERT models Model name Number of layers Attention Heads Embedding Dimension Total Parameters : : : : : : : : : : flaubert small cased 6 8 512 54 M flaubert base uncased 12 12 768 137 M flaubert base cased 12 12 768 138 M flaubert large cased 24 16 1024 373 M Note: flaubert small cased is partially trained so performance is not guaranteed. Consider using it for debugging purpose only. Using FlauBERT with Hugging Face's Transformers Notes: if your transformers version is <=2.10.0, modelname should take one of the following values: References If you use FlauBERT or the FLUE Benchmark for your scientific publication, or if…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy