CamemBERT: a Tasty French Language Model Introduction CamemBERT is a state of the art language model for French based on the RoBERTa model. It is now available on Hugging Face in 6 different versions with varying number of parameters, amount of pretraining data and pretraining data source domains. Pre trained models Model params Arch. Training data camembert base 110M Base OSCAR (138 GB of text) camembert/camembert large 335M Large CCNet (135 GB of text) camembert/camembert base ccnet 110M Base CCNet (135 GB of text) camembert/camembert base wikipedia 4gb 110M Base Wikipedia (4 GB of text) camembert/camembert base oscar 4gb 110M Base Subsample of OSCAR (4 GB of text) camembert/camembert base ccnet 4gb 110M Base Subsample of CCNet (4 GB of text) How to use CamemBERT with HuggingFace Load CamemBERT and its sub word tokenizer : Filling masks using pipeline Extract contextual embedding features from Camembert output Extract contextual embedding features from all Camembert layers Authors CamemBERT was trained and evaluated by Louis Martin\ , Benjamin Muller\ , Pedro Javier Ortiz Suárez\ , Yoann Dupont, Laurent Romary, Éric Villemonte de la Clergerie, Djamé Seddah and Benoît Sagot. Citat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy