QAmembert Model Description We present QAmemBERT , which is a CamemBERT base fine tuned for the Question Answering task for the French language on four French Q&A datasets composed of contexts and questions with their answers inside the context (= SQuAD 1.0 format) but also contexts and questions with their answers not inside the context (= SQuAD 2.0 format). All these datasets were concatenated into a single dataset that we called frenchQA. This represents a total of over 221,348 context/question/answer triplets used to finetune this model and 6,376 to test it . Our methodology is described in a blog post available in English or French. Datasets Dataset Format Train split Dev split Test split piaf SQuAD 1.0 9 224 Q & A X X piaf v2 SQuAD 2.0 9 224 Q & A X X fquad SQuAD 1.0 20 731 Q & A 3 188 Q & A (not used in training because it serves as a test dataset) 2 189 Q & A (not used in our work because not freely available) fquad v2 SQuAD 2.0 20 731 Q & A 3 188 Q & A (not used in training because it serves as a test dataset) X lincoln/newsquadfr SQuAD 1.0 1 650 Q & A 455 Q & A (not used in our work) X lincoln/newsquadfr v2 SQuAD 2.0 1 650 Q & A 455 Q & A (not used in our work) X pragnaka…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy