(BERT base) NER model in the legal domain in Portuguese (LeNER Br) ner bert base portuguese cased lenerbr is a NER model (token classification) in the legal domain in Portuguese that was finetuned on 20/12/2021 in Google Colab from the model pierreguillou/bert base cased pt lenerbr on the dataset LeNER br by using a NER objective. Due to the small size of BERTimbau base and finetuning dataset, the model overfitted before to reach the end of training. Here are the overall final metrics on the validation dataset ( note: see the paragraph "Validation metrics by Named Entity" to get detailed metrics ): f1 : 0.8926146010186757 precision : 0.8810222036028488 recall : 0.9045161290322581 accuracy : 0.9759397808828684 loss : 0.18803243339061737 Check as well the large version of this model with a f1 of 0.908. Note : the model pierreguillou/bert base cased pt lenerbr is a language model that was created through the finetuning of the model BERTimbau base on the dataset LeNER Br language modeling by using a MASK objective. This first specialization of the language model before finetuning on the NER task improved a bit the model quality. To prove it, here are the results of the NER model finetu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy