Custom Legal BERT Model and tokenizer files for Custom Legal BERT model from When Does Pretraining Help? Assessing Self Supervised Learning for Law and the CaseHOLD Dataset. Training Data The pretraining corpus was constructed by ingesting the entire Harvard Law case corpus from 1965 to the present (https://case.law/). The size of this corpus (37GB) is substantial, representing 3,446,187 legal decisions across all federal and state courts, and is larger than the size of the BookCorpus/Wikipedia corpus originally used to train BERT (15GB). Training Objective This model is pretrained from scratch for 2M steps on the MLM and NSP objective, with tokenization and sentence segmentation adapted for legal text (cf. the paper). The model also uses a custom domain specific legal vocabulary. The vocabulary set is constructed using SentencePiece on a subsample (approx. 13M) of sentences from our pretraining corpus, with the number of tokens fixed to 32,000. Usage Please see the casehold repository for scripts that support computing pretrain loss and finetuning on Custom Legal BERT for classification and multiple choice tasks described in the paper: Overruling, Terms of Service, CaseHOLD. Citat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy