LEGAL BERT: The Muppets straight out of Law School LEGAL BERT is a family of BERT models for the legal domain, intended to assist legal NLP research, computational law, and legal technology applications. To pre train the different variations of LEGAL BERT, we collected 12 GB of diverse English legal text from several fields (e.g., legislation, court cases, contracts) scraped from publicly available resources. Sub domain variants (CONTRACTS , EURLEX , ECHR ) and/or general LEGAL BERT perform better than using BERT out of the box for domain specific tasks. This is the sub domain variant pre trained on US contracts. I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras and I. Androutsopoulos. "LEGAL BERT: The Muppets straight out of Law School". In Findings of Empirical Methods in Natural Language Processing (EMNLP 2020) (Short Papers), to be held online, 2020. (https://aclanthology.org/2020.findings emnlp.261) Pre training corpora The pre training corpora of LEGAL BERT include: 116,062 documents of EU legislation, publicly available from EURLEX (http://eur lex.europa.eu), the repository of EU Law running under the EU Publication Office. 61,826 documents of UK legislation, publicl…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy