LEGAL BERT: The Muppets straight out of Law School LEGAL BERT is a family of BERT models for the legal domain, intended to assist legal NLP research, computational law, and legal technology applications. To pre train the different variations of LEGAL BERT, we collected 12 GB of diverse English legal text from several fields (e.g., legislation, court cases, contracts) scraped from publicly available resources. Sub domain variants (CONTRACTS , EURLEX , ECHR ) and/or general LEGAL BERT perform better than using BERT out of the box for domain specific tasks. A light weight model (33% the size of BERT BASE) pre trained from scratch on legal data with competitive performance is also available. I. Chalkidis, M. Fergadiotis, P. Malakasiotis, N. Aletras and I. Androutsopoulos. "LEGAL BERT: The Muppets straight out of Law School". In Findings of Empirical Methods in Natural Language Processing (EMNLP 2020) (Short Papers), to be held online, 2020. (https://aclanthology.org/2020.findings emnlp.261) Pre training corpora The pre training corpora of LEGAL BERT include: 116,062 documents of EU legislation, publicly available from EURLEX (http://eur lex.europa.eu), the repository of EU Law running…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy