MEL: Legal Spanish Language Model Model Name: MEL (Modelo de Español Legal) Model Type: Encoder only Transformer Language: Spanish Domain: Legal Texts Paper: Link to paper Overview MEL is a transformer based language model designed specifically for processing and understanding Spanish legal texts. Built upon XLM RoBERTa large , it is further pre trained on a large corpus of legal documents , including the Boletín Oficial del Estado (BOE), parliamentary transcripts, court rulings, and other legislative texts . MEL significantly improves the performance of legal NLP tasks, such as legal text classification and named entity recognition (NER) . Model Description Architecture Base Model: XLM RoBERTa large Training Objective: Masked Language Modeling (MLM) Pre training Strategy: Continued pre training on Spanish legal texts Context Window: 512 tokens Training Data MEL is trained on a curated corpus of 5.52 million legal texts (~92.7GB) sourced from: BOE (Boletín Oficial del Estado) Parliamentary records Court rulings Legal statutes To ensure high quality text processing, documents were preprocessed by removing unwanted characters, normalizing spacing, chunking texts, and filtering non Sp…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy