Dataset Card for MultiLegalPile Wikipedia Filtered: A filtered version of the MultiLegalPile dataset, together with wikipedia articles Table of Contents Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: Repository: Paper: Leaderboard: Point of Contact: Joel Niklaus Dataset Summary The Multi Legal Pile is a large scale multilingual legal dataset suited for pretraining language models. It spans over 24 languages and four legal text types. Supported Tasks and Leaderboards The dataset supports the tasks of fill mask. Languages The following languages are supported: bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv Dataset Structure It is structured in the following format: {language} {text type} {shard}.jsonl.xz text ty…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy