We introduce BERTurk Legal which is a transformer based language model to retrieve prior legal cases. BERTurk Legal is pre trained on a dataset from the Turkish legal domain. This dataset does not contain any labels related to the prior court case retrieval task. Masked language modeling is used to train BERTurk Legal in a self supervised manner. With zero shot classification, BERTurk Legal provides state of the art results on the dataset consisting of legal cases of the Court of Cassation of Turkey. The results of the experiments show the necessity of developing language models specific to the Turkish law domain. Details of BERTurk Legal can be found in the paper mentioned in the Citation section below. Test dataset can be accessed from the following link: https://github.com/koc lab/yargitay retrieval dataset The model can be loaded and used to create document embeddings as follows. Then, the document embeddings can be utilized for retrieval. Citation If you use the model, please cite the following conference paper.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy