LLMLingua 2 Bert base Multilingual Cased MeetingBank This model was introduced in the paper LLMLingua 2: Data Distillation for Efficient and Faithful Task Agnostic Prompt Compression (Pan et al, 2024). It is a XLM RoBERTa (large sized model) finetuned to perform token classification for task agnostic prompt compression. The probability $p {preserve}$ of each token $x i$ is used as the metric for compression. This model is trained on the extractive text compression dataset constructed with the methodology proposed in the LLMLingua 2 , using training examples from MeetingBank (Hu et al, 2023) as the seed data. You can evaluate the model on downstream tasks such as question answering (QA) and summarization over compressed meeting transcripts using this dataset. For more details, please check the home page of LLMLingua 2 and LLMLingua Series. Usage Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy