Multilingual IPTC Media Topic Classifier News topic classification model based on xlm roberta large and fine tuned on a news corpus in 4 languages (Croatian, Slovenian, Catalan and Greek), annotated with the top level IPTC Media Topic NewsCodes labels. The development and evaluation of the model is described in the paper LLM Teacher Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification (Kuzman and Ljubešić, 2025). The model can be used for classification into topic labels from the IPTC NewsCodes schema and can be applied to any news text in a language, supported by the xlm roberta large . Based on a manually annotated test set (in Croatian, Slovenian, Catalan and Greek), the model achieves macro F1 score of 0.746, micro F1 score of 0.734, and accuracy of 0.734, and outperforms the GPT 4o model (version gpt 4o 2024 05 13 ) used in a zero shot setting. If we use only labels that are predicted with a confidence score equal or higher than 0.90, the model achieves micro F1 and macro F1 of 0.80. Intended use and limitations For reliable results, the classifier should be applied to documents of sufficient length (the rule…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy