FPSC — Faroese Parliament Speech Corpus FPSC is a large scale Faroese parliamentary speech corpus constructed from publicly available recordings from Løgtingið , the Parliament of the Faroe Islands. The dataset contains approximately: 1,600 hours of speech 89,000+ parliamentary speeches 368 parliamentary meetings 75 unique speakers speaker demographic metadata dialect metadata machine generated weak transcripts ROVER voting metadata audio aligned at speech level The corpus was created as part of the paper: FPSC: A Sustainable Pipeline for Building a Faroese Parliamentary Speech Corpus Dávid í Lág, Barbara Scalvini, Carlos Mena, Jón Guðnason LREC 2026 The dataset represents the first large scale corpus of natural spoken Faroese and is intended for: Automatic Speech Recognition (ASR) Low resource speech technology Parliamentary speech analysis Sociolinguistic research Dialect research Weakly supervised ASR training Continual pretraining Multilingual transfer learning Dataset Structure The dataset follows the Hugging Face Audio dataset format and contains one row per parliamentary speech segment. Each entry contains: segmented WAV audio machine generated transcript parliamentary metad…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy