MUSE News MUSE is a comprehensive machine unlearning evaluation benchmark that assesses six key properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. MUSE focuses on two types of textual data that commonly require unlearning: news articles (News) and novels (Books). This repository contains the News corpus of MUSE (MUSE News), which comprises BBC articles collected post August 2023 . Details on Subsets & Splits MUSE News consists of 7 subsets: raw , verbmem , knowmem , privleak , scal , sust , and train . raw : A raw corpus from which all subsets except scal and sust are derived. The splits are: forget : Data intended to be forgotten retain1 : Data used optionally as a calibrator for unlearning retain2 : Retain set, i.e. data seen by the target model and used for evaluation holdout : Data never seen by the target model during pre training and unlearning verbmem : Evaluates verbatim memorization (C1) . It contains a single split forget with 1…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy