TOFU: Task of Fictitious Unlearning 🍢 The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT 4 model. The goal of the task is to unlearn a fine tuned model on various fractions of the forget set. Quick Links Website : The landing page for TOFU arXiv Paper : Detailed information about the TOFU dataset and its significance in unlearning tasks. GitHub Repository : Access the source code, fine tuning scripts, and additional resources for the TOFU dataset. Dataset on Hugging Face : Direct link to download the TOFU dataset. Leaderboard on Hugging Face Spaces : Current rankings and submissions for the TOFU dataset challenges. Summary on Twitter : A concise summary and key takeaways from the project. Applicability 🚀 The dataset is in QA format, making it ideal for use with popular chat models such as Llama2, Mistral, or Qwen. However, it also works for any other large language model. The corresponding code base is written for the Llama2 chat, and Phi 1.5 model…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy