LMD AI Generated Music Detection Benchmark (Note: The corresponding research paper will be released later.) Dataset Description The rapid advancement of AI music generation has raised growing concerns about the authenticity of digital music. While deepfake detection has been extensively studied in the audio domain, symbolic music (MIDI) remains largely unexplored. This dataset presents a comprehensive benchmark for AI generated symbolic music detection , examining how input representations, model architectures, and training compositions affect detection performance and generalizability. We evaluate three input representations — statistical features, piano roll, and event sequences — across diverse model structures. Dataset Sources We constructed a dataset of 5,355 human composed tracks (De duplicated Lakh MIDI) and over 14,000 AI generated MIDI and MP3 files from diverse pipelines, including: Text to MIDI models: MIDI LLM, Text2MIDI Audio to MIDI transcriptions of AI generated audio: Suno v4, Suno v5, Yue Dataset Structure The data is provided in its raw file format to support various MIR research pipelines. All files are organized within the data/ directory, maintaining their orig…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy