Voice "Cloning" is Style Transfer — Audio Dataset Companion dataset for the preprint "Voice 'Cloning' is Style Transfer" (Zhou, Bianchi, Bartelds, Pot, Kwon, Zou; 2026). Code, notebooks, and reproduction figures live at github.com/kzhou cloud/voice cloning public. 🎧 Listen to a small set of paired examples on the project page. What's in here Split files Description : original 699 QC validated human recordings of the Grandfather Passage from 86 non native English speakers, split into sentence level clips cloned 2,270 Step 0 voice clones generated from the sources using ChatterBox, Coqui XTTS, and ElevenLabs V3 (cross sentence cloning paradigm) cloned styles 9,048 Chatterbox clones generated under four different style settings (high/low similarity, expressiveness) — ablation for §4.1 of the paper cloned iterative chatterbox 18,963 Clones of clones: 50 rounds of repeated Chatterbox cloning for 43 speakers × 9 sentences Total 30,980 wavs / ≈10 GB All speakers and annotators are referenced by anonymized IDs ( speaker 001..speaker 086 ). The mapping between anonymized IDs and the raw upstream Prolific IDs is not distributed. Directory layout {model} ∈ {chatterbox, coqui xtts, elevenlabs…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy