The LJ Speech Dataset Version 1.0 July 5, 2017 https://keithito.com/LJ Speech Dataset Overview This is a public domain speech dataset consisting of 13,100 short audio clips of a single speaker reading passages from 7 non fiction books. A transcription is provided for each clip. Clips vary in length from 1 to 10 seconds and have a total length of approximately 24 hours. The texts were published between 1884 and 1964, and are in the public domain. The audio was recorded in 2016 17 by the LibriVox project and is also in the public domain. The following files provide raw lavels for the train/validation/test split train.txt valid.txt test.txt Friendly metadata with the split is provided in the following files: ljspeech train.json ljspeech test.json ljspeech valid.json The JSON files are formatted as follows: The dataset is also usable as a HuggingFace Arrow dataset: https://huggingface.co/docs/datasets/ FILE FORMAT Original metadata is provided in metadata.csv. This file consists of one record per line, delimited by the pipe character (0x7c). The fields are: 1. ID: this is the name of the corresponding .wav file 2. Transcription: words spoken by the reader (UTF 8) 3. Normalized Transcri…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy