video vec2wav2 tokenizer 2 Version 2 continuation shard of the video to AI dataset tokenizer project. Production ready pipeline (Python package video vec2wav2 tokenizer , CLI command video2dataset ) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text to speech (TTS) . Video processing — recursive scan of mp4 / mkv / avi / mov / webm , FFmpeg audio extraction to mono · 16 kHz · 16 bit PCM WAV. Speech recognition — faster whisper, CPU & CUDA, automatic language detection, word level timestamps. Segmentation — cut audio by transcript timestamps into dataset/audio/000001.wav … . Dataset generation — metadata.csv , dataset.jsonl , tts metadata.csv . Feature extraction (optional) — streaming features/train.bin + train.dat with float32 samples, mel spectrograms, duration and sample rate. Statistics — report.json with totals, durations and language distribution. Training — train wav2vec2.py : HuggingFace Wav2Vec2 CTC with resume, multi GPU, mixed precision and checkpointing. Performance — multiprocessing, batch processing, tqdm progress bars and memory efficient streaming that scales past 1 TB of source media. Installation Requires Python 3.1…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy