Whisper Large V3 Russian Podlodka This repository contains a fine tuned Whisper Large V3 model for Russian speech recognition. It serves as the core transcription component of the Pisets system, specifically optimized for long audio recordings such as lectures and interviews. The model was presented in the paper Pisets: A Robust Speech Recognition System for Lectures and Interviews. System Architecture The Pisets system implements a three component architecture to improve recognition accuracy while minimizing hallucinations: 1. Wav2Vec2 : For primary recognition and segmentation. 2. Audio Spectrogram Transformer (AST) : For filtering non speech segments. 3. Whisper (this model) : For the final high quality transcription. Implementation The complete source code and instructions for using the system (including generation of SRT and DocX files) can be found in the GitHub repository: GitHub: https://github.com/bond005/pisets Citation If you use this model or the Pisets system in your research, please cite:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy