Developed by: Anuran Roy Model type: Seq2Seq Speech to Text (LoRA fine tuned) Language(s) (NLP): Hindi ( hi ), English ( en ), Bengali ( bn ), Tamil ( ta ), Telugu ( te ), Marathi ( mr ), Gujarati ( gu ), Kannada ( kn ), Malayalam ( ml ), Punjabi ( pa ), Urdu ( ur ), Odia ( or ) License: CC BY SA 4.0 Finetuned from model: openai/whisper small Uses Direct Use The model can be used directly for speech to text transcription of Indic language audio files. It supports both explicit language specification and automatic language detection. Inputs are taken at 16 kHz frequency (pcm16) Downstream Use The model can be integrated into voice agent pipelines, real time transcription services, or any application requiring Indic language speech recognition. A WebSocket based FastAPI server is provided for real time inference. Out of Scope Use Languages not listed in the supported languages above. Noisy or very low quality audio recordings may produce poor results. Not intended for speaker identification or diarization. How to Get Started with the Model Python Inference WebSocket Server A FastAPI based WebSocket server is included for real time transcription: The server exposes: GET /health — Heal…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy