Swecha Gonthuka ASR (Telugu) Telugu automatic speech recognition model (wav2vec2 based), trained on the Swecha Gonthuka dataset. It is evaluated on Telugu only test sets with Character Error Rate (CER). Model details Training data: Swecha Gonthuka dataset Language: Telugu (te) Metric: CER (Character Error Rate) — text normalized to Telugu script + spaces before scoring. Evaluation results Dataset Test samples CER (%) : : FLEURS (te in) 304 6.32 OpenSLR66 420 9.00 Common Voice 22 (te) 58 11.92 Note: For evaluation we used only those samples that contain no English words Telugu text only for each dataset, to allow a fair evaluation of model capability. Usage Python (Transformers) Responsible and ethical use Intended use: This model is intended for Telugu automatic speech recognition in applications such as transcription, accessibility, and language preservation. Use it in accordance with applicable laws and platform policies. Limitations: Performance may vary with accent, dialect, noise, and recording quality. Do not rely on it as the sole source for critical or legal transcriptions without human review. Misuse: Do not use this model to transcribe private conversations without consen…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy