IISc ARTPARK Project VAANI NeMo Conformer-TDT

ಕನ್ನಡ ವಾಕ್-ಗುರುತಿಸುವಿಕೆ FastConformer Acoustic Model

Trained on thousands of hours of authentic Karnataka dialect recordings under Project VAANI by ARTPARK & the Indian Institute of Science (IISc). Quantized to INT8 and packaged for zero-friction local ONNX inference by OpenVoiceOS.

Architecture
FastConformer-TDT
Duration Transducer
Subsampling
8x Factor
128 Mel filterbanks
INT8 Quantized
476 MB + 5.4 MB
Encoder + Decoder Joint
Target Language
Kannada (kn-IN)
Project VAANI corpus
WebML-Kit On-Device Engine (WASM SIMD / WebGPU)
OpenVoiceOS/artpark-iisc-vaani-fastconformer-kn-onnx · 100% Client-Side
⚡ Detecting Hardware... Checking Cache...
Zero cloud calls. Audio stays strictly local in your browser. Model weights (~476 MB INT8) are streamed directly from Hugging Face and saved in the browser's Cache API for instant subsequent sessions.

Acoustic Input Studio

Ready
00:00.0
Project VAANI Verified Dialects 4 Samples
Urban · Bengaluru ಬೆಂಗಳೂರು ಕರ್ನಾಟಕದ ರಾಜಧಾನಿ "Bengaluru is the capital of Karnataka"
Science · Technical ವಿಜ್ಞಾನ ಮತ್ತು ತಂತ್ರಜ್ಞಾನ ನಮ್ಮ ಭವಿಷ್ಯ "Science & technology shape our future"
Conversational ನಾಳೆ ಬೆಳಗ್ಗೆ ಹತ್ತು ಗಂಟೆಗೆ ಸಭೆ ನಡೆಯಲಿದೆ "The meeting will occur tomorrow at 10 AM"
Agriculture · Malnad ಈ ವರ್ಷ ಮುಂಗಾರು ಮಳೆ ಉತ್ತಮವಾಗಿ ಸುರಿದಿದೆ "This year monsoon rains were abundant"

Transcribed Kannada (ಲಿಪ್ಯಂತರಣ)

ಬೆಂಗಳೂರು ಕರ್ನಾಟಕದ ರಾಜಧಾನಿ
"Bengaluru is the capital of Karnataka"
Kannada Akshara & Subword Decomposition
▁ಬೆಂ #1024 ಗಳೂ #2491 ರು #6 ▁ಕರ್ನಾ #891 ಟಕದ #421 ▁ರಾಜ #310 ಧಾನಿ #772
Inference Time
142 ms
Audio Length
2.40 s
RTF (Real-Time Factor)
0.059x
ONNX Inference Pipeline

NeMo FastConformer-TDT Architecture Flow

Stage 01
Audio Preprocessing
16kHz audio input converted to 128-channel Mel filterbanks using 25ms windows with 10ms frame stride.
128 Mel bins · 16,000 Hz
Stage 02 · INT8
FastConformer Encoder
Depthwise separable convolutions + relative multi-head self-attention with 8x subsampling.
encoder-model.int8.onnx (476 MB)
Stage 03 · TDT
Decoder Joint Network
Joint network emits both token probabilities and duration predictions, skipping blank frames in O(1).
decoder_joint-model.int8.onnx (5.4 MB)
Stage 04
BPE Detokenization
Subword tokens mapped through SentencePiece vocabulary back to Unicode Kannada text strings.
vocab.txt · Max 10 tokens/step
Deployment & SDK

Zero-Friction Inference Snippets

# 1. Install official onnx-asr engine pip install onnx-asr # 2. Recognize Kannada speech in 3 lines of Python import onnx_asr model = onnx_asr.load_model("OpenVoiceOS/artpark-iisc-vaani-fastconformer-kn-onnx") result = model.recognize("kannada_audio.wav") print(f"Kannada: {result}") # Output: "ಬೆಂಗಳೂರು ಕರ್ನಾಟಕದ ರಾಜಧಾನಿ"