WebML-Kit On-Device Engine
(WASM SIMD / WebGPU)
OpenVoiceOS/artpark-iisc-vaani-fastconformer-kn-onnx · 100% Client-Side
Zero cloud calls. Audio stays strictly local in your browser. Model weights (~476 MB INT8) are streamed directly from Hugging Face and saved in the browser's Cache API for instant subsequent sessions.
Acoustic Input Studio
16kHz Mono PCM
Ready
00:00.0
Project VAANI Verified Dialects
4 Samples
Urban · Bengaluru
ಬೆಂಗಳೂರು ಕರ್ನಾಟಕದ ರಾಜಧಾನಿ
"Bengaluru is the capital of Karnataka"
Science · Technical
ವಿಜ್ಞಾನ ಮತ್ತು ತಂತ್ರಜ್ಞಾನ ನಮ್ಮ ಭವಿಷ್ಯ
"Science & technology shape our future"
Conversational
ನಾಳೆ ಬೆಳಗ್ಗೆ ಹತ್ತು ಗಂಟೆಗೆ ಸಭೆ ನಡೆಯಲಿದೆ
"The meeting will occur tomorrow at 10 AM"
Agriculture · Malnad
ಈ ವರ್ಷ ಮುಂಗಾರು ಮಳೆ ಉತ್ತಮವಾಗಿ ಸುರಿದಿದೆ
"This year monsoon rains were abundant"
Transcribed Kannada (ಲಿಪ್ಯಂತರಣ)
BPE Tokenizedಬೆಂಗಳೂರು ಕರ್ನಾಟಕದ ರಾಜಧಾನಿ
"Bengaluru is the capital of Karnataka"
Kannada Akshara & Subword Decomposition
▁ಬೆಂ #1024
ಗಳೂ #2491
ರು #6
▁ಕರ್ನಾ #891
ಟಕದ #421
▁ರಾಜ #310
ಧಾನಿ #772
Inference Time
142 ms
Audio Length
2.40 s
RTF (Real-Time Factor)
0.059x
ONNX Inference Pipeline
NeMo FastConformer-TDT Architecture Flow
Stage 01
Audio Preprocessing
16kHz audio input converted to 128-channel Mel filterbanks using 25ms windows with 10ms frame stride.
128 Mel bins · 16,000 Hz
Stage 02 · INT8
FastConformer Encoder
Depthwise separable convolutions + relative multi-head self-attention with 8x subsampling.
encoder-model.int8.onnx (476 MB)
Stage 03 · TDT
Decoder Joint Network
Joint network emits both token probabilities and duration predictions, skipping blank frames in O(1).
decoder_joint-model.int8.onnx (5.4 MB)
Stage 04
BPE Detokenization
Subword tokens mapped through SentencePiece vocabulary back to Unicode Kannada text strings.
vocab.txt · Max 10 tokens/step
Deployment & SDK
Zero-Friction Inference Snippets
# 1. Install official onnx-asr engine
pip install onnx-asr
# 2. Recognize Kannada speech in 3 lines of Python
import onnx_asr
model = onnx_asr.load_model("OpenVoiceOS/artpark-iisc-vaani-fastconformer-kn-onnx")
result = model.recognize("kannada_audio.wav")
print(f"Kannada: {result}")
# Output: "ಬೆಂಗಳೂರು ಕರ್ನಾಟಕದ ರಾಜಧಾನಿ"