# Speech-model evaluation Compares on-device speech recognition candidates on the same 16 kHz mono clips: - the multilingual Whisper tiny/base models that RECCoon already uses, through the exact sherpa-onnx 1.13.8 version the Android app bundles; - [NVIDIA Parakeet TDT 0.6B v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3) (25 European languages) via its sherpa-onnx int8 export; - [Cactus Whistle](https://huggingface.co/mrfakename/whistle-ONNX), a custom *Needle* encoder/decoder run through its reference Node pipeline. Whistle is **not** a Whisper model and cannot be loaded with sherpa-onnx. Parakeet, unlike Whistle, **is** supported by sherpa-onnx. ## Setup Requires Node 18+, Python 3, `ffmpeg`, and `curl`. ```sh cd tools/whistle-eval # 1. Whistle pack + ONNX graphs + reference JS + onnxruntime-node # and the Whisper tiny/base int8 models. ./setup.sh # 2. Parakeet v3 int8 (~660 MB, optional) mkdir -p models/parakeet-v3-int8 BASE=https://huggingface.co/csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8/resolve/main for f in encoder.int8.onnx decoder.int8.onnx joiner.int8.onnx tokens.txt; do curl -L --fail -o "models/parakeet-v3-int8/$f" "$BASE/$f" done # 3. Python environment for the sherpa-onnx baselines. python3 -m venv .venv .venv/bin/pip install sherpa-onnx numpy ``` `setup.sh` downloads the 17 MB `whistle.pack` and rebuilds the fp32 `encoder.onnx.data` / `decoder.onnx.data` files locally with `build-data.mjs`. They are **not** committed (see `.gitignore`); the pack is enough to recreate them byte-exactly. ## Run ```sh # One clip (language inferred: `it*` -> Italian, otherwise English) .venv/bin/python evaluate.py samples/jfk.wav # A whole folder .venv/bin/python evaluate.py "samples/*.wav" ``` `evaluate.py` prints the transcript and the word error rate (WER) against `samples/.txt` when a reference exists. Parakeet is included automatically when `models/parakeet-v3-int8/` is present. | clip | Whisper tiny | Whisper base | Whistle | **Parakeet v3** | |---|---|---|---|---| | `jfk.wav` | 0% | 4.5% | 0% | 0% | | `en1.wav` | 0% | 0% | 8.7% | 0% | | `en2.wav` | 13.6% | 13.6% | 13.6% | 13.6% | | `it1.wav` | 24% | 12% | 20% | **4%** | | `it1_noisy.wav` | 24% | 12% | 24% | **0%** | | `it2.wav` | 21.7% | 26.1% | 4.3% | 4.3% | Average WER over these six clips (lower is better): | model | English (3) | Italian (3) | All (6) | int8 size | |---|---|---|---|---| | Whisper tiny | 4.5% | 23.2% | 13.9% | ~103 MB | | Whisper base | 6.0% | 16.7% | 11.4% | ~160 MB | | Whistle | 7.4% | 16.1% | 11.8% | ~17 MB pack (desktop) | | **Parakeet v3** | 4.5% | **2.8%** | **3.7%** | ~670 MB | ### Reading the result - **Parakeet v3 is the clear winner**, especially for Italian and noisy audio: it was perfect on the noisy Italian clip that fooled every other model. - The only Parakeet "errors" are number formatting ("four fifteen" → "4:15"), which WER counts even though the transcript is correct. - The cost is size: ~670 MB int8, about four times Whisper base. It is usable on a modern phone but is a large download and a lot of RAM. - Whisper base stays the best small model; Whistle is competitive but not a clear win, and needs a custom runtime. - These are small neural-TTS clips, not real recordings. Run the harness on real RECCoon recordings before tuning further. ## Android notes - Parakeet can be added through the existing sherpa-onnx AAR with `model_type="nemo_transducer"` (Java: `OfflineTransducerModelConfig`). RECCoon 1.5.0-alpha exposes it as the "Parakeet v3 (best)" model in the player. - Whistle would need `onnxruntime-android` plus the log-mel frontend, BPE tokenizer, engram features and greedy decoder from `vendor/js/whistle.js`.