1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
|
# Speech-model evaluation
Compares on-device speech recognition candidates on the same 16 kHz mono clips:
- the multilingual Whisper tiny/base models that RECCoon already uses,
through the exact sherpa-onnx 1.13.8 version the Android app bundles;
- [NVIDIA Parakeet TDT 0.6B v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3)
(25 European languages) via its sherpa-onnx int8 export;
- [Cactus Whistle](https://huggingface.co/mrfakename/whistle-ONNX), a custom
*Needle* encoder/decoder run through its reference Node pipeline.
Whistle is **not** a Whisper model and cannot be loaded with sherpa-onnx.
Parakeet, unlike Whistle, **is** supported by sherpa-onnx.
## Setup
Requires Node 18+, Python 3, `ffmpeg`, and `curl`.
```sh
cd tools/whistle-eval
# 1. Whistle pack + ONNX graphs + reference JS + onnxruntime-node
# and the Whisper tiny/base int8 models.
./setup.sh
# 2. Parakeet v3 int8 (~660 MB, optional)
mkdir -p models/parakeet-v3-int8
BASE=https://huggingface.co/csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8/resolve/main
for f in encoder.int8.onnx decoder.int8.onnx joiner.int8.onnx tokens.txt; do
curl -L --fail -o "models/parakeet-v3-int8/$f" "$BASE/$f"
done
# 3. Python environment for the sherpa-onnx baselines.
python3 -m venv .venv
.venv/bin/pip install sherpa-onnx numpy
```
`setup.sh` downloads the 17 MB `whistle.pack` and rebuilds the fp32
`encoder.onnx.data` / `decoder.onnx.data` files locally with `build-data.mjs`.
They are **not** committed (see `.gitignore`); the pack is enough to recreate
them byte-exactly.
## Run
```sh
# One clip (language inferred: `it*` -> Italian, otherwise English)
.venv/bin/python evaluate.py samples/jfk.wav
# A whole folder
.venv/bin/python evaluate.py "samples/*.wav"
```
`evaluate.py` prints the transcript and the word error rate (WER) against
`samples/<name>.txt` when a reference exists. Parakeet is included
automatically when `models/parakeet-v3-int8/` is present.
| clip | Whisper tiny | Whisper base | Whistle | **Parakeet v3** |
|---|---|---|---|---|
| `jfk.wav` | 0% | 4.5% | 0% | 0% |
| `en1.wav` | 0% | 0% | 8.7% | 0% |
| `en2.wav` | 13.6% | 13.6% | 13.6% | 13.6% |
| `it1.wav` | 24% | 12% | 20% | **4%** |
| `it1_noisy.wav` | 24% | 12% | 24% | **0%** |
| `it2.wav` | 21.7% | 26.1% | 4.3% | 4.3% |
Average WER over these six clips (lower is better):
| model | English (3) | Italian (3) | All (6) | int8 size |
|---|---|---|---|---|
| Whisper tiny | 4.5% | 23.2% | 13.9% | ~103 MB |
| Whisper base | 6.0% | 16.7% | 11.4% | ~160 MB |
| Whistle | 7.4% | 16.1% | 11.8% | ~17 MB pack (desktop) |
| **Parakeet v3** | 4.5% | **2.8%** | **3.7%** | ~670 MB |
### Reading the result
- **Parakeet v3 is the clear winner**, especially for Italian and noisy audio:
it was perfect on the noisy Italian clip that fooled every other model.
- The only Parakeet "errors" are number formatting ("four fifteen" → "4:15"),
which WER counts even though the transcript is correct.
- The cost is size: ~670 MB int8, about four times Whisper base. It is usable
on a modern phone but is a large download and a lot of RAM.
- Whisper base stays the best small model; Whistle is competitive but not a
clear win, and needs a custom runtime.
- These are small neural-TTS clips, not real recordings. Run the harness on
real RECCoon recordings before tuning further.
## Android notes
- Parakeet can be added through the existing sherpa-onnx AAR with
`model_type="nemo_transducer"` (Java: `OfflineTransducerModelConfig`). RECCoon
1.5.0-alpha exposes it as the "Parakeet v3 (best)" model in the player.
- Whistle would need `onnxruntime-android` plus the log-mel frontend, BPE
tokenizer, engram features and greedy decoder from `vendor/js/whistle.js`.
|