summaryrefslogtreecommitdiff
path: root/tools/whistle-eval/README.md
blob: ed6e71e664099721b5333de84eda7cfc1f156183 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
# Speech-model evaluation

Compares on-device speech recognition candidates on the same 16 kHz mono clips:

- the multilingual Whisper tiny/base models that RECCoon already uses,
  through the exact sherpa-onnx 1.13.8 version the Android app bundles;
- [NVIDIA Parakeet TDT 0.6B v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3)
  (25 European languages) via its sherpa-onnx int8 export;
- [Cactus Whistle](https://huggingface.co/mrfakename/whistle-ONNX), a custom
  *Needle* encoder/decoder run through its reference Node pipeline.

Whistle is **not** a Whisper model and cannot be loaded with sherpa-onnx.
Parakeet, unlike Whistle, **is** supported by sherpa-onnx.

## Setup

Requires Node 18+, Python 3, `ffmpeg`, and `curl`.

```sh
cd tools/whistle-eval

# 1. Whistle pack + ONNX graphs + reference JS + onnxruntime-node
#    and the Whisper tiny/base int8 models.
./setup.sh

# 2. Parakeet v3 int8 (~660 MB, optional)
mkdir -p models/parakeet-v3-int8
BASE=https://huggingface.co/csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8/resolve/main
for f in encoder.int8.onnx decoder.int8.onnx joiner.int8.onnx tokens.txt; do
  curl -L --fail -o "models/parakeet-v3-int8/$f" "$BASE/$f"
done

# 3. Python environment for the sherpa-onnx baselines.
python3 -m venv .venv
.venv/bin/pip install sherpa-onnx numpy
```

`setup.sh` downloads the 17 MB `whistle.pack` and rebuilds the fp32
`encoder.onnx.data` / `decoder.onnx.data` files locally with `build-data.mjs`.
They are **not** committed (see `.gitignore`); the pack is enough to recreate
them byte-exactly.

## Run

```sh
# One clip (language inferred: `it*` -> Italian, otherwise English)
.venv/bin/python evaluate.py samples/jfk.wav

# A whole folder
.venv/bin/python evaluate.py "samples/*.wav"
```

`evaluate.py` prints the transcript and the word error rate (WER) against
`samples/<name>.txt` when a reference exists. Parakeet is included
automatically when `models/parakeet-v3-int8/` is present.

| clip | Whisper tiny | Whisper base | Whistle | **Parakeet v3** |
|---|---|---|---|---|
| `jfk.wav` | 0% | 4.5% | 0% | 0% |
| `en1.wav` | 0% | 0% | 8.7% | 0% |
| `en2.wav` | 13.6% | 13.6% | 13.6% | 13.6% |
| `it1.wav` | 24% | 12% | 20% | **4%** |
| `it1_noisy.wav` | 24% | 12% | 24% | **0%** |
| `it2.wav` | 21.7% | 26.1% | 4.3% | 4.3% |

Average WER over these six clips (lower is better):

| model | English (3) | Italian (3) | All (6) | int8 size |
|---|---|---|---|---|
| Whisper tiny | 4.5% | 23.2% | 13.9% | ~103 MB |
| Whisper base | 6.0% | 16.7% | 11.4% | ~160 MB |
| Whistle | 7.4% | 16.1% | 11.8% | ~17 MB pack (desktop) |
| **Parakeet v3** | 4.5% | **2.8%** | **3.7%** | ~670 MB |

### Reading the result

- **Parakeet v3 is the clear winner**, especially for Italian and noisy audio:
  it was perfect on the noisy Italian clip that fooled every other model.
- The only Parakeet "errors" are number formatting ("four fifteen" → "4:15"),
  which WER counts even though the transcript is correct.
- The cost is size: ~670 MB int8, about four times Whisper base. It is usable
  on a modern phone but is a large download and a lot of RAM.
- Whisper base stays the best small model; Whistle is competitive but not a
  clear win, and needs a custom runtime.
- These are small neural-TTS clips, not real recordings. Run the harness on
  real RECCoon recordings before tuning further.

## Android notes

- Parakeet can be added through the existing sherpa-onnx AAR with
  `model_type="nemo_transducer"` (Java: `OfflineTransducerModelConfig`). RECCoon
  1.5.0-alpha exposes it as the "Parakeet v3 (best)" model in the player.
- Whistle would need `onnxruntime-android` plus the log-mel frontend, BPE
  tokenizer, engram features and greedy decoder from `vendor/js/whistle.js`.