diff options
Diffstat (limited to 'tools/whistle-eval/README.md')
| -rw-r--r-- | tools/whistle-eval/README.md | 94 |
1 files changed, 94 insertions, 0 deletions
diff --git a/tools/whistle-eval/README.md b/tools/whistle-eval/README.md new file mode 100644 index 0000000..ed6e71e --- /dev/null +++ b/tools/whistle-eval/README.md @@ -0,0 +1,94 @@ +# Speech-model evaluation + +Compares on-device speech recognition candidates on the same 16 kHz mono clips: + +- the multilingual Whisper tiny/base models that RECCoon already uses, + through the exact sherpa-onnx 1.13.8 version the Android app bundles; +- [NVIDIA Parakeet TDT 0.6B v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3) + (25 European languages) via its sherpa-onnx int8 export; +- [Cactus Whistle](https://huggingface.co/mrfakename/whistle-ONNX), a custom + *Needle* encoder/decoder run through its reference Node pipeline. + +Whistle is **not** a Whisper model and cannot be loaded with sherpa-onnx. +Parakeet, unlike Whistle, **is** supported by sherpa-onnx. + +## Setup + +Requires Node 18+, Python 3, `ffmpeg`, and `curl`. + +```sh +cd tools/whistle-eval + +# 1. Whistle pack + ONNX graphs + reference JS + onnxruntime-node +# and the Whisper tiny/base int8 models. +./setup.sh + +# 2. Parakeet v3 int8 (~660 MB, optional) +mkdir -p models/parakeet-v3-int8 +BASE=https://huggingface.co/csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8/resolve/main +for f in encoder.int8.onnx decoder.int8.onnx joiner.int8.onnx tokens.txt; do + curl -L --fail -o "models/parakeet-v3-int8/$f" "$BASE/$f" +done + +# 3. Python environment for the sherpa-onnx baselines. +python3 -m venv .venv +.venv/bin/pip install sherpa-onnx numpy +``` + +`setup.sh` downloads the 17 MB `whistle.pack` and rebuilds the fp32 +`encoder.onnx.data` / `decoder.onnx.data` files locally with `build-data.mjs`. +They are **not** committed (see `.gitignore`); the pack is enough to recreate +them byte-exactly. + +## Run + +```sh +# One clip (language inferred: `it*` -> Italian, otherwise English) +.venv/bin/python evaluate.py samples/jfk.wav + +# A whole folder +.venv/bin/python evaluate.py "samples/*.wav" +``` + +`evaluate.py` prints the transcript and the word error rate (WER) against +`samples/<name>.txt` when a reference exists. Parakeet is included +automatically when `models/parakeet-v3-int8/` is present. + +| clip | Whisper tiny | Whisper base | Whistle | **Parakeet v3** | +|---|---|---|---|---| +| `jfk.wav` | 0% | 4.5% | 0% | 0% | +| `en1.wav` | 0% | 0% | 8.7% | 0% | +| `en2.wav` | 13.6% | 13.6% | 13.6% | 13.6% | +| `it1.wav` | 24% | 12% | 20% | **4%** | +| `it1_noisy.wav` | 24% | 12% | 24% | **0%** | +| `it2.wav` | 21.7% | 26.1% | 4.3% | 4.3% | + +Average WER over these six clips (lower is better): + +| model | English (3) | Italian (3) | All (6) | int8 size | +|---|---|---|---|---| +| Whisper tiny | 4.5% | 23.2% | 13.9% | ~103 MB | +| Whisper base | 6.0% | 16.7% | 11.4% | ~160 MB | +| Whistle | 7.4% | 16.1% | 11.8% | ~17 MB pack (desktop) | +| **Parakeet v3** | 4.5% | **2.8%** | **3.7%** | ~670 MB | + +### Reading the result + +- **Parakeet v3 is the clear winner**, especially for Italian and noisy audio: + it was perfect on the noisy Italian clip that fooled every other model. +- The only Parakeet "errors" are number formatting ("four fifteen" → "4:15"), + which WER counts even though the transcript is correct. +- The cost is size: ~670 MB int8, about four times Whisper base. It is usable + on a modern phone but is a large download and a lot of RAM. +- Whisper base stays the best small model; Whistle is competitive but not a + clear win, and needs a custom runtime. +- These are small neural-TTS clips, not real recordings. Run the harness on + real RECCoon recordings before tuning further. + +## Android notes + +- Parakeet can be added through the existing sherpa-onnx AAR with + `model_type="nemo_transducer"` (Java: `OfflineTransducerModelConfig`). RECCoon + 1.5.0-alpha exposes it as the "Parakeet v3 (best)" model in the player. +- Whistle would need `onnxruntime-android` plus the log-mel frontend, BPE + tokenizer, engram features and greedy decoder from `vendor/js/whistle.js`. |
