diff options
| author | Tom Cooks <tommasogagliardi+github@gmail.com> | 2026-10-10 15:42:04 -0400 |
|---|---|---|
| committer | Tom Cooks <tommasogagliardi+github@gmail.com> | 2026-10-10 15:42:04 -0400 |
| commit | 901bf58c5cf9151942f8c4cd4e1427a02dec4b52 (patch) | |
| tree | e9e312fb9893881f142a6f79e30417cd0c3f2be8 /TODO.md | |
| parent | 6db1714d0354d1e1105e330ce84c46a54487c3bb (diff) | |
| download | reccoon-901bf58c5cf9151942f8c4cd4e1427a02dec4b52.tar.gz | |
RECCoon 1.0.0: rebrand to Tom Cooks, Parakeet-only, GPLv3, F-Droid-ready
Diffstat (limited to 'TODO.md')
| -rw-r--r-- | TODO.md | 113 |
1 files changed, 113 insertions, 0 deletions
@@ -1,5 +1,118 @@ # RECCoon TODO +## Requested (2026-10-05) + +### 1. Compare transcription models (Whisper / Parakeet / Whistle) + +#### 1a. Parakeet TDT 0.6B v3 — **done, integrated** + +- [x] **Test NVIDIA Parakeet v3** (`csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8`). + - [x] Added `parakeet_baseline.py` to `tools/whistle-eval/` and wired it into + `evaluate.py`. + - [x] Result (same six clips): average WER **3.7%** vs Whisper base 11.4% and + tiny 13.9%. Italian **2.8%** (vs base 16.7%) and **0%** on the noisy + Italian clip. The only errors are number formatting. + - [x] sherpa-onnx already supports it (`model_type="nemo_transducer"`), so no new + runtime is needed. + - [x] Integrated into the app: `ModelRepository.Model.PARAKEET_V3`, + `TranscriptionEngine` builds an `OfflineTransducerModelConfig`, and the + player model spinner lists "Parakeet v3 (best)". + - [ ] Verify decoding and memory use on a real phone; the int8 model is + ~670 MB (~652 MB encoder) and needs a large-heap device. + - [x] Removed the Whisper tiny/base options: Parakeet v3 is now the only model + (and therefore the default). The language picker was removed too, since + Parakeet handles language by itself. The `Model` enum keeps room for more + models later. + - [ ] Consider a smaller multilingual Parakeet/Canary variant if device RAM is tight. + +#### 1b. Cactus Whistle — tested, not integrated + +- [~] **Check whether Cactus Whistle transcribes better than multilingual Whisper.** + Reference: <https://huggingface.co/mrfakename/whistle-ONNX> (ONNX export of + `Cactus-Compute/whistle`; supports en/de/fr/es/it/nl/pl). Motivation: Whisper + makes many mistakes, and a smaller model could replace the ~50 MB sherpa-onnx + AAR and fix the git push problem. + - [x] Build a desktop evaluation harness under `tools/whistle-eval/` (`setup.sh`, + `evaluate.py`, `whisper_baseline.py`, `build-data.mjs`). + - [x] Baseline + metric: the harness runs Whisper tiny/base through the same + sherpa-onnx 1.13.8 version the app bundles and reports WER. + - [x] First result (2026-10-05, six TTS clips, EN + IT): Whistle is **competitive + but not a clear win** — average WER 11.8% vs Whisper base 11.4% and tiny + 13.9%. On Italian Whistle was best (16.1% vs base 16.7%, tiny 23.2%); on + English tiny was best. See `tools/whistle-eval/README.md`. + - [ ] Decide the Android port after running the harness on **real RECCoon + recordings** — the TTS result is too close to call. + - [x] Feasibility: Whistle is a custom *Needle* encoder/decoder, **not** Whisper, so + sherpa-onnx cannot load it. An Android port needs `onnxruntime-android` plus + the mel filterbank, BPE tokenizer and engram features from `js/` / + `whistle.pack`. + - [x] Distribution: `whistle.pack` is ~17 MB (fits GitHub); the fp32 ONNX data is + ~145 MB. `build-data.mjs` reconstructs the `.onnx.data` files from the pack + byte-exactly, so only the pack needs to ship. + - [x] Decision: **do not add Whistle to the model picker** — Parakeet v3 is both + more accurate and supported by sherpa-onnx, so the custom-architecture port + is not worth it. + +### 2. Compact (compressed) recordings + +- [~] **Switch on the main screen to save very small compressed files instead of WAV.** + - [x] Codec: AAC-LC in MP4 (`.m4a`) for universal `MediaCodec` support; bitrate is + 48 kbps per channel (`AudioCompressor`). + - [x] Added a **Compressed (small M4A)** checkbox under the record button, persisted + as `compress_recordings`. + - [x] The temporary WAV is transcoded to `.m4a` after stopping and then deleted; the + M4A is saved to Downloads instead of the WAV. + - [x] Markers still work through the existing JSON sidecar (M4A has no `cue ` chunk). + - [x] `TranscriptionEngine` now streams through `PcmAudioSource`, which decodes + compressed files with `MediaExtractor`/`MediaCodec`, so transcription still + works on M4A. + - [ ] Measure the size win and the transcription-quality impact on a real device and + write the numbers here. Expected size: ~22 MB/hour mono instead of ~317 MB. + +### 3. Forced microphone input (BT-Mic-Force) + +- [~] **Integrate the BT-Mic-Force mic-forcing feature into RECCoon.** + - [x] Added a **Mic** spinner (Auto / Phone mic / detected inputs) to the main + screen. + - [x] The choice is routed into `WavRecorder` via `AudioRecord.setPreferredDevice` + and the `AudioManager` SCO mode (`setMode`, `startBluetoothSco`, + `setBluetoothScoOn`). + - [x] Added `MODIFY_AUDIO_SETTINGS`, `BLUETOOTH_CONNECT` and (pre-31) + `BLUETOOTH`/`BLUETOOTH_ADMIN` permissions; `BLUETOOTH_CONNECT` is requested + only when a Bluetooth input is selected. + - [x] Reused the existing `AudioRecord` capture path; no second foreground service. + - [ ] Verify that Bluetooth SCO actually routes audio on a real device. + +### 4. Live monitoring, output picker, and quick discard + +- [x] **Live monitoring** (`Monitor` checkbox + output picker). `LiveMonitor` plays the + captured PCM through `AudioTrack` on a dedicated thread with a bounded queue, so + a slow output never stalls the recording. Output devices are listed like the mic + picker (Auto / phone speaker / wired / USB / Bluetooth). Feedback warning when + monitoring is enabled. +- [x] **Stop and discard.** A **Discard** button stops the recording and deletes the + temporary file without saving. +- [ ] Verify latency and routing on real devices (wired, Bluetooth A2DP, BT LE). + Bluetooth A2DP adds unavoidable latency; note whether it is usable. +- [ ] Consider requesting audio focus and muting monitoring automatically when the + selected output is the built-in speaker. + +### 5. Rotation feedback (Alberto, from 1.4.0-alpha) + +- [x] **Level meters look upside down when the phone is turned.** Cause: both + activities were locked to `screenOrientation="portrait"`, so nothing followed + the device. + - [x] Removed the portrait lock so the UI rotates with the device. + - [x] Added `configChanges="orientation|screenSize|..."` so a rotation does not + recreate the activity and interrupt an in-progress recording. + - [x] **L/R swapped in landscape.** Android exposes the two mic channels in a fixed + device order; it does not tell the app which physical mic is which, so this is + not reliably auto-detectable. Added a persisted **Swap L/R** checkbox that + exchanges the channels in `WavRecorder`, so both the meters and the saved WAV + match the user's left/right. + - [ ] Confirm on a real device that rotating mid-recording keeps the recording and + the meters upright, and that Swap L/R fixes the inverted meters. + ## Build / distribution - [ ] **The sherpa-onnx binaries are too large to keep in git.** The Android |
