1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
|
# RECCoon TODO
## Requested (2026-10-05)
### 1. Compare transcription models (Whisper / Parakeet / Whistle)
#### 1a. Parakeet TDT 0.6B v3 — **done, integrated**
- [x] **Test NVIDIA Parakeet v3** (`csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8`).
- [x] Added `parakeet_baseline.py` to `tools/whistle-eval/` and wired it into
`evaluate.py`.
- [x] Result (same six clips): average WER **3.7%** vs Whisper base 11.4% and
tiny 13.9%. Italian **2.8%** (vs base 16.7%) and **0%** on the noisy
Italian clip. The only errors are number formatting.
- [x] sherpa-onnx already supports it (`model_type="nemo_transducer"`), so no new
runtime is needed.
- [x] Integrated into the app: `ModelRepository.Model.PARAKEET_V3`,
`TranscriptionEngine` builds an `OfflineTransducerModelConfig`, and the
player model spinner lists "Parakeet v3 (best)".
- [ ] Verify decoding and memory use on a real phone; the int8 model is
~670 MB (~652 MB encoder) and needs a large-heap device.
- [x] Removed the Whisper tiny/base options: Parakeet v3 is now the only model
(and therefore the default). The language picker was removed too, since
Parakeet handles language by itself. The `Model` enum keeps room for more
models later.
- [ ] Consider a smaller multilingual Parakeet/Canary variant if device RAM is tight.
#### 1b. Cactus Whistle — tested, not integrated
- [~] **Check whether Cactus Whistle transcribes better than multilingual Whisper.**
Reference: <https://huggingface.co/mrfakename/whistle-ONNX> (ONNX export of
`Cactus-Compute/whistle`; supports en/de/fr/es/it/nl/pl). Motivation: Whisper
makes many mistakes, and a smaller model could replace the ~50 MB sherpa-onnx
AAR and fix the git push problem.
- [x] Build a desktop evaluation harness under `tools/whistle-eval/` (`setup.sh`,
`evaluate.py`, `whisper_baseline.py`, `build-data.mjs`).
- [x] Baseline + metric: the harness runs Whisper tiny/base through the same
sherpa-onnx 1.13.8 version the app bundles and reports WER.
- [x] First result (2026-10-05, six TTS clips, EN + IT): Whistle is **competitive
but not a clear win** — average WER 11.8% vs Whisper base 11.4% and tiny
13.9%. On Italian Whistle was best (16.1% vs base 16.7%, tiny 23.2%); on
English tiny was best. See `tools/whistle-eval/README.md`.
- [ ] Decide the Android port after running the harness on **real RECCoon
recordings** — the TTS result is too close to call.
- [x] Feasibility: Whistle is a custom *Needle* encoder/decoder, **not** Whisper, so
sherpa-onnx cannot load it. An Android port needs `onnxruntime-android` plus
the mel filterbank, BPE tokenizer and engram features from `js/` /
`whistle.pack`.
- [x] Distribution: `whistle.pack` is ~17 MB (fits GitHub); the fp32 ONNX data is
~145 MB. `build-data.mjs` reconstructs the `.onnx.data` files from the pack
byte-exactly, so only the pack needs to ship.
- [x] Decision: **do not add Whistle to the model picker** — Parakeet v3 is both
more accurate and supported by sherpa-onnx, so the custom-architecture port
is not worth it.
### 2. Compact (compressed) recordings
- [~] **Switch on the main screen to save very small compressed files instead of WAV.**
- [x] Codec: AAC-LC in MP4 (`.m4a`) for universal `MediaCodec` support; bitrate is
48 kbps per channel (`AudioCompressor`).
- [x] Added a **Compressed (small M4A)** checkbox under the record button, persisted
as `compress_recordings`.
- [x] The temporary WAV is transcoded to `.m4a` after stopping and then deleted; the
M4A is saved to Downloads instead of the WAV.
- [x] Markers still work through the existing JSON sidecar (M4A has no `cue ` chunk).
- [x] `TranscriptionEngine` now streams through `PcmAudioSource`, which decodes
compressed files with `MediaExtractor`/`MediaCodec`, so transcription still
works on M4A.
- [ ] Measure the size win and the transcription-quality impact on a real device and
write the numbers here. Expected size: ~22 MB/hour mono instead of ~317 MB.
### 3. Forced microphone input (BT-Mic-Force)
- [~] **Integrate the BT-Mic-Force mic-forcing feature into RECCoon.**
- [x] Added a **Mic** spinner (Auto / Phone mic / detected inputs) to the main
screen.
- [x] The choice is routed into `WavRecorder` via `AudioRecord.setPreferredDevice`
and the `AudioManager` SCO mode (`setMode`, `startBluetoothSco`,
`setBluetoothScoOn`).
- [x] Added `MODIFY_AUDIO_SETTINGS`, `BLUETOOTH_CONNECT` and (pre-31)
`BLUETOOTH`/`BLUETOOTH_ADMIN` permissions; `BLUETOOTH_CONNECT` is requested
only when a Bluetooth input is selected.
- [x] Reused the existing `AudioRecord` capture path; no second foreground service.
- [ ] Verify that Bluetooth SCO actually routes audio on a real device.
### 4. Live monitoring, output picker, and quick discard
- [x] **Live monitoring** (`Monitor` checkbox + output picker). `LiveMonitor` plays the
captured PCM through `AudioTrack` on a dedicated thread with a bounded queue, so
a slow output never stalls the recording. Output devices are listed like the mic
picker (Auto / phone speaker / wired / USB / Bluetooth). Feedback warning when
monitoring is enabled.
- [x] **Stop and discard.** A **Discard** button stops the recording and deletes the
temporary file without saving.
- [ ] Verify latency and routing on real devices (wired, Bluetooth A2DP, BT LE).
Bluetooth A2DP adds unavoidable latency; note whether it is usable.
- [ ] Consider requesting audio focus and muting monitoring automatically when the
selected output is the built-in speaker.
### 5. Rotation feedback (Alberto, from 1.4.0-alpha)
- [x] **Level meters look upside down when the phone is turned.** Cause: both
activities were locked to `screenOrientation="portrait"`, so nothing followed
the device.
- [x] Removed the portrait lock so the UI rotates with the device.
- [x] Added `configChanges="orientation|screenSize|..."` so a rotation does not
recreate the activity and interrupt an in-progress recording.
- [x] **L/R swapped in landscape.** Android exposes the two mic channels in a fixed
device order; it does not tell the app which physical mic is which, so this is
not reliably auto-detectable. Added a persisted **Swap L/R** checkbox that
exchanges the channels in `WavRecorder`, so both the meters and the saved WAV
match the user's left/right.
- [ ] Confirm on a real device that rotating mid-recording keeps the recording and
the meters upright, and that Swap L/R fixes the inverted meters.
## Build / distribution
- [ ] **The sherpa-onnx binaries are too large to keep in git.** The Android
AAR is ~50 MB and the onnxruntime `.so` files push the APK to ~72 MB. The
AAR and the built APK are currently gitignored. Decide how to distribute
them later: Git LFS, a small Maven repository, a Gradle task that
downloads the AAR on first build, or Play Asset/Feature Delivery for the
Whisper models.
## Requested
- [x] **Separate left/right volume meters** — done: stereo capture with a
mono fallback, plus a two-channel peak meter (`LevelMeterView`).
- [x] **Mono switch** — done: the **Mono** checkbox records a single channel.
- [x] **Option to disable notifications and vibrations while recording** — done
via Do Not Disturb total silence (`ACCESS_NOTIFICATION_POLICY`), restored
on stop. Persisted as a checkbox.
- [x] **Live waveform** — done: scrolling `WaveformView` with red clipped
columns.
- [x] **WAV markers** — done: the **Mark** button embeds `cue `/`adtl` chunks in
the WAV and stores a JSON sidecar; the player shows marker chips.
- [x] **Recordings list on the main screen** — done: the top-right *Recordings*
link is replaced by a scrollable list at the bottom of the main screen,
newest first. Tapping a row opens the player. The player has a **Delete**
button and no longer autoplays.
## Backlog
- [x] Export transcripts as `.txt` / `.srt` (saved to Downloads, plus share).
- [x] Foreground service so recording continues when the app is in the
background.
- [x] Library actions: delete, rename, share, "open with".
- [x] Word-level tap-to-seek inside a transcript segment.
- [x] Read embedded markers back from the WAV when there is no JSON sidecar.
- [x] Add a JUnit regression test for `StreamingResampler`.
|