summaryrefslogtreecommitdiff
path: root/TODO.md
blob: 76c4ea6c2a7d8e9f2abce41a1c4ecb0e5af39aa9 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
# RECCoon TODO

## Requested (2026-10-05)

### 1. Compare transcription models (Whisper / Parakeet / Whistle)

#### 1a. Parakeet TDT 0.6B v3 — **done, integrated**

- [x] **Test NVIDIA Parakeet v3** (`csukuangfj/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8`).
  - [x] Added `parakeet_baseline.py` to `tools/whistle-eval/` and wired it into
        `evaluate.py`.
  - [x] Result (same six clips): average WER **3.7%** vs Whisper base 11.4% and
        tiny 13.9%. Italian **2.8%** (vs base 16.7%) and **0%** on the noisy
        Italian clip. The only errors are number formatting.
  - [x] sherpa-onnx already supports it (`model_type="nemo_transducer"`), so no new
        runtime is needed.
  - [x] Integrated into the app: `ModelRepository.Model.PARAKEET_V3`,
        `TranscriptionEngine` builds an `OfflineTransducerModelConfig`, and the
        player model spinner lists "Parakeet v3 (best)".
  - [ ] Verify decoding and memory use on a real phone; the int8 model is
        ~670 MB (~652 MB encoder) and needs a large-heap device.
  - [x] Removed the Whisper tiny/base options: Parakeet v3 is now the only model
        (and therefore the default). The language picker was removed too, since
        Parakeet handles language by itself. The `Model` enum keeps room for more
        models later.
  - [ ] Consider a smaller multilingual Parakeet/Canary variant if device RAM is tight.

#### 1b. Cactus Whistle — tested, not integrated

- [~] **Check whether Cactus Whistle transcribes better than multilingual Whisper.**
      Reference: <https://huggingface.co/mrfakename/whistle-ONNX> (ONNX export of
      `Cactus-Compute/whistle`; supports en/de/fr/es/it/nl/pl). Motivation: Whisper
      makes many mistakes, and a smaller model could replace the ~50 MB sherpa-onnx
      AAR and fix the git push problem.
  - [x] Build a desktop evaluation harness under `tools/whistle-eval/` (`setup.sh`,
        `evaluate.py`, `whisper_baseline.py`, `build-data.mjs`).
  - [x] Baseline + metric: the harness runs Whisper tiny/base through the same
        sherpa-onnx 1.13.8 version the app bundles and reports WER.
  - [x] First result (2026-10-05, six TTS clips, EN + IT): Whistle is **competitive
        but not a clear win** — average WER 11.8% vs Whisper base 11.4% and tiny
        13.9%. On Italian Whistle was best (16.1% vs base 16.7%, tiny 23.2%); on
        English tiny was best. See `tools/whistle-eval/README.md`.
  - [ ] Decide the Android port after running the harness on **real RECCoon
        recordings** — the TTS result is too close to call.
  - [x] Feasibility: Whistle is a custom *Needle* encoder/decoder, **not** Whisper, so
        sherpa-onnx cannot load it. An Android port needs `onnxruntime-android` plus
        the mel filterbank, BPE tokenizer and engram features from `js/` /
        `whistle.pack`.
  - [x] Distribution: `whistle.pack` is ~17 MB (fits GitHub); the fp32 ONNX data is
        ~145 MB. `build-data.mjs` reconstructs the `.onnx.data` files from the pack
        byte-exactly, so only the pack needs to ship.
  - [x] Decision: **do not add Whistle to the model picker** — Parakeet v3 is both
        more accurate and supported by sherpa-onnx, so the custom-architecture port
        is not worth it.

### 2. Compact (compressed) recordings

- [~] **Switch on the main screen to save very small compressed files instead of WAV.**
  - [x] Codec: AAC-LC in MP4 (`.m4a`) for universal `MediaCodec` support; bitrate is
        48 kbps per channel (`AudioCompressor`).
  - [x] Added a **Compressed (small M4A)** checkbox under the record button, persisted
        as `compress_recordings`.
  - [x] The temporary WAV is transcoded to `.m4a` after stopping and then deleted; the
        M4A is saved to Downloads instead of the WAV.
  - [x] Markers still work through the existing JSON sidecar (M4A has no `cue ` chunk).
  - [x] `TranscriptionEngine` now streams through `PcmAudioSource`, which decodes
        compressed files with `MediaExtractor`/`MediaCodec`, so transcription still
        works on M4A.
  - [ ] Measure the size win and the transcription-quality impact on a real device and
        write the numbers here. Expected size: ~22 MB/hour mono instead of ~317 MB.

### 3. Forced microphone input (BT-Mic-Force)

- [~] **Integrate the BT-Mic-Force mic-forcing feature into RECCoon.**
  - [x] Added a **Mic** spinner (Auto / Phone mic / detected inputs) to the main
        screen.
  - [x] The choice is routed into `WavRecorder` via `AudioRecord.setPreferredDevice`
        and the `AudioManager` SCO mode (`setMode`, `startBluetoothSco`,
        `setBluetoothScoOn`).
  - [x] Added `MODIFY_AUDIO_SETTINGS`, `BLUETOOTH_CONNECT` and (pre-31)
        `BLUETOOTH`/`BLUETOOTH_ADMIN` permissions; `BLUETOOTH_CONNECT` is requested
        only when a Bluetooth input is selected.
  - [x] Reused the existing `AudioRecord` capture path; no second foreground service.
  - [ ] Verify that Bluetooth SCO actually routes audio on a real device.

### 4. Live monitoring, output picker, and quick discard

- [x] **Live monitoring** (`Monitor` checkbox + output picker). `LiveMonitor` plays the
      captured PCM through `AudioTrack` on a dedicated thread with a bounded queue, so
      a slow output never stalls the recording. Output devices are listed like the mic
      picker (Auto / phone speaker / wired / USB / Bluetooth). Feedback warning when
      monitoring is enabled.
- [x] **Stop and discard.** A **Discard** button stops the recording and deletes the
      temporary file without saving.
- [ ] Verify latency and routing on real devices (wired, Bluetooth A2DP, BT LE).
      Bluetooth A2DP adds unavoidable latency; note whether it is usable.
- [ ] Consider requesting audio focus and muting monitoring automatically when the
      selected output is the built-in speaker.

### 5. Rotation feedback (Alberto, from 1.4.0-alpha)

- [x] **Level meters look upside down when the phone is turned.** Cause: both
      activities were locked to `screenOrientation="portrait"`, so nothing followed
      the device.
  - [x] Removed the portrait lock so the UI rotates with the device.
  - [x] Added `configChanges="orientation|screenSize|..."` so a rotation does not
        recreate the activity and interrupt an in-progress recording.
  - [x] **L/R swapped in landscape.** Android exposes the two mic channels in a fixed
        device order; it does not tell the app which physical mic is which, so this is
        not reliably auto-detectable. Added a persisted **Swap L/R** checkbox that
        exchanges the channels in `WavRecorder`, so both the meters and the saved WAV
        match the user's left/right.
  - [ ] Confirm on a real device that rotating mid-recording keeps the recording and
        the meters upright, and that Swap L/R fixes the inverted meters.

## Build / distribution

- [ ] **The sherpa-onnx binaries are too large to keep in git.** The Android
      AAR is ~50 MB and the onnxruntime `.so` files push the APK to ~72 MB. The
      AAR and the built APK are currently gitignored. Decide how to distribute
      them later: Git LFS, a small Maven repository, a Gradle task that
      downloads the AAR on first build, or Play Asset/Feature Delivery for the
      Whisper models.

## Requested

- [x] **Separate left/right volume meters** — done: stereo capture with a
      mono fallback, plus a two-channel peak meter (`LevelMeterView`).
- [x] **Mono switch** — done: the **Mono** checkbox records a single channel.
- [x] **Option to disable notifications and vibrations while recording** — done
      via Do Not Disturb total silence (`ACCESS_NOTIFICATION_POLICY`), restored
      on stop. Persisted as a checkbox.
- [x] **Live waveform** — done: scrolling `WaveformView` with red clipped
      columns.
- [x] **WAV markers** — done: the **Mark** button embeds `cue `/`adtl` chunks in
      the WAV and stores a JSON sidecar; the player shows marker chips.
- [x] **Recordings list on the main screen** — done: the top-right *Recordings*
      link is replaced by a scrollable list at the bottom of the main screen,
      newest first. Tapping a row opens the player. The player has a **Delete**
      button and no longer autoplays.

## Backlog

- [x] Export transcripts as `.txt` / `.srt` (saved to Downloads, plus share).
- [x] Foreground service so recording continues when the app is in the
      background.
- [x] Library actions: delete, rename, share, "open with".
- [x] Word-level tap-to-seek inside a transcript segment.
- [x] Read embedded markers back from the WAV when there is no JSON sidecar.
- [x] Add a JUnit regression test for `StreamingResampler`.