From 901bf58c5cf9151942f8c4cd4e1427a02dec4b52 Mon Sep 17 00:00:00 2001 From: Tom Cooks Date: Sat, 10 Oct 2026 15:42:04 -0400 Subject: RECCoon 1.0.0: rebrand to Tom Cooks, Parakeet-only, GPLv3, F-Droid-ready --- README.md | 102 +++++++++++++++++++++++++++++++++++++++++++++----------------- 1 file changed, 75 insertions(+), 27 deletions(-) (limited to 'README.md') diff --git a/README.md b/README.md index 8785d1c..812afd2 100644 --- a/README.md +++ b/README.md @@ -4,22 +4,34 @@ A minimal Android app that records uncompressed WAV audio and transcribes it on device. - **Format:** 44.1 kHz, 16-bit PCM (stereo when the device supports it), no compression +- **Compact mode:** an optional switch saves a small AAC-LC `.m4a` instead of the WAV + (~22 MB/hour mono instead of ~317 MB); transcription still works because the + decoder streams the M4A back to PCM (`PcmAudioSource`) - **Controls:** one big round button — tap to start recording (timer starts), tap again to stop and save - **Meters:** live scrolling waveform plus left/right peak level meters with a clip indicator while recording - **Mono switch:** optionally record a single channel (half the file size) +- **Swap L/R:** exchange the stereo channels when a device reports them inverted in + landscape; applies to the meters and the saved file +- **Mic picker:** choose Auto, the built-in mic, or a specific input (Bluetooth SCO, + USB, wired headset ...) — the BT-Mic-Force feature folded into RECCoon +- **Live monitoring:** an optional **Monitor** switch plays the microphone through a + chosen output (built-in speaker, wired, USB, Bluetooth ...). Use headphones — the + speaker causes feedback. A dedicated thread and a bounded queue keep a slow output + from ever stalling the recording +- **Discard:** stop a recording and throw it away without saving - **Markers:** place labelled cue points while recording; they are embedded in the WAV (`cue ` + `LIST/adtl`) and shown in the player - **Silence:** optional Do Not Disturb total silence (no notifications, no vibration) while recording - **Recordings list:** scrollable list at the bottom of the main screen, newest first - **Player:** scrubber, ±10 s jumps, speed control, delete / rename / share / export, no autoplay - **Transcription:** on-device, offline speech recognition with [sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) using multilingual - Whisper — works with **English and Italian** -- **Name:** RECCoon · **by** wuhei · **version** 1.4.0-alpha + **NVIDIA Parakeet v3** — works with English, Italian and 23 other languages +- **Name:** RECCoon · **by** Tom Cooks · **version** 1.0.0 - **Icon:** raccoon with a microphone ## Requirements -- Android 8.0 (API 26) or newer, arm64-v8a or x86_64 +- Android 8.0 (API 26) or newer, arm64-v8a - `RECORD_AUDIO` permission - On Android 9 and older, `WRITE_EXTERNAL_STORAGE` is also requested - Internet access the first time a transcription model is downloaded @@ -37,16 +49,19 @@ Outputs: - `app/build/outputs/apk/debug/app-debug.apk` - `app/build/outputs/apk/release/app-release.apk` -A prebuilt copy is available at `dist/RECCoon-1.4.0-alpha.apk`. Note that the -sherpa-onnx AAR in `app/libs/` and the APK in `dist/` are **not** tracked in -git because they are large; see `TODO.md`. +A prebuilt copy is available at `dist/RECCoon-1.0.0.apk` (release, signed with +the debug key). The APK in `dist/` is **not** tracked in git. The sherpa-onnx +runtime comes from [JitPack](https://jitpack.io) (an F-Droid-trusted Maven +repository), so no prebuilt binary lives in this repository. ## Recording -- `WavRecorder` uses `AudioRecord` (MIC source) to capture 16-bit PCM at - 44100 Hz into a temporary file, then rewrites a correct 44-byte RIFF/WAVE - header when recording stops. It uses stereo when available and falls back to - mono, and reports per-channel peak levels to the UI. +- `WavRecorder` uses `AudioRecord` (MIC source, or `VOICE_COMMUNICATION` for + Bluetooth SCO) to capture 16-bit PCM at 44100 Hz into a temporary file, then + rewrites a correct 44-byte RIFF/WAVE header when recording stops. It uses + stereo when available and falls back to mono, and reports per-channel peak + levels to the UI. A selected input device is applied with + `AudioRecord.setPreferredDevice`. - `RecordingService` is a microphone-typed foreground service started while recording, so capture continues when the app is in the background. - `MainActivity` handles the runtime permissions, the timer, the level meters @@ -79,32 +94,65 @@ git because they are large; see `TODO.md`. (MediaStore or file), **Export** (writes `.txt` and `.srt` to Downloads) and **Delete** (removes the WAV and its sidecars). - Transcript words are clickable: tapping a word seeks the player to it. Word - timings come from Whisper token timestamps when available, otherwise the + timings come from the model's token timestamps when available, otherwise the segment is divided between its words. ## On-device transcription -- `TranscriptionEngine` streams the WAV, resamples it to 16 kHz and segments it - with Silero VAD, then decodes each speech segment with a multilingual Whisper - model through sherpa-onnx. Streaming keeps memory bounded for long - recordings (offline Whisper only handles 30 s at a time). -- `ModelRepository` downloads the models on first use into app-private storage: +- `TranscriptionEngine` reads the recording through `PcmAudioSource`, which + parses WAV directly and decodes compressed files (AAC/M4A) with + `MediaExtractor` + `MediaCodec`, resamples to 16 kHz and segments the audio + with Silero VAD, then decodes each speech segment with the Parakeet v3 + transducer through sherpa-onnx. Streaming keeps memory bounded for long + recordings. +- `ModelRepository` downloads the model on first use: | Model | Size | Notes | |---|---|---| - | Whisper tiny (int8, multilingual) | ~99 MB | faster, lower accuracy | - | Whisper base (int8, multilingual) | ~154 MB | default, better Italian | - - Both include the small Silero VAD model (~0.6 MB). Files come from the - sherpa-onnx Whisper model repos on Hugging Face and the sherpa-onnx GitHub - release. -- Language can be forced to English (`en`) or Italian (`it`), or left on - **Auto detect**. + | Parakeet v3 (int8, 25 languages) | ~670 MB | only model, multilingual, large download | + + It includes the small Silero VAD model (~0.6 MB). Files come from the + sherpa-onnx model repo on Hugging Face and the sherpa-onnx GitHub release. + Models are stored in the app-specific external folder + `/sdcard/Android/data/com.tomcooks.reccoon/files/models/` so they can be + side-loaded with `adb push` or copied over USB instead of downloaded in-app. +- Parakeet v3 is a transducer, so it is loaded with `model_type="nemo_transducer"` + and handles language detection itself (no language picker). It needs ~670 MB of + storage and a large-heap device. - Transcripts are stored as JSON sidecars in app-private storage (`filesDir/transcripts/.json`). +## Compact (compressed) recordings + +- The **Compressed (small M4A)** checkbox transcodes the finished WAV to + AAC-LC (48 kbps per channel) with `AudioCompressor` and saves that instead. + Markers are kept as a JSON sidecar. Use it when storage matters more than + bit-exact audio. +- `tools/whistle-eval/` contains a desktop harness that compares Whisper, + Parakeet and Cactus Whistle; see its README for the results. + +## Live monitoring and discard + +- **Monitor** plays the incoming audio back through the selected output device while + recording. The output picker mirrors the mic picker (Auto, phone speaker, wired, + USB, Bluetooth). `LiveMonitor` owns an `AudioTrack` on its own thread with a small + bounded queue, so a slow output (for example Bluetooth A2DP) drops monitor chunks + instead of blocking the capture loop and corrupting the recording. Headphones are + strongly recommended — the phone speaker feeds back. +- **Discard** stops the recording and deletes the temporary file without saving it, + for takes that are obviously bad. + +## License + +RECCoon is free software, released under the **GNU General Public License, +version 3 or later** (see `LICENSE`). The Parakeet v3 model is +[NVIDIA Parakeet TDT 0.6B v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3), +licensed CC-BY-4.0. + ## Native libraries and APK size -The APK bundles the sherpa-onnx/onnxruntime native libraries for `arm64-v8a` -and `x86_64`, so it is around 72 MB. Removing `x86_64` from the `abiFilters` -in `app/build.gradle.kts` roughly halves the APK size for phone-only builds. +The APK bundles the sherpa-onnx/onnxruntime native libraries only for +`arm64-v8a` (`abiFilters` in `app/build.gradle.kts`), which keeps it around +35–40 MB. Including `x86_64` as well would roughly double it; that is only +needed for x86 emulators. The ~670 MB Parakeet model is downloaded at runtime, +not bundled. -- cgit v1.2.3