summaryrefslogtreecommitdiff
path: root/README.md
diff options
context:
space:
mode:
authorTom Cooks <tommasogagliardi+github@gmail.com>2026-10-10 15:42:04 -0400
committerTom Cooks <tommasogagliardi+github@gmail.com>2026-10-10 15:42:04 -0400
commit901bf58c5cf9151942f8c4cd4e1427a02dec4b52 (patch)
treee9e312fb9893881f142a6f79e30417cd0c3f2be8 /README.md
parent6db1714d0354d1e1105e330ce84c46a54487c3bb (diff)
downloadreccoon-901bf58c5cf9151942f8c4cd4e1427a02dec4b52.tar.gz
RECCoon 1.0.0: rebrand to Tom Cooks, Parakeet-only, GPLv3, F-Droid-ready
Diffstat (limited to 'README.md')
-rw-r--r--README.md102
1 files changed, 75 insertions, 27 deletions
diff --git a/README.md b/README.md
index 8785d1c..812afd2 100644
--- a/README.md
+++ b/README.md
@@ -4,22 +4,34 @@ A minimal Android app that records uncompressed WAV audio and transcribes it
on device.
- **Format:** 44.1 kHz, 16-bit PCM (stereo when the device supports it), no compression
+- **Compact mode:** an optional switch saves a small AAC-LC `.m4a` instead of the WAV
+ (~22 MB/hour mono instead of ~317 MB); transcription still works because the
+ decoder streams the M4A back to PCM (`PcmAudioSource`)
- **Controls:** one big round button — tap to start recording (timer starts), tap again to stop and save
- **Meters:** live scrolling waveform plus left/right peak level meters with a clip indicator while recording
- **Mono switch:** optionally record a single channel (half the file size)
+- **Swap L/R:** exchange the stereo channels when a device reports them inverted in
+ landscape; applies to the meters and the saved file
+- **Mic picker:** choose Auto, the built-in mic, or a specific input (Bluetooth SCO,
+ USB, wired headset ...) — the BT-Mic-Force feature folded into RECCoon
+- **Live monitoring:** an optional **Monitor** switch plays the microphone through a
+ chosen output (built-in speaker, wired, USB, Bluetooth ...). Use headphones — the
+ speaker causes feedback. A dedicated thread and a bounded queue keep a slow output
+ from ever stalling the recording
+- **Discard:** stop a recording and throw it away without saving
- **Markers:** place labelled cue points while recording; they are embedded in the WAV (`cue ` + `LIST/adtl`) and shown in the player
- **Silence:** optional Do Not Disturb total silence (no notifications, no vibration) while recording
- **Recordings list:** scrollable list at the bottom of the main screen, newest first
- **Player:** scrubber, ±10 s jumps, speed control, delete / rename / share / export, no autoplay
- **Transcription:** on-device, offline speech recognition with
[sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) using multilingual
- Whisper — works with **English and Italian**
-- **Name:** RECCoon · **by** wuhei · **version** 1.4.0-alpha
+ **NVIDIA Parakeet v3** — works with English, Italian and 23 other languages
+- **Name:** RECCoon · **by** Tom Cooks · **version** 1.0.0
- **Icon:** raccoon with a microphone
## Requirements
-- Android 8.0 (API 26) or newer, arm64-v8a or x86_64
+- Android 8.0 (API 26) or newer, arm64-v8a
- `RECORD_AUDIO` permission
- On Android 9 and older, `WRITE_EXTERNAL_STORAGE` is also requested
- Internet access the first time a transcription model is downloaded
@@ -37,16 +49,19 @@ Outputs:
- `app/build/outputs/apk/debug/app-debug.apk`
- `app/build/outputs/apk/release/app-release.apk`
-A prebuilt copy is available at `dist/RECCoon-1.4.0-alpha.apk`. Note that the
-sherpa-onnx AAR in `app/libs/` and the APK in `dist/` are **not** tracked in
-git because they are large; see `TODO.md`.
+A prebuilt copy is available at `dist/RECCoon-1.0.0.apk` (release, signed with
+the debug key). The APK in `dist/` is **not** tracked in git. The sherpa-onnx
+runtime comes from [JitPack](https://jitpack.io) (an F-Droid-trusted Maven
+repository), so no prebuilt binary lives in this repository.
## Recording
-- `WavRecorder` uses `AudioRecord` (MIC source) to capture 16-bit PCM at
- 44100 Hz into a temporary file, then rewrites a correct 44-byte RIFF/WAVE
- header when recording stops. It uses stereo when available and falls back to
- mono, and reports per-channel peak levels to the UI.
+- `WavRecorder` uses `AudioRecord` (MIC source, or `VOICE_COMMUNICATION` for
+ Bluetooth SCO) to capture 16-bit PCM at 44100 Hz into a temporary file, then
+ rewrites a correct 44-byte RIFF/WAVE header when recording stops. It uses
+ stereo when available and falls back to mono, and reports per-channel peak
+ levels to the UI. A selected input device is applied with
+ `AudioRecord.setPreferredDevice`.
- `RecordingService` is a microphone-typed foreground service started while
recording, so capture continues when the app is in the background.
- `MainActivity` handles the runtime permissions, the timer, the level meters
@@ -79,32 +94,65 @@ git because they are large; see `TODO.md`.
(MediaStore or file), **Export** (writes `.txt` and `.srt` to Downloads) and
**Delete** (removes the WAV and its sidecars).
- Transcript words are clickable: tapping a word seeks the player to it. Word
- timings come from Whisper token timestamps when available, otherwise the
+ timings come from the model's token timestamps when available, otherwise the
segment is divided between its words.
## On-device transcription
-- `TranscriptionEngine` streams the WAV, resamples it to 16 kHz and segments it
- with Silero VAD, then decodes each speech segment with a multilingual Whisper
- model through sherpa-onnx. Streaming keeps memory bounded for long
- recordings (offline Whisper only handles 30 s at a time).
-- `ModelRepository` downloads the models on first use into app-private storage:
+- `TranscriptionEngine` reads the recording through `PcmAudioSource`, which
+ parses WAV directly and decodes compressed files (AAC/M4A) with
+ `MediaExtractor` + `MediaCodec`, resamples to 16 kHz and segments the audio
+ with Silero VAD, then decodes each speech segment with the Parakeet v3
+ transducer through sherpa-onnx. Streaming keeps memory bounded for long
+ recordings.
+- `ModelRepository` downloads the model on first use:
| Model | Size | Notes |
|---|---|---|
- | Whisper tiny (int8, multilingual) | ~99 MB | faster, lower accuracy |
- | Whisper base (int8, multilingual) | ~154 MB | default, better Italian |
-
- Both include the small Silero VAD model (~0.6 MB). Files come from the
- sherpa-onnx Whisper model repos on Hugging Face and the sherpa-onnx GitHub
- release.
-- Language can be forced to English (`en`) or Italian (`it`), or left on
- **Auto detect**.
+ | Parakeet v3 (int8, 25 languages) | ~670 MB | only model, multilingual, large download |
+
+ It includes the small Silero VAD model (~0.6 MB). Files come from the
+ sherpa-onnx model repo on Hugging Face and the sherpa-onnx GitHub release.
+ Models are stored in the app-specific external folder
+ `/sdcard/Android/data/com.tomcooks.reccoon/files/models/` so they can be
+ side-loaded with `adb push` or copied over USB instead of downloaded in-app.
+- Parakeet v3 is a transducer, so it is loaded with `model_type="nemo_transducer"`
+ and handles language detection itself (no language picker). It needs ~670 MB of
+ storage and a large-heap device.
- Transcripts are stored as JSON sidecars in app-private storage
(`filesDir/transcripts/<recording>.json`).
+## Compact (compressed) recordings
+
+- The **Compressed (small M4A)** checkbox transcodes the finished WAV to
+ AAC-LC (48 kbps per channel) with `AudioCompressor` and saves that instead.
+ Markers are kept as a JSON sidecar. Use it when storage matters more than
+ bit-exact audio.
+- `tools/whistle-eval/` contains a desktop harness that compares Whisper,
+ Parakeet and Cactus Whistle; see its README for the results.
+
+## Live monitoring and discard
+
+- **Monitor** plays the incoming audio back through the selected output device while
+ recording. The output picker mirrors the mic picker (Auto, phone speaker, wired,
+ USB, Bluetooth). `LiveMonitor` owns an `AudioTrack` on its own thread with a small
+ bounded queue, so a slow output (for example Bluetooth A2DP) drops monitor chunks
+ instead of blocking the capture loop and corrupting the recording. Headphones are
+ strongly recommended — the phone speaker feeds back.
+- **Discard** stops the recording and deletes the temporary file without saving it,
+ for takes that are obviously bad.
+
+## License
+
+RECCoon is free software, released under the **GNU General Public License,
+version 3 or later** (see `LICENSE`). The Parakeet v3 model is
+[NVIDIA Parakeet TDT 0.6B v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3),
+licensed CC-BY-4.0.
+
## Native libraries and APK size
-The APK bundles the sherpa-onnx/onnxruntime native libraries for `arm64-v8a`
-and `x86_64`, so it is around 72 MB. Removing `x86_64` from the `abiFilters`
-in `app/build.gradle.kts` roughly halves the APK size for phone-only builds.
+The APK bundles the sherpa-onnx/onnxruntime native libraries only for
+`arm64-v8a` (`abiFilters` in `app/build.gradle.kts`), which keeps it around
+35–40 MB. Including `x86_64` as well would roughly double it; that is only
+needed for x86 emulators. The ~670 MB Parakeet model is downloaded at runtime,
+not bundled.