# RECCoon A minimal Android app that records uncompressed WAV audio and transcribes it on device. - **Format:** 44.1 kHz, 16-bit PCM (stereo when the device supports it), no compression - **Compact mode:** an optional switch saves a small AAC-LC `.m4a` instead of the WAV (~22 MB/hour mono instead of ~317 MB); transcription still works because the decoder streams the M4A back to PCM (`PcmAudioSource`) - **Controls:** one big round button — tap to start recording (timer starts), tap again to stop and save - **Meters:** live scrolling waveform plus left/right peak level meters with a clip indicator while recording - **Mono switch:** optionally record a single channel (half the file size) - **Swap L/R:** exchange the stereo channels when a device reports them inverted in landscape; applies to the meters and the saved file - **Mic picker:** choose Auto, the built-in mic, or a specific input (Bluetooth SCO, USB, wired headset ...) — the BT-Mic-Force feature folded into RECCoon - **Live monitoring:** an optional **Monitor** switch plays the microphone through a chosen output (built-in speaker, wired, USB, Bluetooth ...). Use headphones — the speaker causes feedback. A dedicated thread and a bounded queue keep a slow output from ever stalling the recording - **Discard:** stop a recording and throw it away without saving - **Markers:** place labelled cue points while recording; they are embedded in the WAV (`cue ` + `LIST/adtl`) and shown in the player - **Silence:** optional Do Not Disturb total silence (no notifications, no vibration) while recording - **Recordings list:** scrollable list at the bottom of the main screen, newest first - **Player:** scrubber, ±10 s jumps, speed control, delete / rename / share / export, no autoplay - **Transcription:** on-device, offline speech recognition with [sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) using multilingual **NVIDIA Parakeet v3** — works with English, Italian and 23 other languages - **Name:** RECCoon · **by** Tom Cooks · **version** 1.0.0 - **Icon:** raccoon with a microphone ## Requirements - Android 8.0 (API 26) or newer, arm64-v8a - `RECORD_AUDIO` permission - On Android 9 and older, `WRITE_EXTERNAL_STORAGE` is also requested - Internet access the first time a transcription model is downloaded ## Build ```sh ./gradlew :app:assembleDebug # debug APK ./gradlew :app:assembleRelease # release APK (signed with the debug key) ./gradlew :app:testDebugUnitTest # JVM unit tests ``` Outputs: - `app/build/outputs/apk/debug/app-debug.apk` - `app/build/outputs/apk/release/app-release.apk` A prebuilt copy is available at `dist/RECCoon-1.0.0.apk` (release, signed with the debug key). The APK in `dist/` is **not** tracked in git. The sherpa-onnx runtime comes from [JitPack](https://jitpack.io) (an F-Droid-trusted Maven repository), so no prebuilt binary lives in this repository. ## Recording - `WavRecorder` uses `AudioRecord` (MIC source, or `VOICE_COMMUNICATION` for Bluetooth SCO) to capture 16-bit PCM at 44100 Hz into a temporary file, then rewrites a correct 44-byte RIFF/WAVE header when recording stops. It uses stereo when available and falls back to mono, and reports per-channel peak levels to the UI. A selected input device is applied with `AudioRecord.setPreferredDevice`. - `RecordingService` is a microphone-typed foreground service started while recording, so capture continues when the app is in the background. - `MainActivity` handles the runtime permissions, the timer, the level meters and copying the finished file into Downloads. On Android 10+ this uses `MediaStore`; on older versions it writes directly to the public Downloads directory and notifies the media scanner. - Files are named `reccoon-YYYY-MM-DD-HHMMSS-XXXXX.wav`, where `XXXXX` is a persistent 5-digit progressive counter. ## Waveform, meters and markers - `LevelMeterView` draws a two-channel peak meter (dBFS scale) with a decaying peak hold and a red `CLIP` warning when a sample reaches full scale. - `WaveformView` draws a scrolling min/max waveform of the left/mono channel; clipped columns are drawn in red. - The **Mark** button (enabled while recording) adds a labelled cue point at the current position. Markers are written into the WAV as standard `cue ` and `LIST/adtl` chunks (readable by Audacity/Reaper) by `WavRecorder`, and also stored as a JSON sidecar by `MarkerStore`. The player reads the sidecar and falls back to `WavMarkers`, which parses the embedded chunks. ## Recordings list and player - The main screen shows a scrollable list of recordings, newest first. Tapping a row opens `PlayerActivity`. - `PlayerActivity` plays a file with the platform `MediaPlayer`, a scrubber (`SeekBar`), ±10 s jumps and a playback speed control (0.5×–2×). Playback does **not** start automatically. - Library actions: **Share** (FileProvider on Android 9 and older), **Rename** (MediaStore or file), **Export** (writes `.txt` and `.srt` to Downloads) and **Delete** (removes the WAV and its sidecars). - Transcript words are clickable: tapping a word seeks the player to it. Word timings come from the model's token timestamps when available, otherwise the segment is divided between its words. ## On-device transcription - `TranscriptionEngine` reads the recording through `PcmAudioSource`, which parses WAV directly and decodes compressed files (AAC/M4A) with `MediaExtractor` + `MediaCodec`, resamples to 16 kHz and segments the audio with Silero VAD, then decodes each speech segment with the Parakeet v3 transducer through sherpa-onnx. Streaming keeps memory bounded for long recordings. - `ModelRepository` downloads the model on first use: | Model | Size | Notes | |---|---|---| | Parakeet v3 (int8, 25 languages) | ~670 MB | only model, multilingual, large download | It includes the small Silero VAD model (~0.6 MB). Files come from the sherpa-onnx model repo on Hugging Face and the sherpa-onnx GitHub release. Models are stored in the app-specific external folder `/sdcard/Android/data/com.tomcooks.reccoon/files/models/` so they can be side-loaded with `adb push` or copied over USB instead of downloaded in-app. - Parakeet v3 is a transducer, so it is loaded with `model_type="nemo_transducer"` and handles language detection itself (no language picker). It needs ~670 MB of storage and a large-heap device. - Transcripts are stored as JSON sidecars in app-private storage (`filesDir/transcripts/.json`). ## Compact (compressed) recordings - The **Compressed (small M4A)** checkbox transcodes the finished WAV to AAC-LC (48 kbps per channel) with `AudioCompressor` and saves that instead. Markers are kept as a JSON sidecar. Use it when storage matters more than bit-exact audio. - `tools/whistle-eval/` contains a desktop harness that compares Whisper, Parakeet and Cactus Whistle; see its README for the results. ## Live monitoring and discard - **Monitor** plays the incoming audio back through the selected output device while recording. The output picker mirrors the mic picker (Auto, phone speaker, wired, USB, Bluetooth). `LiveMonitor` owns an `AudioTrack` on its own thread with a small bounded queue, so a slow output (for example Bluetooth A2DP) drops monitor chunks instead of blocking the capture loop and corrupting the recording. Headphones are strongly recommended — the phone speaker feeds back. - **Discard** stops the recording and deletes the temporary file without saving it, for takes that are obviously bad. ## License RECCoon is free software, released under the **GNU General Public License, version 3 or later** (see `LICENSE`). The Parakeet v3 model is [NVIDIA Parakeet TDT 0.6B v3](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3), licensed CC-BY-4.0. ## Native libraries and APK size The APK bundles the sherpa-onnx/onnxruntime native libraries only for `arm64-v8a` (`abiFilters` in `app/build.gradle.kts`), which keeps it around 35–40 MB. Including `x86_64` as well would roughly double it; that is only needed for x86 emulators. The ~670 MB Parakeet model is downloaded at runtime, not bundled.