# RECCoon A minimal Android app that records uncompressed WAV audio and transcribes it on device. - **Format:** 44.1 kHz, 16-bit PCM (stereo when the device supports it), no compression - **Controls:** one big round button — tap to start recording (timer starts), tap again to stop and save - **Meters:** live scrolling waveform plus left/right peak level meters with a clip indicator while recording - **Mono switch:** optionally record a single channel (half the file size) - **Markers:** place labelled cue points while recording; they are embedded in the WAV (`cue ` + `LIST/adtl`) and shown in the player - **Silence:** optional Do Not Disturb total silence (no notifications, no vibration) while recording - **Recordings list:** scrollable list at the bottom of the main screen, newest first - **Player:** scrubber, ±10 s jumps, speed control, delete / rename / share / export, no autoplay - **Transcription:** on-device, offline speech recognition with [sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) using multilingual Whisper — works with **English and Italian** - **Name:** RECCoon · **by** wuhei · **version** 1.4.0-alpha - **Icon:** raccoon with a microphone ## Requirements - Android 8.0 (API 26) or newer, arm64-v8a or x86_64 - `RECORD_AUDIO` permission - On Android 9 and older, `WRITE_EXTERNAL_STORAGE` is also requested - Internet access the first time a transcription model is downloaded ## Build ```sh ./gradlew :app:assembleDebug # debug APK ./gradlew :app:assembleRelease # release APK (signed with the debug key) ./gradlew :app:testDebugUnitTest # JVM unit tests ``` Outputs: - `app/build/outputs/apk/debug/app-debug.apk` - `app/build/outputs/apk/release/app-release.apk` A prebuilt copy is available at `dist/RECCoon-1.4.0-alpha.apk`. Note that the sherpa-onnx AAR in `app/libs/` and the APK in `dist/` are **not** tracked in git because they are large; see `TODO.md`. ## Recording - `WavRecorder` uses `AudioRecord` (MIC source) to capture 16-bit PCM at 44100 Hz into a temporary file, then rewrites a correct 44-byte RIFF/WAVE header when recording stops. It uses stereo when available and falls back to mono, and reports per-channel peak levels to the UI. - `RecordingService` is a microphone-typed foreground service started while recording, so capture continues when the app is in the background. - `MainActivity` handles the runtime permissions, the timer, the level meters and copying the finished file into Downloads. On Android 10+ this uses `MediaStore`; on older versions it writes directly to the public Downloads directory and notifies the media scanner. - Files are named `reccoon-YYYY-MM-DD-HHMMSS-XXXXX.wav`, where `XXXXX` is a persistent 5-digit progressive counter. ## Waveform, meters and markers - `LevelMeterView` draws a two-channel peak meter (dBFS scale) with a decaying peak hold and a red `CLIP` warning when a sample reaches full scale. - `WaveformView` draws a scrolling min/max waveform of the left/mono channel; clipped columns are drawn in red. - The **Mark** button (enabled while recording) adds a labelled cue point at the current position. Markers are written into the WAV as standard `cue ` and `LIST/adtl` chunks (readable by Audacity/Reaper) by `WavRecorder`, and also stored as a JSON sidecar by `MarkerStore`. The player reads the sidecar and falls back to `WavMarkers`, which parses the embedded chunks. ## Recordings list and player - The main screen shows a scrollable list of recordings, newest first. Tapping a row opens `PlayerActivity`. - `PlayerActivity` plays a file with the platform `MediaPlayer`, a scrubber (`SeekBar`), ±10 s jumps and a playback speed control (0.5×–2×). Playback does **not** start automatically. - Library actions: **Share** (FileProvider on Android 9 and older), **Rename** (MediaStore or file), **Export** (writes `.txt` and `.srt` to Downloads) and **Delete** (removes the WAV and its sidecars). - Transcript words are clickable: tapping a word seeks the player to it. Word timings come from Whisper token timestamps when available, otherwise the segment is divided between its words. ## On-device transcription - `TranscriptionEngine` streams the WAV, resamples it to 16 kHz and segments it with Silero VAD, then decodes each speech segment with a multilingual Whisper model through sherpa-onnx. Streaming keeps memory bounded for long recordings (offline Whisper only handles 30 s at a time). - `ModelRepository` downloads the models on first use into app-private storage: | Model | Size | Notes | |---|---|---| | Whisper tiny (int8, multilingual) | ~99 MB | faster, lower accuracy | | Whisper base (int8, multilingual) | ~154 MB | default, better Italian | Both include the small Silero VAD model (~0.6 MB). Files come from the sherpa-onnx Whisper model repos on Hugging Face and the sherpa-onnx GitHub release. - Language can be forced to English (`en`) or Italian (`it`), or left on **Auto detect**. - Transcripts are stored as JSON sidecars in app-private storage (`filesDir/transcripts/.json`). ## Native libraries and APK size The APK bundles the sherpa-onnx/onnxruntime native libraries for `arm64-v8a` and `x86_64`, so it is around 72 MB. Removing `x86_64` from the `abiFilters` in `app/build.gradle.kts` roughly halves the APK size for phone-only builds.