# RECCoon A minimal Android app that records uncompressed WAV audio and transcribes it on device. - **Format:** 44.1 kHz, 16-bit, mono, PCM (no compression) - **Controls:** one big round button — tap to start recording (timer starts), tap again to stop and save - **Saving:** the WAV is written to the public **Downloads** folder as `reccoon-YYYY-MM-DD-HHMMSS-XXXXX.wav`, where `XXXXX` is a persistent 5-digit progressive counter - **Library:** a list of saved recordings, with a player and a scrubber to jump around long recordings - **Transcription:** on-device, offline speech recognition with [sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) using multilingual Whisper — works with **English and Italian** - **Name:** RECCoon · **by** wuhei · **version** 1.1.0-alpha - **Icon:** raccoon with a microphone ## Requirements - Android 8.0 (API 26) or newer, arm64-v8a or x86_64 - `RECORD_AUDIO` permission - On Android 9 and older, `WRITE_EXTERNAL_STORAGE` is also requested - Internet access the first time a transcription model is downloaded ## Build ```sh ./gradlew :app:assembleDebug # debug APK ./gradlew :app:assembleRelease # release APK (signed with the debug key) ``` Outputs: - `app/build/outputs/apk/debug/app-debug.apk` - `app/build/outputs/apk/release/app-release.apk` A prebuilt copy is available at `dist/RECCoon-1.1.0-alpha.apk`. ## Recording - `WavRecorder` uses `AudioRecord` (MIC source) to capture 16-bit PCM mono at 44100 Hz into a temporary file, then rewrites a correct 44-byte RIFF/WAVE header when recording stops. - `MainActivity` handles the runtime permissions, the timer, and copying the finished file into Downloads. On Android 10+ this uses `MediaStore`; on older versions it writes directly to the public Downloads directory and notifies the media scanner. ## Library and player - `RecordingsActivity` lists the `reccoon-*.wav` files in Downloads. - `PlayerActivity` plays a file with the platform `MediaPlayer`, a scrubber (`SeekBar`), ±10 s jumps and a playback speed control (0.5×–2×). Tapping a transcript line seeks the player to that moment. ## On-device transcription - `TranscriptionEngine` streams the WAV, resamples it to 16 kHz and segments it with Silero VAD, then decodes each speech segment with a multilingual Whisper model through sherpa-onnx. Streaming keeps memory bounded for long recordings (offline Whisper only handles 30 s at a time). - `ModelRepository` downloads the models on first use into app-private storage: | Model | Size | Notes | |---|---|---| | Whisper tiny (int8, multilingual) | ~99 MB | faster, lower accuracy | | Whisper base (int8, multilingual) | ~154 MB | default, better Italian | Both include the small Silero VAD model (~0.6 MB). Files come from the sherpa-onnx Whisper model repos on Hugging Face and the sherpa-onnx GitHub release. - Language can be forced to English (`en`) or Italian (`it`), or left on **Auto detect**. - Transcripts are stored as JSON sidecars in app-private storage (`filesDir/transcripts/.json`) and are shown in the player. A `TXT` badge marks recordings that already have a transcript. ## Native libraries and APK size The APK bundles the sherpa-onnx/onnxruntime native libraries for `arm64-v8a` and `x86_64`, so it is around 70 MB. Removing `x86_64` from the `abiFilters` in `app/build.gradle.kts` roughly halves the APK size for phone-only builds.