1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
|
# RECCoon
A minimal Android app that records uncompressed WAV audio and transcribes it
on device.
- **Format:** 44.1 kHz, 16-bit PCM (stereo when the device supports it), no compression
- **Controls:** one big round button — tap to start recording (timer starts), tap again to stop and save
- **Meters:** live scrolling waveform plus left/right peak level meters with a clip indicator while recording
- **Mono switch:** optionally record a single channel (half the file size)
- **Markers:** place labelled cue points while recording; they are embedded in the WAV (`cue ` + `LIST/adtl`) and shown in the player
- **Silence:** optional Do Not Disturb total silence (no notifications, no vibration) while recording
- **Saving:** the WAV is written to the public **Downloads** folder as
`reccoon-YYYY-MM-DD-HHMMSS-XXXXX.wav`, where `XXXXX` is a persistent
5-digit progressive counter
- **Library:** a list of saved recordings, with a player and a scrubber to jump
around long recordings
- **Transcription:** on-device, offline speech recognition with
[sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) using multilingual
Whisper — works with **English and Italian**
- **Name:** RECCoon · **by** wuhei · **version** 1.3.1-alpha
- **Icon:** raccoon with a microphone
## Requirements
- Android 8.0 (API 26) or newer, arm64-v8a or x86_64
- `RECORD_AUDIO` permission
- On Android 9 and older, `WRITE_EXTERNAL_STORAGE` is also requested
- Internet access the first time a transcription model is downloaded
## Build
```sh
./gradlew :app:assembleDebug # debug APK
./gradlew :app:assembleRelease # release APK (signed with the debug key)
```
Outputs:
- `app/build/outputs/apk/debug/app-debug.apk`
- `app/build/outputs/apk/release/app-release.apk`
A prebuilt copy is available at `dist/RECCoon-1.3.1-alpha.apk`. Note that the
sherpa-onnx AAR in `app/libs/` and the APK in `dist/` are **not** tracked in
git because they are large; see `TODO.md`.
## Recording
- `WavRecorder` uses `AudioRecord` (MIC source) to capture 16-bit PCM at
44100 Hz into a temporary file, then rewrites a correct 44-byte RIFF/WAVE
header when recording stops. It uses stereo when available and falls back to
mono, and reports per-channel peak levels to the UI.
- `MainActivity` handles the runtime permissions, the timer, and copying the
finished file into Downloads. On Android 10+ this uses `MediaStore`; on
older versions it writes directly to the public Downloads directory and
notifies the media scanner.
## Waveform, meters and markers
- `LevelMeterView` draws a two-channel peak meter (dBFS scale) with a decaying
peak hold and a red `CLIP` warning when a sample reaches full scale.
- `WaveformView` draws a scrolling min/max waveform of the left/mono channel;
clipped columns are drawn in red.
- The **Mark** button (enabled while recording) adds a labelled cue point at the
current position. Markers are written into the WAV as standard `cue ` and
`LIST/adtl` chunks (readable by Audacity/Reaper) by `WavRecorder`, and also
stored as a JSON sidecar by `MarkerStore`. `PlayerActivity` shows them as
chips that seek the player when tapped.
- The **Mono** checkbox records a single channel instead of stereo.
- The **Silence alerts** checkbox switches the phone to Do Not Disturb total
silence while recording and restores the previous setting on stop. It needs
the `ACCESS_NOTIFICATION_POLICY` permission, granted in system settings when
first enabled.
## Library and player
- `RecordingsActivity` lists the `reccoon-*.wav` files in Downloads.
- `PlayerActivity` plays a file with the platform `MediaPlayer`, a scrubber
(`SeekBar`), ±10 s jumps and a playback speed control (0.5×–2×). Tapping a
transcript line seeks the player to that moment.
## On-device transcription
- `TranscriptionEngine` streams the WAV, resamples it to 16 kHz and segments it
with Silero VAD, then decodes each speech segment with a multilingual Whisper
model through sherpa-onnx. Streaming keeps memory bounded for long
recordings (offline Whisper only handles 30 s at a time).
- `ModelRepository` downloads the models on first use into app-private storage:
| Model | Size | Notes |
|---|---|---|
| Whisper tiny (int8, multilingual) | ~99 MB | faster, lower accuracy |
| Whisper base (int8, multilingual) | ~154 MB | default, better Italian |
Both include the small Silero VAD model (~0.6 MB). Files come from the
sherpa-onnx Whisper model repos on Hugging Face and the sherpa-onnx GitHub
release.
- Language can be forced to English (`en`) or Italian (`it`), or left on
**Auto detect**.
- Transcripts are stored as JSON sidecars in app-private storage
(`filesDir/transcripts/<recording>.json`) and are shown in the player. A
`TXT` badge marks recordings that already have a transcript.
## Native libraries and APK size
The APK bundles the sherpa-onnx/onnxruntime native libraries for `arm64-v8a`
and `x86_64`, so it is around 70 MB. Removing `x86_64` from the `abiFilters`
in `app/build.gradle.kts` roughly halves the APK size for phone-only builds.
|