The Life Tracker voice layer
Turning speech into filed work, with the source of every transcript on record
Voice notes from an iCloud drop folder, a call recorder that transcribes on the Mac, and Wispr Flow added as a voice layer, all feeding one system that says what it did with each note and where each transcript came from.
Problem
I talk to my life tracker as often as I type: phone voice notes, calls to transcribe, spoken answers in a revision chat. Each path had to run unattended, name its source and keep a note safe when something stalls.
Approach
Keep the unattended path local and mechanical, and add Wispr Flow where it helps, with the local engines as fallbacks. No program here can start Flow's dictation or read it back, and its meeting server is reachable only from an interactive session, so the design works within that.
Architecture
Dictation intake. A launchd watch on an iCloud drop folder fires within a second, with a 60-second scheduler tick as a backstop. A dispatcher with no model in it, under a file lock, writes one candidate per note, then a small agent run routes each as a reminder, question, task, piece of work or plain note. An audio note is taken only when a structural check of the M4A file and a clean decode both pass.
Call recorder. A Swift watcher waits on Core Audio listeners. When a call app takes the microphone it records both sides on separate channels, transcribes each on the device and merges them, dropping echo. Nothing leaves the Mac without an explicit push command.
Wispr Flow layer. When Flow is running, the revision chat's mic focuses the answer box and Flow types into it. Otherwise local Whisper takes over. For calls, the recorder matches a Flow meeting to a recording by time overlap. Flow's transcript takes the slot, the local one is kept beside it, and every transcript records its source. Only Flow's meeting table is read, never its dictation history, and a self-test checks that.
What I built
Notes once sat for hours on iCloud-evicted files, and a pid file let two dispatchers run at once. Evicted notes are now fetched by a child process with a time limit, and a kernel lock replaced the pid file. A failing item backs off and, at the eighth attempt, stops with one notification. A receipt file in iCloud lists what each recent capture became.
Results
Measured live on 11 Sep, notes went from drop to result in 22 to 74 seconds. The recorder's self-test runs 84 checks, six of its cases on Wispr Flow. A separate test covers the mic with Flow running, stopped or unknown.
What didn't work
An unanswered microphone prompt once hung the recorder, so start-up now has watchdogs. No real Flow meeting result has been seen yet, so the parser is tested only on assumed shapes. A meeting Flow records is transcribed on Wispr's servers, so the local-only rule does not cover it. Audio intake is built but switched off, and the 29 Sep dictation changes have not yet run live.
Stack
- Python
- Swift
- Core Audio
- Apple Speech
- Whisper
- Wispr Flow