Audio Rewind
Wear OS + Android
An audio rewind button on your wrist. Capture a rolling window of recent sound, pull back the last 15–30 seconds, replay it, and read it as text.
The problem
You hear something you need — a name, a door number, a price, an instruction — and by the time you get your phone out, it is gone. Every existing answer to this is a recorder: it writes continuously to disk, it fills storage, it costs battery, and it leaves you scrolling a timeline trying to find the ten seconds that mattered.
The recording is the wrong mental model. What people actually want is not "record everything", it is "let me go back a moment". That is a fundamentally different data structure — not a growing file, but a fixed-size window that slides. A rewind buffer you can forget is running.
And it has to be trustworthy in a way a recording app is not. A recorder that silently missed something is merely annoying. A rewind button that loses the last thirty seconds is worse than not having one, because you trusted it.
What I built
- Two applications: a Wear OS watch app and an Android phone companion, plus the transport layer that keeps them agreeing with each other.
- A foreground microphone service on the watch that maintains a rolling audio buffer entirely in RAM. Nothing is written to disk continuously — the buffer is preallocated, filled, and overwritten in place.
- A rewind controller that freezes the last 15 or 30 seconds out of the moving buffer, hands it to playback, and sends it down an asynchronous transcription pipeline.
- A phone↔watch transport so the capture source can be either device, and so the watch can stay a remote control when the phone is doing the listening.
- Surfaces that make the feature one tap away: a Wear OS Tile, a watch-face complication, the in-app screen, and a local memories/history view on the phone.
- An automated end-to-end test system that discovers the real phone and watch, installs builds, drives the UI, traces messages between devices, injects failures, and asserts the full path from user action to final result.
Key features
Capture & buffering
- Rolling microphone capture that runs as a foreground service
- 15-second and 30-second rewind windows
- RAM-based rolling buffer — no continuous recording to storage
- Preallocated ring buffer, no per-write allocation or GC pauses mid-capture
- Explicit user-controlled start; never passive or always-on
- Visible, honest session state on both devices
Output
- Transcription of rewound segments
- Immediate replay of the captured audio
- Results returned to the watch, not just the phone
- Local memories/history of what was captured
Two-device behaviour
- Capture from either the watch microphone or the phone microphone
- Microphone-source preference synchronized between phone and watch
- Watch requests a rewind from the phone when the phone is the active source
- Transcription results and status flow back to the watch
- Resync after disconnect, reconnect, reboot, or process death
Surfaces
- Wear OS Tile for rewind without opening the app
- Watch-face complication for at-a-glance status
- Full watch UI for session control and playback
- Phone companion UI for source selection and history
Technical architecture
The system is split so that the part which must never fail — the audio buffer — has no dependencies on the part which is allowed to fail — the network between two devices. The watch holds the buffer; the transport is treated as unreliable and is never on the critical path for local capture.
-
Capture
Foreground microphone service on the owning device. Audio arrives in a fixed cadence and is written into the buffer as fast as it lands. Lifecycle is explicit: started by the user, visible while running, torn down on stop.
-
Buffer
A preallocated ring buffer of PCM held in RAM. Writes overwrite the oldest samples; the window length is a parameter (15s / 30s) rather than a property of the storage. Because it is preallocated there is no allocation on the audio path, so no GC pause can drop a buffer overrun.
-
Rewind controller
Freezes a copy of the most recent N seconds out of the moving ring, hands the exact slice to playback and transcription, and returns it to the buffer to be overwritten. This is the only place that knows the window semantics.
-
Transcription
Asynchronous, off the audio path entirely. A rewound segment becomes text without ever blocking capture, and a transcription that fails is a failed transcription — it cannot corrupt or stall the buffer.
-
Transport
Phone↔watch messaging over the Wearable APIs, with request/response identity, timeouts, and explicit failure states. When the phone owns the microphone, the watch becomes a remote control and every rewind is a round trip that can fail — so the transport has to fail loudly and recover quietly.
-
State
Session state, capture source and processing status, with one owner per fact. The preference is synchronized rather than duplicated, so the two devices cannot disagree about who is listening.
-
Persistence
Local memories: the rewound segments worth keeping, plus the metadata needed to render and replay them. Bounded and explicit, because the rolling buffer itself is intentionally not durable.
-
UI
Watch app, Wear OS Tile, watch-face complication, and the phone companion. Every surface renders from the same state, so a rewind triggered from the Tile and one triggered in-app are the same code path.
- Audio capture, transport, and transcription never share a thread or a failure domain.
- The rolling buffer is deliberately non-persistent — losing it is a defined behaviour, not a bug.
Visuals
What was technically hard
A rolling window is a moving target
With a growing recording, "the last 30 seconds" is a simple seek. With a fixed ring buffer, those 30 seconds are being overwritten while you ask for them. The read has to resolve the window atomically against concurrent writes, and the slice handed to playback has to be exact — off-by-a-here means an audible repeat or a truncated clip.
Never allocating on the audio path
The obvious implementation writes into a growing list and takes the tail. It works until a garbage-collection pause lands mid-buffer and you lose audio you were told you had. The buffer is preallocated and written in place, which means the failure mode that matters cannot happen at all — a decision that costs a little memory and buys correctness.
A foreground microphone on a watch is a hostile environment
Screen off, ambient mode, the user switching apps, the OS reclaiming the process — all of it can happen while the buffer is live. The service has to survive the ones that should be survivable, stop cleanly for the ones that should not, and never leave a rolling microphone running after the user believes it stopped.
Two devices, one truth, and a network in the middle
When the phone owns the microphone, a rewind from the watch is a network round trip over a Bluetooth link that can drop between the tap and the answer. Every remote operation therefore needs an identity, a timeout, and a defined failure state. A rewind that hangs is far worse than one that says "could not reach the phone" — especially when the buffer is still running and the user cannot tell whether the tap registered.
Requests arrive twice, late, and out of order
A watch Tile can be re-tapped, a message can be redelivered, a surface can be restored and re-fire what it thought it was doing. Idempotency and request identity are not polish here — without them a single tap can consume two rewinds and a user who loses audio twice never comes back.
Privacy is a design constraint, not a policy page
A continuously-running microphone is something people have every right to be suspicious of. So recording is never passive, always visibly indicated, buffered in RAM rather than accumulated on disk, and the failure mode is silence — nothing is written unless the user asks for it to be.
Testing & reliability
The testing philosophy is that "it works on my desk" is not a test result. A user action is only proven end to end when every link in the chain has been observed — and when a failure can be attributed to a specific link rather than to "something went wrong".
End-to-end assertion chain
- 01 User action
- 02 App handler
- 03 Phone / watch transport
- 04 Processing
- 05 State & database
- 06 UI
- 07 Final user result
A test passes only when every link above is observed. When one breaks, the failure names the link instead of saying the feature is broken.
Physical devices, not emulators
The harness discovers a real Android phone and a real Wear OS watch. This is the point: Bluetooth drops, ambient mode, foreground-service kills, and permission revocation are precisely the failures that do not reproduce on an emulator, and they are most of what this system has to survive.
Device discovery, install, launch
The run resolves the connected phone and watch, installs the current build to both, launches the apps, and waits for each surface to report ready — rather than sleeping and hoping.
UI interaction on both devices
Tests drive the real interfaces on the watch and the phone: starting a session, changing the capture source, triggering a rewind, stopping. The action is performed the way a person performs it, so the test exercises the same path a user does.
Transport tracing
Phone↔watch communication is traced, so a test can prove a rewind actually crossed the link and was not quietly satisfied locally. Inter-device bugs are invisible without this.
Failure injection
Deliberately disconnect the link, revoke the permission, kill the foreground service, re-fire the same request, restore a stale surface, and blow past the timeout. The system is expected to fail gracefully and recover; the tests assert exactly that.
End-to-end assertion
Each test asserts the whole chain. Passing means: the action reached the handler, the request crossed the transport, processing completed, state and persistence agree, the UI reflects it, and the user got the result they asked for. A failure names the broken link.
Lifecycle and state cases
Screen off and on, app backgrounded, watch rebooted, connection lost and regained, and repeated identical requests — because the interesting bugs live at the seams, not in the happy path.
My contribution
Solo — product, architecture, both apps, test harness
Solo. I identified the problem, scoped it to a Wear OS app plus a phone companion, designed the buffer and the device transport, wrote both applications, built the transcription pipeline, and built the automated device test system around it.
The engineering judgment worth calling out: the rolling buffer is in RAM and never written continuously, the audio path never allocates, and capture is decoupled from both the network and transcription so that a link failure can never cost the user audio.