Talk to the agent on your own Mac.
Double-tap your earbuds and Claude Code, running on the Mac at home, starts listening. Have it run the tests, read you what failed and fix it, or start the next task — from a walk, the gym or the car. Your phone stays locked in your pocket; the Mac hears you, works in your own files, and speaks the answer back.
PHONE remote: nextTrack → TRIGGER
PHONE audio focus: EXCLUSIVE
PHONE conversation started
— Audon owns the audio- Speech captured
- 2.03 s
- Transcribed
- 1.43 s
- Whole turn
- 5.55 s
How it works
Two devices you own, and a gesture.
The Mac is the server. The phone is a microphone and a speaker. The agent works where your projects already live — your files, your tools, your running processes. The one part of this that is not in your house is the model, and the next section says exactly what that means.
- 01
The phone captures
A double-tap on the earbud opens the microphone — screen locked, phone in your pocket, no wake word and nothing to look at. It costs one sound to own that gesture, which took four builds to find out.
How that gesture was won → - 02
Your Mac transcribes
Language is detected first, then the engine is chosen: English to Apple's on-device recogniser, everything else to whisper large-v3-turbo. With the default engines your audio goes no further than the Mac — they have no network code at all.
- 03
Claude Code does the work
The words go to a Claude Code session on that Mac, on your own Claude account. It reads, runs and edits with the same tools it has at the keyboard, in the same files — not in a hosted sandbox with a copy of them.
- 04
It speaks back
The reply is read into the same earbuds a sentence at a time, as it is written, and a tap stops it mid-sentence. The microphone is never open while it speaks, so a cough, an echo or the television cannot interrupt it. Only you can.
- It knows your voice
- A voice it does not recognise gets a conversation, not your machine — no commands, no files, no mail — so two people can share the phone. That is a courtesy the model is asked to keep, not a lock, and the limits below say how often it is wrong.
- Two phrases, and you choose them
- “Hold this conversation” shuts the microphone until you tap; “stop this conversation” ends it. Naming a phrase is not saying it: “I want to change the stop phrase” is a request, not a stop.
- It never goes quiet on you
- A soft pulse covers the wait, and tool work is narrated in plain words — “searching the web”, “building it” — because in earphones, 40 s of an agent working is otherwise indistinguishable from a crash.
What leaves the house
Your audio stays home. Your words go to Claude.
A voice interface to an agent cannot be entirely offline, and pretending otherwise is how privacy pages mislead. So here is where each thing goes, including the part most voice products leave out.
- Your voiceStays home
Your phone and your Mac. The Mac transcribes it, and the default engines have no network code. Cloud transcription exists, is labelled where you choose it, and is never picked automatically.
- The words you saidGoes to Anthropic
Anthropic, through Claude Code on your Mac — exactly as if you had typed them.
- What the agent reads to answerGoes to Anthropic
Anthropic too, as in any Claude Code session: the files it opens and the output of the commands it runs.
- The spoken replyStays home
Synthesised on your Mac. Cloud voices exist, and send the reply text only if you pick one.
- Phone to Mac, away from homeEncrypted end to end
Through Audon's relay, as ciphertext it cannot read.
Audon runs one relay, and it cannot read your audio.
On the same wifi, the phone talks to the Mac directly. From anywhere else it goes through a relay — and the two ends run their own handshake through it, so what the relay forwards is ciphertext. A room id is the hash of a host’s public key, so the relay cannot substitute itself for your Mac, and the frames are authenticated, so it cannot edit them.
That is measured, not asserted. A tap standing exactly where the relay stands recorded 93 KB of a real transcription turn. The transcript, the credential, the endpoint names and the account were all absent — the longest printable run in the capture is twelve characters of noise.
Your account is a separate thing again, and holds an email address — no password, because there is none to hold, and never audio. It signs the credential your device presents and then gets out of the way: your Mac checks that signature offline, so the account server can be down for a week and every paired device keeps working.
What the relay does see
- IP addresses, at both ends
- Byte counts and timing
- Which room talked to which
That is metadata, it is real, and it is the honest cost of not asking every person to stand up their own server. What the relay never holds is plaintext, or your account — it stores nothing but live connections.
The target case
Built for sentences that change language halfway through.
Mandarin, English and Malay inside one sentence is ordinary speech where this was built. It is also the pattern that makes mainstream engines drop words without telling you. Code-switching is the case Audon is aimed at, not an edge it tolerates.
Names and jargon are what a transcript is for
They are also exactly what a general engine loses. Feeding the decoder the terms it gets wrong — rather than every term you know — moved term recall on real voice notes:
More biasing is not better. Every engine tried has a knee past which recall falls, so what gets fed in is what the engine gets wrong — never the whole vocabulary.
Three engines that looked right on paper
Each was picked on strong published numbers, and each lost on this audio.
- SenseVoice-Small
Chosen on published code-switching benchmarks. Dropped English words entirely, and ran 5× slower than whisper turbo.
- A Malaysian-tuned whisper
Trained on ms/en/zh/ta and looked ideal. It translates rather than transcribes — Mandarin came back as Malay paraphrase, with hallucinations.
- Forcing the Malay locale
On Mandarin-dominant audio, stock whisper collapsed into “Tidak. Tidak. Tidak.”
Evidence
Nothing here is a projection.
Every figure below was run on the development machine — an M1 Pro on macOS 26 — and can be reproduced by the evaluation harness that produced it. Where a number is an estimate, this project says so; there are none on this page.
| What changed | Before | After |
|---|---|---|
| Time to first spoken word | 9.3 s | 4.9 s |
| Time to first model token | 2.25 s | 0.71 s |
| Post-capture processing | 1.58 s | 1.20 s |
| Device probing, off the critical path | 940 ms | 305 ms |
| Term recall, no hint → full glossary | 48.2% | 75.0% |
| Term recall, Apple, with phrase list | 8/14 | 13/14 |
| Hallucinated transcripts | 22 of 235 | 0 |
| Speaker cut off mid-sentence | 1 pause in 8 | 1 in 16 |
of trailing silence is enough to know you have finished, hands-free
per sentence to synthesise a reply, entirely offline
barge-in false positives — structurally, not statistically
realtime to turn a recorded meeting into who said what, offline
Known limits
Stated here rather than discovered later.
This project keeps a list of what it does badly, in the same file as the list of what it does well. These are the entries that will matter to you.
- The model is not in your house
- Claude Code runs on your Mac; the model behind it runs at Anthropic. If that service is down, or your account hits its usage limit, Audon hears you perfectly and has nothing to ask.
- Knowing your voice is a courtesy, not a lock
- Measured on video calls, 4 of 1,833 pieces of other people’s speech scored as the owner — all four from one colleague — while 12% of the owner’s own fell short. A recording of you would pass. The phone in your hand is what holds the door.
- Endpointing is not semantic
- Pause mid-thought for longer than the threshold and your turn ends. Calibration reduces it. It does not remove it, which is why push-to-talk stays the default.
- The pause calibration flatters itself
- It can only see pauses you got away with — one that did cut you off ends a recording and starts another. The figure is reported as a floor, not a target.
- Offline speech has one emotional contour
- Measured at 5.0–5.4 semitones of pitch variation whatever mood is asked for. Mood renders as pace. Expressive synthesis exists and runs 12× slower than real time.
- Read-aloud accuracy figures flatter every engine
- The ranking between engines holds. The absolute numbers do not transfer to spontaneous speech, and this project’s own harness once scored an engine against its own mistakes before that was caught.
- Cloud engines are genuinely off-device
- They exist, they are labelled at the point of choice, and they are never selected automatically. But they are what they are.
- It needs a Mac that stays awake
- That is a property of what this is, not of how it connects. Pairing is one code typed once; the awake Mac is the real cost.
Writing
The failures, in as much detail as the wins.
It is not finished, and that is written down too.
The round trip works on a real phone. There is no App Store build yet. Leave an address and you get one email when there is something to install — or follow the writing, which is where the progress actually gets posted.