TL;DR:
In September, Meta launched Muse, its personal agent. Alongside it came the Muse Charm:
It ships in December.
The Charm's premise struck me: the agent lives in the cloud, the device is just a friendly way to reach it, and anything that needs reading or paying goes to a bigger screen. That's an architecture, not a product. So I asked a developer's question: how much of this can I build this week, with an off-the-shelf dev board and public APIs?
It turned out to be most of it, minus the fingerprint sensor and the payments, which I didn't want anyway. This post covers:
I'm a photographer and an angler on the San Mateo County coast. The questions I actually ask while my hands are full are narrow and local:
|
Ask |
What Kai does |
|---|---|
|
"When's high tide at Pillar Point tomorrow?" |
NOAA tide predictions, a card with highs and lows |
|
"Note: f/11, quarter second, 10-stop ND at the lighthouse" |
a trip note, filed under the current trip |
|
"Research three sunrise spots near Half Moon Bay for this weekend, tell me later" |
a background task; the answer is spoken ~45 s later |
|
"Remind me at 5:30 to pack the ND filters" |
a reminder; the Stick wakes itself at 5:30 |
|
"Tell me if the wind at Half Moon Bay drops under 10 mph this weekend" |
a watch that checks every few hours and speaks once |
The surprise came from the logs: over four days, fewer than one in six real questions touched those niche tools. The rest were whatever came to mind:
A pocket device gets asked everything. So general knowledge had to be excellent, and the specialised tools had to be reliable, not just present.
Kai runs on an M5StickS3, M5Stack's newest "Stick" dev kit. It costs about $21 and needs no soldering: everything a voice gadget needs is already in the case. The Stick family got famous among developers this spring as the reference hardware for Anthropic's open-source Claude Desktop Buddy, a little desk pet that wakes up when a Claude session starts and lets you approve or deny its permission prompts with a button press. Kai asks its newer sibling to do more: listen, talk back, and live on a keychain.
|
Part |
What's inside |
How Kai uses it |
|---|---|---|
|
SoC |
ESP32-S3 (dual-core Xtensa LX7, 240 MHz), 2.4 GHz Wi-Fi, BLE 5 | |
|
Memory |
8 MB flash, 8 MB Octal PSRAM |
a 2 MB speaker ring; recording before Wi-Fi is up |
|
Display |
1.14" ST7789 LCD, 240×135 |
pixel face, captions, five-line cards |
|
Audio |
ES8311 codec, MEMS mic, AW8737 amp, 1 W speaker |
push-to-talk in, Kai's voice out |
|
Input |
two buttons |
hold the front button to talk |
|
Power |
250 mAh LiPo, M5PM1 power chip, USB-C |
rail measurements, power gating, RTC timer wake |
|
Extras |
BMI270 IMU, IR TX/RX, Grove port, Hat2 header |
unused so far |
|
Size |
48 × 24 × 15 mm, 20 g |
Clips to a lanyard or a bag |
ESP32-S3 power domains: two Xtensa LX7 cores and the radio sprint at up to 240 MHz and are powered off in deep sleep, while an always-on RTC domain with a ULP-RISC-V coprocessor, RTC timer and 16 KB of RTC memory stays on at single-digit microamps and wakes the main cores
Two kinds of core, for two kinds of time. The ESP32-S3 pairs two very different processors. For the few seconds you're talking, two Xtensa LX7 cores run at up to 240 MHz, with vector instructions to spare. That's enough to encode a 20 ms Opus frame in 6.4 ms while the radio streams. For the other 23-odd hours of the day, those cores, the radio, and main memory are switched off.
What stays on is a tiny always-on corner of the chip: an RTC timer, 16 KB of RTC memory, and a RISC-V ultra-low-power (ULP) coprocessor that can run its own small program with the main cores asleep. The gap between the two sides is huge: roughly 300 mA while Wi-Fi transmits, and 7–8 µA asleep with only the timer and RTC memory powered, a ratio of about 40,000 to 1.
Kai uses the always-on side today as an alarm clock and a notepad: the RTC timer wakes it for the next reminder, and RTC memory keeps the Wi-Fi channel, BSSID, and relay IP so the reconnect skips the scan. The RISC-V core is the headroom: at ~190 µA it could watch the buttons or the IMU and wake the big cores only when something happens, without ever starting Wi-Fi.
Why this board and not a bare ESP32:
What it doesn't have matters just as much: no fingerprint sensor, no cellular, no GPS, and no Bluetooth Classic, so no headset profiles. Those gaps shaped the design as much as the parts did.
Each of the Stick's limits forced a decision:
|
Constraint |
Consequence |
|---|---|
|
Voice in, voice out; tiny screen; no keyboard |
Answers are one to three spoken sentences plus a five-line card |
|
~225 mAh usable |
The device deep-sleeps between uses, so nothing can push to it |
|
Realtime voice must feel instant |
Slow work is asynchronous: acknowledge now, deliver later |
|
No fingerprint sensor |
No payments or side effects without a deliberate confirmation (deferred) |
The design rule that fell out: voice first, glance second, everything else somewhere else.
Architecture: the Stick streams Opus over Wi-Fi to a Python relay on a home server, which bridges to Gemini Live (fast path) and to a locked-down Hermes Agent (slow path) that calls back through a localhost MCP server
The Stick is deliberately dumb.
The relay is where everything lives. It's a Python service on an always-on box at home.
Keeping the logic off the device paid for itself on day one: every prompt tweak and new tool shipped with a git pull on the server, not a reflash.
The wire protocol is small:
|
Direction |
Message |
|---|---|
|
Stick → relay |
|
|
relay → Stick |
binary Opus, |
|
relay → Stick |
|
Tools are plain async functions. Gemini gets the declaration. The relay sends the data back to Gemini and the card to the screen:
@tool(
"get_tides",
"High and low tide times near a place from the nearest NOAA station (US coasts). Call it for every "
"tide question, including repeats and follow-ups: it refreshes the card on screen.",
{"place": {"type": "STRING", "description": "Beach, harbor, pier or town. Omit for home."},
"day": DAY_PARAM},
)
async def get_tides(ctx: ToolContext, place: str | None = None, day: str | None = None) -> ToolResult:
...
Two early lessons shaped every tool after that:
|
Weather, mid-answer |
Tides |
Saved note |
|---|---|---|
|
|
|
|
A 250 mAh battery forces you to measure before you optimize. I built an estimated power budget from the power chip's rail measurements and component datasheets. It found something I didn't expect:
Energy per question: in firmware 0.7, 84% of each question's charge went to the tail after the answer; firmware 0.8 cut it from 6.6 to 1.6 mAh
The conversation was cheap. The tail was not. After each answer, the device kept the reply screen lit with Wi-Fi awake, then idled, then dimmed, then finally slept. That tail was 84% of each question's charge, and the idle loop was redrawing a full frame at 30 fps, spending half its time drawing nothing new.
The fixes were mundane and effective:
The result: projected battery life at 20 questions a day went from ~1.3 days to ~6 days.
Then Opus: switching the audio from raw PCM (256 kbit/s up, 384 kbit/s down) to Opus (~24 up, ~32 down) cut Wi-Fi airtime about 11–12×.
Waking fast matters as much as sleeping well:
I seriously considered three places for the relay:
ws:// on a trusted network, and a box I already run.I chose local first, remote later. The architecture doesn't care: the phone bridge (Stick → BLE → phone → tunnel → the same relay) only adds a transport. Keeping the brain in one Python process means every tool exists once, not three times in Python, Kotlin and Swift.
A real-time voice model is superb at conversation and terrible at waiting. A Gemini Live function call blocks the turn, so "research three sunrise spots" can't run inside it. Kai needed a second brain for slow work, memory, and schedules.
I compared three routes:
I chose Hermes, and locked it down hard:
A red-team prompt asking it to read ~/.ssh found no tool to do it with.
Integration is two one-way pipes:
POST /v1/runs starts a background run with a voice-shaped instruction. The reply must be JSON: speak (at most two sentences), a card, and details_md.kai_notify delivers a result;
kai_profile_brief updates what Kai knows about me;
Kai's own tide, weather, and light tools are exposed read-only, so Hermes doesn't re-scrape the web with worse data. It once answered sunrise from a web page (7:33) when Kai's tool said 7:03. Job prompts now require Kai's tools.
The key trick is how results come back. ask_agent returns immediately, and Kai says "On it." When the run finishes, the relay injects the result into the live voice session as a text turn, queued so it never talks over you:
await self._send({"type": "nudge"}) # two-note chime: news you didn't just ask for
await self._live.send_client_content(
turns=types.Content(role="user", parts=[types.Part(text=(
"[Background result for the owner, not something they just said. Tell them now in one or "
f"two short spoken sentences, as your own news: {item.speak}]"))]),
turn_complete=True,
)
If the Stick is asleep, the result waits in an inbox, and a yellow badge shows on the face the next time you pick it up.
Memory is a brief, not a database lookup.
recall covers anything older, with an 8-second budget before it falls back to the inbox.Time is where the sleeping device bites. Nothing can push to a Stick in deep sleep. So:
Timer wake sequence: the relay stores the reminder and sends wake_in_s; the Stick deep-sleeps with its RTC timer set; at 3:00 it wakes, says hello with wake=timer, and the relay fires the reminder with a chime and speech
[SILENT] until its condition holds, then calls kai_notify(watch=…), and the relay deletes the job.This is the most interesting design question in the project. You have ~2 seconds on one path and ~25 seconds on the other. Who decides, and how?
The option I rejected: a router in front. Keyword rules, an embedding classifier, or a small model would decide before the answer starts. Those fit text pipelines. With native audio, there's no transcript until the turn ends, so a gate in front adds latency to every question to save a few.
What Kai does instead: the voice model is the router. ask_agent is just another tool, and choosing a tool is something Gemini Live does anyway. That costs zero added latency. The work goes into making that choice well-informed and measurable.
Routing: Gemini Live chooses between answering now, answering then offering to dig deeper, and handing off to Hermes; relay safety nets catch skipped tools and repeated confirmations
Version one used trigger words ("research", "look into", "later"). The logs graded it:
Before touching the prompt, I built an eval:
fast, tool:<name>, slow, offer or clarify. Three runs per case, because the same prompt routes differently from run to run.Then I changed the rules from words to criteria:
|
104 cases × 3 runs |
Before |
After |
|---|---|---|
|
Overall |
90% |
94% |
|
Needs digging deeper (slow or offer) |
56% |
87% |
|
Explicit "research / tell me later" |
94% |
100% |
|
Quick facts wrongly sent slow |
3 |
0 |
The eval earned its keep in the first iteration. My first draft raised accuracy, but it introduced a worse failure. Twice, Kai said "I'm looking into those road closures for you" without calling the tool. That's a promise with nothing behind it. One hard rule fixed it: "never say you're looking into something without calling ask_agent: nothing happens unless you call it". Without the eval, I'd have shipped a Kai that occasionally lies about doing work.
Some fixes don't belong in the prompt at all. Gemini Live sometimes confirmed a reminder twice: once around the tool call, and again when the result came back. No wording reliably stops that, so the relay does: after a successful action tool and one spoken confirmation, it drops any further turn Gemini starts before you speak again. Use the prompt for judgement, and code for invariants.
Kai's moods: idle, listening, thinking, speaking, happy, sleeping, offline
The Muse Charm will ship with a fingerprint sensor, a payments stack, and Meta's cloud behind it. Kai has a 20-gram dev board, a home server and two brains that know their jobs. For a pocket pal that answers the questions I actually ask, that's been enough. And Kai costs $21; you can mod it however you want, and you don't have to wait for December.
Code: github.com/zbruceli/kai (MIT).