palOMine surfaces
A voice loop that lives on the box.
Voice on palOMine runs the same way the rest of the appliance runs — locally, on the hardware, with no third-party speech API in the loop. Whisper Large v3 Turbo and Moonshine Medium Streaming handle speech-to-text; Kokoro handles text-to-speech. The audio passes through the same bounded-authority agent as every other surface.
The local voice stack
Three models. No cloud.
Every voice-in and voice-out turn on palOMine is processed by one of three on-box models. None of them call a hosted speech API. The full list lives on /models.
STT
Whisper Large v3 Turbo
Full-quality speech-to-text for finished utterances: voice notes from the chat gateways, dictated turns in the web UI, transcriptions of recorded replies. First stage of every voice-in surface.
Multilingual ASR with a good latency vs. quality trade-off, sized for the appliance memory budget.
Streaming STT
Moonshine Medium Streaming
Low-latency streaming speech recognition that listens while the user speaks instead of waiting for a finished utterance. Powers the wake-word path and reactive dictation flows that need to feel turn-taking-friendly.
Incremental transcription — words appear as they are spoken, with end-of-turn detection on silence.
TTS
Kokoro
Speaks agent responses back through every surface — the on-device web UI, Telegram, Discord, Signal voice notes, and hands-free sessions. The voice half of every voice-out turn.
Low-latency text-to-speech, paced for natural conversational reply rather than long audio rendering.
Interaction flow
Wake-word on. Push-to-talk when you want it.
The voice loop offers two entry points, both local. Wake-word mode stays in the background listening on Moonshine’s streaming recogniser; push-to-talk opens the same path on demand. Either way, the captured audio never leaves the appliance’s LAN.
Wake-word
Moonshine streams the audio locally until the wake phrase matches; once it does, the appliance hands off to Whisper for a full-quality utterance capture. The background listener is on-box only — no wake-word service is called.
Push-to-talk
A direct capture path that ignores the wake phrase entirely. Useful when the ambient stream is undesirable — quiet rooms, shared spaces, or when the user wants to be deliberate about when a turn starts.
TTFT — voice
Voice turns share the same agent pipeline as chat. Cached agent-context TTFT lands in the same ~0.3–0.9 second band the chat path achieves on /product#benchmarks; streaming STT adds only the cost of finishing the in-progress utterance before the agent’s first token.
TTS pacing
Outbound speech is streamed through Kokoro as the agent’s tokens arrive, so the first audible response is short of a full reply rendered end-to-end. The model is sized for conversational pacing, not long-form narration.
Composes with the chat gateways
Voice notes across Telegram · Discord · Signal.
The voice loop is the same loop on every surface. The chat gateways — /features/channels — accept voice attachments inbound and can deliver voice notes outbound, all without leaving the appliance for inference.
Telegram voice notes
Inbound voice notes are decoded locally and routed to the same agent pipeline as a typed message. Outbound replies can be returned as a generated voice note — transcribed, then rendered through Kokoro on the appliance before delivery.
Discord voice notes
The Discord gateway accepts voice attachments on the same per-server allow-list policy as text. Transcription runs on-box; the platform only sees what the bot emits, the way every bot emits it.
Signal voice notes
Voice attachments through the Signal bridge follow the same end-to-end on-box path. Transcription is local; outbound replies can be sent as a voice note rendered by Kokoro without leaving the network.
Web surface
Dictate in. Speak back.
The on-device web UI exposes the same voice path as the gateways. A dictate button captures audio through Whisper; replies can be toggled to read aloud through Kokoro. Either control can be turned off, and they default to off on shared installations. See / for the web UI.
Bounded-authority core
The appliance decides what voice does.
Voice is a surface, not a special case. The same bounded-authority agent that handles a typed turn handles a voice turn: the wake phrase and the push-to-talk gate are local controls, STT output is just another input to the agent’s pipeline, and TTS output is just another outbound surface that the appliance’s allow/deny rules apply to.
The voice loop might wake the box, but it cannot make a tool call. Whether the user spoke or typed, the appliance decides what voice I/O is permitted under the same authority rules as every other surface.
Privacy
Audio never leaves your network.
There is no cloud STT API. There is no cloud TTS API. The wake-word listener, the full-utterance recogniser, and the speech synthesiser all run on the appliance. Audio is captured, transcribed, and rendered on-box; it does not egress to a third-party speech service in any direction.
Conversation archives — including voice turns and their transcripts — live on the appliance. If voice is disabled on a surface or on the device as a whole, the audio path stops with it. Nothing the appliance records is replayed through a hosted model.
The voice loop ships with the first run.
Whisper, Moonshine, and Kokoro are part of the production model set on every appliance. Get on the waitlist and we’ll let you know when units are ready to be reserved.