Athena realtime voice bridge
This independent service lets OpenClaw's existing browser Talk UI use Athena
Whisper, OpenClaw's agent, and Athena Qwen3-TTS through the browser's supported
OpenAI-style WebRTC transport. It does not modify OpenClaw or switch an Athena
profile. The existing gateway-relay path remains available.
Flow: OpenClaw's athena-talk plugin signs a single-use 60-second browser token;
the browser posts an SDP offer to this service; microphone audio is sent over
WebRTC; Athena transcribes it; the service asks OpenClaw to run
openclaw_agent_consult; Athena speaks the returned text over the same WebRTC
connection. This is a half-duplex prototype. Barge-in and remote-network TURN
support are not implemented.
For replies, the bridge consumes Athena's existing
/v1/audio/speech/pcm-stream endpoint (24 kHz mono PCM) and enqueues audio
frames as they arrive. Playback can start before Qwen finishes synthesizing a
long reply. This does not make playback start while the OpenClaw agent is still
writing its answer: OpenClaw's browser openclaw_agent_consult path submits the
tool result only after the agent's final chat event, so the bridge receives the
complete text first. In one observed run, this agent phase took 14.1 seconds
before TTS began. Streaming this phase would require an upstream OpenClaw Talk
interface for partial agent text, while retaining agent tool access.
The bounded queue applies backpressure instead of truncating long
audio, and the next turn waits until all frames have played. The two-turn
WebRTC test verifies that audio begins before the fake PCM stream finishes.
The bridge announces each committed user audio item before sending its completed transcript with the same item ID. OpenClaw requires that sequence to persist the transcript; omitting the item caused “Realtime transcript refers to an unknown speech item” in the browser. Spoken assistant replies are also announced as conversation items with matching transcript IDs and explicit predecessor IDs. Without them, OpenClaw's ordered transcript storage can wait on a missing assistant item when the next user turn arrives, leaving later questions unanswered.
Required service environment:
| Name | Purpose |
|---|---|
OPENCLAW_PUBLIC_KEY_URL |
Existing OpenClaw HTTPS origin plus /plugins/athena-talk/realtime/public-key. The service fetches the plugin's public verification key. |
ATHENA_API_BASE_URL |
Router API base, http://router:8081/v1 in the Compose deployment. |
ATHENA_API_KEY |
Router key, if required. |
OPENCLAW_ORIGIN |
Exact HTTPS origin of the UI, e.g. https://oc.casaderoll.de. |
PORT |
HTTP listen port; default 8090. The Compose service shares the existing WireGuard gateway network namespace. |
The plugin provides an HTTPS offer route on the existing OpenClaw origin and
forwards SDP to this service. Configure
talk.realtime.providers.athena-talk.realtimeUpstreamUrl with Athena's internal
HTTP URL ending in /v1/realtime/calls. Set the service's
OPENCLAW_PUBLIC_KEY_URL to the OpenClaw HTTPS public-key route above. No new
secret is needed in OpenClaw: its plugin keeps a private signing key in memory,
and the service receives only the public key. Select
talk.realtime.transport: "webrtc" only after the service is reachable. Keep
gateway-relay as a rollback option.
The WebRTC media connection needs a route from the client device to Athena's
ICE candidate addresses. The service runs directly inside Athena's existing
private WireGuard network namespace, so home/VPN clients can reach it at
192.168.1.212:8090; the OpenClaw HTTPS proxy covers signaling only. Clients
outside the private network need TURN support, which is not implemented.
The deployed service is mike-ai-realtime-voice in compose.yaml. It uses no
GPU and does not switch Athena's active model profile. Deploy it with the
stack's normal environment:
cd /opt/mike-ai/stack
docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice
OpenClaw 2026.9.4 uses the athena-talk plugin version 1.2.1 with
talk.realtime.transport set to webrtc and
talk.realtime.providers.athena-talk.realtimeUpstreamUrl set to
http://192.168.1.212:8090/v1/realtime/calls. On the Mac, turn off
“Echtzeitweiterleitung über Gateway verwenden”, which applies only to the
older gateway-relay mode. The tested Mac UI showed this switch on again
after navigation despite a false native preference; this UI behavior remains
unresolved. A Gateway restart after plugin installation was
needed to register the two new HTTPS routes. The plugin signs short-lived
session tokens with an in-memory Ed25519 key; the service fetches only its
public key. A Gateway restart rotates the key automatically.
Local isolated test (fake STT/TTS, synthetic microphone and speaker audio):
python3.11 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/python -m unittest -v test_smoke.py
The synthetic microphone test passed over the actual OpenClaw HTTPS offer
route and WireGuard media path with production Whisper and Qwen3-TTS on
2026-09-16. Subsequent real browser Talk sessions successfully transcribed
and answered multiple user turns. Browser dictation through the same plugin
also produced text after a correction to its separate transcription path.
The first user report of a second turn becoming stuck was addressed by
matching conversation item and predecessor IDs; the later two-turn test
passed. This is still half-duplex and needs a private route for WebRTC media.
The plugin's additional support for uploaded M4A voice notes is documented
in integrations/openclaw-athena-talk/README.md.
For isolated tests, ATHENA_TALK_REALTIME_SECRET can replace
OPENCLAW_PUBLIC_KEY_URL; do not use that test mode in the deployed setup.