Files
AI-Profile-Router/services/athena-realtime-voice

Athena realtime voice bridge

This independent service lets OpenClaw's existing browser Talk UI use Athena Qwen3-ASR, OpenClaw's agent, and Athena Qwen3-TTS through the browser's supported OpenAI-style WebRTC transport. It does not modify OpenClaw or switch an Athena profile. The existing gateway-relay path remains available.

Flow: OpenClaw's athena-talk plugin signs a single-use 60-second browser token; the browser posts an SDP offer to this service; microphone audio is sent over WebRTC; Athena transcribes it; the service asks OpenClaw to run openclaw_agent_consult; Athena speaks the returned text over the same WebRTC connection. This is a half-duplex prototype. Barge-in and remote-network TURN support are not implemented.

For replies, the bridge consumes Athena's existing /v1/audio/speech/pcm-stream endpoint (24 kHz mono PCM) and enqueues audio frames as they arrive. Playback can start before Qwen finishes synthesizing a long reply. This does not make playback start while the OpenClaw agent is still writing its answer: OpenClaw's browser openclaw_agent_consult path submits the tool result only after the agent's final chat event, so the bridge receives the complete text first. In one observed run, this agent phase took 14.1 seconds before TTS began. Streaming this phase would require an upstream OpenClaw Talk interface for partial agent text, while retaining agent tool access. The bounded queue applies backpressure instead of truncating long audio, and the next turn waits until all frames have played. The two-turn WebRTC test verifies that audio begins before the fake PCM stream finishes.

The bridge announces each committed user audio item before sending its completed transcript with the same item ID. OpenClaw requires that sequence to persist the transcript; omitting the item caused “Realtime transcript refers to an unknown speech item” in the browser. Spoken assistant replies are also announced as conversation items with matching transcript IDs and explicit predecessor IDs. Without them, OpenClaw's ordered transcript storage can wait on a missing assistant item when the next user turn arrives, leaving later questions unanswered.

Required service environment:

Name Purpose
OPENCLAW_PUBLIC_KEY_URL Existing OpenClaw HTTPS origin plus /plugins/athena-talk/realtime/public-key. The service fetches the plugin's public verification key.
ATHENA_API_BASE_URL Router API base, http://router:8081/v1 in the Compose deployment.
ATHENA_API_KEY Router key, if required.
OPENCLAW_ORIGIN Exact HTTPS origin of the UI, e.g. https://oc.casaderoll.de.
PORT HTTP listen port; default 8090. The Compose service shares the existing WireGuard gateway network namespace.

The plugin provides an HTTPS offer route on the existing OpenClaw origin and forwards SDP to this service. Configure talk.realtime.providers.athena-talk.realtimeUpstreamUrl with Athena's internal HTTP URL ending in /v1/realtime/calls. Set the service's OPENCLAW_PUBLIC_KEY_URL to the OpenClaw HTTPS public-key route above. No new secret is needed in OpenClaw: its plugin keeps a private signing key in memory, and the service receives only the public key. Select talk.realtime.transport: "webrtc" only after the service is reachable. Keep gateway-relay as a rollback option.

The WebRTC media connection needs a route from the client device to Athena's ICE candidate addresses. The service runs directly inside Athena's existing private WireGuard network namespace, so home/VPN clients can reach it at 192.168.1.212:8090; the OpenClaw HTTPS proxy covers signaling only. Clients outside the private network need TURN support, which is not implemented.

The deployed service is mike-ai-realtime-voice in compose.yaml. It uses no GPU and does not switch Athena's active model profile. Deploy it with the stack's normal environment:

cd /opt/mike-ai/stack
docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice

OpenClaw uses the athena-talk plugin version 1.3.0 with talk.realtime.transport set to webrtc and talk.realtime.providers.athena-talk.realtimeUpstreamUrl set to http://192.168.1.212:8090/v1/realtime/calls. On the Mac, turn off “Echtzeitweiterleitung über Gateway verwenden”, which applies only to the older gateway-relay mode. The tested Mac UI showed this switch on again after navigation despite a false native preference; this UI behavior remains unresolved. A Gateway restart after plugin installation was needed to register the two new HTTPS routes. The plugin signs short-lived session tokens with an in-memory Ed25519 key; the service fetches only its public key. A Gateway restart rotates the key automatically.

Local isolated test (fake STT/TTS, synthetic microphone and speaker audio):

python3.11 -m venv .venv
.venv/bin/pip install -r requirements.txt
.venv/bin/python -m unittest -v test_smoke.py

The synthetic microphone test passed over the actual OpenClaw HTTPS offer route and WireGuard media path with production Qwen3-ASR and Qwen3-TTS on 2026-09-16. Subsequent real browser Talk sessions successfully transcribed and answered multiple user turns. Browser dictation uses a separate plugin path. Since version 1.3.0, recordings longer than six seconds are pre-transcribed incrementally in overlapping short windows while the microphone remains active. This keeps the final tail inside OpenClaw's five-second completion window. Talk through this WebRTC service still ends and transcribes one utterance at a time. The first user report of a second turn becoming stuck was addressed by matching conversation item and predecessor IDs; the later two-turn test passed. This is still half-duplex and needs a private route for WebRTC media. The plugin's additional support for uploaded M4A voice notes is documented in integrations/openclaw-athena-talk/README.md. For isolated tests, ATHENA_TALK_REALTIME_SECRET can replace OPENCLAW_PUBLIC_KEY_URL; do not use that test mode in the deployed setup.