108 lines
5.9 KiB
Markdown
108 lines
5.9 KiB
Markdown
# Athena realtime voice bridge
|
|
|
|
This independent service lets OpenClaw's existing browser Talk UI use Athena
|
|
Whisper, OpenClaw's agent, and Athena Qwen3-TTS through the browser's supported
|
|
OpenAI-style WebRTC transport. It does not modify OpenClaw or switch an Athena
|
|
profile. The existing `gateway-relay` path remains available.
|
|
|
|
Flow: OpenClaw's `athena-talk` plugin signs a single-use 60-second browser token;
|
|
the browser posts an SDP offer to this service; microphone audio is sent over
|
|
WebRTC; Athena transcribes it; the service asks OpenClaw to run
|
|
`openclaw_agent_consult`; Athena speaks the returned text over the same WebRTC
|
|
connection. This is a half-duplex prototype. Barge-in and remote-network TURN
|
|
support are not implemented.
|
|
|
|
For replies, the bridge consumes Athena's existing
|
|
`/v1/audio/speech/pcm-stream` endpoint (24 kHz mono PCM) and enqueues audio
|
|
frames as they arrive. Playback can start before Qwen finishes synthesizing a
|
|
long reply. This does not make playback start while the OpenClaw agent is still
|
|
writing its answer: OpenClaw's browser `openclaw_agent_consult` path submits the
|
|
tool result only after the agent's final chat event, so the bridge receives the
|
|
complete text first. In one observed run, this agent phase took 14.1 seconds
|
|
before TTS began. Streaming this phase would require an upstream OpenClaw Talk
|
|
interface for partial agent text, while retaining agent tool access.
|
|
The bounded queue applies backpressure instead of truncating long
|
|
audio, and the next turn waits until all frames have played. The two-turn
|
|
WebRTC test verifies that audio begins before the fake PCM stream finishes.
|
|
|
|
The bridge announces each committed user audio item before sending its completed
|
|
transcript with the same item ID. OpenClaw requires that sequence to persist the
|
|
transcript; omitting the item caused “Realtime transcript refers to an unknown
|
|
speech item” in the browser.
|
|
Spoken assistant replies are also announced as conversation items with matching
|
|
transcript IDs and explicit predecessor IDs. Without them, OpenClaw's ordered
|
|
transcript storage can wait on a missing assistant item when the next user turn
|
|
arrives, leaving later questions unanswered.
|
|
|
|
Required service environment:
|
|
|
|
| Name | Purpose |
|
|
| --- | --- |
|
|
| `OPENCLAW_PUBLIC_KEY_URL` | Existing OpenClaw HTTPS origin plus `/plugins/athena-talk/realtime/public-key`. The service fetches the plugin's public verification key. |
|
|
| `ATHENA_API_BASE_URL` | Router API base, `http://router:8081/v1` in the Compose deployment. |
|
|
| `ATHENA_API_KEY` | Router key, if required. |
|
|
| `OPENCLAW_ORIGIN` | Exact HTTPS origin of the UI, e.g. `https://oc.casaderoll.de`. |
|
|
| `PORT` | HTTP listen port; default 8090. The Compose service shares the existing WireGuard gateway network namespace. |
|
|
|
|
The plugin provides an HTTPS offer route on the existing OpenClaw origin and
|
|
forwards SDP to this service. Configure
|
|
`talk.realtime.providers.athena-talk.realtimeUpstreamUrl` with Athena's internal
|
|
HTTP URL ending in `/v1/realtime/calls`. Set the service's
|
|
`OPENCLAW_PUBLIC_KEY_URL` to the OpenClaw HTTPS public-key route above. No new
|
|
secret is needed in OpenClaw: its plugin keeps a private signing key in memory,
|
|
and the service receives only the public key. Select
|
|
`talk.realtime.transport: "webrtc"` only after the service is reachable. Keep
|
|
`gateway-relay` as a rollback option.
|
|
|
|
The WebRTC media connection needs a route from the client device to Athena's
|
|
ICE candidate addresses. The service runs directly inside Athena's existing
|
|
private WireGuard network namespace, so home/VPN clients can reach it at
|
|
`192.168.1.212:8090`; the OpenClaw HTTPS proxy covers signaling only. Clients
|
|
outside the private network need TURN support, which is not implemented.
|
|
|
|
The deployed service is `mike-ai-realtime-voice` in `compose.yaml`. It uses no
|
|
GPU and does not switch Athena's active model profile. Deploy it with the
|
|
stack's normal environment:
|
|
|
|
```sh
|
|
cd /opt/mike-ai/stack
|
|
docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice
|
|
```
|
|
|
|
OpenClaw uses the `athena-talk` plugin version 1.3.0 with
|
|
`talk.realtime.transport` set to `webrtc` and
|
|
`talk.realtime.providers.athena-talk.realtimeUpstreamUrl` set to
|
|
`http://192.168.1.212:8090/v1/realtime/calls`. On the Mac, turn off
|
|
“Echtzeitweiterleitung über Gateway verwenden”, which applies only to the
|
|
older `gateway-relay` mode. The tested Mac UI showed this switch on again
|
|
after navigation despite a false native preference; this UI behavior remains
|
|
unresolved. A Gateway restart after plugin installation was
|
|
needed to register the two new HTTPS routes. The plugin signs short-lived
|
|
session tokens with an in-memory Ed25519 key; the service fetches only its
|
|
public key. A Gateway restart rotates the key automatically.
|
|
|
|
Local isolated test (fake STT/TTS, synthetic microphone and speaker audio):
|
|
|
|
```sh
|
|
python3.11 -m venv .venv
|
|
.venv/bin/pip install -r requirements.txt
|
|
.venv/bin/python -m unittest -v test_smoke.py
|
|
```
|
|
|
|
The synthetic microphone test passed over the actual OpenClaw HTTPS offer
|
|
route and WireGuard media path with production Whisper and Qwen3-TTS on
|
|
2026-09-16. Subsequent real browser Talk sessions successfully transcribed
|
|
and answered multiple user turns. Browser dictation uses a separate plugin
|
|
path. Since version 1.3.0, recordings longer than six seconds are
|
|
pre-transcribed incrementally in overlapping short windows while the
|
|
microphone remains active. This keeps the final tail inside OpenClaw's
|
|
five-second completion window. Talk through this WebRTC service still ends
|
|
and transcribes one utterance at a time.
|
|
The first user report of a second turn becoming stuck was addressed by
|
|
matching conversation item and predecessor IDs; the later two-turn test
|
|
passed. This is still half-duplex and needs a private route for WebRTC media.
|
|
The plugin's additional support for uploaded M4A voice notes is documented
|
|
in `integrations/openclaw-athena-talk/README.md`.
|
|
For isolated tests, `ATHENA_TALK_REALTIME_SECRET` can replace
|
|
`OPENCLAW_PUBLIC_KEY_URL`; do not use that test mode in the deployed setup.
|