7.6 KiB
Athena Local Talk for OpenClaw
This private OpenClaw provider connects Gateway-relay Talk sessions to the
existing Athena speech stack. An experimental browser WebRTC path is available
through the separate services/athena-realtime-voice service:
- local VAD collects a spoken utterance,
- Athena Qwen3-ASR transcribes it,
- OpenClaw's normal agent-consult path answers with its configured model and tools,
- Athena Qwen3-TTS returns PCM audio to the Talk client.
The older gateway-relay path splits completed replies into short TTS clips.
The WebRTC path streams Athena's PCM audio to the browser as it is synthesized.
Both paths wait for OpenClaw's agent-consult result before TTS starts; neither
speaks partial text while the agent is still writing.
The provider intentionally uses half-duplex audio: microphone input is paused while a response is being transcribed, generated, synthesized, or played. This prevents speaker feedback from aborting TTS. Spoken interruption (barge-in) is therefore disabled; wait until playback finishes before speaking again.
No public speech provider is used. modelProvider names an existing OpenClaw
model provider whose Athena base URL and API key are reused at runtime; no
second key copy is required. If it is omitted, the plugin checks athena,
llama-cpp, and openai in that order.
Recommended talk.realtime configuration:
{
"provider": "athena-talk",
"model": "athena-local",
"speakerVoice": "alloy",
"mode": "realtime",
"transport": "webrtc",
"brain": "agent-consult",
"providers": {
"athena-talk": {
"modelProvider": "llama-cpp",
"realtimeUpstreamUrl": "http://192.168.1.212:8090/v1/realtime/calls",
"language": "de",
"vadThreshold": 0.018,
"silenceDurationMs": 750,
"prefixPaddingMs": 300,
"maxSpeechSeconds": 180
}
}
}
The plugin requires OpenClaw 2026.9.4 or newer. Build and validate it before installation:
npm install
npm run build
npm run check
openclaw plugins install . --force --accept-capabilities
openclaw plugins inspect athena-talk --runtime --json
The plugin registers Athena Qwen3-ASR (Diktieren) as a separate
realtime transcription provider through OpenClaw's official plugin API. In the
browser composer, hold the microphone for dictation, then release it to send
the 8 kHz G.711 audio through the Gateway. Short recordings are converted to
PCM WAV and sent to Athena's existing /audio/transcriptions endpoint in one
request. Longer recordings are split while the user is still speaking into
six-second windows with 0.5 seconds of overlap. The plugin sends these windows
sequentially to the persistent Qwen3-ASR service, removes duplicated overlap
words, and caches finished
segments until recording stops. Only the short final tail then remains inside
OpenClaw's fixed five-second final-drain window. Each STT request is capped
at 4.5 seconds.
This is incremental pre-transcription over OpenClaw's official transcription
provider API. Qwen3-ASR still receives complete short WAV segments; it is
not a native token-streaming STT protocol. No OpenClaw core file was patched.
The transcribed text is
returned to the composer; this path does not invoke the agent or TTS. The
provider reuses talk.realtime.providers.athena-talk and the configured model
provider for its URL/key. If that model provider has no key, it reuses
tts.providers.openai.apiKey only when the TTS and STT URLs have the same
origin. No second credential is needed. In talk.catalog, it appears under
transcription.providers.
The production provider allows up to 180 seconds per recording. This limit is a local safety cap shared by dictation and Talk, not an OpenClaw or Qwen3-ASR restriction. Incremental segmentation keeps long dictation bounded while it is being recorded.
Voice-note file attachments
An M4A voice note uploaded as a chat attachment does not use the realtime
dictation provider above. OpenClaw processes it through its built-in
tools.media.audio path. On the Unraid installation, automatic provider
selection hit SsrFBlockedError for the private Athena address. Configure the
existing OpenAI-compatible provider and retain the whisper-1 API alias for
Qwen3-ASR:
{
models: {
providers: {
openai: {
// Keep the existing baseUrl and apiKey SecretRef.
request: { allowPrivateNetwork: true },
},
},
},
tools: {
media: {
models: [
{
provider: "openai",
model: "whisper-1",
baseUrl: "http://192.168.1.212:8081/v1",
capabilities: ["audio"],
},
],
},
},
}
OpenClaw 2026.9.4 accepts request.allowPrivateNetwork under
models.providers.openai, not under a tools.media.models[] entry. These
settings hot-reload without restarting the Gateway. The existing
ATHENA_ROUTER_API_KEY SecretRef is reused; do not add another literal key.
On 2026-09-16, OpenClaw's official audio module transcribed the affected M4A
through Athena's then-active Whisper service (127 characters returned). This confirms the endpoint
and file format; a fresh attachment in the chat is still needed to verify the
full message-to-transcript flow.
Restart the gateway once if the installation does not trigger an automatic reload. The managed plugin copy is stored in OpenClaw's persistent data directory, so normal image updates do not remove it. Hermes remains unchanged; the provider reuses the same OpenAI-compatible Athena speech endpoints.
Browser and Mac app status (OpenClaw 2026.9.4)
The original 1.0.0 provider implements gateway-relay only. A direct Gateway call to
talk.session.create with mode=realtime, transport=gateway-relay, and
brain=agent-consult creates an Athena Talk session successfully. The Control
UI and Mac app first call talk.client.create, which rejects that transport
with talk.client.create is client-owned; use talk.session.create for gateway-relay. Their fallback to talk.session.create is not observed in the
affected installation, so the UI still reports a misleading authentication
error. A successfully created test session is tied to its Gateway connection
and disappears when that connection closes.
The 2026.9.4 browser provider-websocket client accepts only the built-in
google-live-bidi protocol and validates the Google WebSocket hostname.
Version 1.1.0 adds a separate webrtc browser session using the UI's
OpenAI-style WebRTC protocol and the Athena realtime voice bridge. The plugin
proxies the SDP offer through the existing OpenClaw HTTPS origin. On
2026-09-16, the installed plugin created a browser session successfully; a
synthetic spoken request traversed the real OpenClaw offer route, private
WireGuard media connection, Athena Whisper, and Qwen3-TTS and returned audio.
The installed browser code also recognizes the bridge's
openclaw_agent_consult call. A real Chrome microphone-to-agent-to-speaker
call and a second question in the same session were observed on 2026-09-16.
The first browser failures required the bridge to announce user and assistant
transcript items with stable IDs and predecessor ordering. The Mac setting
“Echtzeitweiterleitung über Gateway verwenden” should be off for the
browser-owned WebRTC path. On the tested Mac, its UI switch returned to on
after navigation even though the native preference was set to false; Mac UI
behavior still needs verification. No OpenClaw core files were patched; the
plugin lives in OpenClaw's persistent extensions directory and the provider
is selected through talk.realtime configuration. gateway-relay remains available as a
reversible fallback. Never
select provider-websocket for this service. See
services/athena-realtime-voice/README.md for setup and limits.