Files
AI-Profile-Router/integrations/openclaw-athena-talk/README.md
T

168 lines
7.5 KiB
Markdown

# Athena Local Talk for OpenClaw
This private OpenClaw provider connects Gateway-relay Talk sessions to the
existing Athena speech stack. An experimental browser WebRTC path is available
through the separate `services/athena-realtime-voice` service:
1. local VAD collects a spoken utterance,
2. Athena Whisper transcribes it,
3. OpenClaw's normal agent-consult path answers with its configured model and
tools,
4. Athena Qwen3-TTS returns PCM audio to the Talk client.
The older `gateway-relay` path splits completed replies into short TTS clips.
The WebRTC path streams Athena's PCM audio to the browser as it is synthesized.
Both paths wait for OpenClaw's agent-consult result before TTS starts; neither
speaks partial text while the agent is still writing.
The provider intentionally uses half-duplex audio: microphone input is paused
while a response is being transcribed, generated, synthesized, or played. This
prevents speaker feedback from aborting TTS. Spoken interruption (barge-in) is
therefore disabled; wait until playback finishes before speaking again.
No public speech provider is used. `modelProvider` names an existing OpenClaw
model provider whose Athena base URL and API key are reused at runtime; no
second key copy is required. If it is omitted, the plugin checks `athena`,
`llama-cpp`, and `openai` in that order.
Recommended `talk.realtime` configuration:
```json
{
"provider": "athena-talk",
"model": "athena-local",
"speakerVoice": "alloy",
"mode": "realtime",
"transport": "webrtc",
"brain": "agent-consult",
"providers": {
"athena-talk": {
"modelProvider": "llama-cpp",
"realtimeUpstreamUrl": "http://192.168.1.212:8090/v1/realtime/calls",
"language": "de",
"vadThreshold": 0.018,
"silenceDurationMs": 750,
"prefixPaddingMs": 300,
"maxSpeechSeconds": 45
}
}
}
```
The plugin requires OpenClaw 2026.9.4 or newer. Build and validate it before
installation:
```sh
npm install
npm run build
npm run check
openclaw plugins install . --force --accept-capabilities
openclaw plugins inspect athena-talk --runtime --json
```
Version 1.3.0 also registers **Athena Whisper (Diktieren)** as a separate
realtime transcription provider through OpenClaw's official plugin API. In the
browser composer, hold the microphone for dictation, then release it to send
the 8 kHz G.711 audio through the Gateway. Short recordings are converted to
PCM WAV and sent to Athena's existing `/audio/transcriptions` endpoint in one
request. Longer recordings are split while the user is still speaking into
six-second windows with 0.5 seconds of overlap. The plugin sends these windows
sequentially to the persistent Whisper service, carries a short text prompt
into the next request, removes duplicated overlap words, and caches finished
segments until recording stops. Only the short final tail then remains inside
OpenClaw's fixed five-second final-drain window. Each Whisper request is capped
at 4.5 seconds.
This is incremental pre-transcription over OpenClaw's official transcription
provider API. Whisper.cpp still receives complete short WAV segments; it is
not a native token-streaming STT protocol. No OpenClaw core file was patched
and no additional speech container was introduced. The transcribed text is
returned to the composer; this path does not invoke the agent or TTS. The
provider reuses `talk.realtime.providers.athena-talk` and the configured model
provider for its URL/key. If that model provider has no key, it reuses
`tts.providers.openai.apiKey` only when the TTS and STT URLs have the same
origin. No second credential is needed. In `talk.catalog`, it appears under
`transcription.providers`.
### Voice-note file attachments
An M4A voice note uploaded as a chat attachment does **not** use the realtime
dictation provider above. OpenClaw processes it through its built-in
`tools.media.audio` path. On the Unraid installation, automatic provider
selection hit `SsrFBlockedError` for the private Athena address. Configure the
existing OpenAI-compatible provider and select Whisper explicitly:
```json5
{
models: {
providers: {
openai: {
// Keep the existing baseUrl and apiKey SecretRef.
request: { allowPrivateNetwork: true },
},
},
},
tools: {
media: {
models: [
{
provider: "openai",
model: "whisper-1",
baseUrl: "http://192.168.1.212:8081/v1",
capabilities: ["audio"],
},
],
},
},
}
```
OpenClaw 2026.9.4 accepts `request.allowPrivateNetwork` under
`models.providers.openai`, not under a `tools.media.models[]` entry. These
settings hot-reload without restarting the Gateway. The existing
`ATHENA_ROUTER_API_KEY` SecretRef is reused; do not add another literal key.
On 2026-09-16, OpenClaw's official audio module transcribed the affected M4A
through Athena Whisper (127 characters returned). This confirms the endpoint
and file format; a fresh attachment in the chat is still needed to verify the
full message-to-transcript flow.
Restart the gateway once if the installation does not trigger an automatic
reload. The managed plugin copy is stored in OpenClaw's persistent data
directory, so normal image updates do not remove it. Hermes remains unchanged;
the provider reuses the same OpenAI-compatible Athena speech endpoints.
## Browser and Mac app status (OpenClaw 2026.9.4)
The original 1.0.0 provider implements `gateway-relay` only. A direct Gateway call to
`talk.session.create` with `mode=realtime`, `transport=gateway-relay`, and
`brain=agent-consult` creates an Athena Talk session successfully. The Control
UI and Mac app first call `talk.client.create`, which rejects that transport
with `talk.client.create is client-owned; use talk.session.create for
gateway-relay`. Their fallback to `talk.session.create` is not observed in the
affected installation, so the UI still reports a misleading authentication
error. A successfully created test session is tied to its Gateway connection
and disappears when that connection closes.
The 2026.9.4 browser `provider-websocket` client accepts only the built-in
`google-live-bidi` protocol and validates the Google WebSocket hostname.
Version 1.1.0 adds a separate `webrtc` browser session using the UI's
OpenAI-style WebRTC protocol and the Athena realtime voice bridge. The plugin
proxies the SDP offer through the existing OpenClaw HTTPS origin. On
2026-09-16, the installed plugin created a browser session successfully; a
synthetic spoken request traversed the real OpenClaw offer route, private
WireGuard media connection, Athena Whisper, and Qwen3-TTS and returned audio.
The installed browser code also recognizes the bridge's
`openclaw_agent_consult` call. A real Chrome microphone-to-agent-to-speaker
call and a second question in the same session were observed on 2026-09-16.
The first browser failures required the bridge to announce user and assistant
transcript items with stable IDs and predecessor ordering. The Mac setting
“Echtzeitweiterleitung über Gateway verwenden” should be off for the
browser-owned WebRTC path. On the tested Mac, its UI switch returned to on
after navigation even though the native preference was set to false; Mac UI
behavior still needs verification. No OpenClaw core files were patched; the
plugin lives in OpenClaw's persistent extensions directory and the provider
is selected through `talk.realtime` configuration. `gateway-relay` remains available as a
reversible fallback. Never
select `provider-websocket` for this service. See
`services/athena-realtime-voice/README.md` for setup and limits.