158 lines
6.9 KiB
Markdown
158 lines
6.9 KiB
Markdown
# Athena Local Talk for OpenClaw
|
|
|
|
This private OpenClaw provider connects Gateway-relay Talk sessions to the
|
|
existing Athena speech stack. An experimental browser WebRTC path is available
|
|
through the separate `services/athena-realtime-voice` service:
|
|
|
|
1. local VAD collects a spoken utterance,
|
|
2. Athena Whisper transcribes it,
|
|
3. OpenClaw's normal agent-consult path answers with its configured model and
|
|
tools,
|
|
4. Athena Qwen3-TTS returns PCM audio to the Talk client.
|
|
|
|
The older `gateway-relay` path splits completed replies into short TTS clips.
|
|
The WebRTC path streams Athena's PCM audio to the browser as it is synthesized.
|
|
Both paths wait for OpenClaw's agent-consult result before TTS starts; neither
|
|
speaks partial text while the agent is still writing.
|
|
|
|
The provider intentionally uses half-duplex audio: microphone input is paused
|
|
while a response is being transcribed, generated, synthesized, or played. This
|
|
prevents speaker feedback from aborting TTS. Spoken interruption (barge-in) is
|
|
therefore disabled; wait until playback finishes before speaking again.
|
|
|
|
No public speech provider is used. `modelProvider` names an existing OpenClaw
|
|
model provider whose Athena base URL and API key are reused at runtime; no
|
|
second key copy is required. If it is omitted, the plugin checks `athena`,
|
|
`llama-cpp`, and `openai` in that order.
|
|
|
|
Recommended `talk.realtime` configuration:
|
|
|
|
```json
|
|
{
|
|
"provider": "athena-talk",
|
|
"model": "athena-local",
|
|
"speakerVoice": "alloy",
|
|
"mode": "realtime",
|
|
"transport": "webrtc",
|
|
"brain": "agent-consult",
|
|
"providers": {
|
|
"athena-talk": {
|
|
"modelProvider": "llama-cpp",
|
|
"realtimeUpstreamUrl": "http://192.168.1.212:8090/v1/realtime/calls",
|
|
"language": "de",
|
|
"vadThreshold": 0.018,
|
|
"silenceDurationMs": 750,
|
|
"prefixPaddingMs": 300,
|
|
"maxSpeechSeconds": 45
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
The plugin requires OpenClaw 2026.9.4 or newer. Build and validate it before
|
|
installation:
|
|
|
|
```sh
|
|
npm install
|
|
npm run build
|
|
npm run check
|
|
openclaw plugins install . --force --accept-capabilities
|
|
openclaw plugins inspect athena-talk --runtime --json
|
|
```
|
|
|
|
Version 1.2.1 also registers **Athena Whisper (Diktieren)** as a separate
|
|
realtime transcription provider through OpenClaw's official plugin API. In the
|
|
browser composer, hold the microphone for dictation, then release it to send
|
|
the 8 kHz G.711 audio through the Gateway. The plugin converts it to PCM WAV
|
|
and calls the same Athena `/audio/transcriptions` endpoint used by Talk. The
|
|
transcribed text is returned to the composer; this path does not invoke the
|
|
agent or TTS. The transcription provider reuses `talk.realtime.providers.athena-talk`
|
|
and the configured model provider for its URL/key. If that model provider has
|
|
no key, it reuses `tts.providers.openai.apiKey` only when the TTS and STT URLs
|
|
have the same origin. No second credential is needed. In `talk.catalog`, it
|
|
appears under `transcription.providers`. OpenClaw
|
|
currently gives a transcription provider five seconds to return its final text
|
|
after recording stops; the plugin caps its Whisper request at 4.5 seconds.
|
|
|
|
### Voice-note file attachments
|
|
|
|
An M4A voice note uploaded as a chat attachment does **not** use the realtime
|
|
dictation provider above. OpenClaw processes it through its built-in
|
|
`tools.media.audio` path. On the Unraid installation, automatic provider
|
|
selection hit `SsrFBlockedError` for the private Athena address. Configure the
|
|
existing OpenAI-compatible provider and select Whisper explicitly:
|
|
|
|
```json5
|
|
{
|
|
models: {
|
|
providers: {
|
|
openai: {
|
|
// Keep the existing baseUrl and apiKey SecretRef.
|
|
request: { allowPrivateNetwork: true },
|
|
},
|
|
},
|
|
},
|
|
tools: {
|
|
media: {
|
|
models: [
|
|
{
|
|
provider: "openai",
|
|
model: "whisper-1",
|
|
baseUrl: "http://192.168.1.212:8081/v1",
|
|
capabilities: ["audio"],
|
|
},
|
|
],
|
|
},
|
|
},
|
|
}
|
|
```
|
|
|
|
OpenClaw 2026.9.4 accepts `request.allowPrivateNetwork` under
|
|
`models.providers.openai`, not under a `tools.media.models[]` entry. These
|
|
settings hot-reload without restarting the Gateway. The existing
|
|
`ATHENA_ROUTER_API_KEY` SecretRef is reused; do not add another literal key.
|
|
On 2026-09-16, OpenClaw's official audio module transcribed the affected M4A
|
|
through Athena Whisper (127 characters returned). This confirms the endpoint
|
|
and file format; a fresh attachment in the chat is still needed to verify the
|
|
full message-to-transcript flow.
|
|
|
|
Restart the gateway once if the installation does not trigger an automatic
|
|
reload. The managed plugin copy is stored in OpenClaw's persistent data
|
|
directory, so normal image updates do not remove it. Hermes remains unchanged;
|
|
the provider reuses the same OpenAI-compatible Athena speech endpoints.
|
|
|
|
## Browser and Mac app status (OpenClaw 2026.9.4)
|
|
|
|
The original 1.0.0 provider implements `gateway-relay` only. A direct Gateway call to
|
|
`talk.session.create` with `mode=realtime`, `transport=gateway-relay`, and
|
|
`brain=agent-consult` creates an Athena Talk session successfully. The Control
|
|
UI and Mac app first call `talk.client.create`, which rejects that transport
|
|
with `talk.client.create is client-owned; use talk.session.create for
|
|
gateway-relay`. Their fallback to `talk.session.create` is not observed in the
|
|
affected installation, so the UI still reports a misleading authentication
|
|
error. A successfully created test session is tied to its Gateway connection
|
|
and disappears when that connection closes.
|
|
|
|
The 2026.9.4 browser `provider-websocket` client accepts only the built-in
|
|
`google-live-bidi` protocol and validates the Google WebSocket hostname.
|
|
Version 1.1.0 adds a separate `webrtc` browser session using the UI's
|
|
OpenAI-style WebRTC protocol and the Athena realtime voice bridge. The plugin
|
|
proxies the SDP offer through the existing OpenClaw HTTPS origin. On
|
|
2026-09-16, the installed plugin created a browser session successfully; a
|
|
synthetic spoken request traversed the real OpenClaw offer route, private
|
|
WireGuard media connection, Athena Whisper, and Qwen3-TTS and returned audio.
|
|
The installed browser code also recognizes the bridge's
|
|
`openclaw_agent_consult` call. A real Chrome microphone-to-agent-to-speaker
|
|
call and a second question in the same session were observed on 2026-09-16.
|
|
The first browser failures required the bridge to announce user and assistant
|
|
transcript items with stable IDs and predecessor ordering. The Mac setting
|
|
“Echtzeitweiterleitung über Gateway verwenden” should be off for the
|
|
browser-owned WebRTC path. On the tested Mac, its UI switch returned to on
|
|
after navigation even though the native preference was set to false; Mac UI
|
|
behavior still needs verification. No OpenClaw core files were patched; the
|
|
plugin lives in OpenClaw's persistent extensions directory and the provider
|
|
is selected through `talk.realtime` configuration. `gateway-relay` remains available as a
|
|
reversible fallback. Never
|
|
select `provider-websocket` for this service. See
|
|
`services/athena-realtime-voice/README.md` for setup and limits.
|