48 lines
1.6 KiB
Markdown
48 lines
1.6 KiB
Markdown
# Athena Local Talk for OpenClaw
|
|
|
|
This private OpenClaw provider connects browser/Desktop Talk to the existing
|
|
Athena speech stack:
|
|
|
|
1. local VAD collects a spoken utterance,
|
|
2. Athena Whisper transcribes it,
|
|
3. OpenClaw's normal agent-consult path answers with its configured model and
|
|
tools,
|
|
4. Athena XTTS/Piper returns PCM audio to the Talk client.
|
|
|
|
Long replies are synthesized incrementally. The first short phrase starts
|
|
playing as soon as it is ready while the next phrase is generated in parallel.
|
|
|
|
The provider intentionally uses half-duplex audio: microphone input is paused
|
|
while a response is being transcribed, generated, synthesized, or played. This
|
|
prevents speaker feedback from aborting TTS. Spoken interruption (barge-in) is
|
|
therefore disabled; wait until playback finishes before speaking again.
|
|
|
|
No public speech provider is used. The provider reuses the already configured
|
|
`models.providers.athena` base URL and API key; no second key copy is required.
|
|
|
|
Recommended `talk.realtime` configuration:
|
|
|
|
```json
|
|
{
|
|
"provider": "athena-talk",
|
|
"model": "athena-local",
|
|
"speakerVoice": "alloy",
|
|
"language": "de",
|
|
"mode": "realtime",
|
|
"transport": "gateway-relay",
|
|
"brain": "agent-consult",
|
|
"providers": {
|
|
"athena-talk": {
|
|
"vadThreshold": 0.018,
|
|
"silenceDurationMs": 750,
|
|
"prefixPaddingMs": 300,
|
|
"maxSpeechSeconds": 45
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
Install from the OpenClaw container with `openclaw plugins install <path>` and
|
|
restart the gateway once. The plugin is stored in OpenClaw's persistent data
|
|
directory, so normal image updates do not remove it.
|