diff --git a/services/athena-realtime-voice/README.md b/services/athena-realtime-voice/README.md index e48e58b..37d29e7 100644 --- a/services/athena-realtime-voice/README.md +++ b/services/athena-realtime-voice/README.md @@ -15,7 +15,13 @@ support are not implemented. For replies, the bridge consumes Athena's existing `/v1/audio/speech/pcm-stream` endpoint (24 kHz mono PCM) and enqueues audio frames as they arrive. Playback can start before Qwen finishes synthesizing a -long reply. The bounded queue applies backpressure instead of truncating long +long reply. This does not make playback start while the OpenClaw agent is still +writing its answer: OpenClaw's browser `openclaw_agent_consult` path submits the +tool result only after the agent's final chat event, so the bridge receives the +complete text first. In one observed run, this agent phase took 14.1 seconds +before TTS began. Streaming this phase would require an upstream OpenClaw Talk +interface for partial agent text, while retaining agent tool access. +The bounded queue applies backpressure instead of truncating long audio, and the next turn waits until all frames have played. The two-turn WebRTC test verifies that audio begins before the fake PCM stream finishes.