From c0815ad29212ea2599a86433d7b1bb319024c3f6 Mon Sep 17 00:00:00 2001 From: Mikei386 <44135113+Mikei386@users.noreply.github.com> Date: Wed, 16 Sep 2026 14:58:29 +0200 Subject: [PATCH] Document OpenClaw agent completion latency before TTS --- services/athena-realtime-voice/README.md | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/services/athena-realtime-voice/README.md b/services/athena-realtime-voice/README.md index e48e58b..37d29e7 100644 --- a/services/athena-realtime-voice/README.md +++ b/services/athena-realtime-voice/README.md @@ -15,7 +15,13 @@ support are not implemented. For replies, the bridge consumes Athena's existing `/v1/audio/speech/pcm-stream` endpoint (24 kHz mono PCM) and enqueues audio frames as they arrive. Playback can start before Qwen finishes synthesizing a -long reply. The bounded queue applies backpressure instead of truncating long +long reply. This does not make playback start while the OpenClaw agent is still +writing its answer: OpenClaw's browser `openclaw_agent_consult` path submits the +tool result only after the agent's final chat event, so the bridge receives the +complete text first. In one observed run, this agent phase took 14.1 seconds +before TTS began. Streaming this phase would require an upstream OpenClaw Talk +interface for partial agent text, while retaining agent tool access. +The bounded queue applies backpressure instead of truncating long audio, and the next turn waits until all frames have played. The two-turn WebRTC test verifies that audio begins before the fake PCM stream finishes.