Document OpenClaw agent completion latency before TTS

This commit is contained in:
Mikei386
2026-09-16 14:58:29 +02:00
parent 3f684b7325
commit c0815ad292
+7 -1
View File
@@ -15,7 +15,13 @@ support are not implemented.
For replies, the bridge consumes Athena's existing
`/v1/audio/speech/pcm-stream` endpoint (24 kHz mono PCM) and enqueues audio
frames as they arrive. Playback can start before Qwen finishes synthesizing a
long reply. The bounded queue applies backpressure instead of truncating long
long reply. This does not make playback start while the OpenClaw agent is still
writing its answer: OpenClaw's browser `openclaw_agent_consult` path submits the
tool result only after the agent's final chat event, so the bridge receives the
complete text first. In one observed run, this agent phase took 14.1 seconds
before TTS began. Streaming this phase would require an upstream OpenClaw Talk
interface for partial agent text, while retaining agent tool access.
The bounded queue applies backpressure instead of truncating long
audio, and the next turn waits until all frames have played. The two-turn
WebRTC test verifies that audio begins before the fake PCM stream finishes.