Document OpenClaw agent completion latency before TTS
This commit is contained in:
@@ -15,7 +15,13 @@ support are not implemented.
|
|||||||
For replies, the bridge consumes Athena's existing
|
For replies, the bridge consumes Athena's existing
|
||||||
`/v1/audio/speech/pcm-stream` endpoint (24 kHz mono PCM) and enqueues audio
|
`/v1/audio/speech/pcm-stream` endpoint (24 kHz mono PCM) and enqueues audio
|
||||||
frames as they arrive. Playback can start before Qwen finishes synthesizing a
|
frames as they arrive. Playback can start before Qwen finishes synthesizing a
|
||||||
long reply. The bounded queue applies backpressure instead of truncating long
|
long reply. This does not make playback start while the OpenClaw agent is still
|
||||||
|
writing its answer: OpenClaw's browser `openclaw_agent_consult` path submits the
|
||||||
|
tool result only after the agent's final chat event, so the bridge receives the
|
||||||
|
complete text first. In one observed run, this agent phase took 14.1 seconds
|
||||||
|
before TTS began. Streaming this phase would require an upstream OpenClaw Talk
|
||||||
|
interface for partial agent text, while retaining agent tool access.
|
||||||
|
The bounded queue applies backpressure instead of truncating long
|
||||||
audio, and the next turn waits until all frames have played. The two-turn
|
audio, and the next turn waits until all frames have played. The two-turn
|
||||||
WebRTC test verifies that audio begins before the fake PCM stream finishes.
|
WebRTC test verifies that audio begins before the fake PCM stream finishes.
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user