Stream realtime TTS audio as it is synthesized

This commit is contained in:
Mikei386
2026-09-16 14:55:12 +02:00
parent 97c8072b3d
commit 3f684b7325
3 changed files with 60 additions and 31 deletions
+7
View File
@@ -12,6 +12,13 @@ WebRTC; Athena transcribes it; the service asks OpenClaw to run
connection. This is a half-duplex prototype. Barge-in and remote-network TURN
support are not implemented.
For replies, the bridge consumes Athena's existing
`/v1/audio/speech/pcm-stream` endpoint (24 kHz mono PCM) and enqueues audio
frames as they arrive. Playback can start before Qwen finishes synthesizing a
long reply. The bounded queue applies backpressure instead of truncating long
audio, and the next turn waits until all frames have played. The two-turn
WebRTC test verifies that audio begins before the fake PCM stream finishes.
The bridge announces each committed user audio item before sending its completed
transcript with the same item ID. OpenClaw requires that sequence to persist the
transcript; omitting the item caused “Realtime transcript refers to an unknown