Stream long OpenClaw dictation in segments

This commit is contained in:
Mikei386
2026-09-21 16:06:15 +02:00
parent e577c55489
commit 46e5bbdf7f
10 changed files with 313 additions and 72 deletions
+7 -3
View File
@@ -69,7 +69,7 @@ cd /opt/mike-ai/stack
docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice
```
OpenClaw 2026.9.4 uses the `athena-talk` plugin version 1.2.1 with
OpenClaw uses the `athena-talk` plugin version 1.3.0 with
`talk.realtime.transport` set to `webrtc` and
`talk.realtime.providers.athena-talk.realtimeUpstreamUrl` set to
`http://192.168.1.212:8090/v1/realtime/calls`. On the Mac, turn off
@@ -92,8 +92,12 @@ python3.11 -m venv .venv
The synthetic microphone test passed over the actual OpenClaw HTTPS offer
route and WireGuard media path with production Whisper and Qwen3-TTS on
2026-09-16. Subsequent real browser Talk sessions successfully transcribed
and answered multiple user turns. Browser dictation through the same plugin
also produced text after a correction to its separate transcription path.
and answered multiple user turns. Browser dictation uses a separate plugin
path. Since version 1.3.0, recordings longer than six seconds are
pre-transcribed incrementally in overlapping short windows while the
microphone remains active. This keeps the final tail inside OpenClaw's
five-second completion window. Talk through this WebRTC service still ends
and transcribes one utterance at a time.
The first user report of a second turn becoming stuck was addressed by
matching conversation item and predecessor IDs; the later two-turn test
passed. This is still half-duplex and needs a private route for WebRTC media.