Stream long OpenClaw dictation in segments

This commit is contained in:
Mikei386
2026-09-21 16:06:15 +02:00
parent e577c55489
commit 46e5bbdf7f
10 changed files with 313 additions and 72 deletions
+21 -11
View File
@@ -60,19 +60,29 @@ openclaw plugins install . --force --accept-capabilities
openclaw plugins inspect athena-talk --runtime --json
```
Version 1.2.1 also registers **Athena Whisper (Diktieren)** as a separate
Version 1.3.0 also registers **Athena Whisper (Diktieren)** as a separate
realtime transcription provider through OpenClaw's official plugin API. In the
browser composer, hold the microphone for dictation, then release it to send
the 8 kHz G.711 audio through the Gateway. The plugin converts it to PCM WAV
and calls the same Athena `/audio/transcriptions` endpoint used by Talk. The
transcribed text is returned to the composer; this path does not invoke the
agent or TTS. The transcription provider reuses `talk.realtime.providers.athena-talk`
and the configured model provider for its URL/key. If that model provider has
no key, it reuses `tts.providers.openai.apiKey` only when the TTS and STT URLs
have the same origin. No second credential is needed. In `talk.catalog`, it
appears under `transcription.providers`. OpenClaw
currently gives a transcription provider five seconds to return its final text
after recording stops; the plugin caps its Whisper request at 4.5 seconds.
the 8 kHz G.711 audio through the Gateway. Short recordings are converted to
PCM WAV and sent to Athena's existing `/audio/transcriptions` endpoint in one
request. Longer recordings are split while the user is still speaking into
six-second windows with 0.5 seconds of overlap. The plugin sends these windows
sequentially to the persistent Whisper service, carries a short text prompt
into the next request, removes duplicated overlap words, and caches finished
segments until recording stops. Only the short final tail then remains inside
OpenClaw's fixed five-second final-drain window. Each Whisper request is capped
at 4.5 seconds.
This is incremental pre-transcription over OpenClaw's official transcription
provider API. Whisper.cpp still receives complete short WAV segments; it is
not a native token-streaming STT protocol. No OpenClaw core file was patched
and no additional speech container was introduced. The transcribed text is
returned to the composer; this path does not invoke the agent or TTS. The
provider reuses `talk.realtime.providers.athena-talk` and the configured model
provider for its URL/key. If that model provider has no key, it reuses
`tts.providers.openai.apiKey` only when the TTS and STT URLs have the same
origin. No second credential is needed. In `talk.catalog`, it appears under
`transcription.providers`.
### Voice-note file attachments