Stream long OpenClaw dictation in segments
This commit is contained in:
@@ -60,19 +60,29 @@ openclaw plugins install . --force --accept-capabilities
|
||||
openclaw plugins inspect athena-talk --runtime --json
|
||||
```
|
||||
|
||||
Version 1.2.1 also registers **Athena Whisper (Diktieren)** as a separate
|
||||
Version 1.3.0 also registers **Athena Whisper (Diktieren)** as a separate
|
||||
realtime transcription provider through OpenClaw's official plugin API. In the
|
||||
browser composer, hold the microphone for dictation, then release it to send
|
||||
the 8 kHz G.711 audio through the Gateway. The plugin converts it to PCM WAV
|
||||
and calls the same Athena `/audio/transcriptions` endpoint used by Talk. The
|
||||
transcribed text is returned to the composer; this path does not invoke the
|
||||
agent or TTS. The transcription provider reuses `talk.realtime.providers.athena-talk`
|
||||
and the configured model provider for its URL/key. If that model provider has
|
||||
no key, it reuses `tts.providers.openai.apiKey` only when the TTS and STT URLs
|
||||
have the same origin. No second credential is needed. In `talk.catalog`, it
|
||||
appears under `transcription.providers`. OpenClaw
|
||||
currently gives a transcription provider five seconds to return its final text
|
||||
after recording stops; the plugin caps its Whisper request at 4.5 seconds.
|
||||
the 8 kHz G.711 audio through the Gateway. Short recordings are converted to
|
||||
PCM WAV and sent to Athena's existing `/audio/transcriptions` endpoint in one
|
||||
request. Longer recordings are split while the user is still speaking into
|
||||
six-second windows with 0.5 seconds of overlap. The plugin sends these windows
|
||||
sequentially to the persistent Whisper service, carries a short text prompt
|
||||
into the next request, removes duplicated overlap words, and caches finished
|
||||
segments until recording stops. Only the short final tail then remains inside
|
||||
OpenClaw's fixed five-second final-drain window. Each Whisper request is capped
|
||||
at 4.5 seconds.
|
||||
|
||||
This is incremental pre-transcription over OpenClaw's official transcription
|
||||
provider API. Whisper.cpp still receives complete short WAV segments; it is
|
||||
not a native token-streaming STT protocol. No OpenClaw core file was patched
|
||||
and no additional speech container was introduced. The transcribed text is
|
||||
returned to the composer; this path does not invoke the agent or TTS. The
|
||||
provider reuses `talk.realtime.providers.athena-talk` and the configured model
|
||||
provider for its URL/key. If that model provider has no key, it reuses
|
||||
`tts.providers.openai.apiKey` only when the TTS and STT URLs have the same
|
||||
origin. No second credential is needed. In `talk.catalog`, it appears under
|
||||
`transcription.providers`.
|
||||
|
||||
### Voice-note file attachments
|
||||
|
||||
|
||||
Reference in New Issue
Block a user