diff --git a/README.md b/README.md index d9ea175..c250d4c 100644 --- a/README.md +++ b/README.md @@ -151,6 +151,11 @@ verbindet Mikrofon → Athena Whisper → normalen OpenClaw-Agenten → aktives Athena-TTS, sodass Modell, Werkzeuge und Memory auch im Sprachmodus erhalten bleiben. Die Installation landet in OpenClaws persistentem Datenverzeichnis und bleibt deshalb bei normalen Container-Updates bestehen. +Die separate Diktierfunktion verarbeitet seit Plugin-Version 1.3.0 längere +Aufnahmen bereits während des Sprechens in überlappenden Sechs-Sekunden- +Abschnitten. Beim Loslassen bleibt nur der kurze Rest für OpenClaws festes +Fünf-Sekunden-Abschlussfenster. Dafür wurden weder OpenClaw selbst verändert +noch ein weiterer Container angelegt. OpenClaw wird über den Provider **llama.cpp → Existing llama-server** mit `http://192.168.1.212:8081/v1` verbunden. Der Router beantwortet sowohl diff --git a/docs/LIVE_STATE.md b/docs/LIVE_STATE.md index 35b52e6..0043c65 100644 --- a/docs/LIVE_STATE.md +++ b/docs/LIVE_STATE.md @@ -64,7 +64,10 @@ nicht mehr den aktuellen Containerzustand. Er braucht keine GPU und keinen Profilwechsel. Das Plugin liegt in `integrations/openclaw-athena-talk` und läuft auf Unraid, nicht auf Athena. Diktat und hochgeladene M4A-Sprachnachrichten nutzen den separaten - Transkriptionspfad des Plugins. OpenClaw liefert den fertigen Agententext + Transkriptionspfad des Plugins. Plugin 1.3.0 zerlegt längere Browser-Diktate + während der Aufnahme in überlappende Sechs-Sekunden-Abschnitte und hält so + den Abschluss innerhalb von OpenClaws festem Fünf-Sekunden-Fenster. OpenClaw + selbst wurde dafür nicht gepatcht. OpenClaw liefert den fertigen Agententext an die Brücke; das Sprechen beginnt daher erst nach Abschluss der Agentenantwort. Details und Grenzen stehen in der [Voice-Doku](../services/athena-realtime-voice/README.md). - `mike-ai-mikes-applio-ui` läuft gesund und ohne GPU. Die Quelle liegt im diff --git a/docs/RECOVERY.md b/docs/RECOVERY.md index a296d32..1cad98e 100644 --- a/docs/RECOVERY.md +++ b/docs/RECOVERY.md @@ -16,6 +16,12 @@ eingeschlossen: Im Archiv `athena-2026-09-16T13-02-56.tar.gz` wurden sowohl `compose.yaml` als auch `services/athena-realtime-voice/server.py` geprüft. Das OpenClaw-Plugin auf Unraid liegt außerhalb dieses Athena-Backups; seine Quelle ist im Git-Repository unter `integrations/openclaw-athena-talk` erfasst. +Die produktiv installierte Version 1.3.0 liegt zusätzlich im persistenten +OpenClaw-Appdata. Vor ihrer Installation wurde die bisherige Version als +`/mnt/nvme-storage/appdata/OpenClaw/config/plugin-backups/athena-talk-1.2.1-before-streaming.tar.gz` +gesichert. Für eine Neuinstallation ist der im Git dokumentierte Build mit +`openclaw plugins install --force --accept-capabilities` zu +installieren; eine Änderung an OpenClaw-Core-Dateien ist nicht erforderlich. Piper-Daten sind kein aktueller Sicherungsbestand. Modellgewichte unter `/data/models` und das reproduzierbare Whisper-Volume gehören nicht zu diesen diff --git a/integrations/openclaw-athena-talk/README.md b/integrations/openclaw-athena-talk/README.md index df78cef..1534da2 100644 --- a/integrations/openclaw-athena-talk/README.md +++ b/integrations/openclaw-athena-talk/README.md @@ -60,19 +60,29 @@ openclaw plugins install . --force --accept-capabilities openclaw plugins inspect athena-talk --runtime --json ``` -Version 1.2.1 also registers **Athena Whisper (Diktieren)** as a separate +Version 1.3.0 also registers **Athena Whisper (Diktieren)** as a separate realtime transcription provider through OpenClaw's official plugin API. In the browser composer, hold the microphone for dictation, then release it to send -the 8 kHz G.711 audio through the Gateway. The plugin converts it to PCM WAV -and calls the same Athena `/audio/transcriptions` endpoint used by Talk. The -transcribed text is returned to the composer; this path does not invoke the -agent or TTS. The transcription provider reuses `talk.realtime.providers.athena-talk` -and the configured model provider for its URL/key. If that model provider has -no key, it reuses `tts.providers.openai.apiKey` only when the TTS and STT URLs -have the same origin. No second credential is needed. In `talk.catalog`, it -appears under `transcription.providers`. OpenClaw -currently gives a transcription provider five seconds to return its final text -after recording stops; the plugin caps its Whisper request at 4.5 seconds. +the 8 kHz G.711 audio through the Gateway. Short recordings are converted to +PCM WAV and sent to Athena's existing `/audio/transcriptions` endpoint in one +request. Longer recordings are split while the user is still speaking into +six-second windows with 0.5 seconds of overlap. The plugin sends these windows +sequentially to the persistent Whisper service, carries a short text prompt +into the next request, removes duplicated overlap words, and caches finished +segments until recording stops. Only the short final tail then remains inside +OpenClaw's fixed five-second final-drain window. Each Whisper request is capped +at 4.5 seconds. + +This is incremental pre-transcription over OpenClaw's official transcription +provider API. Whisper.cpp still receives complete short WAV segments; it is +not a native token-streaming STT protocol. No OpenClaw core file was patched +and no additional speech container was introduced. The transcribed text is +returned to the composer; this path does not invoke the agent or TTS. The +provider reuses `talk.realtime.providers.athena-talk` and the configured model +provider for its URL/key. If that model provider has no key, it reuses +`tts.providers.openai.apiKey` only when the TTS and STT URLs have the same +origin. No second credential is needed. In `talk.catalog`, it appears under +`transcription.providers`. ### Voice-note file attachments diff --git a/integrations/openclaw-athena-talk/dist/index.js b/integrations/openclaw-athena-talk/dist/index.js index 6f48ff9..7b81683 100644 --- a/integrations/openclaw-athena-talk/dist/index.js +++ b/integrations/openclaw-athena-talk/dist/index.js @@ -4,6 +4,10 @@ const AUDIO_FORMAT = { encoding: "pcm16", sampleRateHz: 24000, channels: 1 }; const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls"; const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key"; const MAX_OFFER_BYTES = 64 * 1024; +const DICTATION_SAMPLE_RATE_HZ = 8000; +const DICTATION_SEGMENT_BYTES = DICTATION_SAMPLE_RATE_HZ * 6; +const DICTATION_OVERLAP_BYTES = DICTATION_SAMPLE_RATE_HZ / 2; +const DICTATION_REQUEST_TIMEOUT_MS = 4500; const browserKeys = generateKeyPairSync("ed25519"); const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString(); function base64url(value) { @@ -149,13 +153,38 @@ function resolveTranscriptionConfig(cfg, rawConfig) { const talkProvider = record(record(talkConfig.providers)["athena-talk"]); return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } }); } +function comparableWord(value) { + return value.toLocaleLowerCase("de-DE").replace(/[^\p{L}\p{N}]+/gu, ""); +} +function removeTranscriptOverlap(previous, current) { + const priorWords = previous.trim().split(/\s+/).filter(Boolean); + const currentWords = current.trim().split(/\s+/).filter(Boolean); + const maximum = Math.min(12, priorWords.length, currentWords.length); + for (let count = maximum; count >= 1; count -= 1) { + const left = priorWords.slice(-count).map(comparableWord); + const right = currentWords.slice(0, count).map(comparableWord); + if (!left.every((word, index) => word && word === right[index])) + continue; + // A single short word is too ambiguous to remove safely. Longer words are + // sufficient because the audio overlap is only half a second. + if (count === 1 && left[0].length < 5) + continue; + return currentWords.slice(count).join(" "); + } + return currentWords.join(" "); +} class AthenaTranscriptionSession { req; config; connected = false; closed = false; audio = []; - bytes = 0; + bufferedBytes = 0; + totalBytes = 0; + processing = Promise.resolve(); + completedTranscripts = []; + emittedTranscripts = 0; + processingError = null; constructor(req, config) { this.req = req; this.config = config; @@ -165,50 +194,88 @@ class AthenaTranscriptionSession { sendAudio(audio) { if (!this.isConnected() || audio.length === 0) return; - if (this.bytes === 0) + if (this.totalBytes === 0) this.req.onSpeechStart?.(); const maxBytes = this.config.maxSpeechSeconds * 8000; - if (this.bytes + audio.length > maxBytes) { + if (this.totalBytes + audio.length > maxBytes) { this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`)); this.close(); return; } this.audio.push(Buffer.from(audio)); - this.bytes += audio.length; + this.bufferedBytes += audio.length; + this.totalBytes += audio.length; + while (this.bufferedBytes >= DICTATION_SEGMENT_BYTES) { + this.queueFullSegment(); + } } close() { if (this.closed) return; this.closed = true; this.connected = false; - if (!this.bytes) + if (!this.totalBytes) return; - const audio = Buffer.concat(this.audio); - this.audio.length = 0; - void this.transcribe(audio); + this.emitCompletedTranscripts(); + const tail = Buffer.concat(this.audio); + this.audio = []; + this.bufferedBytes = 0; + if (tail.length > DICTATION_OVERLAP_BYTES || this.completedTranscripts.length === 0) { + this.queueTranscription(tail); + } + void this.processing.finally(() => { + this.emitCompletedTranscripts(); + if (this.processingError && this.completedTranscripts.length === 0) { + this.req.onError?.(this.processingError); + } + }); } - async transcribe(audio) { - try { - const form = new FormData(); - form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav"); - form.append("model", "whisper-1"); - form.append("language", this.config.language); - const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, { - method: "POST", - headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {}, - body: form, - signal: AbortSignal.timeout(4500), - }); - if (!response.ok) - throw new Error(`Athena STT failed (HTTP ${response.status})`); - const text = String(record(await response.json()).text || "").trim(); - if (text) - this.req.onTranscript?.(text); - } - catch (error) { - this.req.onError?.(error instanceof Error ? error : new Error(String(error))); + queueFullSegment() { + const buffered = Buffer.concat(this.audio); + const segment = Buffer.from(buffered.subarray(0, DICTATION_SEGMENT_BYTES)); + const retained = Buffer.from(buffered.subarray(DICTATION_SEGMENT_BYTES - DICTATION_OVERLAP_BYTES)); + this.audio = retained.length ? [retained] : []; + this.bufferedBytes = retained.length; + this.queueTranscription(segment); + } + queueTranscription(audio) { + if (!audio.length) + return; + this.processing = this.processing.then(async () => { + const previous = this.completedTranscripts.join(" "); + const text = await this.transcribe(audio, previous.slice(-240)); + const novel = removeTranscriptOverlap(previous, text); + if (novel) + this.completedTranscripts.push(novel); + if (this.closed) + this.emitCompletedTranscripts(); + }).catch((error) => { + this.processingError = error instanceof Error ? error : new Error(String(error)); + }); + } + emitCompletedTranscripts() { + while (this.emittedTranscripts < this.completedTranscripts.length) { + this.req.onTranscript?.(this.completedTranscripts[this.emittedTranscripts]); + this.emittedTranscripts += 1; } } + async transcribe(audio, prompt) { + const form = new FormData(); + form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav"); + form.append("model", "whisper-1"); + form.append("language", this.config.language); + if (prompt) + form.append("prompt", prompt); + const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, { + method: "POST", + headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {}, + body: form, + signal: AbortSignal.timeout(DICTATION_REQUEST_TIMEOUT_MS), + }); + if (!response.ok) + throw new Error(`Athena STT failed (HTTP ${response.status})`); + return String(record(await response.json()).text || "").trim(); + } } function pcmRms(pcm) { if (pcm.length < 2) diff --git a/integrations/openclaw-athena-talk/index.ts b/integrations/openclaw-athena-talk/index.ts index ec652a5..8a37ca9 100644 --- a/integrations/openclaw-athena-talk/index.ts +++ b/integrations/openclaw-athena-talk/index.ts @@ -20,6 +20,10 @@ type ProviderConfig = { const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls"; const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key"; const MAX_OFFER_BYTES = 64 * 1024; +const DICTATION_SAMPLE_RATE_HZ = 8000; +const DICTATION_SEGMENT_BYTES = DICTATION_SAMPLE_RATE_HZ * 6; +const DICTATION_OVERLAP_BYTES = DICTATION_SAMPLE_RATE_HZ / 2; +const DICTATION_REQUEST_TIMEOUT_MS = 4500; const browserKeys = generateKeyPairSync("ed25519"); const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString(); @@ -171,11 +175,36 @@ function resolveTranscriptionConfig(cfg: unknown, rawConfig: unknown): Required< return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } }); } +function comparableWord(value: string): string { + return value.toLocaleLowerCase("de-DE").replace(/[^\p{L}\p{N}]+/gu, ""); +} + +function removeTranscriptOverlap(previous: string, current: string): string { + const priorWords = previous.trim().split(/\s+/).filter(Boolean); + const currentWords = current.trim().split(/\s+/).filter(Boolean); + const maximum = Math.min(12, priorWords.length, currentWords.length); + for (let count = maximum; count >= 1; count -= 1) { + const left = priorWords.slice(-count).map(comparableWord); + const right = currentWords.slice(0, count).map(comparableWord); + if (!left.every((word, index) => word && word === right[index])) continue; + // A single short word is too ambiguous to remove safely. Longer words are + // sufficient because the audio overlap is only half a second. + if (count === 1 && left[0].length < 5) continue; + return currentWords.slice(count).join(" "); + } + return currentWords.join(" "); +} + class AthenaTranscriptionSession { private connected = false; private closed = false; - private readonly audio: Buffer[] = []; - private bytes = 0; + private audio: Buffer[] = []; + private bufferedBytes = 0; + private totalBytes = 0; + private processing: Promise = Promise.resolve(); + private readonly completedTranscripts: string[] = []; + private emittedTranscripts = 0; + private processingError: Error | null = null; constructor( private readonly req: { @@ -193,46 +222,85 @@ class AthenaTranscriptionSession { sendAudio(audio: Buffer): void { if (!this.isConnected() || audio.length === 0) return; - if (this.bytes === 0) this.req.onSpeechStart?.(); + if (this.totalBytes === 0) this.req.onSpeechStart?.(); const maxBytes = this.config.maxSpeechSeconds * 8000; - if (this.bytes + audio.length > maxBytes) { + if (this.totalBytes + audio.length > maxBytes) { this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`)); this.close(); return; } this.audio.push(Buffer.from(audio)); - this.bytes += audio.length; + this.bufferedBytes += audio.length; + this.totalBytes += audio.length; + while (this.bufferedBytes >= DICTATION_SEGMENT_BYTES) { + this.queueFullSegment(); + } } close(): void { if (this.closed) return; this.closed = true; this.connected = false; - if (!this.bytes) return; - const audio = Buffer.concat(this.audio); - this.audio.length = 0; - void this.transcribe(audio); + if (!this.totalBytes) return; + this.emitCompletedTranscripts(); + const tail = Buffer.concat(this.audio); + this.audio = []; + this.bufferedBytes = 0; + if (tail.length > DICTATION_OVERLAP_BYTES || this.completedTranscripts.length === 0) { + this.queueTranscription(tail); + } + void this.processing.finally(() => { + this.emitCompletedTranscripts(); + if (this.processingError && this.completedTranscripts.length === 0) { + this.req.onError?.(this.processingError); + } + }); } - private async transcribe(audio: Buffer): Promise { - try { - const form = new FormData(); - form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav"); - form.append("model", "whisper-1"); - form.append("language", this.config.language); - const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, { - method: "POST", - headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {}, - body: form, - signal: AbortSignal.timeout(4500), - }); - if (!response.ok) throw new Error(`Athena STT failed (HTTP ${response.status})`); - const text = String(record(await response.json()).text || "").trim(); - if (text) this.req.onTranscript?.(text); - } catch (error) { - this.req.onError?.(error instanceof Error ? error : new Error(String(error))); + private queueFullSegment(): void { + const buffered = Buffer.concat(this.audio); + const segment = Buffer.from(buffered.subarray(0, DICTATION_SEGMENT_BYTES)); + const retained = Buffer.from(buffered.subarray(DICTATION_SEGMENT_BYTES - DICTATION_OVERLAP_BYTES)); + this.audio = retained.length ? [retained] : []; + this.bufferedBytes = retained.length; + this.queueTranscription(segment); + } + + private queueTranscription(audio: Buffer): void { + if (!audio.length) return; + this.processing = this.processing.then(async () => { + const previous = this.completedTranscripts.join(" "); + const text = await this.transcribe(audio, previous.slice(-240)); + const novel = removeTranscriptOverlap(previous, text); + if (novel) this.completedTranscripts.push(novel); + if (this.closed) this.emitCompletedTranscripts(); + }).catch((error: unknown) => { + this.processingError = error instanceof Error ? error : new Error(String(error)); + }); + } + + private emitCompletedTranscripts(): void { + while (this.emittedTranscripts < this.completedTranscripts.length) { + this.req.onTranscript?.(this.completedTranscripts[this.emittedTranscripts]); + this.emittedTranscripts += 1; } } + + private async transcribe(audio: Buffer, prompt: string): Promise { + const form = new FormData(); + form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav"); + form.append("model", "whisper-1"); + form.append("language", this.config.language); + if (prompt) form.append("prompt", prompt); + const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, { + method: "POST", + headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {}, + body: form, + signal: AbortSignal.timeout(DICTATION_REQUEST_TIMEOUT_MS), + }); + if (!response.ok) throw new Error(`Athena STT failed (HTTP ${response.status})`); + return String(record(await response.json()).text || "").trim(); + } } function pcmRms(pcm: Buffer): number { diff --git a/integrations/openclaw-athena-talk/package-lock.json b/integrations/openclaw-athena-talk/package-lock.json index 9930a78..20025a5 100644 --- a/integrations/openclaw-athena-talk/package-lock.json +++ b/integrations/openclaw-athena-talk/package-lock.json @@ -1,12 +1,12 @@ { "name": "@casaderoll/openclaw-athena-talk", - "version": "1.2.1", + "version": "1.3.0", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "@casaderoll/openclaw-athena-talk", - "version": "1.2.1", + "version": "1.3.0", "devDependencies": { "@types/node": "^24.0.0", "openclaw": "2026.9.4", diff --git a/integrations/openclaw-athena-talk/package.json b/integrations/openclaw-athena-talk/package.json index 59892a1..478d303 100644 --- a/integrations/openclaw-athena-talk/package.json +++ b/integrations/openclaw-athena-talk/package.json @@ -1,6 +1,6 @@ { "name": "@casaderoll/openclaw-athena-talk", - "version": "1.2.1", + "version": "1.3.0", "private": true, "description": "Local OpenClaw Talk provider backed by Athena Whisper and Qwen3-TTS", "type": "module", diff --git a/integrations/openclaw-athena-talk/test-transcription.mjs b/integrations/openclaw-athena-talk/test-transcription.mjs index bf5bce9..45968db 100644 --- a/integrations/openclaw-athena-talk/test-transcription.mjs +++ b/integrations/openclaw-athena-talk/test-transcription.mjs @@ -3,6 +3,14 @@ import { createServer } from "node:http"; import { test } from "node:test"; import plugin from "./dist/index.js"; +async function waitFor(predicate, timeoutMs = 2000) { + const deadline = Date.now() + timeoutMs; + while (!predicate()) { + if (Date.now() >= deadline) throw new Error("timed out waiting for condition"); + await new Promise((resolve) => setTimeout(resolve, 10)); + } +} + test("dictation registers separately and sends G.711 audio to Athena Whisper", async () => { let transcription; plugin.register({ @@ -79,3 +87,73 @@ test("dictation reuses only a TTS key for the same Athena origin", () => { cfg: withTts("http://other:8081/v1"), rawConfig: {}, }).apiKey, ""); }); + +test("long dictation is transcribed incrementally before the recording closes", async () => { + let transcription; + plugin.register({ + registerRealtimeTranscriptionProvider: (value) => { transcription = value; }, + registerRealtimeVoiceProvider: () => {}, + registerHttpRoute: () => {}, + }); + + const answers = [ + "Dies ist ein langer Abschnitt", + "langer Abschnitt mit einer Fortsetzung", + "einer Fortsetzung und einem Ende.", + ]; + let uploads = 0; + const server = createServer(async (req, res) => { + const index = uploads++; + const form = await new Request("http://localhost", { + method: "POST", + headers: { "Content-Type": req.headers["content-type"] }, + body: req, + duplex: "half", + }).formData(); + const wav = Buffer.from(await form.get("file").arrayBuffer()); + assert.equal(wav.readUInt32LE(24), 8000); + if (index > 0) assert.ok(String(form.get("prompt") || "").length > 0); + res.writeHead(200, { "Content-Type": "application/json" }) + .end(JSON.stringify({ text: answers[index] })); + }); + await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve)); + const cfg = { talk: { realtime: { providers: { "athena-talk": { + baseUrl: `http://127.0.0.1:${server.address().port}/v1`, language: "de", + } } } } }; + const providerConfig = transcription.resolveConfig({ cfg, rawConfig: {} }); + const transcripts = []; + const errors = []; + try { + const session = transcription.createSession({ + cfg, + providerConfig, + onTranscript: (text) => transcripts.push(text), + onError: (error) => errors.push(error), + }); + await session.connect(); + + // Six seconds start the first request while dictation is still active. + session.sendAudio(Buffer.alloc(48_000, 0xff)); + await waitFor(() => uploads === 1); + assert.deepEqual(transcripts, []); + + // Another 5.5 seconds form the next overlapping segment. The remaining + // 2 seconds are finalized only when the user stops dictation. + session.sendAudio(Buffer.alloc(44_000, 0xff)); + await waitFor(() => uploads === 2); + session.sendAudio(Buffer.alloc(16_000, 0xff)); + session.close(); + + await waitFor(() => transcripts.length === 3); + assert.deepEqual(transcripts, [ + "Dies ist ein langer Abschnitt", + "mit einer Fortsetzung", + "und einem Ende.", + ]); + assert.equal(uploads, 3); + assert.deepEqual(errors, []); + } finally { + server.closeAllConnections(); + await new Promise((resolve) => server.close(resolve)); + } +}); diff --git a/services/athena-realtime-voice/README.md b/services/athena-realtime-voice/README.md index 464970a..ddae30d 100644 --- a/services/athena-realtime-voice/README.md +++ b/services/athena-realtime-voice/README.md @@ -69,7 +69,7 @@ cd /opt/mike-ai/stack docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice ``` -OpenClaw 2026.9.4 uses the `athena-talk` plugin version 1.2.1 with +OpenClaw uses the `athena-talk` plugin version 1.3.0 with `talk.realtime.transport` set to `webrtc` and `talk.realtime.providers.athena-talk.realtimeUpstreamUrl` set to `http://192.168.1.212:8090/v1/realtime/calls`. On the Mac, turn off @@ -92,8 +92,12 @@ python3.11 -m venv .venv The synthetic microphone test passed over the actual OpenClaw HTTPS offer route and WireGuard media path with production Whisper and Qwen3-TTS on 2026-09-16. Subsequent real browser Talk sessions successfully transcribed -and answered multiple user turns. Browser dictation through the same plugin -also produced text after a correction to its separate transcription path. +and answered multiple user turns. Browser dictation uses a separate plugin +path. Since version 1.3.0, recordings longer than six seconds are +pre-transcribed incrementally in overlapping short windows while the +microphone remains active. This keeps the final tail inside OpenClaw's +five-second completion window. Talk through this WebRTC service still ends +and transcribes one utterance at a time. The first user report of a second turn becoming stuck was addressed by matching conversation item and predecessor IDs; the later two-turn test passed. This is still half-duplex and needs a private route for WebRTC media.