Stream long OpenClaw dictation in segments

This commit is contained in:
Mikei386
2026-09-21 16:06:15 +02:00
parent e577c55489
commit 46e5bbdf7f
10 changed files with 313 additions and 72 deletions
+5
View File
@@ -151,6 +151,11 @@ verbindet Mikrofon → Athena Whisper → normalen OpenClaw-Agenten → aktives
Athena-TTS, sodass Modell, Werkzeuge und Memory auch im Sprachmodus erhalten Athena-TTS, sodass Modell, Werkzeuge und Memory auch im Sprachmodus erhalten
bleiben. Die Installation landet in OpenClaws persistentem Datenverzeichnis bleiben. Die Installation landet in OpenClaws persistentem Datenverzeichnis
und bleibt deshalb bei normalen Container-Updates bestehen. und bleibt deshalb bei normalen Container-Updates bestehen.
Die separate Diktierfunktion verarbeitet seit Plugin-Version 1.3.0 längere
Aufnahmen bereits während des Sprechens in überlappenden Sechs-Sekunden-
Abschnitten. Beim Loslassen bleibt nur der kurze Rest für OpenClaws festes
Fünf-Sekunden-Abschlussfenster. Dafür wurden weder OpenClaw selbst verändert
noch ein weiterer Container angelegt.
OpenClaw wird über den Provider **llama.cpp → Existing llama-server** mit OpenClaw wird über den Provider **llama.cpp → Existing llama-server** mit
`http://192.168.1.212:8081/v1` verbunden. Der Router beantwortet sowohl `http://192.168.1.212:8081/v1` verbunden. Der Router beantwortet sowohl
+4 -1
View File
@@ -64,7 +64,10 @@ nicht mehr den aktuellen Containerzustand.
Er braucht keine GPU und keinen Profilwechsel. Das Plugin liegt in Er braucht keine GPU und keinen Profilwechsel. Das Plugin liegt in
`integrations/openclaw-athena-talk` und läuft auf Unraid, nicht auf Athena. `integrations/openclaw-athena-talk` und läuft auf Unraid, nicht auf Athena.
Diktat und hochgeladene M4A-Sprachnachrichten nutzen den separaten Diktat und hochgeladene M4A-Sprachnachrichten nutzen den separaten
Transkriptionspfad des Plugins. OpenClaw liefert den fertigen Agententext Transkriptionspfad des Plugins. Plugin 1.3.0 zerlegt längere Browser-Diktate
während der Aufnahme in überlappende Sechs-Sekunden-Abschnitte und hält so
den Abschluss innerhalb von OpenClaws festem Fünf-Sekunden-Fenster. OpenClaw
selbst wurde dafür nicht gepatcht. OpenClaw liefert den fertigen Agententext
an die Brücke; das Sprechen beginnt daher erst nach Abschluss der an die Brücke; das Sprechen beginnt daher erst nach Abschluss der
Agentenantwort. Details und Grenzen stehen in der [Voice-Doku](../services/athena-realtime-voice/README.md). Agentenantwort. Details und Grenzen stehen in der [Voice-Doku](../services/athena-realtime-voice/README.md).
- `mike-ai-mikes-applio-ui` läuft gesund und ohne GPU. Die Quelle liegt im - `mike-ai-mikes-applio-ui` läuft gesund und ohne GPU. Die Quelle liegt im
+6
View File
@@ -16,6 +16,12 @@ eingeschlossen: Im Archiv `athena-2026-09-16T13-02-56.tar.gz` wurden sowohl
`compose.yaml` als auch `services/athena-realtime-voice/server.py` geprüft. `compose.yaml` als auch `services/athena-realtime-voice/server.py` geprüft.
Das OpenClaw-Plugin auf Unraid liegt außerhalb dieses Athena-Backups; seine Das OpenClaw-Plugin auf Unraid liegt außerhalb dieses Athena-Backups; seine
Quelle ist im Git-Repository unter `integrations/openclaw-athena-talk` erfasst. Quelle ist im Git-Repository unter `integrations/openclaw-athena-talk` erfasst.
Die produktiv installierte Version 1.3.0 liegt zusätzlich im persistenten
OpenClaw-Appdata. Vor ihrer Installation wurde die bisherige Version als
`/mnt/nvme-storage/appdata/OpenClaw/config/plugin-backups/athena-talk-1.2.1-before-streaming.tar.gz`
gesichert. Für eine Neuinstallation ist der im Git dokumentierte Build mit
`openclaw plugins install <paket.tgz> --force --accept-capabilities` zu
installieren; eine Änderung an OpenClaw-Core-Dateien ist nicht erforderlich.
Piper-Daten sind kein aktueller Sicherungsbestand. Modellgewichte unter Piper-Daten sind kein aktueller Sicherungsbestand. Modellgewichte unter
`/data/models` und das reproduzierbare Whisper-Volume gehören nicht zu diesen `/data/models` und das reproduzierbare Whisper-Volume gehören nicht zu diesen
+21 -11
View File
@@ -60,19 +60,29 @@ openclaw plugins install . --force --accept-capabilities
openclaw plugins inspect athena-talk --runtime --json openclaw plugins inspect athena-talk --runtime --json
``` ```
Version 1.2.1 also registers **Athena Whisper (Diktieren)** as a separate Version 1.3.0 also registers **Athena Whisper (Diktieren)** as a separate
realtime transcription provider through OpenClaw's official plugin API. In the realtime transcription provider through OpenClaw's official plugin API. In the
browser composer, hold the microphone for dictation, then release it to send browser composer, hold the microphone for dictation, then release it to send
the 8 kHz G.711 audio through the Gateway. The plugin converts it to PCM WAV the 8 kHz G.711 audio through the Gateway. Short recordings are converted to
and calls the same Athena `/audio/transcriptions` endpoint used by Talk. The PCM WAV and sent to Athena's existing `/audio/transcriptions` endpoint in one
transcribed text is returned to the composer; this path does not invoke the request. Longer recordings are split while the user is still speaking into
agent or TTS. The transcription provider reuses `talk.realtime.providers.athena-talk` six-second windows with 0.5 seconds of overlap. The plugin sends these windows
and the configured model provider for its URL/key. If that model provider has sequentially to the persistent Whisper service, carries a short text prompt
no key, it reuses `tts.providers.openai.apiKey` only when the TTS and STT URLs into the next request, removes duplicated overlap words, and caches finished
have the same origin. No second credential is needed. In `talk.catalog`, it segments until recording stops. Only the short final tail then remains inside
appears under `transcription.providers`. OpenClaw OpenClaw's fixed five-second final-drain window. Each Whisper request is capped
currently gives a transcription provider five seconds to return its final text at 4.5 seconds.
after recording stops; the plugin caps its Whisper request at 4.5 seconds.
This is incremental pre-transcription over OpenClaw's official transcription
provider API. Whisper.cpp still receives complete short WAV segments; it is
not a native token-streaming STT protocol. No OpenClaw core file was patched
and no additional speech container was introduced. The transcribed text is
returned to the composer; this path does not invoke the agent or TTS. The
provider reuses `talk.realtime.providers.athena-talk` and the configured model
provider for its URL/key. If that model provider has no key, it reuses
`tts.providers.openai.apiKey` only when the TTS and STT URLs have the same
origin. No second credential is needed. In `talk.catalog`, it appears under
`transcription.providers`.
### Voice-note file attachments ### Voice-note file attachments
+95 -28
View File
@@ -4,6 +4,10 @@ const AUDIO_FORMAT = { encoding: "pcm16", sampleRateHz: 24000, channels: 1 };
const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls"; const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls";
const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key"; const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key";
const MAX_OFFER_BYTES = 64 * 1024; const MAX_OFFER_BYTES = 64 * 1024;
const DICTATION_SAMPLE_RATE_HZ = 8000;
const DICTATION_SEGMENT_BYTES = DICTATION_SAMPLE_RATE_HZ * 6;
const DICTATION_OVERLAP_BYTES = DICTATION_SAMPLE_RATE_HZ / 2;
const DICTATION_REQUEST_TIMEOUT_MS = 4500;
const browserKeys = generateKeyPairSync("ed25519"); const browserKeys = generateKeyPairSync("ed25519");
const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString(); const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString();
function base64url(value) { function base64url(value) {
@@ -149,13 +153,38 @@ function resolveTranscriptionConfig(cfg, rawConfig) {
const talkProvider = record(record(talkConfig.providers)["athena-talk"]); const talkProvider = record(record(talkConfig.providers)["athena-talk"]);
return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } }); return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } });
} }
function comparableWord(value) {
return value.toLocaleLowerCase("de-DE").replace(/[^\p{L}\p{N}]+/gu, "");
}
function removeTranscriptOverlap(previous, current) {
const priorWords = previous.trim().split(/\s+/).filter(Boolean);
const currentWords = current.trim().split(/\s+/).filter(Boolean);
const maximum = Math.min(12, priorWords.length, currentWords.length);
for (let count = maximum; count >= 1; count -= 1) {
const left = priorWords.slice(-count).map(comparableWord);
const right = currentWords.slice(0, count).map(comparableWord);
if (!left.every((word, index) => word && word === right[index]))
continue;
// A single short word is too ambiguous to remove safely. Longer words are
// sufficient because the audio overlap is only half a second.
if (count === 1 && left[0].length < 5)
continue;
return currentWords.slice(count).join(" ");
}
return currentWords.join(" ");
}
class AthenaTranscriptionSession { class AthenaTranscriptionSession {
req; req;
config; config;
connected = false; connected = false;
closed = false; closed = false;
audio = []; audio = [];
bytes = 0; bufferedBytes = 0;
totalBytes = 0;
processing = Promise.resolve();
completedTranscripts = [];
emittedTranscripts = 0;
processingError = null;
constructor(req, config) { constructor(req, config) {
this.req = req; this.req = req;
this.config = config; this.config = config;
@@ -165,50 +194,88 @@ class AthenaTranscriptionSession {
sendAudio(audio) { sendAudio(audio) {
if (!this.isConnected() || audio.length === 0) if (!this.isConnected() || audio.length === 0)
return; return;
if (this.bytes === 0) if (this.totalBytes === 0)
this.req.onSpeechStart?.(); this.req.onSpeechStart?.();
const maxBytes = this.config.maxSpeechSeconds * 8000; const maxBytes = this.config.maxSpeechSeconds * 8000;
if (this.bytes + audio.length > maxBytes) { if (this.totalBytes + audio.length > maxBytes) {
this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`)); this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`));
this.close(); this.close();
return; return;
} }
this.audio.push(Buffer.from(audio)); this.audio.push(Buffer.from(audio));
this.bytes += audio.length; this.bufferedBytes += audio.length;
this.totalBytes += audio.length;
while (this.bufferedBytes >= DICTATION_SEGMENT_BYTES) {
this.queueFullSegment();
}
} }
close() { close() {
if (this.closed) if (this.closed)
return; return;
this.closed = true; this.closed = true;
this.connected = false; this.connected = false;
if (!this.bytes) if (!this.totalBytes)
return; return;
const audio = Buffer.concat(this.audio); this.emitCompletedTranscripts();
this.audio.length = 0; const tail = Buffer.concat(this.audio);
void this.transcribe(audio); this.audio = [];
this.bufferedBytes = 0;
if (tail.length > DICTATION_OVERLAP_BYTES || this.completedTranscripts.length === 0) {
this.queueTranscription(tail);
}
void this.processing.finally(() => {
this.emitCompletedTranscripts();
if (this.processingError && this.completedTranscripts.length === 0) {
this.req.onError?.(this.processingError);
}
});
} }
async transcribe(audio) { queueFullSegment() {
try { const buffered = Buffer.concat(this.audio);
const form = new FormData(); const segment = Buffer.from(buffered.subarray(0, DICTATION_SEGMENT_BYTES));
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav"); const retained = Buffer.from(buffered.subarray(DICTATION_SEGMENT_BYTES - DICTATION_OVERLAP_BYTES));
form.append("model", "whisper-1"); this.audio = retained.length ? [retained] : [];
form.append("language", this.config.language); this.bufferedBytes = retained.length;
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, { this.queueTranscription(segment);
method: "POST", }
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {}, queueTranscription(audio) {
body: form, if (!audio.length)
signal: AbortSignal.timeout(4500), return;
}); this.processing = this.processing.then(async () => {
if (!response.ok) const previous = this.completedTranscripts.join(" ");
throw new Error(`Athena STT failed (HTTP ${response.status})`); const text = await this.transcribe(audio, previous.slice(-240));
const text = String(record(await response.json()).text || "").trim(); const novel = removeTranscriptOverlap(previous, text);
if (text) if (novel)
this.req.onTranscript?.(text); this.completedTranscripts.push(novel);
} if (this.closed)
catch (error) { this.emitCompletedTranscripts();
this.req.onError?.(error instanceof Error ? error : new Error(String(error))); }).catch((error) => {
this.processingError = error instanceof Error ? error : new Error(String(error));
});
}
emitCompletedTranscripts() {
while (this.emittedTranscripts < this.completedTranscripts.length) {
this.req.onTranscript?.(this.completedTranscripts[this.emittedTranscripts]);
this.emittedTranscripts += 1;
} }
} }
async transcribe(audio, prompt) {
const form = new FormData();
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
form.append("model", "whisper-1");
form.append("language", this.config.language);
if (prompt)
form.append("prompt", prompt);
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
method: "POST",
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
body: form,
signal: AbortSignal.timeout(DICTATION_REQUEST_TIMEOUT_MS),
});
if (!response.ok)
throw new Error(`Athena STT failed (HTTP ${response.status})`);
return String(record(await response.json()).text || "").trim();
}
} }
function pcmRms(pcm) { function pcmRms(pcm) {
if (pcm.length < 2) if (pcm.length < 2)
+94 -26
View File
@@ -20,6 +20,10 @@ type ProviderConfig = {
const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls"; const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls";
const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key"; const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key";
const MAX_OFFER_BYTES = 64 * 1024; const MAX_OFFER_BYTES = 64 * 1024;
const DICTATION_SAMPLE_RATE_HZ = 8000;
const DICTATION_SEGMENT_BYTES = DICTATION_SAMPLE_RATE_HZ * 6;
const DICTATION_OVERLAP_BYTES = DICTATION_SAMPLE_RATE_HZ / 2;
const DICTATION_REQUEST_TIMEOUT_MS = 4500;
const browserKeys = generateKeyPairSync("ed25519"); const browserKeys = generateKeyPairSync("ed25519");
const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString(); const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString();
@@ -171,11 +175,36 @@ function resolveTranscriptionConfig(cfg: unknown, rawConfig: unknown): Required<
return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } }); return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } });
} }
function comparableWord(value: string): string {
return value.toLocaleLowerCase("de-DE").replace(/[^\p{L}\p{N}]+/gu, "");
}
function removeTranscriptOverlap(previous: string, current: string): string {
const priorWords = previous.trim().split(/\s+/).filter(Boolean);
const currentWords = current.trim().split(/\s+/).filter(Boolean);
const maximum = Math.min(12, priorWords.length, currentWords.length);
for (let count = maximum; count >= 1; count -= 1) {
const left = priorWords.slice(-count).map(comparableWord);
const right = currentWords.slice(0, count).map(comparableWord);
if (!left.every((word, index) => word && word === right[index])) continue;
// A single short word is too ambiguous to remove safely. Longer words are
// sufficient because the audio overlap is only half a second.
if (count === 1 && left[0].length < 5) continue;
return currentWords.slice(count).join(" ");
}
return currentWords.join(" ");
}
class AthenaTranscriptionSession { class AthenaTranscriptionSession {
private connected = false; private connected = false;
private closed = false; private closed = false;
private readonly audio: Buffer[] = []; private audio: Buffer[] = [];
private bytes = 0; private bufferedBytes = 0;
private totalBytes = 0;
private processing: Promise<void> = Promise.resolve();
private readonly completedTranscripts: string[] = [];
private emittedTranscripts = 0;
private processingError: Error | null = null;
constructor( constructor(
private readonly req: { private readonly req: {
@@ -193,46 +222,85 @@ class AthenaTranscriptionSession {
sendAudio(audio: Buffer): void { sendAudio(audio: Buffer): void {
if (!this.isConnected() || audio.length === 0) return; if (!this.isConnected() || audio.length === 0) return;
if (this.bytes === 0) this.req.onSpeechStart?.(); if (this.totalBytes === 0) this.req.onSpeechStart?.();
const maxBytes = this.config.maxSpeechSeconds * 8000; const maxBytes = this.config.maxSpeechSeconds * 8000;
if (this.bytes + audio.length > maxBytes) { if (this.totalBytes + audio.length > maxBytes) {
this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`)); this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`));
this.close(); this.close();
return; return;
} }
this.audio.push(Buffer.from(audio)); this.audio.push(Buffer.from(audio));
this.bytes += audio.length; this.bufferedBytes += audio.length;
this.totalBytes += audio.length;
while (this.bufferedBytes >= DICTATION_SEGMENT_BYTES) {
this.queueFullSegment();
}
} }
close(): void { close(): void {
if (this.closed) return; if (this.closed) return;
this.closed = true; this.closed = true;
this.connected = false; this.connected = false;
if (!this.bytes) return; if (!this.totalBytes) return;
const audio = Buffer.concat(this.audio); this.emitCompletedTranscripts();
this.audio.length = 0; const tail = Buffer.concat(this.audio);
void this.transcribe(audio); this.audio = [];
this.bufferedBytes = 0;
if (tail.length > DICTATION_OVERLAP_BYTES || this.completedTranscripts.length === 0) {
this.queueTranscription(tail);
}
void this.processing.finally(() => {
this.emitCompletedTranscripts();
if (this.processingError && this.completedTranscripts.length === 0) {
this.req.onError?.(this.processingError);
}
});
} }
private async transcribe(audio: Buffer): Promise<void> { private queueFullSegment(): void {
try { const buffered = Buffer.concat(this.audio);
const form = new FormData(); const segment = Buffer.from(buffered.subarray(0, DICTATION_SEGMENT_BYTES));
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav"); const retained = Buffer.from(buffered.subarray(DICTATION_SEGMENT_BYTES - DICTATION_OVERLAP_BYTES));
form.append("model", "whisper-1"); this.audio = retained.length ? [retained] : [];
form.append("language", this.config.language); this.bufferedBytes = retained.length;
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, { this.queueTranscription(segment);
method: "POST", }
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
body: form, private queueTranscription(audio: Buffer): void {
signal: AbortSignal.timeout(4500), if (!audio.length) return;
}); this.processing = this.processing.then(async () => {
if (!response.ok) throw new Error(`Athena STT failed (HTTP ${response.status})`); const previous = this.completedTranscripts.join(" ");
const text = String(record(await response.json()).text || "").trim(); const text = await this.transcribe(audio, previous.slice(-240));
if (text) this.req.onTranscript?.(text); const novel = removeTranscriptOverlap(previous, text);
} catch (error) { if (novel) this.completedTranscripts.push(novel);
this.req.onError?.(error instanceof Error ? error : new Error(String(error))); if (this.closed) this.emitCompletedTranscripts();
}).catch((error: unknown) => {
this.processingError = error instanceof Error ? error : new Error(String(error));
});
}
private emitCompletedTranscripts(): void {
while (this.emittedTranscripts < this.completedTranscripts.length) {
this.req.onTranscript?.(this.completedTranscripts[this.emittedTranscripts]);
this.emittedTranscripts += 1;
} }
} }
private async transcribe(audio: Buffer, prompt: string): Promise<string> {
const form = new FormData();
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
form.append("model", "whisper-1");
form.append("language", this.config.language);
if (prompt) form.append("prompt", prompt);
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
method: "POST",
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
body: form,
signal: AbortSignal.timeout(DICTATION_REQUEST_TIMEOUT_MS),
});
if (!response.ok) throw new Error(`Athena STT failed (HTTP ${response.status})`);
return String(record(await response.json()).text || "").trim();
}
} }
function pcmRms(pcm: Buffer): number { function pcmRms(pcm: Buffer): number {
+2 -2
View File
@@ -1,12 +1,12 @@
{ {
"name": "@casaderoll/openclaw-athena-talk", "name": "@casaderoll/openclaw-athena-talk",
"version": "1.2.1", "version": "1.3.0",
"lockfileVersion": 3, "lockfileVersion": 3,
"requires": true, "requires": true,
"packages": { "packages": {
"": { "": {
"name": "@casaderoll/openclaw-athena-talk", "name": "@casaderoll/openclaw-athena-talk",
"version": "1.2.1", "version": "1.3.0",
"devDependencies": { "devDependencies": {
"@types/node": "^24.0.0", "@types/node": "^24.0.0",
"openclaw": "2026.9.4", "openclaw": "2026.9.4",
@@ -1,6 +1,6 @@
{ {
"name": "@casaderoll/openclaw-athena-talk", "name": "@casaderoll/openclaw-athena-talk",
"version": "1.2.1", "version": "1.3.0",
"private": true, "private": true,
"description": "Local OpenClaw Talk provider backed by Athena Whisper and Qwen3-TTS", "description": "Local OpenClaw Talk provider backed by Athena Whisper and Qwen3-TTS",
"type": "module", "type": "module",
@@ -3,6 +3,14 @@ import { createServer } from "node:http";
import { test } from "node:test"; import { test } from "node:test";
import plugin from "./dist/index.js"; import plugin from "./dist/index.js";
async function waitFor(predicate, timeoutMs = 2000) {
const deadline = Date.now() + timeoutMs;
while (!predicate()) {
if (Date.now() >= deadline) throw new Error("timed out waiting for condition");
await new Promise((resolve) => setTimeout(resolve, 10));
}
}
test("dictation registers separately and sends G.711 audio to Athena Whisper", async () => { test("dictation registers separately and sends G.711 audio to Athena Whisper", async () => {
let transcription; let transcription;
plugin.register({ plugin.register({
@@ -79,3 +87,73 @@ test("dictation reuses only a TTS key for the same Athena origin", () => {
cfg: withTts("http://other:8081/v1"), rawConfig: {}, cfg: withTts("http://other:8081/v1"), rawConfig: {},
}).apiKey, ""); }).apiKey, "");
}); });
test("long dictation is transcribed incrementally before the recording closes", async () => {
let transcription;
plugin.register({
registerRealtimeTranscriptionProvider: (value) => { transcription = value; },
registerRealtimeVoiceProvider: () => {},
registerHttpRoute: () => {},
});
const answers = [
"Dies ist ein langer Abschnitt",
"langer Abschnitt mit einer Fortsetzung",
"einer Fortsetzung und einem Ende.",
];
let uploads = 0;
const server = createServer(async (req, res) => {
const index = uploads++;
const form = await new Request("http://localhost", {
method: "POST",
headers: { "Content-Type": req.headers["content-type"] },
body: req,
duplex: "half",
}).formData();
const wav = Buffer.from(await form.get("file").arrayBuffer());
assert.equal(wav.readUInt32LE(24), 8000);
if (index > 0) assert.ok(String(form.get("prompt") || "").length > 0);
res.writeHead(200, { "Content-Type": "application/json" })
.end(JSON.stringify({ text: answers[index] }));
});
await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve));
const cfg = { talk: { realtime: { providers: { "athena-talk": {
baseUrl: `http://127.0.0.1:${server.address().port}/v1`, language: "de",
} } } } };
const providerConfig = transcription.resolveConfig({ cfg, rawConfig: {} });
const transcripts = [];
const errors = [];
try {
const session = transcription.createSession({
cfg,
providerConfig,
onTranscript: (text) => transcripts.push(text),
onError: (error) => errors.push(error),
});
await session.connect();
// Six seconds start the first request while dictation is still active.
session.sendAudio(Buffer.alloc(48_000, 0xff));
await waitFor(() => uploads === 1);
assert.deepEqual(transcripts, []);
// Another 5.5 seconds form the next overlapping segment. The remaining
// 2 seconds are finalized only when the user stops dictation.
session.sendAudio(Buffer.alloc(44_000, 0xff));
await waitFor(() => uploads === 2);
session.sendAudio(Buffer.alloc(16_000, 0xff));
session.close();
await waitFor(() => transcripts.length === 3);
assert.deepEqual(transcripts, [
"Dies ist ein langer Abschnitt",
"mit einer Fortsetzung",
"und einem Ende.",
]);
assert.equal(uploads, 3);
assert.deepEqual(errors, []);
} finally {
server.closeAllConnections();
await new Promise((resolve) => server.close(resolve));
}
});
+7 -3
View File
@@ -69,7 +69,7 @@ cd /opt/mike-ai/stack
docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice
``` ```
OpenClaw 2026.9.4 uses the `athena-talk` plugin version 1.2.1 with OpenClaw uses the `athena-talk` plugin version 1.3.0 with
`talk.realtime.transport` set to `webrtc` and `talk.realtime.transport` set to `webrtc` and
`talk.realtime.providers.athena-talk.realtimeUpstreamUrl` set to `talk.realtime.providers.athena-talk.realtimeUpstreamUrl` set to
`http://192.168.1.212:8090/v1/realtime/calls`. On the Mac, turn off `http://192.168.1.212:8090/v1/realtime/calls`. On the Mac, turn off
@@ -92,8 +92,12 @@ python3.11 -m venv .venv
The synthetic microphone test passed over the actual OpenClaw HTTPS offer The synthetic microphone test passed over the actual OpenClaw HTTPS offer
route and WireGuard media path with production Whisper and Qwen3-TTS on route and WireGuard media path with production Whisper and Qwen3-TTS on
2026-09-16. Subsequent real browser Talk sessions successfully transcribed 2026-09-16. Subsequent real browser Talk sessions successfully transcribed
and answered multiple user turns. Browser dictation through the same plugin and answered multiple user turns. Browser dictation uses a separate plugin
also produced text after a correction to its separate transcription path. path. Since version 1.3.0, recordings longer than six seconds are
pre-transcribed incrementally in overlapping short windows while the
microphone remains active. This keeps the final tail inside OpenClaw's
five-second completion window. Talk through this WebRTC service still ends
and transcribes one utterance at a time.
The first user report of a second turn becoming stuck was addressed by The first user report of a second turn becoming stuck was addressed by
matching conversation item and predecessor IDs; the later two-turn test matching conversation item and predecessor IDs; the later two-turn test
passed. This is still half-duplex and needs a private route for WebRTC media. passed. This is still half-duplex and needs a private route for WebRTC media.