Stream long OpenClaw dictation in segments
This commit is contained in:
@@ -151,6 +151,11 @@ verbindet Mikrofon → Athena Whisper → normalen OpenClaw-Agenten → aktives
|
|||||||
Athena-TTS, sodass Modell, Werkzeuge und Memory auch im Sprachmodus erhalten
|
Athena-TTS, sodass Modell, Werkzeuge und Memory auch im Sprachmodus erhalten
|
||||||
bleiben. Die Installation landet in OpenClaws persistentem Datenverzeichnis
|
bleiben. Die Installation landet in OpenClaws persistentem Datenverzeichnis
|
||||||
und bleibt deshalb bei normalen Container-Updates bestehen.
|
und bleibt deshalb bei normalen Container-Updates bestehen.
|
||||||
|
Die separate Diktierfunktion verarbeitet seit Plugin-Version 1.3.0 längere
|
||||||
|
Aufnahmen bereits während des Sprechens in überlappenden Sechs-Sekunden-
|
||||||
|
Abschnitten. Beim Loslassen bleibt nur der kurze Rest für OpenClaws festes
|
||||||
|
Fünf-Sekunden-Abschlussfenster. Dafür wurden weder OpenClaw selbst verändert
|
||||||
|
noch ein weiterer Container angelegt.
|
||||||
|
|
||||||
OpenClaw wird über den Provider **llama.cpp → Existing llama-server** mit
|
OpenClaw wird über den Provider **llama.cpp → Existing llama-server** mit
|
||||||
`http://192.168.1.212:8081/v1` verbunden. Der Router beantwortet sowohl
|
`http://192.168.1.212:8081/v1` verbunden. Der Router beantwortet sowohl
|
||||||
|
|||||||
+4
-1
@@ -64,7 +64,10 @@ nicht mehr den aktuellen Containerzustand.
|
|||||||
Er braucht keine GPU und keinen Profilwechsel. Das Plugin liegt in
|
Er braucht keine GPU und keinen Profilwechsel. Das Plugin liegt in
|
||||||
`integrations/openclaw-athena-talk` und läuft auf Unraid, nicht auf Athena.
|
`integrations/openclaw-athena-talk` und läuft auf Unraid, nicht auf Athena.
|
||||||
Diktat und hochgeladene M4A-Sprachnachrichten nutzen den separaten
|
Diktat und hochgeladene M4A-Sprachnachrichten nutzen den separaten
|
||||||
Transkriptionspfad des Plugins. OpenClaw liefert den fertigen Agententext
|
Transkriptionspfad des Plugins. Plugin 1.3.0 zerlegt längere Browser-Diktate
|
||||||
|
während der Aufnahme in überlappende Sechs-Sekunden-Abschnitte und hält so
|
||||||
|
den Abschluss innerhalb von OpenClaws festem Fünf-Sekunden-Fenster. OpenClaw
|
||||||
|
selbst wurde dafür nicht gepatcht. OpenClaw liefert den fertigen Agententext
|
||||||
an die Brücke; das Sprechen beginnt daher erst nach Abschluss der
|
an die Brücke; das Sprechen beginnt daher erst nach Abschluss der
|
||||||
Agentenantwort. Details und Grenzen stehen in der [Voice-Doku](../services/athena-realtime-voice/README.md).
|
Agentenantwort. Details und Grenzen stehen in der [Voice-Doku](../services/athena-realtime-voice/README.md).
|
||||||
- `mike-ai-mikes-applio-ui` läuft gesund und ohne GPU. Die Quelle liegt im
|
- `mike-ai-mikes-applio-ui` läuft gesund und ohne GPU. Die Quelle liegt im
|
||||||
|
|||||||
@@ -16,6 +16,12 @@ eingeschlossen: Im Archiv `athena-2026-09-16T13-02-56.tar.gz` wurden sowohl
|
|||||||
`compose.yaml` als auch `services/athena-realtime-voice/server.py` geprüft.
|
`compose.yaml` als auch `services/athena-realtime-voice/server.py` geprüft.
|
||||||
Das OpenClaw-Plugin auf Unraid liegt außerhalb dieses Athena-Backups; seine
|
Das OpenClaw-Plugin auf Unraid liegt außerhalb dieses Athena-Backups; seine
|
||||||
Quelle ist im Git-Repository unter `integrations/openclaw-athena-talk` erfasst.
|
Quelle ist im Git-Repository unter `integrations/openclaw-athena-talk` erfasst.
|
||||||
|
Die produktiv installierte Version 1.3.0 liegt zusätzlich im persistenten
|
||||||
|
OpenClaw-Appdata. Vor ihrer Installation wurde die bisherige Version als
|
||||||
|
`/mnt/nvme-storage/appdata/OpenClaw/config/plugin-backups/athena-talk-1.2.1-before-streaming.tar.gz`
|
||||||
|
gesichert. Für eine Neuinstallation ist der im Git dokumentierte Build mit
|
||||||
|
`openclaw plugins install <paket.tgz> --force --accept-capabilities` zu
|
||||||
|
installieren; eine Änderung an OpenClaw-Core-Dateien ist nicht erforderlich.
|
||||||
|
|
||||||
Piper-Daten sind kein aktueller Sicherungsbestand. Modellgewichte unter
|
Piper-Daten sind kein aktueller Sicherungsbestand. Modellgewichte unter
|
||||||
`/data/models` und das reproduzierbare Whisper-Volume gehören nicht zu diesen
|
`/data/models` und das reproduzierbare Whisper-Volume gehören nicht zu diesen
|
||||||
|
|||||||
@@ -60,19 +60,29 @@ openclaw plugins install . --force --accept-capabilities
|
|||||||
openclaw plugins inspect athena-talk --runtime --json
|
openclaw plugins inspect athena-talk --runtime --json
|
||||||
```
|
```
|
||||||
|
|
||||||
Version 1.2.1 also registers **Athena Whisper (Diktieren)** as a separate
|
Version 1.3.0 also registers **Athena Whisper (Diktieren)** as a separate
|
||||||
realtime transcription provider through OpenClaw's official plugin API. In the
|
realtime transcription provider through OpenClaw's official plugin API. In the
|
||||||
browser composer, hold the microphone for dictation, then release it to send
|
browser composer, hold the microphone for dictation, then release it to send
|
||||||
the 8 kHz G.711 audio through the Gateway. The plugin converts it to PCM WAV
|
the 8 kHz G.711 audio through the Gateway. Short recordings are converted to
|
||||||
and calls the same Athena `/audio/transcriptions` endpoint used by Talk. The
|
PCM WAV and sent to Athena's existing `/audio/transcriptions` endpoint in one
|
||||||
transcribed text is returned to the composer; this path does not invoke the
|
request. Longer recordings are split while the user is still speaking into
|
||||||
agent or TTS. The transcription provider reuses `talk.realtime.providers.athena-talk`
|
six-second windows with 0.5 seconds of overlap. The plugin sends these windows
|
||||||
and the configured model provider for its URL/key. If that model provider has
|
sequentially to the persistent Whisper service, carries a short text prompt
|
||||||
no key, it reuses `tts.providers.openai.apiKey` only when the TTS and STT URLs
|
into the next request, removes duplicated overlap words, and caches finished
|
||||||
have the same origin. No second credential is needed. In `talk.catalog`, it
|
segments until recording stops. Only the short final tail then remains inside
|
||||||
appears under `transcription.providers`. OpenClaw
|
OpenClaw's fixed five-second final-drain window. Each Whisper request is capped
|
||||||
currently gives a transcription provider five seconds to return its final text
|
at 4.5 seconds.
|
||||||
after recording stops; the plugin caps its Whisper request at 4.5 seconds.
|
|
||||||
|
This is incremental pre-transcription over OpenClaw's official transcription
|
||||||
|
provider API. Whisper.cpp still receives complete short WAV segments; it is
|
||||||
|
not a native token-streaming STT protocol. No OpenClaw core file was patched
|
||||||
|
and no additional speech container was introduced. The transcribed text is
|
||||||
|
returned to the composer; this path does not invoke the agent or TTS. The
|
||||||
|
provider reuses `talk.realtime.providers.athena-talk` and the configured model
|
||||||
|
provider for its URL/key. If that model provider has no key, it reuses
|
||||||
|
`tts.providers.openai.apiKey` only when the TTS and STT URLs have the same
|
||||||
|
origin. No second credential is needed. In `talk.catalog`, it appears under
|
||||||
|
`transcription.providers`.
|
||||||
|
|
||||||
### Voice-note file attachments
|
### Voice-note file attachments
|
||||||
|
|
||||||
|
|||||||
+85
-18
@@ -4,6 +4,10 @@ const AUDIO_FORMAT = { encoding: "pcm16", sampleRateHz: 24000, channels: 1 };
|
|||||||
const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls";
|
const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls";
|
||||||
const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key";
|
const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key";
|
||||||
const MAX_OFFER_BYTES = 64 * 1024;
|
const MAX_OFFER_BYTES = 64 * 1024;
|
||||||
|
const DICTATION_SAMPLE_RATE_HZ = 8000;
|
||||||
|
const DICTATION_SEGMENT_BYTES = DICTATION_SAMPLE_RATE_HZ * 6;
|
||||||
|
const DICTATION_OVERLAP_BYTES = DICTATION_SAMPLE_RATE_HZ / 2;
|
||||||
|
const DICTATION_REQUEST_TIMEOUT_MS = 4500;
|
||||||
const browserKeys = generateKeyPairSync("ed25519");
|
const browserKeys = generateKeyPairSync("ed25519");
|
||||||
const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString();
|
const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString();
|
||||||
function base64url(value) {
|
function base64url(value) {
|
||||||
@@ -149,13 +153,38 @@ function resolveTranscriptionConfig(cfg, rawConfig) {
|
|||||||
const talkProvider = record(record(talkConfig.providers)["athena-talk"]);
|
const talkProvider = record(record(talkConfig.providers)["athena-talk"]);
|
||||||
return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } });
|
return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } });
|
||||||
}
|
}
|
||||||
|
function comparableWord(value) {
|
||||||
|
return value.toLocaleLowerCase("de-DE").replace(/[^\p{L}\p{N}]+/gu, "");
|
||||||
|
}
|
||||||
|
function removeTranscriptOverlap(previous, current) {
|
||||||
|
const priorWords = previous.trim().split(/\s+/).filter(Boolean);
|
||||||
|
const currentWords = current.trim().split(/\s+/).filter(Boolean);
|
||||||
|
const maximum = Math.min(12, priorWords.length, currentWords.length);
|
||||||
|
for (let count = maximum; count >= 1; count -= 1) {
|
||||||
|
const left = priorWords.slice(-count).map(comparableWord);
|
||||||
|
const right = currentWords.slice(0, count).map(comparableWord);
|
||||||
|
if (!left.every((word, index) => word && word === right[index]))
|
||||||
|
continue;
|
||||||
|
// A single short word is too ambiguous to remove safely. Longer words are
|
||||||
|
// sufficient because the audio overlap is only half a second.
|
||||||
|
if (count === 1 && left[0].length < 5)
|
||||||
|
continue;
|
||||||
|
return currentWords.slice(count).join(" ");
|
||||||
|
}
|
||||||
|
return currentWords.join(" ");
|
||||||
|
}
|
||||||
class AthenaTranscriptionSession {
|
class AthenaTranscriptionSession {
|
||||||
req;
|
req;
|
||||||
config;
|
config;
|
||||||
connected = false;
|
connected = false;
|
||||||
closed = false;
|
closed = false;
|
||||||
audio = [];
|
audio = [];
|
||||||
bytes = 0;
|
bufferedBytes = 0;
|
||||||
|
totalBytes = 0;
|
||||||
|
processing = Promise.resolve();
|
||||||
|
completedTranscripts = [];
|
||||||
|
emittedTranscripts = 0;
|
||||||
|
processingError = null;
|
||||||
constructor(req, config) {
|
constructor(req, config) {
|
||||||
this.req = req;
|
this.req = req;
|
||||||
this.config = config;
|
this.config = config;
|
||||||
@@ -165,49 +194,87 @@ class AthenaTranscriptionSession {
|
|||||||
sendAudio(audio) {
|
sendAudio(audio) {
|
||||||
if (!this.isConnected() || audio.length === 0)
|
if (!this.isConnected() || audio.length === 0)
|
||||||
return;
|
return;
|
||||||
if (this.bytes === 0)
|
if (this.totalBytes === 0)
|
||||||
this.req.onSpeechStart?.();
|
this.req.onSpeechStart?.();
|
||||||
const maxBytes = this.config.maxSpeechSeconds * 8000;
|
const maxBytes = this.config.maxSpeechSeconds * 8000;
|
||||||
if (this.bytes + audio.length > maxBytes) {
|
if (this.totalBytes + audio.length > maxBytes) {
|
||||||
this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`));
|
this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`));
|
||||||
this.close();
|
this.close();
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
this.audio.push(Buffer.from(audio));
|
this.audio.push(Buffer.from(audio));
|
||||||
this.bytes += audio.length;
|
this.bufferedBytes += audio.length;
|
||||||
|
this.totalBytes += audio.length;
|
||||||
|
while (this.bufferedBytes >= DICTATION_SEGMENT_BYTES) {
|
||||||
|
this.queueFullSegment();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
close() {
|
close() {
|
||||||
if (this.closed)
|
if (this.closed)
|
||||||
return;
|
return;
|
||||||
this.closed = true;
|
this.closed = true;
|
||||||
this.connected = false;
|
this.connected = false;
|
||||||
if (!this.bytes)
|
if (!this.totalBytes)
|
||||||
return;
|
return;
|
||||||
const audio = Buffer.concat(this.audio);
|
this.emitCompletedTranscripts();
|
||||||
this.audio.length = 0;
|
const tail = Buffer.concat(this.audio);
|
||||||
void this.transcribe(audio);
|
this.audio = [];
|
||||||
|
this.bufferedBytes = 0;
|
||||||
|
if (tail.length > DICTATION_OVERLAP_BYTES || this.completedTranscripts.length === 0) {
|
||||||
|
this.queueTranscription(tail);
|
||||||
}
|
}
|
||||||
async transcribe(audio) {
|
void this.processing.finally(() => {
|
||||||
try {
|
this.emitCompletedTranscripts();
|
||||||
|
if (this.processingError && this.completedTranscripts.length === 0) {
|
||||||
|
this.req.onError?.(this.processingError);
|
||||||
|
}
|
||||||
|
});
|
||||||
|
}
|
||||||
|
queueFullSegment() {
|
||||||
|
const buffered = Buffer.concat(this.audio);
|
||||||
|
const segment = Buffer.from(buffered.subarray(0, DICTATION_SEGMENT_BYTES));
|
||||||
|
const retained = Buffer.from(buffered.subarray(DICTATION_SEGMENT_BYTES - DICTATION_OVERLAP_BYTES));
|
||||||
|
this.audio = retained.length ? [retained] : [];
|
||||||
|
this.bufferedBytes = retained.length;
|
||||||
|
this.queueTranscription(segment);
|
||||||
|
}
|
||||||
|
queueTranscription(audio) {
|
||||||
|
if (!audio.length)
|
||||||
|
return;
|
||||||
|
this.processing = this.processing.then(async () => {
|
||||||
|
const previous = this.completedTranscripts.join(" ");
|
||||||
|
const text = await this.transcribe(audio, previous.slice(-240));
|
||||||
|
const novel = removeTranscriptOverlap(previous, text);
|
||||||
|
if (novel)
|
||||||
|
this.completedTranscripts.push(novel);
|
||||||
|
if (this.closed)
|
||||||
|
this.emitCompletedTranscripts();
|
||||||
|
}).catch((error) => {
|
||||||
|
this.processingError = error instanceof Error ? error : new Error(String(error));
|
||||||
|
});
|
||||||
|
}
|
||||||
|
emitCompletedTranscripts() {
|
||||||
|
while (this.emittedTranscripts < this.completedTranscripts.length) {
|
||||||
|
this.req.onTranscript?.(this.completedTranscripts[this.emittedTranscripts]);
|
||||||
|
this.emittedTranscripts += 1;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
async transcribe(audio, prompt) {
|
||||||
const form = new FormData();
|
const form = new FormData();
|
||||||
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
||||||
form.append("model", "whisper-1");
|
form.append("model", "whisper-1");
|
||||||
form.append("language", this.config.language);
|
form.append("language", this.config.language);
|
||||||
|
if (prompt)
|
||||||
|
form.append("prompt", prompt);
|
||||||
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
|
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
|
||||||
method: "POST",
|
method: "POST",
|
||||||
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
|
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
|
||||||
body: form,
|
body: form,
|
||||||
signal: AbortSignal.timeout(4500),
|
signal: AbortSignal.timeout(DICTATION_REQUEST_TIMEOUT_MS),
|
||||||
});
|
});
|
||||||
if (!response.ok)
|
if (!response.ok)
|
||||||
throw new Error(`Athena STT failed (HTTP ${response.status})`);
|
throw new Error(`Athena STT failed (HTTP ${response.status})`);
|
||||||
const text = String(record(await response.json()).text || "").trim();
|
return String(record(await response.json()).text || "").trim();
|
||||||
if (text)
|
|
||||||
this.req.onTranscript?.(text);
|
|
||||||
}
|
|
||||||
catch (error) {
|
|
||||||
this.req.onError?.(error instanceof Error ? error : new Error(String(error)));
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
function pcmRms(pcm) {
|
function pcmRms(pcm) {
|
||||||
|
|||||||
@@ -20,6 +20,10 @@ type ProviderConfig = {
|
|||||||
const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls";
|
const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls";
|
||||||
const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key";
|
const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key";
|
||||||
const MAX_OFFER_BYTES = 64 * 1024;
|
const MAX_OFFER_BYTES = 64 * 1024;
|
||||||
|
const DICTATION_SAMPLE_RATE_HZ = 8000;
|
||||||
|
const DICTATION_SEGMENT_BYTES = DICTATION_SAMPLE_RATE_HZ * 6;
|
||||||
|
const DICTATION_OVERLAP_BYTES = DICTATION_SAMPLE_RATE_HZ / 2;
|
||||||
|
const DICTATION_REQUEST_TIMEOUT_MS = 4500;
|
||||||
const browserKeys = generateKeyPairSync("ed25519");
|
const browserKeys = generateKeyPairSync("ed25519");
|
||||||
const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString();
|
const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString();
|
||||||
|
|
||||||
@@ -171,11 +175,36 @@ function resolveTranscriptionConfig(cfg: unknown, rawConfig: unknown): Required<
|
|||||||
return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } });
|
return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } });
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function comparableWord(value: string): string {
|
||||||
|
return value.toLocaleLowerCase("de-DE").replace(/[^\p{L}\p{N}]+/gu, "");
|
||||||
|
}
|
||||||
|
|
||||||
|
function removeTranscriptOverlap(previous: string, current: string): string {
|
||||||
|
const priorWords = previous.trim().split(/\s+/).filter(Boolean);
|
||||||
|
const currentWords = current.trim().split(/\s+/).filter(Boolean);
|
||||||
|
const maximum = Math.min(12, priorWords.length, currentWords.length);
|
||||||
|
for (let count = maximum; count >= 1; count -= 1) {
|
||||||
|
const left = priorWords.slice(-count).map(comparableWord);
|
||||||
|
const right = currentWords.slice(0, count).map(comparableWord);
|
||||||
|
if (!left.every((word, index) => word && word === right[index])) continue;
|
||||||
|
// A single short word is too ambiguous to remove safely. Longer words are
|
||||||
|
// sufficient because the audio overlap is only half a second.
|
||||||
|
if (count === 1 && left[0].length < 5) continue;
|
||||||
|
return currentWords.slice(count).join(" ");
|
||||||
|
}
|
||||||
|
return currentWords.join(" ");
|
||||||
|
}
|
||||||
|
|
||||||
class AthenaTranscriptionSession {
|
class AthenaTranscriptionSession {
|
||||||
private connected = false;
|
private connected = false;
|
||||||
private closed = false;
|
private closed = false;
|
||||||
private readonly audio: Buffer[] = [];
|
private audio: Buffer[] = [];
|
||||||
private bytes = 0;
|
private bufferedBytes = 0;
|
||||||
|
private totalBytes = 0;
|
||||||
|
private processing: Promise<void> = Promise.resolve();
|
||||||
|
private readonly completedTranscripts: string[] = [];
|
||||||
|
private emittedTranscripts = 0;
|
||||||
|
private processingError: Error | null = null;
|
||||||
|
|
||||||
constructor(
|
constructor(
|
||||||
private readonly req: {
|
private readonly req: {
|
||||||
@@ -193,45 +222,84 @@ class AthenaTranscriptionSession {
|
|||||||
|
|
||||||
sendAudio(audio: Buffer): void {
|
sendAudio(audio: Buffer): void {
|
||||||
if (!this.isConnected() || audio.length === 0) return;
|
if (!this.isConnected() || audio.length === 0) return;
|
||||||
if (this.bytes === 0) this.req.onSpeechStart?.();
|
if (this.totalBytes === 0) this.req.onSpeechStart?.();
|
||||||
const maxBytes = this.config.maxSpeechSeconds * 8000;
|
const maxBytes = this.config.maxSpeechSeconds * 8000;
|
||||||
if (this.bytes + audio.length > maxBytes) {
|
if (this.totalBytes + audio.length > maxBytes) {
|
||||||
this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`));
|
this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`));
|
||||||
this.close();
|
this.close();
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
this.audio.push(Buffer.from(audio));
|
this.audio.push(Buffer.from(audio));
|
||||||
this.bytes += audio.length;
|
this.bufferedBytes += audio.length;
|
||||||
|
this.totalBytes += audio.length;
|
||||||
|
while (this.bufferedBytes >= DICTATION_SEGMENT_BYTES) {
|
||||||
|
this.queueFullSegment();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
close(): void {
|
close(): void {
|
||||||
if (this.closed) return;
|
if (this.closed) return;
|
||||||
this.closed = true;
|
this.closed = true;
|
||||||
this.connected = false;
|
this.connected = false;
|
||||||
if (!this.bytes) return;
|
if (!this.totalBytes) return;
|
||||||
const audio = Buffer.concat(this.audio);
|
this.emitCompletedTranscripts();
|
||||||
this.audio.length = 0;
|
const tail = Buffer.concat(this.audio);
|
||||||
void this.transcribe(audio);
|
this.audio = [];
|
||||||
|
this.bufferedBytes = 0;
|
||||||
|
if (tail.length > DICTATION_OVERLAP_BYTES || this.completedTranscripts.length === 0) {
|
||||||
|
this.queueTranscription(tail);
|
||||||
|
}
|
||||||
|
void this.processing.finally(() => {
|
||||||
|
this.emitCompletedTranscripts();
|
||||||
|
if (this.processingError && this.completedTranscripts.length === 0) {
|
||||||
|
this.req.onError?.(this.processingError);
|
||||||
|
}
|
||||||
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
private async transcribe(audio: Buffer): Promise<void> {
|
private queueFullSegment(): void {
|
||||||
try {
|
const buffered = Buffer.concat(this.audio);
|
||||||
|
const segment = Buffer.from(buffered.subarray(0, DICTATION_SEGMENT_BYTES));
|
||||||
|
const retained = Buffer.from(buffered.subarray(DICTATION_SEGMENT_BYTES - DICTATION_OVERLAP_BYTES));
|
||||||
|
this.audio = retained.length ? [retained] : [];
|
||||||
|
this.bufferedBytes = retained.length;
|
||||||
|
this.queueTranscription(segment);
|
||||||
|
}
|
||||||
|
|
||||||
|
private queueTranscription(audio: Buffer): void {
|
||||||
|
if (!audio.length) return;
|
||||||
|
this.processing = this.processing.then(async () => {
|
||||||
|
const previous = this.completedTranscripts.join(" ");
|
||||||
|
const text = await this.transcribe(audio, previous.slice(-240));
|
||||||
|
const novel = removeTranscriptOverlap(previous, text);
|
||||||
|
if (novel) this.completedTranscripts.push(novel);
|
||||||
|
if (this.closed) this.emitCompletedTranscripts();
|
||||||
|
}).catch((error: unknown) => {
|
||||||
|
this.processingError = error instanceof Error ? error : new Error(String(error));
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
private emitCompletedTranscripts(): void {
|
||||||
|
while (this.emittedTranscripts < this.completedTranscripts.length) {
|
||||||
|
this.req.onTranscript?.(this.completedTranscripts[this.emittedTranscripts]);
|
||||||
|
this.emittedTranscripts += 1;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
private async transcribe(audio: Buffer, prompt: string): Promise<string> {
|
||||||
const form = new FormData();
|
const form = new FormData();
|
||||||
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
||||||
form.append("model", "whisper-1");
|
form.append("model", "whisper-1");
|
||||||
form.append("language", this.config.language);
|
form.append("language", this.config.language);
|
||||||
|
if (prompt) form.append("prompt", prompt);
|
||||||
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
|
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
|
||||||
method: "POST",
|
method: "POST",
|
||||||
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
|
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
|
||||||
body: form,
|
body: form,
|
||||||
signal: AbortSignal.timeout(4500),
|
signal: AbortSignal.timeout(DICTATION_REQUEST_TIMEOUT_MS),
|
||||||
});
|
});
|
||||||
if (!response.ok) throw new Error(`Athena STT failed (HTTP ${response.status})`);
|
if (!response.ok) throw new Error(`Athena STT failed (HTTP ${response.status})`);
|
||||||
const text = String(record(await response.json()).text || "").trim();
|
return String(record(await response.json()).text || "").trim();
|
||||||
if (text) this.req.onTranscript?.(text);
|
|
||||||
} catch (error) {
|
|
||||||
this.req.onError?.(error instanceof Error ? error : new Error(String(error)));
|
|
||||||
}
|
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
+2
-2
@@ -1,12 +1,12 @@
|
|||||||
{
|
{
|
||||||
"name": "@casaderoll/openclaw-athena-talk",
|
"name": "@casaderoll/openclaw-athena-talk",
|
||||||
"version": "1.2.1",
|
"version": "1.3.0",
|
||||||
"lockfileVersion": 3,
|
"lockfileVersion": 3,
|
||||||
"requires": true,
|
"requires": true,
|
||||||
"packages": {
|
"packages": {
|
||||||
"": {
|
"": {
|
||||||
"name": "@casaderoll/openclaw-athena-talk",
|
"name": "@casaderoll/openclaw-athena-talk",
|
||||||
"version": "1.2.1",
|
"version": "1.3.0",
|
||||||
"devDependencies": {
|
"devDependencies": {
|
||||||
"@types/node": "^24.0.0",
|
"@types/node": "^24.0.0",
|
||||||
"openclaw": "2026.9.4",
|
"openclaw": "2026.9.4",
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
{
|
{
|
||||||
"name": "@casaderoll/openclaw-athena-talk",
|
"name": "@casaderoll/openclaw-athena-talk",
|
||||||
"version": "1.2.1",
|
"version": "1.3.0",
|
||||||
"private": true,
|
"private": true,
|
||||||
"description": "Local OpenClaw Talk provider backed by Athena Whisper and Qwen3-TTS",
|
"description": "Local OpenClaw Talk provider backed by Athena Whisper and Qwen3-TTS",
|
||||||
"type": "module",
|
"type": "module",
|
||||||
|
|||||||
@@ -3,6 +3,14 @@ import { createServer } from "node:http";
|
|||||||
import { test } from "node:test";
|
import { test } from "node:test";
|
||||||
import plugin from "./dist/index.js";
|
import plugin from "./dist/index.js";
|
||||||
|
|
||||||
|
async function waitFor(predicate, timeoutMs = 2000) {
|
||||||
|
const deadline = Date.now() + timeoutMs;
|
||||||
|
while (!predicate()) {
|
||||||
|
if (Date.now() >= deadline) throw new Error("timed out waiting for condition");
|
||||||
|
await new Promise((resolve) => setTimeout(resolve, 10));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
test("dictation registers separately and sends G.711 audio to Athena Whisper", async () => {
|
test("dictation registers separately and sends G.711 audio to Athena Whisper", async () => {
|
||||||
let transcription;
|
let transcription;
|
||||||
plugin.register({
|
plugin.register({
|
||||||
@@ -79,3 +87,73 @@ test("dictation reuses only a TTS key for the same Athena origin", () => {
|
|||||||
cfg: withTts("http://other:8081/v1"), rawConfig: {},
|
cfg: withTts("http://other:8081/v1"), rawConfig: {},
|
||||||
}).apiKey, "");
|
}).apiKey, "");
|
||||||
});
|
});
|
||||||
|
|
||||||
|
test("long dictation is transcribed incrementally before the recording closes", async () => {
|
||||||
|
let transcription;
|
||||||
|
plugin.register({
|
||||||
|
registerRealtimeTranscriptionProvider: (value) => { transcription = value; },
|
||||||
|
registerRealtimeVoiceProvider: () => {},
|
||||||
|
registerHttpRoute: () => {},
|
||||||
|
});
|
||||||
|
|
||||||
|
const answers = [
|
||||||
|
"Dies ist ein langer Abschnitt",
|
||||||
|
"langer Abschnitt mit einer Fortsetzung",
|
||||||
|
"einer Fortsetzung und einem Ende.",
|
||||||
|
];
|
||||||
|
let uploads = 0;
|
||||||
|
const server = createServer(async (req, res) => {
|
||||||
|
const index = uploads++;
|
||||||
|
const form = await new Request("http://localhost", {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "Content-Type": req.headers["content-type"] },
|
||||||
|
body: req,
|
||||||
|
duplex: "half",
|
||||||
|
}).formData();
|
||||||
|
const wav = Buffer.from(await form.get("file").arrayBuffer());
|
||||||
|
assert.equal(wav.readUInt32LE(24), 8000);
|
||||||
|
if (index > 0) assert.ok(String(form.get("prompt") || "").length > 0);
|
||||||
|
res.writeHead(200, { "Content-Type": "application/json" })
|
||||||
|
.end(JSON.stringify({ text: answers[index] }));
|
||||||
|
});
|
||||||
|
await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve));
|
||||||
|
const cfg = { talk: { realtime: { providers: { "athena-talk": {
|
||||||
|
baseUrl: `http://127.0.0.1:${server.address().port}/v1`, language: "de",
|
||||||
|
} } } } };
|
||||||
|
const providerConfig = transcription.resolveConfig({ cfg, rawConfig: {} });
|
||||||
|
const transcripts = [];
|
||||||
|
const errors = [];
|
||||||
|
try {
|
||||||
|
const session = transcription.createSession({
|
||||||
|
cfg,
|
||||||
|
providerConfig,
|
||||||
|
onTranscript: (text) => transcripts.push(text),
|
||||||
|
onError: (error) => errors.push(error),
|
||||||
|
});
|
||||||
|
await session.connect();
|
||||||
|
|
||||||
|
// Six seconds start the first request while dictation is still active.
|
||||||
|
session.sendAudio(Buffer.alloc(48_000, 0xff));
|
||||||
|
await waitFor(() => uploads === 1);
|
||||||
|
assert.deepEqual(transcripts, []);
|
||||||
|
|
||||||
|
// Another 5.5 seconds form the next overlapping segment. The remaining
|
||||||
|
// 2 seconds are finalized only when the user stops dictation.
|
||||||
|
session.sendAudio(Buffer.alloc(44_000, 0xff));
|
||||||
|
await waitFor(() => uploads === 2);
|
||||||
|
session.sendAudio(Buffer.alloc(16_000, 0xff));
|
||||||
|
session.close();
|
||||||
|
|
||||||
|
await waitFor(() => transcripts.length === 3);
|
||||||
|
assert.deepEqual(transcripts, [
|
||||||
|
"Dies ist ein langer Abschnitt",
|
||||||
|
"mit einer Fortsetzung",
|
||||||
|
"und einem Ende.",
|
||||||
|
]);
|
||||||
|
assert.equal(uploads, 3);
|
||||||
|
assert.deepEqual(errors, []);
|
||||||
|
} finally {
|
||||||
|
server.closeAllConnections();
|
||||||
|
await new Promise((resolve) => server.close(resolve));
|
||||||
|
}
|
||||||
|
});
|
||||||
|
|||||||
@@ -69,7 +69,7 @@ cd /opt/mike-ai/stack
|
|||||||
docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice
|
docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice
|
||||||
```
|
```
|
||||||
|
|
||||||
OpenClaw 2026.9.4 uses the `athena-talk` plugin version 1.2.1 with
|
OpenClaw uses the `athena-talk` plugin version 1.3.0 with
|
||||||
`talk.realtime.transport` set to `webrtc` and
|
`talk.realtime.transport` set to `webrtc` and
|
||||||
`talk.realtime.providers.athena-talk.realtimeUpstreamUrl` set to
|
`talk.realtime.providers.athena-talk.realtimeUpstreamUrl` set to
|
||||||
`http://192.168.1.212:8090/v1/realtime/calls`. On the Mac, turn off
|
`http://192.168.1.212:8090/v1/realtime/calls`. On the Mac, turn off
|
||||||
@@ -92,8 +92,12 @@ python3.11 -m venv .venv
|
|||||||
The synthetic microphone test passed over the actual OpenClaw HTTPS offer
|
The synthetic microphone test passed over the actual OpenClaw HTTPS offer
|
||||||
route and WireGuard media path with production Whisper and Qwen3-TTS on
|
route and WireGuard media path with production Whisper and Qwen3-TTS on
|
||||||
2026-09-16. Subsequent real browser Talk sessions successfully transcribed
|
2026-09-16. Subsequent real browser Talk sessions successfully transcribed
|
||||||
and answered multiple user turns. Browser dictation through the same plugin
|
and answered multiple user turns. Browser dictation uses a separate plugin
|
||||||
also produced text after a correction to its separate transcription path.
|
path. Since version 1.3.0, recordings longer than six seconds are
|
||||||
|
pre-transcribed incrementally in overlapping short windows while the
|
||||||
|
microphone remains active. This keeps the final tail inside OpenClaw's
|
||||||
|
five-second completion window. Talk through this WebRTC service still ends
|
||||||
|
and transcribes one utterance at a time.
|
||||||
The first user report of a second turn becoming stuck was addressed by
|
The first user report of a second turn becoming stuck was addressed by
|
||||||
matching conversation item and predecessor IDs; the later two-turn test
|
matching conversation item and predecessor IDs; the later two-turn test
|
||||||
passed. This is still half-duplex and needs a private route for WebRTC media.
|
passed. This is still half-duplex and needs a private route for WebRTC media.
|
||||||
|
|||||||
Reference in New Issue
Block a user