Compare commits
3
Commits
33c04150e3
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
da10f6b48d | ||
|
|
8a323e5b9e | ||
|
|
6378b50086 |
No files matched your search
@@ -4,6 +4,8 @@ Stand: **20. September 2026**, auf Athena geprüft. Die fünf Textprofil-Images
|
|||||||
tragen **llama.cpp b29c606** (0.4.1).
|
tragen **llama.cpp b29c606** (0.4.1).
|
||||||
Der [geprüfte Live-Stand](docs/LIVE_STATE.md) beschreibt Profile, GPUs und die
|
Der [geprüfte Live-Stand](docs/LIVE_STATE.md) beschreibt Profile, GPUs und die
|
||||||
Abweichung zwischen dem bereitgestellten Stack und dem Git-Checkout.
|
Abweichung zwischen dem bereitgestellten Stack und dem Git-Checkout.
|
||||||
|
Seit dem 24. September verarbeitet auch Ultra Bilder; sein Vision-Projektor
|
||||||
|
läuft auf der CPU. [Änderung und Test](docs/ULTRA_CPU_VISION_20260924.md).
|
||||||
Der [Updatebericht vom 15. September](docs/UPDATE_AUDIT_20260915.md)
|
Der [Updatebericht vom 15. September](docs/UPDATE_AUDIT_20260915.md)
|
||||||
enthält die aktuellen Build- und Testbelege. Der ältere
|
enthält die aktuellen Build- und Testbelege. Der ältere
|
||||||
[b10930-Bericht](docs/LLAMA_B10930_UPDATE_20260912.md) dokumentiert einen
|
[b10930-Bericht](docs/LLAMA_B10930_UPDATE_20260912.md) dokumentiert einen
|
||||||
@@ -22,7 +24,7 @@ Sie betreibt:
|
|||||||
- den OpenAI-kompatiblen Profile Router,
|
- den OpenAI-kompatiblen Profile Router,
|
||||||
- Qwen-Image-2.1 INT8 für Textbilder und Referenzbild-Bearbeitung,
|
- Qwen-Image-2.1 INT8 für Textbilder und Referenzbild-Bearbeitung,
|
||||||
- Qwen3-TTS und TTS-Gateway für Sprache,
|
- Qwen3-TTS und TTS-Gateway für Sprache,
|
||||||
- Whisper.cpp und die WebRTC-Brücke für OpenClaw Talk,
|
- Qwen3-ASR auf der CPU und die WebRTC-Brücke für OpenClaw Talk,
|
||||||
- EmbeddingGemma auf der CPU für OpenClaws semantische Memory-Suche,
|
- EmbeddingGemma auf der CPU für OpenClaws semantische Memory-Suche,
|
||||||
- die GPU-lose Mikes-Applio-UI als gesonderten Checkout,
|
- die GPU-lose Mikes-Applio-UI als gesonderten Checkout,
|
||||||
- das Athena-Dashboard,
|
- das Athena-Dashboard,
|
||||||
@@ -93,6 +95,8 @@ nicht direkt. Kein automatischer Host-Neustart ist vorgesehen.
|
|||||||
- Qwen-Image-2.1 INT8: Bildgenerierung und Editing auf der RTX 5080; das Textmodell wird dafür kurz entladen und danach
|
- Qwen-Image-2.1 INT8: Bildgenerierung und Editing auf der RTX 5080; das Textmodell wird dafür kurz entladen und danach
|
||||||
automatisch wiederhergestellt
|
automatisch wiederhergestellt
|
||||||
- Qwen3-TTS 1.7B: RTX 3060; kein Piper-Fallback
|
- Qwen3-TTS 1.7B: RTX 3060; kein Piper-Fallback
|
||||||
|
- Qwen3-ASR 0.6B Q8: CPU, hinter dem OpenAI-kompatiblen
|
||||||
|
Transkriptionsendpunkt mit dem Modellnamen `qwen3-asr`.
|
||||||
- EmbeddingGemma 300M Q8: CPU, OpenAI-kompatibel auf Port 8082; kein GPU-Zugriff
|
- EmbeddingGemma 300M Q8: CPU, OpenAI-kompatibel auf Port 8082; kein GPU-Zugriff
|
||||||
|
|
||||||
Die geprüften Live-Werte stehen in [docs/LIVE_STATE.md](docs/LIVE_STATE.md).
|
Die geprüften Live-Werte stehen in [docs/LIVE_STATE.md](docs/LIVE_STATE.md).
|
||||||
|
|||||||
@@ -26,7 +26,7 @@ dokumentieren den ausgerollten Stand, reduzierte Statuslatenzen und Tests.
|
|||||||
Q5-Worker auf der RTX 3060
|
Q5-Worker auf der RTX 3060
|
||||||
- FLUX.2 Klein 9B FP8 Beta als gestoppter Rückfallcontainer
|
- FLUX.2 Klein 9B FP8 Beta als gestoppter Rückfallcontainer
|
||||||
- Qwen3-TTS 1.7B auf der RTX 3060 hinter dem TTS-Gateway; kein Piper-Fallback
|
- Qwen3-TTS 1.7B auf der RTX 3060 hinter dem TTS-Gateway; kein Piper-Fallback
|
||||||
- Whisper.cpp `small` auf der CPU für lokale deutsche Spracherkennung
|
- Qwen3-ASR 0.6B Q8 auf der CPU für lokale deutsche Spracherkennung
|
||||||
- EmbeddingGemma 300M Q8 auf der CPU für OpenClaws hybride Memory-Suche
|
- EmbeddingGemma 300M Q8 auf der CPU für OpenClaws hybride Memory-Suche
|
||||||
- Live-Dashboard mit 21 Tagen Detailhistorie für GPUs und Slot-Kontextbelegung auf Port 8099
|
- Live-Dashboard mit 21 Tagen Detailhistorie für GPUs und Slot-Kontextbelegung auf Port 8099
|
||||||
- Portainer CE als optionale Container-Ansicht auf Port 9443
|
- Portainer CE als optionale Container-Ansicht auf Port 9443
|
||||||
@@ -142,12 +142,16 @@ Prefills konkurrieren weiterhin um Rechenleistung und den gemeinsamen KV-Pool.
|
|||||||
- Athena-Dashboard: `http://192.168.1.212:8099`
|
- Athena-Dashboard: `http://192.168.1.212:8099`
|
||||||
|
|
||||||
Der Router stellt Sprache OpenAI-kompatibel bereit: Sprachausgabe über
|
Der Router stellt Sprache OpenAI-kompatibel bereit: Sprachausgabe über
|
||||||
`/v1/audio/speech` und Spracherkennung über `/v1/audio/transcriptions`. Das
|
`/v1/audio/speech` und Spracherkennung über `/v1/audio/transcriptions`.
|
||||||
Whisper-Modell liegt persistent im Docker-Volume `whisper-data`; Audiodaten
|
Spracherkennung nutzt Qwen3-ASR-0.6B Q8 auf der CPU; das Modell liegt unter
|
||||||
werden lokal auf Athena verarbeitet. Für OpenClaw Talk liegt der lokale
|
`/data/models/qwen3-asr-0.6b-q8`. Der Adapter liefert reinen Text unter
|
||||||
|
`qwen3-asr`; OpenClaw und die Voice-Brücke verwenden denselben Modellnamen.
|
||||||
|
Whisper-Container, Image und
|
||||||
|
Modellvolume sind entfernt. Audiodaten werden lokal auf Athena verarbeitet.
|
||||||
|
Für OpenClaw Talk liegt der lokale
|
||||||
Realtime-Provider unter
|
Realtime-Provider unter
|
||||||
[`integrations/openclaw-athena-talk`](integrations/openclaw-athena-talk). Er
|
[`integrations/openclaw-athena-talk`](integrations/openclaw-athena-talk). Er
|
||||||
verbindet Mikrofon → Athena Whisper → normalen OpenClaw-Agenten → aktives
|
verbindet Mikrofon → Athena STT → normalen OpenClaw-Agenten → aktives
|
||||||
Athena-TTS, sodass Modell, Werkzeuge und Memory auch im Sprachmodus erhalten
|
Athena-TTS, sodass Modell, Werkzeuge und Memory auch im Sprachmodus erhalten
|
||||||
bleiben. Die Installation landet in OpenClaws persistentem Datenverzeichnis
|
bleiben. Die Installation landet in OpenClaws persistentem Datenverzeichnis
|
||||||
und bleibt deshalb bei normalen Container-Updates bestehen.
|
und bleibt deshalb bei normalen Container-Updates bestehen.
|
||||||
|
|||||||
+69
-31
@@ -394,9 +394,8 @@ services:
|
|||||||
- --spec-draft-type-v
|
- --spec-draft-type-v
|
||||||
- f16
|
- f16
|
||||||
|
|
||||||
# Text-only maximum-context profile. This exact IQ4_XS-pure / 256K / 80:20
|
# Maximum-context profile. Keep the vision projector on CPU so images work
|
||||||
# combination completed the 220K fill test on RTX 5080 + RTX 3060.
|
# without consuming the tightly budgeted GPU memory of the 256K context.
|
||||||
# Deliberately no vision projector: Ultra prioritizes maximum usable context.
|
|
||||||
llama-ultra:
|
llama-ultra:
|
||||||
<<: *llama-common
|
<<: *llama-common
|
||||||
container_name: mike-ai-llama-ultra
|
container_name: mike-ai-llama-ultra
|
||||||
@@ -408,6 +407,9 @@ services:
|
|||||||
command:
|
command:
|
||||||
- --model
|
- --model
|
||||||
- "/models/${ULTRA_MODEL_FILE:?ULTRA_MODEL_FILE is required}"
|
- "/models/${ULTRA_MODEL_FILE:?ULTRA_MODEL_FILE is required}"
|
||||||
|
- --mmproj
|
||||||
|
- "/models/${VISION_PROJECTOR_FILE:?VISION_PROJECTOR_FILE is required}"
|
||||||
|
- --no-mmproj-offload
|
||||||
- --alias
|
- --alias
|
||||||
- qwen-ultra
|
- qwen-ultra
|
||||||
- --ctx-size
|
- --ctx-size
|
||||||
@@ -656,7 +658,7 @@ services:
|
|||||||
MUSIC_START_TIMEOUT: "600"
|
MUSIC_START_TIMEOUT: "600"
|
||||||
VOICE_CHANGE_START_TIMEOUT: "600"
|
VOICE_CHANGE_START_TIMEOUT: "600"
|
||||||
APPLIO_START_TIMEOUT: "900"
|
APPLIO_START_TIMEOUT: "900"
|
||||||
STT_WORKER_URL: http://whisper:8084
|
STT_WORKER_URL: http://qwen-asr-worker:8084
|
||||||
STT_TIMEOUT: "300"
|
STT_TIMEOUT: "300"
|
||||||
networks: [frontend, control, inference]
|
networks: [frontend, control, inference]
|
||||||
security_opt: ["no-new-privileges:true"]
|
security_opt: ["no-new-privileges:true"]
|
||||||
@@ -679,7 +681,7 @@ services:
|
|||||||
condition: service_healthy
|
condition: service_healthy
|
||||||
tts-gateway:
|
tts-gateway:
|
||||||
condition: service_healthy
|
condition: service_healthy
|
||||||
whisper:
|
qwen-asr-worker:
|
||||||
condition: service_healthy
|
condition: service_healthy
|
||||||
|
|
||||||
image-worker:
|
image-worker:
|
||||||
@@ -974,40 +976,77 @@ services:
|
|||||||
retries: 12
|
retries: 12
|
||||||
start_period: 10s
|
start_period: 10s
|
||||||
|
|
||||||
whisper:
|
qwen-asr:
|
||||||
build:
|
image: ${LLAMA_CPU_IMAGE:-mike-ai/llama.cpp-cpu:local}
|
||||||
context: .
|
container_name: mike-ai-qwen-asr
|
||||||
dockerfile: platform/docker/whisper/Dockerfile
|
|
||||||
args:
|
|
||||||
WHISPER_CPP_VERSION: ${WHISPER_CPP_VERSION:-v1.9.1}
|
|
||||||
image: mike-ai/whisper:local
|
|
||||||
container_name: mike-ai-whisper
|
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
read_only: true
|
read_only: true
|
||||||
tmpfs:
|
tmpfs:
|
||||||
- /tmp:size=2g,mode=1777
|
- /tmp:size=256m,mode=1777
|
||||||
volumes:
|
volumes:
|
||||||
- whisper-data:/models
|
- "${QWEN_ASR_MODEL_DIR:-/data/models/qwen3-asr-0.6b-q8}:/models:ro"
|
||||||
environment:
|
command:
|
||||||
WHISPER_HOST: 0.0.0.0
|
- --model
|
||||||
WHISPER_PORT: "8084"
|
- /models/Qwen3-ASR-0.6B-Q8_0.gguf
|
||||||
WHISPER_CLI: /opt/whisper.cpp/build/bin/whisper-cli
|
- --mmproj
|
||||||
WHISPER_MODEL: /models/ggml-small.bin
|
- /models/mmproj-Qwen3-ASR-0.6B-Q8_0.gguf
|
||||||
WHISPER_MODEL_URL: ${WHISPER_MODEL_URL:-https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-small.bin}
|
- --no-mmproj-offload
|
||||||
WHISPER_SERVER_PORT: "8085"
|
- --n-gpu-layers
|
||||||
WHISPER_THREADS: ${WHISPER_THREADS:-8}
|
- "0"
|
||||||
WHISPER_LANGUAGE: ${WHISPER_LANGUAGE:-de}
|
- --alias
|
||||||
networks: [frontend, inference]
|
- qwen3-asr-0.6b
|
||||||
|
- --ctx-size
|
||||||
|
- "4096"
|
||||||
|
- --threads
|
||||||
|
- "6"
|
||||||
|
- --parallel
|
||||||
|
- "1"
|
||||||
|
- --host
|
||||||
|
- 0.0.0.0
|
||||||
|
- --port
|
||||||
|
- "8080"
|
||||||
|
- --no-ui
|
||||||
|
- --fit
|
||||||
|
- "off"
|
||||||
|
cpus: 6
|
||||||
|
mem_limit: 6g
|
||||||
|
networks: [inference]
|
||||||
security_opt: ["no-new-privileges:true"]
|
security_opt: ["no-new-privileges:true"]
|
||||||
cap_drop: [ALL]
|
cap_drop: [ALL]
|
||||||
# The entrypoint supervises whisper-server after dropping it to uid 10004.
|
|
||||||
cap_add: [CHOWN, SETUID, SETGID, KILL]
|
|
||||||
healthcheck:
|
healthcheck:
|
||||||
test: [CMD, curl, -fsS, "http://127.0.0.1:8084/status"]
|
test: [CMD, curl, -fsS, "http://127.0.0.1:8080/health"]
|
||||||
interval: 10s
|
interval: 10s
|
||||||
timeout: 5s
|
timeout: 5s
|
||||||
retries: 90
|
retries: 12
|
||||||
start_period: 20m
|
start_period: 30s
|
||||||
|
|
||||||
|
qwen-asr-worker:
|
||||||
|
build:
|
||||||
|
context: .
|
||||||
|
dockerfile: platform/docker/qwen-asr-worker/Dockerfile
|
||||||
|
image: mike-ai/qwen-asr-worker:local
|
||||||
|
container_name: mike-ai-qwen-asr-worker
|
||||||
|
restart: unless-stopped
|
||||||
|
read_only: true
|
||||||
|
tmpfs:
|
||||||
|
- /tmp:size=256m,mode=1777
|
||||||
|
environment:
|
||||||
|
QWEN_ASR_HOST: 0.0.0.0
|
||||||
|
QWEN_ASR_PORT: "8084"
|
||||||
|
QWEN_ASR_LANGUAGE: de
|
||||||
|
QWEN_ASR_SERVER_URL: http://qwen-asr:8080
|
||||||
|
networks: [inference]
|
||||||
|
security_opt: ["no-new-privileges:true"]
|
||||||
|
cap_drop: [ALL]
|
||||||
|
depends_on:
|
||||||
|
qwen-asr:
|
||||||
|
condition: service_healthy
|
||||||
|
healthcheck:
|
||||||
|
test: [CMD, python, -c, "import json,urllib.request; assert json.load(urllib.request.urlopen('http://127.0.0.1:8084/status', timeout=2))['ready']"]
|
||||||
|
interval: 10s
|
||||||
|
timeout: 5s
|
||||||
|
retries: 12
|
||||||
|
start_period: 15s
|
||||||
|
|
||||||
llama-dashboard:
|
llama-dashboard:
|
||||||
build: ./platform/llama-dashboard
|
build: ./platform/llama-dashboard
|
||||||
@@ -1122,7 +1161,6 @@ networks:
|
|||||||
name: mike-ai-tools-egress
|
name: mike-ai-tools-egress
|
||||||
|
|
||||||
volumes:
|
volumes:
|
||||||
whisper-data:
|
|
||||||
router-state:
|
router-state:
|
||||||
router-images:
|
router-images:
|
||||||
portainer-data:
|
portainer-data:
|
||||||
|
|||||||
@@ -47,9 +47,9 @@
|
|||||||
"model_env": "ULTRA_MODEL_FILE",
|
"model_env": "ULTRA_MODEL_FILE",
|
||||||
"model_family": "Qwen3.8-27B IQ4 XS Pure",
|
"model_family": "Qwen3.8-27B IQ4 XS Pure",
|
||||||
"gpu_split": "80:20",
|
"gpu_split": "80:20",
|
||||||
"vision": false,
|
"vision": true,
|
||||||
"mtp": 2,
|
"mtp": 2,
|
||||||
"description": "Maximaler Textkontext; bewusst ohne Vision-Projektor."
|
"description": "Maximaler Kontext mit Vision-Projektor auf der CPU."
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"id": "uncensored",
|
"id": "uncensored",
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
#!/usr/bin/env python3
|
#!/usr/bin/env python3
|
||||||
"""Mock-STT-Worker für lokale Tests.
|
"""Mock-STT-Worker für lokale Tests.
|
||||||
|
|
||||||
Simuliert den Whisper-STT-Worker:
|
Simuliert den Qwen3-ASR-Adapter:
|
||||||
GET /status → ready: true
|
GET /status → ready: true
|
||||||
POST /transcribe → liefert festes Transkript
|
POST /transcribe → liefert festes Transkript
|
||||||
|
|
||||||
|
|||||||
+8
-8
@@ -220,7 +220,7 @@ def main() -> None:
|
|||||||
# Test 1: Multipart + Content-Length (bestehender Pfad)
|
# Test 1: Multipart + Content-Length (bestehender Pfad)
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 1: Multipart + Content-Length")
|
print("Test 1: Multipart + Content-Length")
|
||||||
mp = build_multipart({"model": "whisper-1", "language": "de"},
|
mp = build_multipart({"model": "qwen3-asr", "language": "de"},
|
||||||
file_data=fake_webm)
|
file_data=fake_webm)
|
||||||
status, body = http_request(
|
status, body = http_request(
|
||||||
"POST", PORTS["router"], "/v1/audio/transcriptions",
|
"POST", PORTS["router"], "/v1/audio/transcriptions",
|
||||||
@@ -235,7 +235,7 @@ def main() -> None:
|
|||||||
# Test 2: Multipart + Transfer-Encoding chunked (einfach)
|
# Test 2: Multipart + Transfer-Encoding chunked (einfach)
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 2: Multipart + chunked (einfach)")
|
print("Test 2: Multipart + chunked (einfach)")
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=fake_webm)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=fake_webm)
|
||||||
chunked = build_chunked_body([mp])
|
chunked = build_chunked_body([mp])
|
||||||
raw = (
|
raw = (
|
||||||
b"POST /v1/audio/transcriptions HTTP/1.1\r\n"
|
b"POST /v1/audio/transcriptions HTTP/1.1\r\n"
|
||||||
@@ -254,7 +254,7 @@ def main() -> None:
|
|||||||
# Test 3: Mehrere unterschiedlich große Chunks
|
# Test 3: Mehrere unterschiedlich große Chunks
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 3: Mehrere unterschiedlich große Chunks")
|
print("Test 3: Mehrere unterschiedlich große Chunks")
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=fake_webm)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=fake_webm)
|
||||||
# In 5 Chunks aufteilen (unterschiedlich groß)
|
# In 5 Chunks aufteilen (unterschiedlich groß)
|
||||||
chunks = []
|
chunks = []
|
||||||
sizes = [10, 50, 7, 100, 33]
|
sizes = [10, 50, 7, 100, 33]
|
||||||
@@ -283,7 +283,7 @@ def main() -> None:
|
|||||||
# Test 4: Boundary über Chunk-Grenzen verteilt
|
# Test 4: Boundary über Chunk-Grenzen verteilt
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 4: Boundary über Chunk-Grenzen verteilt")
|
print("Test 4: Boundary über Chunk-Grenzen verteilt")
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=fake_webm)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=fake_webm)
|
||||||
# Boundary-String finden und Chunk-Grenze genau dorthin setzen
|
# Boundary-String finden und Chunk-Grenze genau dorthin setzen
|
||||||
boundary_str = b"--testboundary123"
|
boundary_str = b"--testboundary123"
|
||||||
idx = mp.find(boundary_str, 10) # zweite Boundary (vor file)
|
idx = mp.find(boundary_str, 10) # zweite Boundary (vor file)
|
||||||
@@ -310,7 +310,7 @@ def main() -> None:
|
|||||||
# Test 5: Chunk Extensions
|
# Test 5: Chunk Extensions
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 5: Chunk Extensions")
|
print("Test 5: Chunk Extensions")
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=fake_webm)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=fake_webm)
|
||||||
chunks_ext = [
|
chunks_ext = [
|
||||||
(mp[:20], "ext1=value1"),
|
(mp[:20], "ext1=value1"),
|
||||||
(mp[20:60], None),
|
(mp[20:60], None),
|
||||||
@@ -335,7 +335,7 @@ def main() -> None:
|
|||||||
# hier explizit mit Trailer)
|
# hier explizit mit Trailer)
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 6: 0-Chunk mit Trailer")
|
print("Test 6: 0-Chunk mit Trailer")
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=fake_webm)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=fake_webm)
|
||||||
chunked = build_chunked_body([mp])
|
chunked = build_chunked_body([mp])
|
||||||
# Trailer hinzufügen
|
# Trailer hinzufügen
|
||||||
chunked_with_trailer = chunked.replace(
|
chunked_with_trailer = chunked.replace(
|
||||||
@@ -376,7 +376,7 @@ def main() -> None:
|
|||||||
print("Test 8: Uploadgrößenlimit")
|
print("Test 8: Uploadgrößenlimit")
|
||||||
# MAX_UPLOAD_SIZE = 1 MB, also 2 MB senden
|
# MAX_UPLOAD_SIZE = 1 MB, also 2 MB senden
|
||||||
big_data = b"A" * (2 * 1024 * 1024)
|
big_data = b"A" * (2 * 1024 * 1024)
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=big_data)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=big_data)
|
||||||
chunked = build_chunked_body([mp[:1024 * 1024], mp[1024 * 1024:]])
|
chunked = build_chunked_body([mp[:1024 * 1024], mp[1024 * 1024:]])
|
||||||
raw = (
|
raw = (
|
||||||
b"POST /v1/audio/transcriptions HTTP/1.1\r\n"
|
b"POST /v1/audio/transcriptions HTTP/1.1\r\n"
|
||||||
@@ -402,7 +402,7 @@ def main() -> None:
|
|||||||
+ b"WEBM_OPUS_AUDIO_DATA" * 100)
|
+ b"WEBM_OPUS_AUDIO_DATA" * 100)
|
||||||
boundary = "950bd961b24c4a32801e31b128c85e09"
|
boundary = "950bd961b24c4a32801e31b128c85e09"
|
||||||
mp = build_multipart(
|
mp = build_multipart(
|
||||||
{"model": "whisper-1", "language": "de"},
|
{"model": "qwen3-asr", "language": "de"},
|
||||||
file_data=webm_data,
|
file_data=webm_data,
|
||||||
filename="recording.webm",
|
filename="recording.webm",
|
||||||
boundary=boundary,
|
boundary=boundary,
|
||||||
|
|||||||
+12
-12
@@ -71,13 +71,13 @@ def test_quoted_boundary():
|
|||||||
boundary = "----WebKitFormBoundary7MA4YWxkTrZu0gW"
|
boundary = "----WebKitFormBoundary7MA4YWxkTrZu0gW"
|
||||||
webm = b"\x1a\x45\xdf\xa3" + b"\x00\x01\x02\x03\xff\xfe\xfd" * 50
|
webm = b"\x1a\x45\xdf\xa3" + b"\x00\x01\x02\x03\xff\xfe\xfd" * 50
|
||||||
body, ct = build(
|
body, ct = build(
|
||||||
[("model", "whisper-1", None), ("file", webm, "t.webm")],
|
[("model", "qwen3-asr", None), ("file", webm, "t.webm")],
|
||||||
boundary, quoted=True,
|
boundary, quoted=True,
|
||||||
)
|
)
|
||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == webm, "file_data mismatch"
|
assert fd == webm, "file_data mismatch"
|
||||||
assert fn == "t.webm", f"filename mismatch: {fn!r}"
|
assert fn == "t.webm", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
print(" quoted boundary: OK")
|
print(" quoted boundary: OK")
|
||||||
|
|
||||||
|
|
||||||
@@ -109,7 +109,7 @@ def test_openwebui_style():
|
|||||||
f"\r\n--{boundary}\r\n"
|
f"\r\n--{boundary}\r\n"
|
||||||
f'Content-Disposition: form-data; name="model"\r\n'
|
f'Content-Disposition: form-data; name="model"\r\n'
|
||||||
f"\r\n"
|
f"\r\n"
|
||||||
f"whisper-1\r\n"
|
f"qwen3-asr\r\n"
|
||||||
f"--{boundary}\r\n"
|
f"--{boundary}\r\n"
|
||||||
f'Content-Disposition: form-data; name="temperature"\r\n'
|
f'Content-Disposition: form-data; name="temperature"\r\n'
|
||||||
f"\r\n"
|
f"\r\n"
|
||||||
@@ -119,7 +119,7 @@ def test_openwebui_style():
|
|||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == webm, "file_data mismatch"
|
assert fd == webm, "file_data mismatch"
|
||||||
assert fn == "rec.webm", f"filename mismatch: {fn!r}"
|
assert fn == "rec.webm", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
assert fl["temperature"] == "0.0", f"temperature mismatch: {fl!r}"
|
assert fl["temperature"] == "0.0", f"temperature mismatch: {fl!r}"
|
||||||
print(" Open-WebUI-artig: OK")
|
print(" Open-WebUI-artig: OK")
|
||||||
|
|
||||||
@@ -128,13 +128,13 @@ def test_file_before_model():
|
|||||||
"""File-Feld vor model-Feld."""
|
"""File-Feld vor model-Feld."""
|
||||||
boundary = "boundary123"
|
boundary = "boundary123"
|
||||||
body, ct = build(
|
body, ct = build(
|
||||||
[("file", b"DATA", "f.wav"), ("model", "whisper-1", None)],
|
[("file", b"DATA", "f.wav"), ("model", "qwen3-asr", None)],
|
||||||
boundary, quoted=False,
|
boundary, quoted=False,
|
||||||
)
|
)
|
||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == b"DATA", "file_data mismatch"
|
assert fd == b"DATA", "file_data mismatch"
|
||||||
assert fn == "f.wav", f"filename mismatch: {fn!r}"
|
assert fn == "f.wav", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
print(" File vor model: OK")
|
print(" File vor model: OK")
|
||||||
|
|
||||||
|
|
||||||
@@ -142,13 +142,13 @@ def test_file_after_model():
|
|||||||
"""model-Feld vor File-Feld."""
|
"""model-Feld vor File-Feld."""
|
||||||
boundary = "boundary456"
|
boundary = "boundary456"
|
||||||
body, ct = build(
|
body, ct = build(
|
||||||
[("model", "whisper-1", None), ("file", b"DATA", "g.wav")],
|
[("model", "qwen3-asr", None), ("file", b"DATA", "g.wav")],
|
||||||
boundary, quoted=False,
|
boundary, quoted=False,
|
||||||
)
|
)
|
||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == b"DATA", "file_data mismatch"
|
assert fd == b"DATA", "file_data mismatch"
|
||||||
assert fn == "g.wav", f"filename mismatch: {fn!r}"
|
assert fn == "g.wav", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
print(" File nach model: OK")
|
print(" File nach model: OK")
|
||||||
|
|
||||||
|
|
||||||
@@ -182,13 +182,13 @@ def test_extra_headers_ignored():
|
|||||||
f"\r\n--{boundary}\r\n"
|
f"\r\n--{boundary}\r\n"
|
||||||
f'Content-Disposition: form-data; name="model"\r\n'
|
f'Content-Disposition: form-data; name="model"\r\n'
|
||||||
f"\r\n"
|
f"\r\n"
|
||||||
f"whisper-1\r\n"
|
f"qwen3-asr\r\n"
|
||||||
f"--{boundary}--\r\n"
|
f"--{boundary}--\r\n"
|
||||||
).encode()
|
).encode()
|
||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == webm, "file_data mismatch"
|
assert fd == webm, "file_data mismatch"
|
||||||
assert fn == "x.webm", f"filename mismatch: {fn!r}"
|
assert fn == "x.webm", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
print(" Extra-Header ignoriert: OK")
|
print(" Extra-Header ignoriert: OK")
|
||||||
|
|
||||||
|
|
||||||
@@ -198,7 +198,7 @@ def test_all_fields():
|
|||||||
body, ct = build(
|
body, ct = build(
|
||||||
[
|
[
|
||||||
("file", b"AUDIO", "a.webm"),
|
("file", b"AUDIO", "a.webm"),
|
||||||
("model", "whisper-1", None),
|
("model", "qwen3-asr", None),
|
||||||
("language", "de", None),
|
("language", "de", None),
|
||||||
("prompt", "Kontext", None),
|
("prompt", "Kontext", None),
|
||||||
("response_format", "verbose_json", None),
|
("response_format", "verbose_json", None),
|
||||||
@@ -209,7 +209,7 @@ def test_all_fields():
|
|||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == b"AUDIO", "file_data mismatch"
|
assert fd == b"AUDIO", "file_data mismatch"
|
||||||
assert fn == "a.webm", f"filename mismatch: {fn!r}"
|
assert fn == "a.webm", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
assert fl["language"] == "de", f"language mismatch: {fl!r}"
|
assert fl["language"] == "de", f"language mismatch: {fl!r}"
|
||||||
assert fl["prompt"] == "Kontext", f"prompt mismatch: {fl!r}"
|
assert fl["prompt"] == "Kontext", f"prompt mismatch: {fl!r}"
|
||||||
assert fl["response_format"] == "verbose_json", f"response_format mismatch: {fl!r}"
|
assert fl["response_format"] == "verbose_json", f"response_format mismatch: {fl!r}"
|
||||||
|
|||||||
@@ -22,7 +22,7 @@ nicht automatisch ein ungenutzter Rest.
|
|||||||
| `mike-ai-llama-fast` | Qwen3.8-27B `IQ4-MIX`, Qwen-MMProj BF16 | Schnelles Q4-Text-/Vision-Profil mit 76.800 Token Kontext. |
|
| `mike-ai-llama-fast` | Qwen3.8-27B `IQ4-MIX`, Qwen-MMProj BF16 | Schnelles Q4-Text-/Vision-Profil mit 76.800 Token Kontext. |
|
||||||
| `mike-ai-llama-large` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Q4-Text-/Vision-Profil mit 192.000 Token Kontext und Verteilung auf beide GPUs. |
|
| `mike-ai-llama-large` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Q4-Text-/Vision-Profil mit 192.000 Token Kontext und Verteilung auf beide GPUs. |
|
||||||
| `mike-ai-llama-medium` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Standard-Q4-Text-/Vision-Profil mit gemeinsamem 160.000-Token-KV-Pool, aktuell zwei Slots und Verteilung auf beide GPUs. |
|
| `mike-ai-llama-medium` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Standard-Q4-Text-/Vision-Profil mit gemeinsamem 160.000-Token-KV-Pool, aktuell zwei Slots und Verteilung auf beide GPUs. |
|
||||||
| `mike-ai-llama-ultra` | Qwen3.8-27B `IQ4_XS-pure`, ohne Vision-Projektor | Maximales Langkontextprofil mit 262.144 Token Kontext und Verteilung auf beide GPUs. |
|
| `mike-ai-llama-ultra` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 auf CPU | Text-/Vision-Profil mit 262.144 Token Kontext und Verteilung des Textmodells auf beide GPUs. |
|
||||||
| `mike-ai-llama-uncensored` | Qwen3.8-27B Abliterated `Q4_K_M`, eigener MMProj F16 | Spezialprofil mit 80.000 Token Kontext und gelockerten Modellgrenzen. |
|
| `mike-ai-llama-uncensored` | Qwen3.8-27B Abliterated `Q4_K_M`, eigener MMProj F16 | Spezialprofil mit 80.000 Token Kontext und gelockerten Modellgrenzen. |
|
||||||
| `mike-ai-ltx2-studio` | LTX-Video-Backend | GPU-Worker für lokale Videogenerierung; beim Abgleich gestoppt. |
|
| `mike-ai-ltx2-studio` | LTX-Video-Backend | GPU-Worker für lokale Videogenerierung; beim Abgleich gestoppt. |
|
||||||
| `mike-ai-mcp-athena-operator` | kein Modell | Stellt Hermes begrenzte Werkzeuge zum Prüfen, Ändern, Testen, Sichern und Versionieren von Athena bereit. |
|
| `mike-ai-mcp-athena-operator` | kein Modell | Stellt Hermes begrenzte Werkzeuge zum Prüfen, Ändern, Testen, Sichern und Versionieren von Athena bereit. |
|
||||||
@@ -32,13 +32,14 @@ nicht automatisch ein ungenutzter Rest.
|
|||||||
| `mike-ai-portainer` | kein Modell; Portainer CE | Optionale Docker-Verwaltungsoberfläche. |
|
| `mike-ai-portainer` | kein Modell; Portainer CE | Optionale Docker-Verwaltungsoberfläche. |
|
||||||
| `mike-ai-profile-controller` | kein Modell | Startet und stoppt ausschließlich freigegebene Modellprofile und Spezialworker in einer sicheren Reihenfolge. |
|
| `mike-ai-profile-controller` | kein Modell | Startet und stoppt ausschließlich freigegebene Modellprofile und Spezialworker in einer sicheren Reihenfolge. |
|
||||||
| `mike-ai-qwen-cron-test` | kleines Qwen-Testmodell | Gestoppter CPU-/Cron-Worker-Versuch; nicht produktiv eingesetzt. |
|
| `mike-ai-qwen-cron-test` | kleines Qwen-Testmodell | Gestoppter CPU-/Cron-Worker-Versuch; nicht produktiv eingesetzt. |
|
||||||
| `mike-ai-realtime-voice` | kein Modell | Laufende, gesunde WebRTC-Brücke für OpenClaw Talk; verbindet über den privaten WireGuard-Pfad Athena Whisper, OpenClaw-Agent und Qwen3-TTS ohne Profilwechsel. |
|
| `mike-ai-realtime-voice` | kein Modell | Laufende, gesunde WebRTC-Brücke für OpenClaw Talk; verbindet über den privaten WireGuard-Pfad Athena Qwen3-ASR, OpenClaw-Agent und Qwen3-TTS ohne Profilwechsel. |
|
||||||
| `mike-ai-qwen3-tts` | `Qwen/Qwen3-TTS-12Hz-1.7B-Base`, Stimme Serena | Hochwertige deutsche Sprachausgabe auf der RTX 3060 im LLM-Betrieb. |
|
| `mike-ai-qwen3-tts` | `Qwen/Qwen3-TTS-12Hz-1.7B-Base`, Stimme Serena | Hochwertige deutsche Sprachausgabe auf der RTX 3060 im LLM-Betrieb. |
|
||||||
| `mike-ai-router` | kein eigenes Modell | Einzige OpenAI-kompatible Modelladresse; koordiniert Profile, Bildaufträge, Sprache und Betriebsarten. |
|
| `mike-ai-router` | kein eigenes Modell | Einzige OpenAI-kompatible Modelladresse; koordiniert Profile, Bildaufträge, Sprache und Betriebsarten. |
|
||||||
| `mike-ai-stem-separator` | BS-RoFormer Viperx 1297, `htdemucs_ft`, `htdemucs_6s`, `MossFormer2_SE_48K` | Trennt Gesang, Instrumente oder Sprache/Hintergrundgeräusche im exklusiven Separationsmodus. |
|
| `mike-ai-stem-separator` | BS-RoFormer Viperx 1297, `htdemucs_ft`, `htdemucs_6s`, `MossFormer2_SE_48K` | Trennt Gesang, Instrumente oder Sprache/Hintergrundgeräusche im exklusiven Separationsmodus. |
|
||||||
| `mike-ai-tts-gateway` | kein eigenes Modell | Normalisiert Text, konvertiert Ausgabeformate und stellt Qwen3-TTS sowie natives PCM-Streaming über eine stabile interne API bereit. |
|
| `mike-ai-tts-gateway` | kein eigenes Modell | Normalisiert Text, konvertiert Ausgabeformate und stellt Qwen3-TTS sowie natives PCM-Streaming über eine stabile interne API bereit. |
|
||||||
| `mike-ai-voice-studio` | `k2-fsa/OmniVoice` 0.2.1 mit Whisper-ASR | Erzeugt Text-to-Speech mit einer Referenzstimme; kein Audio-to-Audio-Voice-Changer. |
|
| `mike-ai-voice-studio` | `k2-fsa/OmniVoice` 0.2.1 mit Whisper-ASR | Erzeugt Text-to-Speech mit einer Referenzstimme; kein Audio-to-Audio-Voice-Changer. |
|
||||||
| `mike-ai-whisper` | Whisper.cpp 1.9.4, `ggml-small` | Lokale deutsche Spracherkennung auf der CPU über `/v1/audio/transcriptions`. |
|
| `mike-ai-qwen-asr` | Qwen3-ASR 0.6B Q8 | CPU-Inferenz für lokale deutsche Spracherkennung. |
|
||||||
|
| `mike-ai-qwen-asr-worker` | kein eigenes Modell | Audio-Adapter für `/v1/audio/transcriptions` mit dem Modellnamen `qwen3-asr`. |
|
||||||
| `mike-ai-wireguard-gateway` | kein Modell | Veröffentlicht Dashboard und Fachoberflächen ausschließlich über den privaten WireGuard-Pfad. |
|
| `mike-ai-wireguard-gateway` | kein Modell | Veröffentlicht Dashboard und Fachoberflächen ausschließlich über den privaten WireGuard-Pfad. |
|
||||||
| `mike-ai-xvc-studio` | `chenxie95/X-VC`, GLM-4-Voice-Tokenizer und optional Resemble Enhance | Wandelt eine vorhandene Sprachaufnahme anhand einer Referenzstimme in Audio zu Audio um; gibt das native 16-kHz-Ergebnis und optional eine neural restaurierte 44,1-kHz-Fassung aus. |
|
| `mike-ai-xvc-studio` | `chenxie95/X-VC`, GLM-4-Voice-Tokenizer und optional Resemble Enhance | Wandelt eine vorhandene Sprachaufnahme anhand einer Referenzstimme in Audio zu Audio um; gibt das native 16-kHz-Ergebnis und optional eine neural restaurierte 44,1-kHz-Fassung aus. |
|
||||||
| `mike-ai-yue2-playground` | Image `mike-ai/yue2:3b-0.1.6` | Vorhandener, gestoppter Playground; in diesem Abgleich nicht funktional getestet. |
|
| `mike-ai-yue2-playground` | Image `mike-ai/yue2:3b-0.1.6` | Vorhandener, gestoppter Playground; in diesem Abgleich nicht funktional getestet. |
|
||||||
|
|||||||
@@ -1,5 +1,19 @@
|
|||||||
# Geprüfter Live-Stand auf Athena
|
# Geprüfter Live-Stand auf Athena
|
||||||
|
|
||||||
|
Nachtrag vom 25. September 2026: Die produktive Spracherkennung läuft über
|
||||||
|
Qwen3-ASR 0.6B Q8 auf der CPU (`mike-ai-qwen-asr` und
|
||||||
|
`mike-ai-qwen-asr-worker`). Der Router bietet nur `qwen3-asr` als STT-Modell
|
||||||
|
an; OpenClaw und die Voice-Brücke verwenden denselben Namen. Eine M4A-Aufnahme
|
||||||
|
wurde über `/v1/audio/transcriptions` geprüft. Whisper-Container,
|
||||||
|
Images und Modellvolume wurden entfernt. Das aktive Ultra-Profil wurde bei
|
||||||
|
dieser Umstellung nicht gewechselt. Die Angaben zu Whisper weiter unten
|
||||||
|
beschreiben den historischen Stand vom 21. September.
|
||||||
|
|
||||||
|
Nachtrag vom 24. September 2026: Ultra verarbeitet nun Bilder mit einem
|
||||||
|
CPU-seitigen BF16-Vision-Projektor. Der Stand und der Funktionstest sind in
|
||||||
|
[Ultra-Vision mit CPU-Projektor](ULTRA_CPU_VISION_20260924.md) dokumentiert.
|
||||||
|
Die folgende Tabelle bildet weiterhin den historischen Stand vom 21. September ab.
|
||||||
|
|
||||||
Stand: **21. September 2026**. Quelle: lesender SSH-Abgleich von Docker,
|
Stand: **21. September 2026**. Quelle: lesender SSH-Abgleich von Docker,
|
||||||
Compose-Dateien, Git-Inhalten, Images, Health-Endpunkten und Backup-Timern.
|
Compose-Dateien, Git-Inhalten, Images, Health-Endpunkten und Backup-Timern.
|
||||||
Containerzustände sind Momentaufnahmen; der Controller darf Profile danach
|
Containerzustände sind Momentaufnahmen; der Controller darf Profile danach
|
||||||
|
|||||||
+4
-2
@@ -24,8 +24,10 @@ gesichert. Für eine Neuinstallation ist der im Git dokumentierte Build mit
|
|||||||
installieren; eine Änderung an OpenClaw-Core-Dateien ist nicht erforderlich.
|
installieren; eine Änderung an OpenClaw-Core-Dateien ist nicht erforderlich.
|
||||||
|
|
||||||
Piper-Daten sind kein aktueller Sicherungsbestand. Modellgewichte unter
|
Piper-Daten sind kein aktueller Sicherungsbestand. Modellgewichte unter
|
||||||
`/data/models` und das reproduzierbare Whisper-Volume gehören nicht zu diesen
|
`/data/models`, einschließlich Qwen3-ASR unter
|
||||||
Backup-Mounts. Ein Backup ausschließlich auf `/data` schützt nicht vor deren Ausfall.
|
`/data/models/qwen3-asr-0.6b-q8`, gehören nicht zu diesen Backup-Mounts.
|
||||||
|
Das frühere Whisper-Volume wurde am 25. September 2026 entfernt.
|
||||||
|
Ein Backup ausschließlich auf `/data` schützt nicht vor einem Ausfall der Datenplatte.
|
||||||
Das gilt auch für das reproduzierbare EmbeddingGemma-Gewicht unter
|
Das gilt auch für das reproduzierbare EmbeddingGemma-Gewicht unter
|
||||||
`/data/models/embeddinggemma`; URL und SHA-256 stehen in
|
`/data/models/embeddinggemma`; URL und SHA-256 stehen in
|
||||||
`config/install.env.example`, sodass der Installer es erneut laden und prüfen kann.
|
`config/install.env.example`, sodass der Installer es erneut laden und prüfen kann.
|
||||||
|
|||||||
@@ -9,7 +9,7 @@ Standardprofil: **medium** · globales Ausgabelimit: **8192 Token**
|
|||||||
| fast | `qwen-fast` | 76,800 | 1 | Qwen3.8-27B IQ4 Mix | 5080 only | ja | 2 |
|
| fast | `qwen-fast` | 76,800 | 1 | Qwen3.8-27B IQ4 Mix | 5080 only | ja | 2 |
|
||||||
| medium | `qwen-medium` | 160,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 85:15 | ja | 2 |
|
| medium | `qwen-medium` | 160,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 85:15 | ja | 2 |
|
||||||
| large | `qwen-large` | 192,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 86:14 | ja | 2 |
|
| large | `qwen-large` | 192,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 86:14 | ja | 2 |
|
||||||
| ultra | `qwen-ultra` | 262,144 | 1 | Qwen3.8-27B IQ4 XS Pure | 80:20 | nein | 2 |
|
| ultra | `qwen-ultra` | 262,144 | 1 | Qwen3.8-27B IQ4 XS Pure | 80:20 | ja | 2 |
|
||||||
| uncensored | `qwen-uncensored` | 80,000 | 1 | Qwen3.8-27B Abliterated Q4_K_M | 90:10 | ja | 2 |
|
| uncensored | `qwen-uncensored` | 80,000 | 1 | Qwen3.8-27B Abliterated Q4_K_M | 90:10 | ja | 2 |
|
||||||
|
|
||||||
## Zweck
|
## Zweck
|
||||||
@@ -17,5 +17,5 @@ Standardprofil: **medium** · globales Ausgabelimit: **8192 Token**
|
|||||||
- **fast**: Schnelles Profil für kurze Chats und zügige Werkzeugaufgaben.
|
- **fast**: Schnelles Profil für kurze Chats und zügige Werkzeugaufgaben.
|
||||||
- **medium**: Ausgewogenes Standardprofil für Alltag und lange agentische Aufgaben.
|
- **medium**: Ausgewogenes Standardprofil für Alltag und lange agentische Aufgaben.
|
||||||
- **large**: Großes Profil für umfangreiche Dokumente und lange technische Arbeiten.
|
- **large**: Großes Profil für umfangreiche Dokumente und lange technische Arbeiten.
|
||||||
- **ultra**: Maximaler Textkontext; bewusst ohne Vision-Projektor.
|
- **ultra**: Maximaler Kontext mit Vision-Projektor auf der CPU.
|
||||||
- **uncensored**: Weniger restriktives Spezialprofil; Werkzeugrechte bleiben unverändert.
|
- **uncensored**: Weniger restriktives Spezialprofil; Werkzeugrechte bleiben unverändert.
|
||||||
@@ -0,0 +1,23 @@
|
|||||||
|
# Ultra-Vision mit CPU-Projektor
|
||||||
|
|
||||||
|
Stand: 24. September 2026. Ultra verwendet weiterhin Qwen3.8-27B IQ4_XS Pure
|
||||||
|
mit 262.144 Token Kontext. Der BF16-Vision-Projektor wird mit `--mmproj`
|
||||||
|
geladen und durch `--no-mmproj-offload` auf der CPU gehalten. So kann Ultra
|
||||||
|
Bilder verarbeiten, ohne den knapp bemessenen GPU-Speicher des 256K-Profils
|
||||||
|
zusätzlich mit dem Projektor zu belegen. Router und Profilmatrix melden Ultra
|
||||||
|
als vision-fähig.
|
||||||
|
|
||||||
|
Der frühere text-only-Modus war die Ursache dafür, dass der Router Bildanfragen
|
||||||
|
an `qwen-ultra` abwies. Ein Test über den Router nach dem Deployment lieferte
|
||||||
|
HTTP 200 für ein einzelnes synthetisches Farbbild (`Red`) und für fünf
|
||||||
|
synthetische Bilder (`5`). Der Ultra-Start lud den multimodalen Projektor.
|
||||||
|
Gemessen wurden 11,47 Sekunden bis zur Betriebsbereitschaft und 1,66 bzw.
|
||||||
|
0,9 Sekunden für die beiden kleinen Testanfragen. Große Screenshots und lange
|
||||||
|
Kontexte wurden damit nicht vermessen; deren Laufzeit kann deutlich höher
|
||||||
|
sein. Nach dem Test wurde Medium wieder aktiviert, Ultra ist gestoppt.
|
||||||
|
|
||||||
|
Vor der Änderung wurden die drei produktiven Dateien auf Athena unter
|
||||||
|
`/var/tmp/ultra-vision-cpu-20260924/` gesichert. Ein Rückbau muss Compose,
|
||||||
|
Router-Profilregister und Profilmatrix gemeinsam auf den vorigen Stand setzen
|
||||||
|
und den Router neu laden. Das Verzeichnis liegt nur temporär auf dem Host;
|
||||||
|
die dauerhafte Versionierung erfolgt im Git-Repository.
|
||||||
@@ -5,7 +5,7 @@ existing Athena speech stack. An experimental browser WebRTC path is available
|
|||||||
through the separate `services/athena-realtime-voice` service:
|
through the separate `services/athena-realtime-voice` service:
|
||||||
|
|
||||||
1. local VAD collects a spoken utterance,
|
1. local VAD collects a spoken utterance,
|
||||||
2. Athena Whisper transcribes it,
|
2. Athena Qwen3-ASR transcribes it,
|
||||||
3. OpenClaw's normal agent-consult path answers with its configured model and
|
3. OpenClaw's normal agent-consult path answers with its configured model and
|
||||||
tools,
|
tools,
|
||||||
4. Athena Qwen3-TTS returns PCM audio to the Talk client.
|
4. Athena Qwen3-TTS returns PCM audio to the Talk client.
|
||||||
@@ -60,23 +60,24 @@ openclaw plugins install . --force --accept-capabilities
|
|||||||
openclaw plugins inspect athena-talk --runtime --json
|
openclaw plugins inspect athena-talk --runtime --json
|
||||||
```
|
```
|
||||||
|
|
||||||
Version 1.3.0 also registers **Athena Whisper (Diktieren)** as a separate
|
Version 1.3.1 sends `qwen3-asr` throughout dictation and Talk. The plugin
|
||||||
|
registers **Athena Qwen3-ASR (Diktieren)** as a separate
|
||||||
realtime transcription provider through OpenClaw's official plugin API. In the
|
realtime transcription provider through OpenClaw's official plugin API. In the
|
||||||
browser composer, hold the microphone for dictation, then release it to send
|
browser composer, hold the microphone for dictation, then release it to send
|
||||||
the 8 kHz G.711 audio through the Gateway. Short recordings are converted to
|
the 8 kHz G.711 audio through the Gateway. Short recordings are converted to
|
||||||
PCM WAV and sent to Athena's existing `/audio/transcriptions` endpoint in one
|
PCM WAV and sent to Athena's existing `/audio/transcriptions` endpoint in one
|
||||||
request. Longer recordings are split while the user is still speaking into
|
request. Longer recordings are split while the user is still speaking into
|
||||||
six-second windows with 0.5 seconds of overlap. The plugin sends these windows
|
six-second windows with 0.5 seconds of overlap. The plugin sends these windows
|
||||||
sequentially to the persistent Whisper service, carries a short text prompt
|
sequentially to the persistent Qwen3-ASR service, removes duplicated overlap
|
||||||
into the next request, removes duplicated overlap words, and caches finished
|
words, and caches finished
|
||||||
segments until recording stops. Only the short final tail then remains inside
|
segments until recording stops. Only the short final tail then remains inside
|
||||||
OpenClaw's fixed five-second final-drain window. Each Whisper request is capped
|
OpenClaw's fixed five-second final-drain window. Each STT request is capped
|
||||||
at 4.5 seconds.
|
at 4.5 seconds.
|
||||||
|
|
||||||
This is incremental pre-transcription over OpenClaw's official transcription
|
This is incremental pre-transcription over OpenClaw's official transcription
|
||||||
provider API. Whisper.cpp still receives complete short WAV segments; it is
|
provider API. Qwen3-ASR still receives complete short WAV segments; it is
|
||||||
not a native token-streaming STT protocol. No OpenClaw core file was patched
|
not a native token-streaming STT protocol. No OpenClaw core file was patched.
|
||||||
and no additional speech container was introduced. The transcribed text is
|
The transcribed text is
|
||||||
returned to the composer; this path does not invoke the agent or TTS. The
|
returned to the composer; this path does not invoke the agent or TTS. The
|
||||||
provider reuses `talk.realtime.providers.athena-talk` and the configured model
|
provider reuses `talk.realtime.providers.athena-talk` and the configured model
|
||||||
provider for its URL/key. If that model provider has no key, it reuses
|
provider for its URL/key. If that model provider has no key, it reuses
|
||||||
@@ -85,7 +86,7 @@ origin. No second credential is needed. In `talk.catalog`, it appears under
|
|||||||
`transcription.providers`.
|
`transcription.providers`.
|
||||||
|
|
||||||
The production provider allows up to 180 seconds per recording. This limit is
|
The production provider allows up to 180 seconds per recording. This limit is
|
||||||
a local safety cap shared by dictation and Talk, not an OpenClaw or Whisper
|
a local safety cap shared by dictation and Talk, not an OpenClaw or Qwen3-ASR
|
||||||
restriction. Incremental segmentation keeps long dictation bounded while it
|
restriction. Incremental segmentation keeps long dictation bounded while it
|
||||||
is being recorded.
|
is being recorded.
|
||||||
|
|
||||||
@@ -95,7 +96,7 @@ An M4A voice note uploaded as a chat attachment does **not** use the realtime
|
|||||||
dictation provider above. OpenClaw processes it through its built-in
|
dictation provider above. OpenClaw processes it through its built-in
|
||||||
`tools.media.audio` path. On the Unraid installation, automatic provider
|
`tools.media.audio` path. On the Unraid installation, automatic provider
|
||||||
selection hit `SsrFBlockedError` for the private Athena address. Configure the
|
selection hit `SsrFBlockedError` for the private Athena address. Configure the
|
||||||
existing OpenAI-compatible provider and select Whisper explicitly:
|
existing OpenAI-compatible provider with the `qwen3-asr` model:
|
||||||
|
|
||||||
```json5
|
```json5
|
||||||
{
|
{
|
||||||
@@ -112,7 +113,7 @@ existing OpenAI-compatible provider and select Whisper explicitly:
|
|||||||
models: [
|
models: [
|
||||||
{
|
{
|
||||||
provider: "openai",
|
provider: "openai",
|
||||||
model: "whisper-1",
|
model: "qwen3-asr",
|
||||||
baseUrl: "http://192.168.1.212:8081/v1",
|
baseUrl: "http://192.168.1.212:8081/v1",
|
||||||
capabilities: ["audio"],
|
capabilities: ["audio"],
|
||||||
},
|
},
|
||||||
@@ -127,7 +128,7 @@ OpenClaw 2026.9.4 accepts `request.allowPrivateNetwork` under
|
|||||||
settings hot-reload without restarting the Gateway. The existing
|
settings hot-reload without restarting the Gateway. The existing
|
||||||
`ATHENA_ROUTER_API_KEY` SecretRef is reused; do not add another literal key.
|
`ATHENA_ROUTER_API_KEY` SecretRef is reused; do not add another literal key.
|
||||||
On 2026-09-16, OpenClaw's official audio module transcribed the affected M4A
|
On 2026-09-16, OpenClaw's official audio module transcribed the affected M4A
|
||||||
through Athena Whisper (127 characters returned). This confirms the endpoint
|
through Athena's then-active Whisper service (127 characters returned). This confirms the endpoint
|
||||||
and file format; a fresh attachment in the chat is still needed to verify the
|
and file format; a fresh attachment in the chat is still needed to verify the
|
||||||
full message-to-transcript flow.
|
full message-to-transcript flow.
|
||||||
|
|
||||||
|
|||||||
+6
-6
@@ -138,7 +138,7 @@ function wavFromPcm16(pcm, sampleRate = 24000) {
|
|||||||
return Buffer.concat([header, pcm]);
|
return Buffer.concat([header, pcm]);
|
||||||
}
|
}
|
||||||
// The browser transcription relay sends 8 kHz G.711 mu-law, while Athena's
|
// The browser transcription relay sends 8 kHz G.711 mu-law, while Athena's
|
||||||
// existing Whisper endpoint accepts PCM WAV uploads.
|
// transcription endpoint accepts PCM WAV uploads.
|
||||||
function wavFromMulaw8k(audio) {
|
function wavFromMulaw8k(audio) {
|
||||||
const pcm = Buffer.allocUnsafe(audio.length * 2);
|
const pcm = Buffer.allocUnsafe(audio.length * 2);
|
||||||
for (let i = 0; i < audio.length; i += 1) {
|
for (let i = 0; i < audio.length; i += 1) {
|
||||||
@@ -262,7 +262,7 @@ class AthenaTranscriptionSession {
|
|||||||
async transcribe(audio, prompt) {
|
async transcribe(audio, prompt) {
|
||||||
const form = new FormData();
|
const form = new FormData();
|
||||||
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
||||||
form.append("model", "whisper-1");
|
form.append("model", "qwen3-asr");
|
||||||
form.append("language", this.config.language);
|
form.append("language", this.config.language);
|
||||||
if (prompt)
|
if (prompt)
|
||||||
form.append("prompt", prompt);
|
form.append("prompt", prompt);
|
||||||
@@ -471,7 +471,7 @@ class AthenaTalkBridge {
|
|||||||
const form = new FormData();
|
const form = new FormData();
|
||||||
const wavBytes = Uint8Array.from(wavFromPcm16(pcm));
|
const wavBytes = Uint8Array.from(wavFromPcm16(pcm));
|
||||||
form.append("file", new Blob([wavBytes], { type: "audio/wav" }), "talk.wav");
|
form.append("file", new Blob([wavBytes], { type: "audio/wav" }), "talk.wav");
|
||||||
form.append("model", "whisper-1");
|
form.append("model", "qwen3-asr");
|
||||||
form.append("language", this.cfg.language);
|
form.append("language", this.cfg.language);
|
||||||
const response = await fetch(`${this.cfg.baseUrl}/audio/transcriptions`, {
|
const response = await fetch(`${this.cfg.baseUrl}/audio/transcriptions`, {
|
||||||
method: "POST",
|
method: "POST",
|
||||||
@@ -567,9 +567,9 @@ export default definePluginEntry({
|
|||||||
register(api) {
|
register(api) {
|
||||||
api.registerRealtimeTranscriptionProvider({
|
api.registerRealtimeTranscriptionProvider({
|
||||||
id: "athena-talk",
|
id: "athena-talk",
|
||||||
label: "Athena Whisper (Diktieren)",
|
label: "Athena Qwen3-ASR (Diktieren)",
|
||||||
defaultModel: "whisper-1",
|
defaultModel: "qwen3-asr",
|
||||||
models: ["whisper-1"],
|
models: ["qwen3-asr"],
|
||||||
autoSelectOrder: 1,
|
autoSelectOrder: 1,
|
||||||
resolveConfig: ({ cfg, rawConfig }) => resolveTranscriptionConfig(cfg, rawConfig),
|
resolveConfig: ({ cfg, rawConfig }) => resolveTranscriptionConfig(cfg, rawConfig),
|
||||||
isConfigured: ({ providerConfig }) => Boolean(record(providerConfig).baseUrl),
|
isConfigured: ({ providerConfig }) => Boolean(record(providerConfig).baseUrl),
|
||||||
|
|||||||
@@ -158,7 +158,7 @@ function wavFromPcm16(pcm: Buffer, sampleRate = 24000): Buffer {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// The browser transcription relay sends 8 kHz G.711 mu-law, while Athena's
|
// The browser transcription relay sends 8 kHz G.711 mu-law, while Athena's
|
||||||
// existing Whisper endpoint accepts PCM WAV uploads.
|
// transcription endpoint accepts PCM WAV uploads.
|
||||||
function wavFromMulaw8k(audio: Buffer): Buffer {
|
function wavFromMulaw8k(audio: Buffer): Buffer {
|
||||||
const pcm = Buffer.allocUnsafe(audio.length * 2);
|
const pcm = Buffer.allocUnsafe(audio.length * 2);
|
||||||
for (let i = 0; i < audio.length; i += 1) {
|
for (let i = 0; i < audio.length; i += 1) {
|
||||||
@@ -289,7 +289,7 @@ class AthenaTranscriptionSession {
|
|||||||
private async transcribe(audio: Buffer, prompt: string): Promise<string> {
|
private async transcribe(audio: Buffer, prompt: string): Promise<string> {
|
||||||
const form = new FormData();
|
const form = new FormData();
|
||||||
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
||||||
form.append("model", "whisper-1");
|
form.append("model", "qwen3-asr");
|
||||||
form.append("language", this.config.language);
|
form.append("language", this.config.language);
|
||||||
if (prompt) form.append("prompt", prompt);
|
if (prompt) form.append("prompt", prompt);
|
||||||
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
|
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
|
||||||
@@ -490,7 +490,7 @@ class AthenaTalkBridge {
|
|||||||
const form = new FormData();
|
const form = new FormData();
|
||||||
const wavBytes = Uint8Array.from(wavFromPcm16(pcm));
|
const wavBytes = Uint8Array.from(wavFromPcm16(pcm));
|
||||||
form.append("file", new Blob([wavBytes], { type: "audio/wav" }), "talk.wav");
|
form.append("file", new Blob([wavBytes], { type: "audio/wav" }), "talk.wav");
|
||||||
form.append("model", "whisper-1");
|
form.append("model", "qwen3-asr");
|
||||||
form.append("language", this.cfg.language);
|
form.append("language", this.cfg.language);
|
||||||
const response = await fetch(`${this.cfg.baseUrl}/audio/transcriptions`, {
|
const response = await fetch(`${this.cfg.baseUrl}/audio/transcriptions`, {
|
||||||
method: "POST",
|
method: "POST",
|
||||||
@@ -578,9 +578,9 @@ export default definePluginEntry({
|
|||||||
register(api) {
|
register(api) {
|
||||||
api.registerRealtimeTranscriptionProvider({
|
api.registerRealtimeTranscriptionProvider({
|
||||||
id: "athena-talk",
|
id: "athena-talk",
|
||||||
label: "Athena Whisper (Diktieren)",
|
label: "Athena Qwen3-ASR (Diktieren)",
|
||||||
defaultModel: "whisper-1",
|
defaultModel: "qwen3-asr",
|
||||||
models: ["whisper-1"],
|
models: ["qwen3-asr"],
|
||||||
autoSelectOrder: 1,
|
autoSelectOrder: 1,
|
||||||
resolveConfig: ({ cfg, rawConfig }) => resolveTranscriptionConfig(cfg, rawConfig),
|
resolveConfig: ({ cfg, rawConfig }) => resolveTranscriptionConfig(cfg, rawConfig),
|
||||||
isConfigured: ({ providerConfig }) => Boolean(record(providerConfig).baseUrl),
|
isConfigured: ({ providerConfig }) => Boolean(record(providerConfig).baseUrl),
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
{
|
{
|
||||||
"id": "athena-talk",
|
"id": "athena-talk",
|
||||||
"name": "Athena Local Talk",
|
"name": "Athena Local Talk",
|
||||||
"description": "Private OpenClaw Talk provider using Athena Whisper and Qwen3-TTS.",
|
"description": "Private OpenClaw Talk provider using Athena Qwen3-ASR and Qwen3-TTS.",
|
||||||
"activation": {
|
"activation": {
|
||||||
"onStartup": true
|
"onStartup": true
|
||||||
},
|
},
|
||||||
|
|||||||
+2
-2
@@ -1,12 +1,12 @@
|
|||||||
{
|
{
|
||||||
"name": "@casaderoll/openclaw-athena-talk",
|
"name": "@casaderoll/openclaw-athena-talk",
|
||||||
"version": "1.3.0",
|
"version": "1.3.1",
|
||||||
"lockfileVersion": 3,
|
"lockfileVersion": 3,
|
||||||
"requires": true,
|
"requires": true,
|
||||||
"packages": {
|
"packages": {
|
||||||
"": {
|
"": {
|
||||||
"name": "@casaderoll/openclaw-athena-talk",
|
"name": "@casaderoll/openclaw-athena-talk",
|
||||||
"version": "1.3.0",
|
"version": "1.3.1",
|
||||||
"devDependencies": {
|
"devDependencies": {
|
||||||
"@types/node": "^24.0.0",
|
"@types/node": "^24.0.0",
|
||||||
"openclaw": "2026.9.4",
|
"openclaw": "2026.9.4",
|
||||||
|
|||||||
@@ -1,8 +1,8 @@
|
|||||||
{
|
{
|
||||||
"name": "@casaderoll/openclaw-athena-talk",
|
"name": "@casaderoll/openclaw-athena-talk",
|
||||||
"version": "1.3.0",
|
"version": "1.3.1",
|
||||||
"private": true,
|
"private": true,
|
||||||
"description": "Local OpenClaw Talk provider backed by Athena Whisper and Qwen3-TTS",
|
"description": "Local OpenClaw Talk provider backed by Athena Qwen3-ASR and Qwen3-TTS",
|
||||||
"type": "module",
|
"type": "module",
|
||||||
"files": [
|
"files": [
|
||||||
"dist",
|
"dist",
|
||||||
|
|||||||
@@ -11,7 +11,7 @@ async function waitFor(predicate, timeoutMs = 2000) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
test("dictation registers separately and sends G.711 audio to Athena Whisper", async () => {
|
test("dictation registers separately and sends G.711 audio to Athena Qwen3-ASR", async () => {
|
||||||
let transcription;
|
let transcription;
|
||||||
plugin.register({
|
plugin.register({
|
||||||
registerRealtimeTranscriptionProvider: (value) => { transcription = value; },
|
registerRealtimeTranscriptionProvider: (value) => { transcription = value; },
|
||||||
@@ -37,7 +37,7 @@ test("dictation registers separately and sends G.711 audio to Athena Whisper", a
|
|||||||
assert.equal(wav.readUInt16LE(34), 16);
|
assert.equal(wav.readUInt16LE(34), 16);
|
||||||
assert.equal(wav.length, 48);
|
assert.equal(wav.length, 48);
|
||||||
assert.equal(form.get("language"), "de");
|
assert.equal(form.get("language"), "de");
|
||||||
assert.equal(form.get("model"), "whisper-1");
|
assert.equal(form.get("model"), "qwen3-asr");
|
||||||
res.writeHead(200, { "Content-Type": "application/json" }).end(JSON.stringify({ text: "Hallo Athena" }));
|
res.writeHead(200, { "Content-Type": "application/json" }).end(JSON.stringify({ text: "Hallo Athena" }));
|
||||||
});
|
});
|
||||||
await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve));
|
await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve));
|
||||||
|
|||||||
@@ -52,7 +52,8 @@ case "$command" in
|
|||||||
[[ $# -eq 0 ]] || { echo "core akzeptiert keine weiteren Services" >&2; exit 2; }
|
[[ $# -eq 0 ]] || { echo "core akzeptiert keine weiteren Services" >&2; exit 2; }
|
||||||
run "$ROOT_DIR/platform/mcp/install-tools.sh"
|
run "$ROOT_DIR/platform/mcp/install-tools.sh"
|
||||||
run "${compose[@]}" up -d --build \
|
run "${compose[@]}" up -d --build \
|
||||||
wireguard-gateway embedding qwen3-tts tts-gateway profile-controller router llama-dashboard portainer backup
|
wireguard-gateway embedding qwen3-tts tts-gateway profile-controller \
|
||||||
|
qwen-asr qwen-asr-worker router llama-dashboard portainer backup
|
||||||
else
|
else
|
||||||
run "${compose[@]}" up -d --build --no-deps "$@"
|
run "${compose[@]}" up -d --build --no-deps "$@"
|
||||||
fi
|
fi
|
||||||
|
|||||||
@@ -0,0 +1,17 @@
|
|||||||
|
FROM python:3.13.7-slim-bookworm
|
||||||
|
|
||||||
|
RUN apt-get update \
|
||||||
|
&& apt-get install -y --no-install-recommends ffmpeg \
|
||||||
|
&& rm -rf /var/lib/apt/lists/* \
|
||||||
|
&& useradd --system --uid 10005 --home-dir /nonexistent --shell /usr/sbin/nologin stt
|
||||||
|
|
||||||
|
WORKDIR /app
|
||||||
|
COPY router/qwen_asr_worker.py /app/qwen_asr_worker.py
|
||||||
|
|
||||||
|
ENV QWEN_ASR_HOST=0.0.0.0 \
|
||||||
|
QWEN_ASR_PORT=8084 \
|
||||||
|
QWEN_ASR_LANGUAGE=de
|
||||||
|
|
||||||
|
USER 10005:10005
|
||||||
|
EXPOSE 8084
|
||||||
|
ENTRYPOINT ["python", "/app/qwen_asr_worker.py"]
|
||||||
@@ -133,8 +133,8 @@ display:
|
|||||||
language: "de"
|
language: "de"
|
||||||
show_reasoning: false
|
show_reasoning: false
|
||||||
|
|
||||||
# Speech input stays local and private. A fixed German hint avoids Whisper
|
# Speech input stays local and private. The fixed German hint helps with
|
||||||
# interpreting short utterances as English while retaining the fast base model.
|
# short utterances and product names.
|
||||||
stt:
|
stt:
|
||||||
enabled: true
|
enabled: true
|
||||||
provider: "local"
|
provider: "local"
|
||||||
|
|||||||
@@ -32,7 +32,7 @@ gateway, UI, CPU-STT, backup and operator containers may remain active.
|
|||||||
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
|
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
|
||||||
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
|
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
|
||||||
- Large: the same Q4 model, 192,000 tokens, vision.
|
- Large: the same Q4 model, 192,000 tokens, vision.
|
||||||
- Ultra: the same Q4 model, 262,144 tokens, no vision projector.
|
- Ultra: the same Q4 model, 262,144 tokens, vision projector on CPU.
|
||||||
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
|
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
|
||||||
|
|
||||||
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
|
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
|
||||||
|
|||||||
@@ -16,10 +16,10 @@ nach Standardbenchmark, Tool-Calling-Test und Kontexttest übernommen.
|
|||||||
|
|
||||||
- Fast: Qwen3.8-27B IQ4-MIX mit MTP2
|
- Fast: Qwen3.8-27B IQ4-MIX mit MTP2
|
||||||
- Medium und Large: Qwen3.8-27B IQ4_XS Pure mit MTP3
|
- Medium und Large: Qwen3.8-27B IQ4_XS Pure mit MTP3
|
||||||
- Ultra: Qwen3.8-27B IQ4_XS Pure mit MTP2 und maximalem Textkontext
|
- Ultra: Qwen3.8-27B IQ4_XS Pure mit MTP2, maximalem Kontext und Vision-Projektor auf CPU
|
||||||
- Uncensored: Blackfrost Qwen3.8-27B Abliterated Q4_K_M mit MTP2
|
- Uncensored: Blackfrost Qwen3.8-27B Abliterated Q4_K_M mit MTP2
|
||||||
- Fast, Medium, Large und Uncensored: integrierte Vision; der jeweilige
|
- Alle fünf Profile: integrierte Vision. Bei Fast, Medium, Large und
|
||||||
Projektor liegt vollständig auf der RTX 3060
|
Uncensored liegt der Projektor auf der RTX 3060; bei Ultra auf der CPU.
|
||||||
|
|
||||||
Die produktive Runtime ist seit dem 12. September 2026 auf llama.cpp
|
Die produktive Runtime ist seit dem 12. September 2026 auf llama.cpp
|
||||||
**Build 10930**, Commit `56381e407c0ccfb3a6f71e668a27a901001d22ce`,
|
**Build 10930**, Commit `56381e407c0ccfb3a6f71e668a27a901001d22ce`,
|
||||||
|
|||||||
@@ -3,4 +3,4 @@ Description=Legacy native Qwen Ultra 256K profile (Docker is the production path
|
|||||||
|
|
||||||
[Service]
|
[Service]
|
||||||
ExecStart=
|
ExecStart=
|
||||||
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --alias qwen-ultra --ctx-size 262144 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --load-mode none --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
|
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-ultra --ctx-size 262144 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --load-mode none --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
|
||||||
@@ -29,12 +29,11 @@ Sprachausgabe (Qwen3-TTS auf RTX 3060):
|
|||||||
POST /v1/audio/speech (OpenAI-kompatibel)
|
POST /v1/audio/speech (OpenAI-kompatibel)
|
||||||
GET /v1/audio/voices (verfügbare Stimmen)
|
GET /v1/audio/voices (verfügbare Stimmen)
|
||||||
|
|
||||||
Spracherkennung (whisper.cpp, deutsch, CPU-only):
|
Spracherkennung (Qwen3-ASR, deutsch, CPU-only):
|
||||||
POST /v1/audio/transcriptions (OpenAI-kompatibel)
|
POST /v1/audio/transcriptions (OpenAI-kompatibel)
|
||||||
GET /v1/audio/models (verfügbare Audio-Modelle)
|
GET /v1/audio/models (verfügbare Audio-Modelle)
|
||||||
|
|
||||||
Der TTS-Worker (mike-ai-xtts.service) und der STT-Worker
|
TTS und Qwen3-ASR laufen als separate, langlebige Dienste.
|
||||||
(mike-ai-whisper.service) laufen als separate, langlebige Prozesse.
|
|
||||||
Der Router leitet /v1/audio/speech und /v1/audio/transcriptions
|
Der Router leitet /v1/audio/speech und /v1/audio/transcriptions
|
||||||
per HTTP an die Worker weiter.
|
per HTTP an die Worker weiter.
|
||||||
|
|
||||||
@@ -216,11 +215,11 @@ TTS_DEFAULT_VOICE = os.environ.get(
|
|||||||
TTS_FORMATS = ("mp3", "wav", "pcm")
|
TTS_FORMATS = ("mp3", "wav", "pcm")
|
||||||
TTS_DEFAULT_FORMAT = "mp3"
|
TTS_DEFAULT_FORMAT = "mp3"
|
||||||
|
|
||||||
# --- Spracherkennung (whisper.cpp, deutsch, CPU-only) ---
|
# --- Spracherkennung (Qwen3-ASR, deutsch, CPU-only) ---
|
||||||
STT_WORKER_URL = os.environ.get("STT_WORKER_URL", "http://127.0.0.1:8084")
|
STT_WORKER_URL = os.environ.get("STT_WORKER_URL", "http://127.0.0.1:8084")
|
||||||
STT_TIMEOUT = float(os.environ.get("STT_TIMEOUT", "120")) # s, pro Transkription
|
STT_TIMEOUT = float(os.environ.get("STT_TIMEOUT", "120")) # s, pro Transkription
|
||||||
STT_CONNECT_TIMEOUT = float(os.environ.get("STT_CONNECT_TIMEOUT", "5"))
|
STT_CONNECT_TIMEOUT = float(os.environ.get("STT_CONNECT_TIMEOUT", "5"))
|
||||||
STT_MODEL = "whisper-1" # virtuelles Modell für /v1/audio/transcriptions
|
STT_MODEL = "qwen3-asr"
|
||||||
|
|
||||||
# Maximale Upload-Größe (Bytes) – verhindert unbegrenzten RAM-Verbrauch.
|
# Maximale Upload-Größe (Bytes) – verhindert unbegrenzten RAM-Verbrauch.
|
||||||
# 50 MB ist für Audio-Dateien (WebM/Opus, WAV, MP3) mehr als ausreichend.
|
# 50 MB ist für Audio-Dateien (WebM/Opus, WAV, MP3) mehr als ausreichend.
|
||||||
@@ -2846,7 +2845,7 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
models.append({
|
models.append({
|
||||||
"id": STT_MODEL,
|
"id": STT_MODEL,
|
||||||
"object": "model",
|
"object": "model",
|
||||||
"owned_by": "whisper.cpp",
|
"owned_by": "qwen3-asr",
|
||||||
"type": "transcription",
|
"type": "transcription",
|
||||||
})
|
})
|
||||||
if tts.get("ready"):
|
if tts.get("ready"):
|
||||||
@@ -2969,7 +2968,7 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
|
|
||||||
# Modell-Validierung
|
# Modell-Validierung
|
||||||
model = fields.get("model", STT_MODEL)
|
model = fields.get("model", STT_MODEL)
|
||||||
if model not in (STT_MODEL, "whisper"):
|
if model != STT_MODEL:
|
||||||
self._send_error(400, f"unbekanntes Modell: {model!r} "
|
self._send_error(400, f"unbekanntes Modell: {model!r} "
|
||||||
f"(erwartet: {STT_MODEL})",
|
f"(erwartet: {STT_MODEL})",
|
||||||
"invalid_request_error", "unknown_model")
|
"invalid_request_error", "unknown_model")
|
||||||
|
|||||||
@@ -0,0 +1,147 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""OpenAI-router STT adapter for the persistent, CPU-only Qwen3-ASR server."""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import subprocess
|
||||||
|
import tempfile
|
||||||
|
import time
|
||||||
|
import urllib.request
|
||||||
|
import uuid
|
||||||
|
from email import policy
|
||||||
|
from email.parser import BytesParser
|
||||||
|
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||||
|
|
||||||
|
|
||||||
|
HOST = os.environ.get("QWEN_ASR_HOST", "0.0.0.0")
|
||||||
|
PORT = int(os.environ.get("QWEN_ASR_PORT", "8084"))
|
||||||
|
SERVER_URL = os.environ.get("QWEN_ASR_SERVER_URL", "http://qwen-asr:8080").rstrip("/")
|
||||||
|
LANGUAGE = os.environ.get("QWEN_ASR_LANGUAGE", "de")
|
||||||
|
MAX_BODY_BYTES = 25 * 1024 * 1024
|
||||||
|
|
||||||
|
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
|
||||||
|
log = logging.getLogger("qwen-asr-worker")
|
||||||
|
|
||||||
|
|
||||||
|
def clean_transcript(value: str) -> str:
|
||||||
|
"""llama.cpp may include a Qwen task marker before the spoken words."""
|
||||||
|
if "<asr_text>" in value:
|
||||||
|
value = value.split("<asr_text>", 1)[1]
|
||||||
|
return value.replace("<|endoftext|>", "").strip()
|
||||||
|
|
||||||
|
|
||||||
|
def transcribe(audio: bytes, filename: str, language: str) -> dict:
|
||||||
|
suffix = os.path.splitext(filename)[1].lower() or ".wav"
|
||||||
|
with tempfile.TemporaryDirectory(prefix="qwen_asr_") as directory:
|
||||||
|
source = os.path.join(directory, "input" + suffix)
|
||||||
|
wav = os.path.join(directory, "audio.wav")
|
||||||
|
with open(source, "wb") as handle:
|
||||||
|
handle.write(audio)
|
||||||
|
result = subprocess.run(
|
||||||
|
["ffmpeg", "-nostdin", "-hide_banner", "-loglevel", "error", "-y",
|
||||||
|
"-i", source, "-ac", "1", "-ar", "16000", "-c:a", "pcm_s16le", wav],
|
||||||
|
capture_output=True, text=True, timeout=30,
|
||||||
|
)
|
||||||
|
if result.returncode:
|
||||||
|
raise ValueError("Audio konnte nicht gelesen werden: " + result.stderr[-300:])
|
||||||
|
with open(wav, "rb") as handle:
|
||||||
|
pcm = handle.read()
|
||||||
|
|
||||||
|
boundary = "athena-qwen-asr-" + uuid.uuid4().hex
|
||||||
|
body = b"".join([
|
||||||
|
f"--{boundary}\r\n".encode(),
|
||||||
|
b'Content-Disposition: form-data; name="file"; filename="audio.wav"\r\n',
|
||||||
|
b"Content-Type: audio/wav\r\n\r\n", pcm, b"\r\n",
|
||||||
|
f"--{boundary}\r\n".encode(),
|
||||||
|
b'Content-Disposition: form-data; name="model"\r\n\r\n',
|
||||||
|
b"qwen3-asr-0.6b\r\n",
|
||||||
|
f"--{boundary}\r\n".encode(),
|
||||||
|
b'Content-Disposition: form-data; name="language"\r\n\r\n',
|
||||||
|
language.encode(), b"\r\n",
|
||||||
|
f"--{boundary}--\r\n".encode(),
|
||||||
|
])
|
||||||
|
request = urllib.request.Request(
|
||||||
|
SERVER_URL + "/v1/audio/transcriptions", data=body,
|
||||||
|
headers={"Content-Type": f"multipart/form-data; boundary={boundary}"},
|
||||||
|
method="POST",
|
||||||
|
)
|
||||||
|
started = time.monotonic()
|
||||||
|
with urllib.request.urlopen(request, timeout=60) as response:
|
||||||
|
payload = json.loads(response.read())
|
||||||
|
if not isinstance(payload, dict) or not isinstance(payload.get("text"), str):
|
||||||
|
raise RuntimeError("Qwen3-ASR returned no transcription")
|
||||||
|
text = clean_transcript(payload["text"])
|
||||||
|
elapsed = int((time.monotonic() - started) * 1000)
|
||||||
|
log.info("Qwen3-ASR transcribed %d characters in %d ms", len(text), elapsed)
|
||||||
|
return {"text": text, "language": language, "duration_ms": elapsed,
|
||||||
|
"engine": "qwen3-asr-0.6b"}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_audio(body: bytes, content_type: str) -> tuple[bytes, str, str]:
|
||||||
|
if "multipart/form-data" not in content_type.lower():
|
||||||
|
return body, "audio.wav", LANGUAGE
|
||||||
|
message = BytesParser(policy=policy.default).parsebytes(
|
||||||
|
b"MIME-Version: 1.0\r\nContent-Type: " + content_type.encode() +
|
||||||
|
b"\r\n\r\n" + body
|
||||||
|
)
|
||||||
|
if not message.is_multipart():
|
||||||
|
raise ValueError("Invalid multipart upload")
|
||||||
|
audio = b""
|
||||||
|
filename = "audio.wav"
|
||||||
|
language = LANGUAGE
|
||||||
|
for part in message.iter_parts():
|
||||||
|
name = part.get_param("name", header="content-disposition")
|
||||||
|
if name == "file":
|
||||||
|
audio = part.get_payload(decode=True) or b""
|
||||||
|
filename = os.path.basename(part.get_filename() or filename)
|
||||||
|
elif name == "language":
|
||||||
|
language = (part.get_payload(decode=True) or b"").decode("utf-8").strip()
|
||||||
|
return audio, filename, language if language and language != "auto" else LANGUAGE
|
||||||
|
|
||||||
|
|
||||||
|
class Handler(BaseHTTPRequestHandler):
|
||||||
|
def send_json(self, status: int, data: dict) -> None:
|
||||||
|
body = json.dumps(data, ensure_ascii=False).encode()
|
||||||
|
self.send_response(status)
|
||||||
|
self.send_header("Content-Type", "application/json; charset=utf-8")
|
||||||
|
self.send_header("Content-Length", str(len(body)))
|
||||||
|
self.end_headers()
|
||||||
|
self.wfile.write(body)
|
||||||
|
|
||||||
|
def do_GET(self) -> None:
|
||||||
|
if self.path != "/status":
|
||||||
|
self.send_json(404, {"error": "not found"})
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
with urllib.request.urlopen(SERVER_URL + "/health", timeout=2) as response:
|
||||||
|
ready = response.status == 200
|
||||||
|
except Exception:
|
||||||
|
ready = False
|
||||||
|
self.send_json(200, {"ready": ready, "model": "qwen3-asr-0.6b",
|
||||||
|
"engine": "qwen3-asr", "language": LANGUAGE})
|
||||||
|
|
||||||
|
def do_POST(self) -> None:
|
||||||
|
if self.path != "/transcribe":
|
||||||
|
self.send_json(404, {"error": "not found"})
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
size = int(self.headers.get("Content-Length", "0"))
|
||||||
|
if not 0 < size <= MAX_BODY_BYTES:
|
||||||
|
self.send_json(413, {"error": "Invalid audio size"})
|
||||||
|
return
|
||||||
|
audio, filename, language = parse_audio(
|
||||||
|
self.rfile.read(size), self.headers.get("Content-Type", "")
|
||||||
|
)
|
||||||
|
if not audio:
|
||||||
|
raise ValueError("Missing audio file")
|
||||||
|
self.send_json(200, transcribe(audio, filename, language))
|
||||||
|
except ValueError as exc:
|
||||||
|
self.send_json(400, {"error": str(exc)})
|
||||||
|
except Exception:
|
||||||
|
log.exception("Transcription failed")
|
||||||
|
self.send_json(503, {"error": "Qwen3-ASR unavailable"})
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
ThreadingHTTPServer((HOST, PORT), Handler).serve_forever()
|
||||||
@@ -18,7 +18,7 @@
|
|||||||
"ultra": {
|
"ultra": {
|
||||||
"context": 262144,
|
"context": 262144,
|
||||||
"model_alias": "qwen-ultra",
|
"model_alias": "qwen-ultra",
|
||||||
"vision": false
|
"vision": true
|
||||||
},
|
},
|
||||||
"uncensored": {
|
"uncensored": {
|
||||||
"context": 80000,
|
"context": 80000,
|
||||||
|
|||||||
@@ -47,7 +47,7 @@ def main() -> int:
|
|||||||
body = (
|
body = (
|
||||||
f"--{boundary}\r\n"
|
f"--{boundary}\r\n"
|
||||||
'Content-Disposition: form-data; name="model"\r\n\r\n'
|
'Content-Disposition: form-data; name="model"\r\n\r\n'
|
||||||
"whisper-1\r\n"
|
"qwen3-asr\r\n"
|
||||||
f"--{boundary}\r\n"
|
f"--{boundary}\r\n"
|
||||||
'Content-Disposition: form-data; name="language"\r\n\r\n'
|
'Content-Disposition: form-data; name="language"\r\n\r\n'
|
||||||
"de\r\n"
|
"de\r\n"
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
# Athena realtime voice bridge
|
# Athena realtime voice bridge
|
||||||
|
|
||||||
This independent service lets OpenClaw's existing browser Talk UI use Athena
|
This independent service lets OpenClaw's existing browser Talk UI use Athena
|
||||||
Whisper, OpenClaw's agent, and Athena Qwen3-TTS through the browser's supported
|
Qwen3-ASR, OpenClaw's agent, and Athena Qwen3-TTS through the browser's supported
|
||||||
OpenAI-style WebRTC transport. It does not modify OpenClaw or switch an Athena
|
OpenAI-style WebRTC transport. It does not modify OpenClaw or switch an Athena
|
||||||
profile. The existing `gateway-relay` path remains available.
|
profile. The existing `gateway-relay` path remains available.
|
||||||
|
|
||||||
@@ -90,7 +90,7 @@ python3.11 -m venv .venv
|
|||||||
```
|
```
|
||||||
|
|
||||||
The synthetic microphone test passed over the actual OpenClaw HTTPS offer
|
The synthetic microphone test passed over the actual OpenClaw HTTPS offer
|
||||||
route and WireGuard media path with production Whisper and Qwen3-TTS on
|
route and WireGuard media path with production Qwen3-ASR and Qwen3-TTS on
|
||||||
2026-09-16. Subsequent real browser Talk sessions successfully transcribed
|
2026-09-16. Subsequent real browser Talk sessions successfully transcribed
|
||||||
and answered multiple user turns. Browser dictation uses a separate plugin
|
and answered multiple user turns. Browser dictation uses a separate plugin
|
||||||
path. Since version 1.3.0, recordings longer than six seconds are
|
path. Since version 1.3.0, recordings longer than six seconds are
|
||||||
|
|||||||
@@ -281,7 +281,7 @@ class RealtimeSession:
|
|||||||
try:
|
try:
|
||||||
form = FormData()
|
form = FormData()
|
||||||
form.add_field("file", make_wav(pcm), filename="talk.wav", content_type="audio/wav")
|
form.add_field("file", make_wav(pcm), filename="talk.wav", content_type="audio/wav")
|
||||||
form.add_field("model", "whisper-1")
|
form.add_field("model", "qwen3-asr")
|
||||||
form.add_field("language", os.environ.get("STT_LANGUAGE", "de"))
|
form.add_field("language", os.environ.get("STT_LANGUAGE", "de"))
|
||||||
async with self.http.post(
|
async with self.http.post(
|
||||||
os.environ["ATHENA_API_BASE_URL"].rstrip("/") + "/audio/transcriptions",
|
os.environ["ATHENA_API_BASE_URL"].rstrip("/") + "/audio/transcriptions",
|
||||||
|
|||||||
+4
-1
@@ -27,6 +27,8 @@ for name in \
|
|||||||
mike-ai-wireguard-gateway \
|
mike-ai-wireguard-gateway \
|
||||||
mike-ai-profile-controller \
|
mike-ai-profile-controller \
|
||||||
mike-ai-embedding \
|
mike-ai-embedding \
|
||||||
|
mike-ai-qwen-asr \
|
||||||
|
mike-ai-qwen-asr-worker \
|
||||||
mike-ai-router \
|
mike-ai-router \
|
||||||
mike-ai-qwen3-tts \
|
mike-ai-qwen3-tts \
|
||||||
mike-ai-tts-gateway \
|
mike-ai-tts-gateway \
|
||||||
@@ -70,7 +72,8 @@ for legacy in \
|
|||||||
mike-ai-mcp-deemix \
|
mike-ai-mcp-deemix \
|
||||||
mike-ai-mcp-github \
|
mike-ai-mcp-github \
|
||||||
mike-ai-mcp-homeassistant \
|
mike-ai-mcp-homeassistant \
|
||||||
mike-ai-mcp-navidrome; do
|
mike-ai-mcp-navidrome \
|
||||||
|
mike-ai-whisper; do
|
||||||
[[ -z $(docker inspect -f '{{.Name}}' "$legacy" 2>/dev/null || true) ]] || \
|
[[ -z $(docker inspect -f '{{.Name}}' "$legacy" 2>/dev/null || true) ]] || \
|
||||||
fail "Altlast existiert noch: $legacy"
|
fail "Altlast existiert noch: $legacy"
|
||||||
done
|
done
|
||||||
|
|||||||
Reference in new issue
Block a user