Compare commits
10
Commits
57309f3722
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
da10f6b48d | ||
|
|
8a323e5b9e | ||
|
|
6378b50086 | ||
|
|
33c04150e3 | ||
|
|
46e5bbdf7f | ||
|
|
e577c55489 | ||
|
|
bb06545956 | ||
|
|
fd3f3f5979 | ||
|
|
5106d0d2ed | ||
|
|
b832157226 |
No files matched your search
@@ -4,6 +4,8 @@ Stand: **20. September 2026**, auf Athena geprüft. Die fünf Textprofil-Images
|
|||||||
tragen **llama.cpp b29c606** (0.4.1).
|
tragen **llama.cpp b29c606** (0.4.1).
|
||||||
Der [geprüfte Live-Stand](docs/LIVE_STATE.md) beschreibt Profile, GPUs und die
|
Der [geprüfte Live-Stand](docs/LIVE_STATE.md) beschreibt Profile, GPUs und die
|
||||||
Abweichung zwischen dem bereitgestellten Stack und dem Git-Checkout.
|
Abweichung zwischen dem bereitgestellten Stack und dem Git-Checkout.
|
||||||
|
Seit dem 24. September verarbeitet auch Ultra Bilder; sein Vision-Projektor
|
||||||
|
läuft auf der CPU. [Änderung und Test](docs/ULTRA_CPU_VISION_20260924.md).
|
||||||
Der [Updatebericht vom 15. September](docs/UPDATE_AUDIT_20260915.md)
|
Der [Updatebericht vom 15. September](docs/UPDATE_AUDIT_20260915.md)
|
||||||
enthält die aktuellen Build- und Testbelege. Der ältere
|
enthält die aktuellen Build- und Testbelege. Der ältere
|
||||||
[b10930-Bericht](docs/LLAMA_B10930_UPDATE_20260912.md) dokumentiert einen
|
[b10930-Bericht](docs/LLAMA_B10930_UPDATE_20260912.md) dokumentiert einen
|
||||||
@@ -22,7 +24,7 @@ Sie betreibt:
|
|||||||
- den OpenAI-kompatiblen Profile Router,
|
- den OpenAI-kompatiblen Profile Router,
|
||||||
- Qwen-Image-2.1 INT8 für Textbilder und Referenzbild-Bearbeitung,
|
- Qwen-Image-2.1 INT8 für Textbilder und Referenzbild-Bearbeitung,
|
||||||
- Qwen3-TTS und TTS-Gateway für Sprache,
|
- Qwen3-TTS und TTS-Gateway für Sprache,
|
||||||
- Whisper.cpp und die WebRTC-Brücke für OpenClaw Talk,
|
- Qwen3-ASR auf der CPU und die WebRTC-Brücke für OpenClaw Talk,
|
||||||
- EmbeddingGemma auf der CPU für OpenClaws semantische Memory-Suche,
|
- EmbeddingGemma auf der CPU für OpenClaws semantische Memory-Suche,
|
||||||
- die GPU-lose Mikes-Applio-UI als gesonderten Checkout,
|
- die GPU-lose Mikes-Applio-UI als gesonderten Checkout,
|
||||||
- das Athena-Dashboard,
|
- das Athena-Dashboard,
|
||||||
@@ -93,6 +95,8 @@ nicht direkt. Kein automatischer Host-Neustart ist vorgesehen.
|
|||||||
- Qwen-Image-2.1 INT8: Bildgenerierung und Editing auf der RTX 5080; das Textmodell wird dafür kurz entladen und danach
|
- Qwen-Image-2.1 INT8: Bildgenerierung und Editing auf der RTX 5080; das Textmodell wird dafür kurz entladen und danach
|
||||||
automatisch wiederhergestellt
|
automatisch wiederhergestellt
|
||||||
- Qwen3-TTS 1.7B: RTX 3060; kein Piper-Fallback
|
- Qwen3-TTS 1.7B: RTX 3060; kein Piper-Fallback
|
||||||
|
- Qwen3-ASR 0.6B Q8: CPU, hinter dem OpenAI-kompatiblen
|
||||||
|
Transkriptionsendpunkt mit dem Modellnamen `qwen3-asr`.
|
||||||
- EmbeddingGemma 300M Q8: CPU, OpenAI-kompatibel auf Port 8082; kein GPU-Zugriff
|
- EmbeddingGemma 300M Q8: CPU, OpenAI-kompatibel auf Port 8082; kein GPU-Zugriff
|
||||||
|
|
||||||
Die geprüften Live-Werte stehen in [docs/LIVE_STATE.md](docs/LIVE_STATE.md).
|
Die geprüften Live-Werte stehen in [docs/LIVE_STATE.md](docs/LIVE_STATE.md).
|
||||||
|
|||||||
@@ -22,9 +22,11 @@ dokumentieren den ausgerollten Stand, reduzierte Statuslatenzen und Tests.
|
|||||||
- genau ein aktives llama.cpp-Profil: Fast, Medium, Large, Ultra oder Uncensored
|
- genau ein aktives llama.cpp-Profil: Fast, Medium, Large, Ultra oder Uncensored
|
||||||
- Profile Router auf Port 8081
|
- Profile Router auf Port 8081
|
||||||
- Qwen-Image-2.1 INT8 als produktiver Bildworker auf der RTX 5080
|
- Qwen-Image-2.1 INT8 als produktiver Bildworker auf der RTX 5080
|
||||||
|
- offizielle Qwen-Image-2.1 Prompt-Enhancer T2I und I2I als kurzlebige
|
||||||
|
Q5-Worker auf der RTX 3060
|
||||||
- FLUX.2 Klein 9B FP8 Beta als gestoppter Rückfallcontainer
|
- FLUX.2 Klein 9B FP8 Beta als gestoppter Rückfallcontainer
|
||||||
- Qwen3-TTS 1.7B auf der RTX 3060 hinter dem TTS-Gateway; kein Piper-Fallback
|
- Qwen3-TTS 1.7B auf der RTX 3060 hinter dem TTS-Gateway; kein Piper-Fallback
|
||||||
- Whisper.cpp `small` auf der CPU für lokale deutsche Spracherkennung
|
- Qwen3-ASR 0.6B Q8 auf der CPU für lokale deutsche Spracherkennung
|
||||||
- EmbeddingGemma 300M Q8 auf der CPU für OpenClaws hybride Memory-Suche
|
- EmbeddingGemma 300M Q8 auf der CPU für OpenClaws hybride Memory-Suche
|
||||||
- Live-Dashboard mit 21 Tagen Detailhistorie für GPUs und Slot-Kontextbelegung auf Port 8099
|
- Live-Dashboard mit 21 Tagen Detailhistorie für GPUs und Slot-Kontextbelegung auf Port 8099
|
||||||
- Portainer CE als optionale Container-Ansicht auf Port 9443
|
- Portainer CE als optionale Container-Ansicht auf Port 9443
|
||||||
@@ -48,8 +50,12 @@ ist nicht mehr Bestandteil der produktiven Architektur.
|
|||||||
Bildanfragen an den Router verwenden seit dem 21. September
|
Bildanfragen an den Router verwenden seit dem 21. September
|
||||||
**Qwen-Image-2.1 INT8** über gepinntes ComfyUI. Der Worker läuft mit Low-VRAM
|
**Qwen-Image-2.1 INT8** über gepinntes ComfyUI. Der Worker läuft mit Low-VRAM
|
||||||
auf der RTX 5080, 25 Schritten, CFG 1 und kann bis zu vier Referenzbilder
|
auf der RTX 5080, 25 Schritten, CFG 1 und kann bis zu vier Referenzbilder
|
||||||
verarbeiten. Das aktive Textprofil und TTS werden für den Auftrag angehalten
|
verarbeiten. Vor dem Rendern schreibt der passende offizielle Q5-Prompt-
|
||||||
und anschließend wiederhergestellt. Der bisherige FLUX.2-Worker und seine
|
Enhancer den kurzen Text automatisch um: PE-T2I für reine Textaufträge und
|
||||||
|
PE-I2I für Referenzbilder. Er läuft dafür vorübergehend auf der RTX 3060.
|
||||||
|
OpenClaw benötigt weder ein neues Werkzeug noch besondere Promptregeln. Das
|
||||||
|
aktive Textprofil und TTS werden für den Auftrag angehalten und anschließend
|
||||||
|
wiederhergestellt. Der bisherige FLUX.2-Worker und seine
|
||||||
Gewichte bleiben als gestoppter, explizit allowlist-beschränkter Rückfallpfad
|
Gewichte bleiben als gestoppter, explizit allowlist-beschränkter Rückfallpfad
|
||||||
erhalten. Reproduzierbarer Stand und Rückschaltung:
|
erhalten. Reproduzierbarer Stand und Rückschaltung:
|
||||||
[Qwen-Image-2.1](docs/QWEN_IMAGE_21.md).
|
[Qwen-Image-2.1](docs/QWEN_IMAGE_21.md).
|
||||||
@@ -136,15 +142,24 @@ Prefills konkurrieren weiterhin um Rechenleistung und den gemeinsamen KV-Pool.
|
|||||||
- Athena-Dashboard: `http://192.168.1.212:8099`
|
- Athena-Dashboard: `http://192.168.1.212:8099`
|
||||||
|
|
||||||
Der Router stellt Sprache OpenAI-kompatibel bereit: Sprachausgabe über
|
Der Router stellt Sprache OpenAI-kompatibel bereit: Sprachausgabe über
|
||||||
`/v1/audio/speech` und Spracherkennung über `/v1/audio/transcriptions`. Das
|
`/v1/audio/speech` und Spracherkennung über `/v1/audio/transcriptions`.
|
||||||
Whisper-Modell liegt persistent im Docker-Volume `whisper-data`; Audiodaten
|
Spracherkennung nutzt Qwen3-ASR-0.6B Q8 auf der CPU; das Modell liegt unter
|
||||||
werden lokal auf Athena verarbeitet. Für OpenClaw Talk liegt der lokale
|
`/data/models/qwen3-asr-0.6b-q8`. Der Adapter liefert reinen Text unter
|
||||||
|
`qwen3-asr`; OpenClaw und die Voice-Brücke verwenden denselben Modellnamen.
|
||||||
|
Whisper-Container, Image und
|
||||||
|
Modellvolume sind entfernt. Audiodaten werden lokal auf Athena verarbeitet.
|
||||||
|
Für OpenClaw Talk liegt der lokale
|
||||||
Realtime-Provider unter
|
Realtime-Provider unter
|
||||||
[`integrations/openclaw-athena-talk`](integrations/openclaw-athena-talk). Er
|
[`integrations/openclaw-athena-talk`](integrations/openclaw-athena-talk). Er
|
||||||
verbindet Mikrofon → Athena Whisper → normalen OpenClaw-Agenten → aktives
|
verbindet Mikrofon → Athena STT → normalen OpenClaw-Agenten → aktives
|
||||||
Athena-TTS, sodass Modell, Werkzeuge und Memory auch im Sprachmodus erhalten
|
Athena-TTS, sodass Modell, Werkzeuge und Memory auch im Sprachmodus erhalten
|
||||||
bleiben. Die Installation landet in OpenClaws persistentem Datenverzeichnis
|
bleiben. Die Installation landet in OpenClaws persistentem Datenverzeichnis
|
||||||
und bleibt deshalb bei normalen Container-Updates bestehen.
|
und bleibt deshalb bei normalen Container-Updates bestehen.
|
||||||
|
Die separate Diktierfunktion verarbeitet seit Plugin-Version 1.3.0 längere
|
||||||
|
Aufnahmen bereits während des Sprechens in überlappenden Sechs-Sekunden-
|
||||||
|
Abschnitten. Beim Loslassen bleibt nur der kurze Rest für OpenClaws festes
|
||||||
|
Fünf-Sekunden-Abschlussfenster. Dafür wurden weder OpenClaw selbst verändert
|
||||||
|
noch ein weiterer Container angelegt.
|
||||||
|
|
||||||
OpenClaw wird über den Provider **llama.cpp → Existing llama-server** mit
|
OpenClaw wird über den Provider **llama.cpp → Existing llama-server** mit
|
||||||
`http://192.168.1.212:8081/v1` verbunden. Der Router beantwortet sowohl
|
`http://192.168.1.212:8081/v1` verbunden. Der Router beantwortet sowohl
|
||||||
|
|||||||
+215
-31
@@ -394,9 +394,8 @@ services:
|
|||||||
- --spec-draft-type-v
|
- --spec-draft-type-v
|
||||||
- f16
|
- f16
|
||||||
|
|
||||||
# Text-only maximum-context profile. This exact IQ4_XS-pure / 256K / 80:20
|
# Maximum-context profile. Keep the vision projector on CPU so images work
|
||||||
# combination completed the 220K fill test on RTX 5080 + RTX 3060.
|
# without consuming the tightly budgeted GPU memory of the 256K context.
|
||||||
# Deliberately no vision projector: Ultra prioritizes maximum usable context.
|
|
||||||
llama-ultra:
|
llama-ultra:
|
||||||
<<: *llama-common
|
<<: *llama-common
|
||||||
container_name: mike-ai-llama-ultra
|
container_name: mike-ai-llama-ultra
|
||||||
@@ -408,6 +407,9 @@ services:
|
|||||||
command:
|
command:
|
||||||
- --model
|
- --model
|
||||||
- "/models/${ULTRA_MODEL_FILE:?ULTRA_MODEL_FILE is required}"
|
- "/models/${ULTRA_MODEL_FILE:?ULTRA_MODEL_FILE is required}"
|
||||||
|
- --mmproj
|
||||||
|
- "/models/${VISION_PROJECTOR_FILE:?VISION_PROJECTOR_FILE is required}"
|
||||||
|
- --no-mmproj-offload
|
||||||
- --alias
|
- --alias
|
||||||
- qwen-ultra
|
- qwen-ultra
|
||||||
- --ctx-size
|
- --ctx-size
|
||||||
@@ -572,6 +574,8 @@ services:
|
|||||||
ALLOWED_PROFILES: fast,medium,large,ultra,uncensored
|
ALLOWED_PROFILES: fast,medium,large,ultra,uncensored
|
||||||
IMAGE_WORKER: image
|
IMAGE_WORKER: image
|
||||||
RESTORE_WORKER: restore
|
RESTORE_WORKER: restore
|
||||||
|
IMAGE_PROMPT_I2I_WORKER: image-prompt-i2i
|
||||||
|
IMAGE_PROMPT_T2I_WORKER: image-prompt-t2i
|
||||||
FLUX_STANDBY_WORKER: flux-standby
|
FLUX_STANDBY_WORKER: flux-standby
|
||||||
TTS_WORKER: qwen3
|
TTS_WORKER: qwen3
|
||||||
MUSIC_WORKER: acestep
|
MUSIC_WORKER: acestep
|
||||||
@@ -602,6 +606,8 @@ services:
|
|||||||
volumes:
|
volumes:
|
||||||
- ./router/router_profiles.json:/etc/mike-ai/router-profiles.json:ro
|
- ./router/router_profiles.json:/etc/mike-ai/router-profiles.json:ro
|
||||||
- ./config/global-system-policy.txt:/etc/mike-ai/global-system-policy.txt:ro
|
- ./config/global-system-policy.txt:/etc/mike-ai/global-system-policy.txt:ro
|
||||||
|
- "${QWEN_IMAGE_PE_I2I_MODEL_DIR:-/data/models/qwen-image-2.1-pe-i2i-q5}/system_prompt.txt:/etc/mike-ai/qwen-image-pe-i2i-system-prompt.txt:ro"
|
||||||
|
- "${QWEN_IMAGE_PE_T2I_MODEL_DIR:-/data/models/qwen-image-2.1-pe-t2i-q5}/system_prompt.txt:/etc/mike-ai/qwen-image-pe-t2i-system-prompt.txt:ro"
|
||||||
- router-state:/var/lib/mike-ai-profile-router
|
- router-state:/var/lib/mike-ai-profile-router
|
||||||
- router-images:/data/images
|
- router-images:/data/images
|
||||||
environment:
|
environment:
|
||||||
@@ -631,6 +637,11 @@ services:
|
|||||||
IMAGE_DIR: /data/images
|
IMAGE_DIR: /data/images
|
||||||
IMAGE_WORKER_URL: http://image-worker:8086
|
IMAGE_WORKER_URL: http://image-worker:8086
|
||||||
IMAGE_WORKER_TOKEN: "${CONTROLLER_TOKEN:?CONTROLLER_TOKEN is required}"
|
IMAGE_WORKER_TOKEN: "${CONTROLLER_TOKEN:?CONTROLLER_TOKEN is required}"
|
||||||
|
IMAGE_PROMPT_I2I_URL: http://image-prompt-enhancer-i2i:8080
|
||||||
|
IMAGE_PROMPT_T2I_URL: http://image-prompt-enhancer-t2i:8080
|
||||||
|
IMAGE_PROMPT_I2I_SYSTEM_FILE: /etc/mike-ai/qwen-image-pe-i2i-system-prompt.txt
|
||||||
|
IMAGE_PROMPT_T2I_SYSTEM_FILE: /etc/mike-ai/qwen-image-pe-t2i-system-prompt.txt
|
||||||
|
IMAGE_PROMPT_ENHANCER_TIMEOUT: "180"
|
||||||
IMAGE_MODEL_NAME: Qwen-Image-2.1-int8
|
IMAGE_MODEL_NAME: Qwen-Image-2.1-int8
|
||||||
IMAGE_INFERENCE_STEPS: "25"
|
IMAGE_INFERENCE_STEPS: "25"
|
||||||
CHAT_IMAGE_ALLOW_REMOTE_URLS: "false"
|
CHAT_IMAGE_ALLOW_REMOTE_URLS: "false"
|
||||||
@@ -647,7 +658,7 @@ services:
|
|||||||
MUSIC_START_TIMEOUT: "600"
|
MUSIC_START_TIMEOUT: "600"
|
||||||
VOICE_CHANGE_START_TIMEOUT: "600"
|
VOICE_CHANGE_START_TIMEOUT: "600"
|
||||||
APPLIO_START_TIMEOUT: "900"
|
APPLIO_START_TIMEOUT: "900"
|
||||||
STT_WORKER_URL: http://whisper:8084
|
STT_WORKER_URL: http://qwen-asr-worker:8084
|
||||||
STT_TIMEOUT: "300"
|
STT_TIMEOUT: "300"
|
||||||
networks: [frontend, control, inference]
|
networks: [frontend, control, inference]
|
||||||
security_opt: ["no-new-privileges:true"]
|
security_opt: ["no-new-privileges:true"]
|
||||||
@@ -670,7 +681,7 @@ services:
|
|||||||
condition: service_healthy
|
condition: service_healthy
|
||||||
tts-gateway:
|
tts-gateway:
|
||||||
condition: service_healthy
|
condition: service_healthy
|
||||||
whisper:
|
qwen-asr-worker:
|
||||||
condition: service_healthy
|
condition: service_healthy
|
||||||
|
|
||||||
image-worker:
|
image-worker:
|
||||||
@@ -716,6 +727,143 @@ services:
|
|||||||
retries: 90
|
retries: 90
|
||||||
start_period: 10s
|
start_period: 10s
|
||||||
|
|
||||||
|
image-prompt-enhancer-i2i:
|
||||||
|
image: mike-ai/llama.cpp:b10930
|
||||||
|
container_name: mike-ai-image-prompt-enhancer-i2i
|
||||||
|
restart: "no"
|
||||||
|
profiles: [image]
|
||||||
|
labels:
|
||||||
|
com.mike-ai.image-worker: image-prompt-i2i
|
||||||
|
deploy:
|
||||||
|
resources:
|
||||||
|
reservations:
|
||||||
|
devices:
|
||||||
|
- driver: nvidia
|
||||||
|
device_ids: ["${QWEN_IMAGE_PE_GPU:-0}"]
|
||||||
|
capabilities: [gpu]
|
||||||
|
read_only: true
|
||||||
|
tmpfs: ["/tmp:size=512m,mode=1777"]
|
||||||
|
volumes:
|
||||||
|
- "${QWEN_IMAGE_PE_I2I_MODEL_DIR:-/data/models/qwen-image-2.1-pe-i2i-q5}:/models:ro"
|
||||||
|
command:
|
||||||
|
- --model
|
||||||
|
- /models/Qwen-Image-2.1-PE-I2I.Q5_K_M.gguf
|
||||||
|
- --mmproj
|
||||||
|
- /models/Qwen-Image-2.1-PE-I2I.mmproj-bf16.gguf
|
||||||
|
- --mmproj-offload
|
||||||
|
- --mmproj-device
|
||||||
|
- CUDA0
|
||||||
|
- --alias
|
||||||
|
- qwen-image-pe-i2i
|
||||||
|
- --ctx-size
|
||||||
|
- "16384"
|
||||||
|
- --flash-attn
|
||||||
|
- "on"
|
||||||
|
- --cache-type-k
|
||||||
|
- q8_0
|
||||||
|
- --cache-type-v
|
||||||
|
- q8_0
|
||||||
|
- --threads
|
||||||
|
- "6"
|
||||||
|
- --threads-batch
|
||||||
|
- "6"
|
||||||
|
- --batch-size
|
||||||
|
- "1024"
|
||||||
|
- --ubatch-size
|
||||||
|
- "128"
|
||||||
|
- --parallel
|
||||||
|
- "1"
|
||||||
|
- --jinja
|
||||||
|
- --reasoning
|
||||||
|
- auto
|
||||||
|
- --host
|
||||||
|
- 0.0.0.0
|
||||||
|
- --port
|
||||||
|
- "8080"
|
||||||
|
- --metrics
|
||||||
|
- --n-gpu-layers
|
||||||
|
- all
|
||||||
|
- --device
|
||||||
|
- CUDA0
|
||||||
|
- --split-mode
|
||||||
|
- none
|
||||||
|
- --no-ui
|
||||||
|
networks: [inference]
|
||||||
|
security_opt: ["no-new-privileges:true"]
|
||||||
|
cap_drop: [ALL]
|
||||||
|
healthcheck:
|
||||||
|
test: [CMD, curl, -fsS, "http://127.0.0.1:8080/health"]
|
||||||
|
interval: 5s
|
||||||
|
timeout: 3s
|
||||||
|
retries: 40
|
||||||
|
start_period: 10s
|
||||||
|
|
||||||
|
image-prompt-enhancer-t2i:
|
||||||
|
image: mike-ai/llama.cpp:b10930
|
||||||
|
container_name: mike-ai-image-prompt-enhancer-t2i
|
||||||
|
restart: "no"
|
||||||
|
profiles: [image]
|
||||||
|
labels:
|
||||||
|
com.mike-ai.image-worker: image-prompt-t2i
|
||||||
|
deploy:
|
||||||
|
resources:
|
||||||
|
reservations:
|
||||||
|
devices:
|
||||||
|
- driver: nvidia
|
||||||
|
device_ids: ["${QWEN_IMAGE_PE_GPU:-0}"]
|
||||||
|
capabilities: [gpu]
|
||||||
|
read_only: true
|
||||||
|
tmpfs: ["/tmp:size=512m,mode=1777"]
|
||||||
|
volumes:
|
||||||
|
- "${QWEN_IMAGE_PE_T2I_MODEL_DIR:-/data/models/qwen-image-2.1-pe-t2i-q5}:/models:ro"
|
||||||
|
command:
|
||||||
|
- --model
|
||||||
|
- /models/Qwen-Image-2.1-PE-T2I.Q5_K_M.gguf
|
||||||
|
- --alias
|
||||||
|
- qwen-image-pe-t2i
|
||||||
|
- --ctx-size
|
||||||
|
- "16384"
|
||||||
|
- --flash-attn
|
||||||
|
- "on"
|
||||||
|
- --cache-type-k
|
||||||
|
- q8_0
|
||||||
|
- --cache-type-v
|
||||||
|
- q8_0
|
||||||
|
- --threads
|
||||||
|
- "6"
|
||||||
|
- --threads-batch
|
||||||
|
- "6"
|
||||||
|
- --batch-size
|
||||||
|
- "1024"
|
||||||
|
- --ubatch-size
|
||||||
|
- "128"
|
||||||
|
- --parallel
|
||||||
|
- "1"
|
||||||
|
- --jinja
|
||||||
|
- --reasoning
|
||||||
|
- auto
|
||||||
|
- --host
|
||||||
|
- 0.0.0.0
|
||||||
|
- --port
|
||||||
|
- "8080"
|
||||||
|
- --metrics
|
||||||
|
- --n-gpu-layers
|
||||||
|
- all
|
||||||
|
- --device
|
||||||
|
- CUDA0
|
||||||
|
- --split-mode
|
||||||
|
- none
|
||||||
|
- --no-ui
|
||||||
|
networks: [inference]
|
||||||
|
security_opt: ["no-new-privileges:true"]
|
||||||
|
cap_drop: [ALL]
|
||||||
|
healthcheck:
|
||||||
|
test: [CMD, curl, -fsS, "http://127.0.0.1:8080/health"]
|
||||||
|
interval: 5s
|
||||||
|
timeout: 3s
|
||||||
|
retries: 40
|
||||||
|
start_period: 10s
|
||||||
|
|
||||||
# Previous production image model, retained as a stopped rollback target.
|
# Previous production image model, retained as a stopped rollback target.
|
||||||
# It is outside the normal router path and can only be started through the
|
# It is outside the normal router path and can only be started through the
|
||||||
# controller's allowlisted flux-standby endpoint.
|
# controller's allowlisted flux-standby endpoint.
|
||||||
@@ -828,40 +976,77 @@ services:
|
|||||||
retries: 12
|
retries: 12
|
||||||
start_period: 10s
|
start_period: 10s
|
||||||
|
|
||||||
whisper:
|
qwen-asr:
|
||||||
build:
|
image: ${LLAMA_CPU_IMAGE:-mike-ai/llama.cpp-cpu:local}
|
||||||
context: .
|
container_name: mike-ai-qwen-asr
|
||||||
dockerfile: platform/docker/whisper/Dockerfile
|
|
||||||
args:
|
|
||||||
WHISPER_CPP_VERSION: ${WHISPER_CPP_VERSION:-v1.9.1}
|
|
||||||
image: mike-ai/whisper:local
|
|
||||||
container_name: mike-ai-whisper
|
|
||||||
restart: unless-stopped
|
restart: unless-stopped
|
||||||
read_only: true
|
read_only: true
|
||||||
tmpfs:
|
tmpfs:
|
||||||
- /tmp:size=2g,mode=1777
|
- /tmp:size=256m,mode=1777
|
||||||
volumes:
|
volumes:
|
||||||
- whisper-data:/models
|
- "${QWEN_ASR_MODEL_DIR:-/data/models/qwen3-asr-0.6b-q8}:/models:ro"
|
||||||
environment:
|
command:
|
||||||
WHISPER_HOST: 0.0.0.0
|
- --model
|
||||||
WHISPER_PORT: "8084"
|
- /models/Qwen3-ASR-0.6B-Q8_0.gguf
|
||||||
WHISPER_CLI: /opt/whisper.cpp/build/bin/whisper-cli
|
- --mmproj
|
||||||
WHISPER_MODEL: /models/ggml-small.bin
|
- /models/mmproj-Qwen3-ASR-0.6B-Q8_0.gguf
|
||||||
WHISPER_MODEL_URL: ${WHISPER_MODEL_URL:-https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-small.bin}
|
- --no-mmproj-offload
|
||||||
WHISPER_SERVER_PORT: "8085"
|
- --n-gpu-layers
|
||||||
WHISPER_THREADS: ${WHISPER_THREADS:-8}
|
- "0"
|
||||||
WHISPER_LANGUAGE: ${WHISPER_LANGUAGE:-de}
|
- --alias
|
||||||
networks: [frontend, inference]
|
- qwen3-asr-0.6b
|
||||||
|
- --ctx-size
|
||||||
|
- "4096"
|
||||||
|
- --threads
|
||||||
|
- "6"
|
||||||
|
- --parallel
|
||||||
|
- "1"
|
||||||
|
- --host
|
||||||
|
- 0.0.0.0
|
||||||
|
- --port
|
||||||
|
- "8080"
|
||||||
|
- --no-ui
|
||||||
|
- --fit
|
||||||
|
- "off"
|
||||||
|
cpus: 6
|
||||||
|
mem_limit: 6g
|
||||||
|
networks: [inference]
|
||||||
security_opt: ["no-new-privileges:true"]
|
security_opt: ["no-new-privileges:true"]
|
||||||
cap_drop: [ALL]
|
cap_drop: [ALL]
|
||||||
# The entrypoint supervises whisper-server after dropping it to uid 10004.
|
|
||||||
cap_add: [CHOWN, SETUID, SETGID, KILL]
|
|
||||||
healthcheck:
|
healthcheck:
|
||||||
test: [CMD, curl, -fsS, "http://127.0.0.1:8084/status"]
|
test: [CMD, curl, -fsS, "http://127.0.0.1:8080/health"]
|
||||||
interval: 10s
|
interval: 10s
|
||||||
timeout: 5s
|
timeout: 5s
|
||||||
retries: 90
|
retries: 12
|
||||||
start_period: 20m
|
start_period: 30s
|
||||||
|
|
||||||
|
qwen-asr-worker:
|
||||||
|
build:
|
||||||
|
context: .
|
||||||
|
dockerfile: platform/docker/qwen-asr-worker/Dockerfile
|
||||||
|
image: mike-ai/qwen-asr-worker:local
|
||||||
|
container_name: mike-ai-qwen-asr-worker
|
||||||
|
restart: unless-stopped
|
||||||
|
read_only: true
|
||||||
|
tmpfs:
|
||||||
|
- /tmp:size=256m,mode=1777
|
||||||
|
environment:
|
||||||
|
QWEN_ASR_HOST: 0.0.0.0
|
||||||
|
QWEN_ASR_PORT: "8084"
|
||||||
|
QWEN_ASR_LANGUAGE: de
|
||||||
|
QWEN_ASR_SERVER_URL: http://qwen-asr:8080
|
||||||
|
networks: [inference]
|
||||||
|
security_opt: ["no-new-privileges:true"]
|
||||||
|
cap_drop: [ALL]
|
||||||
|
depends_on:
|
||||||
|
qwen-asr:
|
||||||
|
condition: service_healthy
|
||||||
|
healthcheck:
|
||||||
|
test: [CMD, python, -c, "import json,urllib.request; assert json.load(urllib.request.urlopen('http://127.0.0.1:8084/status', timeout=2))['ready']"]
|
||||||
|
interval: 10s
|
||||||
|
timeout: 5s
|
||||||
|
retries: 12
|
||||||
|
start_period: 15s
|
||||||
|
|
||||||
llama-dashboard:
|
llama-dashboard:
|
||||||
build: ./platform/llama-dashboard
|
build: ./platform/llama-dashboard
|
||||||
@@ -976,7 +1161,6 @@ networks:
|
|||||||
name: mike-ai-tools-egress
|
name: mike-ai-tools-egress
|
||||||
|
|
||||||
volumes:
|
volumes:
|
||||||
whisper-data:
|
|
||||||
router-state:
|
router-state:
|
||||||
router-images:
|
router-images:
|
||||||
portainer-data:
|
portainer-data:
|
||||||
|
|||||||
@@ -47,9 +47,9 @@
|
|||||||
"model_env": "ULTRA_MODEL_FILE",
|
"model_env": "ULTRA_MODEL_FILE",
|
||||||
"model_family": "Qwen3.8-27B IQ4 XS Pure",
|
"model_family": "Qwen3.8-27B IQ4 XS Pure",
|
||||||
"gpu_split": "80:20",
|
"gpu_split": "80:20",
|
||||||
"vision": false,
|
"vision": true,
|
||||||
"mtp": 2,
|
"mtp": 2,
|
||||||
"description": "Maximaler Textkontext; bewusst ohne Vision-Projektor."
|
"description": "Maximaler Kontext mit Vision-Projektor auf der CPU."
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"id": "uncensored",
|
"id": "uncensored",
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
#!/usr/bin/env python3
|
#!/usr/bin/env python3
|
||||||
"""Mock-STT-Worker für lokale Tests.
|
"""Mock-STT-Worker für lokale Tests.
|
||||||
|
|
||||||
Simuliert den Whisper-STT-Worker:
|
Simuliert den Qwen3-ASR-Adapter:
|
||||||
GET /status → ready: true
|
GET /status → ready: true
|
||||||
POST /transcribe → liefert festes Transkript
|
POST /transcribe → liefert festes Transkript
|
||||||
|
|
||||||
|
|||||||
+8
-8
@@ -220,7 +220,7 @@ def main() -> None:
|
|||||||
# Test 1: Multipart + Content-Length (bestehender Pfad)
|
# Test 1: Multipart + Content-Length (bestehender Pfad)
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 1: Multipart + Content-Length")
|
print("Test 1: Multipart + Content-Length")
|
||||||
mp = build_multipart({"model": "whisper-1", "language": "de"},
|
mp = build_multipart({"model": "qwen3-asr", "language": "de"},
|
||||||
file_data=fake_webm)
|
file_data=fake_webm)
|
||||||
status, body = http_request(
|
status, body = http_request(
|
||||||
"POST", PORTS["router"], "/v1/audio/transcriptions",
|
"POST", PORTS["router"], "/v1/audio/transcriptions",
|
||||||
@@ -235,7 +235,7 @@ def main() -> None:
|
|||||||
# Test 2: Multipart + Transfer-Encoding chunked (einfach)
|
# Test 2: Multipart + Transfer-Encoding chunked (einfach)
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 2: Multipart + chunked (einfach)")
|
print("Test 2: Multipart + chunked (einfach)")
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=fake_webm)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=fake_webm)
|
||||||
chunked = build_chunked_body([mp])
|
chunked = build_chunked_body([mp])
|
||||||
raw = (
|
raw = (
|
||||||
b"POST /v1/audio/transcriptions HTTP/1.1\r\n"
|
b"POST /v1/audio/transcriptions HTTP/1.1\r\n"
|
||||||
@@ -254,7 +254,7 @@ def main() -> None:
|
|||||||
# Test 3: Mehrere unterschiedlich große Chunks
|
# Test 3: Mehrere unterschiedlich große Chunks
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 3: Mehrere unterschiedlich große Chunks")
|
print("Test 3: Mehrere unterschiedlich große Chunks")
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=fake_webm)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=fake_webm)
|
||||||
# In 5 Chunks aufteilen (unterschiedlich groß)
|
# In 5 Chunks aufteilen (unterschiedlich groß)
|
||||||
chunks = []
|
chunks = []
|
||||||
sizes = [10, 50, 7, 100, 33]
|
sizes = [10, 50, 7, 100, 33]
|
||||||
@@ -283,7 +283,7 @@ def main() -> None:
|
|||||||
# Test 4: Boundary über Chunk-Grenzen verteilt
|
# Test 4: Boundary über Chunk-Grenzen verteilt
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 4: Boundary über Chunk-Grenzen verteilt")
|
print("Test 4: Boundary über Chunk-Grenzen verteilt")
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=fake_webm)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=fake_webm)
|
||||||
# Boundary-String finden und Chunk-Grenze genau dorthin setzen
|
# Boundary-String finden und Chunk-Grenze genau dorthin setzen
|
||||||
boundary_str = b"--testboundary123"
|
boundary_str = b"--testboundary123"
|
||||||
idx = mp.find(boundary_str, 10) # zweite Boundary (vor file)
|
idx = mp.find(boundary_str, 10) # zweite Boundary (vor file)
|
||||||
@@ -310,7 +310,7 @@ def main() -> None:
|
|||||||
# Test 5: Chunk Extensions
|
# Test 5: Chunk Extensions
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 5: Chunk Extensions")
|
print("Test 5: Chunk Extensions")
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=fake_webm)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=fake_webm)
|
||||||
chunks_ext = [
|
chunks_ext = [
|
||||||
(mp[:20], "ext1=value1"),
|
(mp[:20], "ext1=value1"),
|
||||||
(mp[20:60], None),
|
(mp[20:60], None),
|
||||||
@@ -335,7 +335,7 @@ def main() -> None:
|
|||||||
# hier explizit mit Trailer)
|
# hier explizit mit Trailer)
|
||||||
# ------------------------------------------------------------------
|
# ------------------------------------------------------------------
|
||||||
print("Test 6: 0-Chunk mit Trailer")
|
print("Test 6: 0-Chunk mit Trailer")
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=fake_webm)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=fake_webm)
|
||||||
chunked = build_chunked_body([mp])
|
chunked = build_chunked_body([mp])
|
||||||
# Trailer hinzufügen
|
# Trailer hinzufügen
|
||||||
chunked_with_trailer = chunked.replace(
|
chunked_with_trailer = chunked.replace(
|
||||||
@@ -376,7 +376,7 @@ def main() -> None:
|
|||||||
print("Test 8: Uploadgrößenlimit")
|
print("Test 8: Uploadgrößenlimit")
|
||||||
# MAX_UPLOAD_SIZE = 1 MB, also 2 MB senden
|
# MAX_UPLOAD_SIZE = 1 MB, also 2 MB senden
|
||||||
big_data = b"A" * (2 * 1024 * 1024)
|
big_data = b"A" * (2 * 1024 * 1024)
|
||||||
mp = build_multipart({"model": "whisper-1"}, file_data=big_data)
|
mp = build_multipart({"model": "qwen3-asr"}, file_data=big_data)
|
||||||
chunked = build_chunked_body([mp[:1024 * 1024], mp[1024 * 1024:]])
|
chunked = build_chunked_body([mp[:1024 * 1024], mp[1024 * 1024:]])
|
||||||
raw = (
|
raw = (
|
||||||
b"POST /v1/audio/transcriptions HTTP/1.1\r\n"
|
b"POST /v1/audio/transcriptions HTTP/1.1\r\n"
|
||||||
@@ -402,7 +402,7 @@ def main() -> None:
|
|||||||
+ b"WEBM_OPUS_AUDIO_DATA" * 100)
|
+ b"WEBM_OPUS_AUDIO_DATA" * 100)
|
||||||
boundary = "950bd961b24c4a32801e31b128c85e09"
|
boundary = "950bd961b24c4a32801e31b128c85e09"
|
||||||
mp = build_multipart(
|
mp = build_multipart(
|
||||||
{"model": "whisper-1", "language": "de"},
|
{"model": "qwen3-asr", "language": "de"},
|
||||||
file_data=webm_data,
|
file_data=webm_data,
|
||||||
filename="recording.webm",
|
filename="recording.webm",
|
||||||
boundary=boundary,
|
boundary=boundary,
|
||||||
|
|||||||
+12
-12
@@ -71,13 +71,13 @@ def test_quoted_boundary():
|
|||||||
boundary = "----WebKitFormBoundary7MA4YWxkTrZu0gW"
|
boundary = "----WebKitFormBoundary7MA4YWxkTrZu0gW"
|
||||||
webm = b"\x1a\x45\xdf\xa3" + b"\x00\x01\x02\x03\xff\xfe\xfd" * 50
|
webm = b"\x1a\x45\xdf\xa3" + b"\x00\x01\x02\x03\xff\xfe\xfd" * 50
|
||||||
body, ct = build(
|
body, ct = build(
|
||||||
[("model", "whisper-1", None), ("file", webm, "t.webm")],
|
[("model", "qwen3-asr", None), ("file", webm, "t.webm")],
|
||||||
boundary, quoted=True,
|
boundary, quoted=True,
|
||||||
)
|
)
|
||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == webm, "file_data mismatch"
|
assert fd == webm, "file_data mismatch"
|
||||||
assert fn == "t.webm", f"filename mismatch: {fn!r}"
|
assert fn == "t.webm", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
print(" quoted boundary: OK")
|
print(" quoted boundary: OK")
|
||||||
|
|
||||||
|
|
||||||
@@ -109,7 +109,7 @@ def test_openwebui_style():
|
|||||||
f"\r\n--{boundary}\r\n"
|
f"\r\n--{boundary}\r\n"
|
||||||
f'Content-Disposition: form-data; name="model"\r\n'
|
f'Content-Disposition: form-data; name="model"\r\n'
|
||||||
f"\r\n"
|
f"\r\n"
|
||||||
f"whisper-1\r\n"
|
f"qwen3-asr\r\n"
|
||||||
f"--{boundary}\r\n"
|
f"--{boundary}\r\n"
|
||||||
f'Content-Disposition: form-data; name="temperature"\r\n'
|
f'Content-Disposition: form-data; name="temperature"\r\n'
|
||||||
f"\r\n"
|
f"\r\n"
|
||||||
@@ -119,7 +119,7 @@ def test_openwebui_style():
|
|||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == webm, "file_data mismatch"
|
assert fd == webm, "file_data mismatch"
|
||||||
assert fn == "rec.webm", f"filename mismatch: {fn!r}"
|
assert fn == "rec.webm", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
assert fl["temperature"] == "0.0", f"temperature mismatch: {fl!r}"
|
assert fl["temperature"] == "0.0", f"temperature mismatch: {fl!r}"
|
||||||
print(" Open-WebUI-artig: OK")
|
print(" Open-WebUI-artig: OK")
|
||||||
|
|
||||||
@@ -128,13 +128,13 @@ def test_file_before_model():
|
|||||||
"""File-Feld vor model-Feld."""
|
"""File-Feld vor model-Feld."""
|
||||||
boundary = "boundary123"
|
boundary = "boundary123"
|
||||||
body, ct = build(
|
body, ct = build(
|
||||||
[("file", b"DATA", "f.wav"), ("model", "whisper-1", None)],
|
[("file", b"DATA", "f.wav"), ("model", "qwen3-asr", None)],
|
||||||
boundary, quoted=False,
|
boundary, quoted=False,
|
||||||
)
|
)
|
||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == b"DATA", "file_data mismatch"
|
assert fd == b"DATA", "file_data mismatch"
|
||||||
assert fn == "f.wav", f"filename mismatch: {fn!r}"
|
assert fn == "f.wav", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
print(" File vor model: OK")
|
print(" File vor model: OK")
|
||||||
|
|
||||||
|
|
||||||
@@ -142,13 +142,13 @@ def test_file_after_model():
|
|||||||
"""model-Feld vor File-Feld."""
|
"""model-Feld vor File-Feld."""
|
||||||
boundary = "boundary456"
|
boundary = "boundary456"
|
||||||
body, ct = build(
|
body, ct = build(
|
||||||
[("model", "whisper-1", None), ("file", b"DATA", "g.wav")],
|
[("model", "qwen3-asr", None), ("file", b"DATA", "g.wav")],
|
||||||
boundary, quoted=False,
|
boundary, quoted=False,
|
||||||
)
|
)
|
||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == b"DATA", "file_data mismatch"
|
assert fd == b"DATA", "file_data mismatch"
|
||||||
assert fn == "g.wav", f"filename mismatch: {fn!r}"
|
assert fn == "g.wav", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
print(" File nach model: OK")
|
print(" File nach model: OK")
|
||||||
|
|
||||||
|
|
||||||
@@ -182,13 +182,13 @@ def test_extra_headers_ignored():
|
|||||||
f"\r\n--{boundary}\r\n"
|
f"\r\n--{boundary}\r\n"
|
||||||
f'Content-Disposition: form-data; name="model"\r\n'
|
f'Content-Disposition: form-data; name="model"\r\n'
|
||||||
f"\r\n"
|
f"\r\n"
|
||||||
f"whisper-1\r\n"
|
f"qwen3-asr\r\n"
|
||||||
f"--{boundary}--\r\n"
|
f"--{boundary}--\r\n"
|
||||||
).encode()
|
).encode()
|
||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == webm, "file_data mismatch"
|
assert fd == webm, "file_data mismatch"
|
||||||
assert fn == "x.webm", f"filename mismatch: {fn!r}"
|
assert fn == "x.webm", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
print(" Extra-Header ignoriert: OK")
|
print(" Extra-Header ignoriert: OK")
|
||||||
|
|
||||||
|
|
||||||
@@ -198,7 +198,7 @@ def test_all_fields():
|
|||||||
body, ct = build(
|
body, ct = build(
|
||||||
[
|
[
|
||||||
("file", b"AUDIO", "a.webm"),
|
("file", b"AUDIO", "a.webm"),
|
||||||
("model", "whisper-1", None),
|
("model", "qwen3-asr", None),
|
||||||
("language", "de", None),
|
("language", "de", None),
|
||||||
("prompt", "Kontext", None),
|
("prompt", "Kontext", None),
|
||||||
("response_format", "verbose_json", None),
|
("response_format", "verbose_json", None),
|
||||||
@@ -209,7 +209,7 @@ def test_all_fields():
|
|||||||
fd, fn, fl = parse_multipart(body, ct)
|
fd, fn, fl = parse_multipart(body, ct)
|
||||||
assert fd == b"AUDIO", "file_data mismatch"
|
assert fd == b"AUDIO", "file_data mismatch"
|
||||||
assert fn == "a.webm", f"filename mismatch: {fn!r}"
|
assert fn == "a.webm", f"filename mismatch: {fn!r}"
|
||||||
assert fl["model"] == "whisper-1", f"model mismatch: {fl!r}"
|
assert fl["model"] == "qwen3-asr", f"model mismatch: {fl!r}"
|
||||||
assert fl["language"] == "de", f"language mismatch: {fl!r}"
|
assert fl["language"] == "de", f"language mismatch: {fl!r}"
|
||||||
assert fl["prompt"] == "Kontext", f"prompt mismatch: {fl!r}"
|
assert fl["prompt"] == "Kontext", f"prompt mismatch: {fl!r}"
|
||||||
assert fl["response_format"] == "verbose_json", f"response_format mismatch: {fl!r}"
|
assert fl["response_format"] == "verbose_json", f"response_format mismatch: {fl!r}"
|
||||||
|
|||||||
@@ -3,6 +3,7 @@ import json
|
|||||||
import os
|
import os
|
||||||
import sys
|
import sys
|
||||||
import threading
|
import threading
|
||||||
|
import tempfile
|
||||||
import unittest
|
import unittest
|
||||||
from pathlib import Path
|
from pathlib import Path
|
||||||
from unittest.mock import patch
|
from unittest.mock import patch
|
||||||
@@ -13,6 +14,26 @@ import ai_profile_router as router
|
|||||||
|
|
||||||
|
|
||||||
class CoordinationTests(unittest.TestCase):
|
class CoordinationTests(unittest.TestCase):
|
||||||
|
def test_prompt_enhancer_accepts_plain_and_fenced_json(self):
|
||||||
|
plain = router._parse_prompt_enhancer_result(
|
||||||
|
'{"rewritten_prompt":"new scene","wh_ratio":"3:2"}')
|
||||||
|
fenced = router._parse_prompt_enhancer_result(
|
||||||
|
'```json\n{"rewritten_prompt":"portrait","wh_ratio":"3:4"}\n```')
|
||||||
|
self.assertEqual(plain['rewritten_prompt'], 'new scene')
|
||||||
|
self.assertEqual(fenced['wh_ratio'], '3:4')
|
||||||
|
|
||||||
|
def test_prompt_enhancer_rejects_missing_rewrite(self):
|
||||||
|
with self.assertRaisesRegex(RuntimeError, 'rewritten_prompt'):
|
||||||
|
router._parse_prompt_enhancer_result('{"wh_ratio":"1:1"}')
|
||||||
|
|
||||||
|
def test_image_data_url_resolves_volume_relative_reference(self):
|
||||||
|
with tempfile.TemporaryDirectory() as directory, \
|
||||||
|
patch.object(router, 'IMAGE_DIR', directory):
|
||||||
|
path = Path(directory) / '.edit-reference.ref'
|
||||||
|
path.write_bytes(b'\x89PNG\r\n\x1a\ncontent')
|
||||||
|
result = router._image_data_url(path.name)
|
||||||
|
self.assertTrue(result.startswith('data:image/png;base64,'))
|
||||||
|
|
||||||
def test_mode_uses_one_consistent_controller_snapshot(self):
|
def test_mode_uses_one_consistent_controller_snapshot(self):
|
||||||
with patch.object(router, 'PROFILE_CONTROL_URL', 'http://controller'), \
|
with patch.object(router, 'PROFILE_CONTROL_URL', 'http://controller'), \
|
||||||
patch.object(router, '_profile_controller_request', return_value={
|
patch.object(router, '_profile_controller_request', return_value={
|
||||||
|
|||||||
@@ -2,7 +2,7 @@
|
|||||||
|
|
||||||
Stand: 21. September 2026
|
Stand: 21. September 2026
|
||||||
|
|
||||||
Athena hat 31 reguläre Docker-Container. Nicht jeder Container enthält
|
Athena hat 33 reguläre Docker-Container. Nicht jeder Container enthält
|
||||||
ein KI-Modell: Router, Oberflächen, Netzwerk, Steuerung und Sicherung sind
|
ein KI-Modell: Router, Oberflächen, Netzwerk, Steuerung und Sicherung sind
|
||||||
gewöhnliche Dienste. Die rechenintensiven GPU-Worker werden absichtlich nur bei
|
gewöhnliche Dienste. Die rechenintensiven GPU-Worker werden absichtlich nur bei
|
||||||
Bedarf gestartet. Ein Container im Zustand `Created` oder `Exited (0)` ist daher
|
Bedarf gestartet. Ein Container im Zustand `Created` oder `Exited (0)` ist daher
|
||||||
@@ -15,12 +15,14 @@ nicht automatisch ein ungenutzter Rest.
|
|||||||
| `mike-ai-bonsai2-ab` | Bonsai-2-Vergleichsmodell | Gestoppter, reproduzierbar dokumentierter A/B-Testcontainer; kein Produktivprofil. |
|
| `mike-ai-bonsai2-ab` | Bonsai-2-Vergleichsmodell | Gestoppter, reproduzierbar dokumentierter A/B-Testcontainer; kein Produktivprofil. |
|
||||||
| `mike-ai-applio-studio` | Applio/RVC; Stimmenmodelle werden nutzerseitig ergänzt | Vollständige RVC-Oberfläche für Inferenz, Modellverwaltung und Training auf der RTX 5080. Für eine Konvertierung ist ein importiertes oder trainiertes `.pth`-Modell nötig; eine Referenzaufnahme allein reicht nicht. |
|
| `mike-ai-applio-studio` | Applio/RVC; Stimmenmodelle werden nutzerseitig ergänzt | Vollständige RVC-Oberfläche für Inferenz, Modellverwaltung und Training auf der RTX 5080. Für eine Konvertierung ist ein importiertes oder trainiertes `.pth`-Modell nötig; eine Referenzaufnahme allein reicht nicht. |
|
||||||
| `mike-ai-image-worker` | Qwen-Image-2.1 INT8, Qwen3-VL-8B INT8 und VAE | Produktiver Bildworker; erzeugt und bearbeitet Bilder transaktional auf der RTX 5080. |
|
| `mike-ai-image-worker` | Qwen-Image-2.1 INT8, Qwen3-VL-8B INT8 und VAE | Produktiver Bildworker; erzeugt und bearbeitet Bilder transaktional auf der RTX 5080. |
|
||||||
|
| `mike-ai-image-prompt-enhancer-t2i` | Qwen-Image-2.1 PE-T2I Q5_K_M | Kurzlebiger offizieller Prompt-Aufbereiter für reine Textaufträge auf der RTX 3060; außerhalb eines Bildauftrags gestoppt. |
|
||||||
|
| `mike-ai-image-prompt-enhancer-i2i` | Qwen-Image-2.1 PE-I2I Q5_K_M mit BF16-MMProj | Kurzlebiger offizieller Prompt-Aufbereiter für bis zu vier Referenzbilder auf der RTX 3060; außerhalb eines Bildauftrags gestoppt. |
|
||||||
| `mike-ai-flux-image-worker` | FLUX.2 Klein 9B FP8, Qwen3-8B NF4 Textencoder und VAE | Gestoppter Rückfallworker; wird vom normalen Router-Bildpfad nicht gestartet. |
|
| `mike-ai-flux-image-worker` | FLUX.2 Klein 9B FP8, Qwen3-8B NF4 Textencoder und VAE | Gestoppter Rückfallworker; wird vom normalen Router-Bildpfad nicht gestartet. |
|
||||||
| `mike-ai-llama-dashboard` | kein Modell | Zeigt Telemetrie, Profile, GPU-Nutzung, Slot-Kontextbelegung mit Verlauf und Betriebsarten an und bietet die Modusumschaltung. |
|
| `mike-ai-llama-dashboard` | kein Modell | Zeigt Telemetrie, Profile, GPU-Nutzung, Slot-Kontextbelegung mit Verlauf und Betriebsarten an und bietet die Modusumschaltung. |
|
||||||
| `mike-ai-llama-fast` | Qwen3.8-27B `IQ4-MIX`, Qwen-MMProj BF16 | Schnelles Q4-Text-/Vision-Profil mit 76.800 Token Kontext. |
|
| `mike-ai-llama-fast` | Qwen3.8-27B `IQ4-MIX`, Qwen-MMProj BF16 | Schnelles Q4-Text-/Vision-Profil mit 76.800 Token Kontext. |
|
||||||
| `mike-ai-llama-large` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Q4-Text-/Vision-Profil mit 192.000 Token Kontext und Verteilung auf beide GPUs. |
|
| `mike-ai-llama-large` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Q4-Text-/Vision-Profil mit 192.000 Token Kontext und Verteilung auf beide GPUs. |
|
||||||
| `mike-ai-llama-medium` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Standard-Q4-Text-/Vision-Profil mit gemeinsamem 160.000-Token-KV-Pool, aktuell zwei Slots und Verteilung auf beide GPUs. |
|
| `mike-ai-llama-medium` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Standard-Q4-Text-/Vision-Profil mit gemeinsamem 160.000-Token-KV-Pool, aktuell zwei Slots und Verteilung auf beide GPUs. |
|
||||||
| `mike-ai-llama-ultra` | Qwen3.8-27B `IQ4_XS-pure`, ohne Vision-Projektor | Maximales Langkontextprofil mit 262.144 Token Kontext und Verteilung auf beide GPUs. |
|
| `mike-ai-llama-ultra` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 auf CPU | Text-/Vision-Profil mit 262.144 Token Kontext und Verteilung des Textmodells auf beide GPUs. |
|
||||||
| `mike-ai-llama-uncensored` | Qwen3.8-27B Abliterated `Q4_K_M`, eigener MMProj F16 | Spezialprofil mit 80.000 Token Kontext und gelockerten Modellgrenzen. |
|
| `mike-ai-llama-uncensored` | Qwen3.8-27B Abliterated `Q4_K_M`, eigener MMProj F16 | Spezialprofil mit 80.000 Token Kontext und gelockerten Modellgrenzen. |
|
||||||
| `mike-ai-ltx2-studio` | LTX-Video-Backend | GPU-Worker für lokale Videogenerierung; beim Abgleich gestoppt. |
|
| `mike-ai-ltx2-studio` | LTX-Video-Backend | GPU-Worker für lokale Videogenerierung; beim Abgleich gestoppt. |
|
||||||
| `mike-ai-mcp-athena-operator` | kein Modell | Stellt Hermes begrenzte Werkzeuge zum Prüfen, Ändern, Testen, Sichern und Versionieren von Athena bereit. |
|
| `mike-ai-mcp-athena-operator` | kein Modell | Stellt Hermes begrenzte Werkzeuge zum Prüfen, Ändern, Testen, Sichern und Versionieren von Athena bereit. |
|
||||||
@@ -30,13 +32,14 @@ nicht automatisch ein ungenutzter Rest.
|
|||||||
| `mike-ai-portainer` | kein Modell; Portainer CE | Optionale Docker-Verwaltungsoberfläche. |
|
| `mike-ai-portainer` | kein Modell; Portainer CE | Optionale Docker-Verwaltungsoberfläche. |
|
||||||
| `mike-ai-profile-controller` | kein Modell | Startet und stoppt ausschließlich freigegebene Modellprofile und Spezialworker in einer sicheren Reihenfolge. |
|
| `mike-ai-profile-controller` | kein Modell | Startet und stoppt ausschließlich freigegebene Modellprofile und Spezialworker in einer sicheren Reihenfolge. |
|
||||||
| `mike-ai-qwen-cron-test` | kleines Qwen-Testmodell | Gestoppter CPU-/Cron-Worker-Versuch; nicht produktiv eingesetzt. |
|
| `mike-ai-qwen-cron-test` | kleines Qwen-Testmodell | Gestoppter CPU-/Cron-Worker-Versuch; nicht produktiv eingesetzt. |
|
||||||
| `mike-ai-realtime-voice` | kein Modell | Laufende, gesunde WebRTC-Brücke für OpenClaw Talk; verbindet über den privaten WireGuard-Pfad Athena Whisper, OpenClaw-Agent und Qwen3-TTS ohne Profilwechsel. |
|
| `mike-ai-realtime-voice` | kein Modell | Laufende, gesunde WebRTC-Brücke für OpenClaw Talk; verbindet über den privaten WireGuard-Pfad Athena Qwen3-ASR, OpenClaw-Agent und Qwen3-TTS ohne Profilwechsel. |
|
||||||
| `mike-ai-qwen3-tts` | `Qwen/Qwen3-TTS-12Hz-1.7B-Base`, Stimme Serena | Hochwertige deutsche Sprachausgabe auf der RTX 3060 im LLM-Betrieb. |
|
| `mike-ai-qwen3-tts` | `Qwen/Qwen3-TTS-12Hz-1.7B-Base`, Stimme Serena | Hochwertige deutsche Sprachausgabe auf der RTX 3060 im LLM-Betrieb. |
|
||||||
| `mike-ai-router` | kein eigenes Modell | Einzige OpenAI-kompatible Modelladresse; koordiniert Profile, Bildaufträge, Sprache und Betriebsarten. |
|
| `mike-ai-router` | kein eigenes Modell | Einzige OpenAI-kompatible Modelladresse; koordiniert Profile, Bildaufträge, Sprache und Betriebsarten. |
|
||||||
| `mike-ai-stem-separator` | BS-RoFormer Viperx 1297, `htdemucs_ft`, `htdemucs_6s`, `MossFormer2_SE_48K` | Trennt Gesang, Instrumente oder Sprache/Hintergrundgeräusche im exklusiven Separationsmodus. |
|
| `mike-ai-stem-separator` | BS-RoFormer Viperx 1297, `htdemucs_ft`, `htdemucs_6s`, `MossFormer2_SE_48K` | Trennt Gesang, Instrumente oder Sprache/Hintergrundgeräusche im exklusiven Separationsmodus. |
|
||||||
| `mike-ai-tts-gateway` | kein eigenes Modell | Normalisiert Text, konvertiert Ausgabeformate und stellt Qwen3-TTS sowie natives PCM-Streaming über eine stabile interne API bereit. |
|
| `mike-ai-tts-gateway` | kein eigenes Modell | Normalisiert Text, konvertiert Ausgabeformate und stellt Qwen3-TTS sowie natives PCM-Streaming über eine stabile interne API bereit. |
|
||||||
| `mike-ai-voice-studio` | `k2-fsa/OmniVoice` 0.2.1 mit Whisper-ASR | Erzeugt Text-to-Speech mit einer Referenzstimme; kein Audio-to-Audio-Voice-Changer. |
|
| `mike-ai-voice-studio` | `k2-fsa/OmniVoice` 0.2.1 mit Whisper-ASR | Erzeugt Text-to-Speech mit einer Referenzstimme; kein Audio-to-Audio-Voice-Changer. |
|
||||||
| `mike-ai-whisper` | Whisper.cpp 1.9.4, `ggml-small` | Lokale deutsche Spracherkennung auf der CPU über `/v1/audio/transcriptions`. |
|
| `mike-ai-qwen-asr` | Qwen3-ASR 0.6B Q8 | CPU-Inferenz für lokale deutsche Spracherkennung. |
|
||||||
|
| `mike-ai-qwen-asr-worker` | kein eigenes Modell | Audio-Adapter für `/v1/audio/transcriptions` mit dem Modellnamen `qwen3-asr`. |
|
||||||
| `mike-ai-wireguard-gateway` | kein Modell | Veröffentlicht Dashboard und Fachoberflächen ausschließlich über den privaten WireGuard-Pfad. |
|
| `mike-ai-wireguard-gateway` | kein Modell | Veröffentlicht Dashboard und Fachoberflächen ausschließlich über den privaten WireGuard-Pfad. |
|
||||||
| `mike-ai-xvc-studio` | `chenxie95/X-VC`, GLM-4-Voice-Tokenizer und optional Resemble Enhance | Wandelt eine vorhandene Sprachaufnahme anhand einer Referenzstimme in Audio zu Audio um; gibt das native 16-kHz-Ergebnis und optional eine neural restaurierte 44,1-kHz-Fassung aus. |
|
| `mike-ai-xvc-studio` | `chenxie95/X-VC`, GLM-4-Voice-Tokenizer und optional Resemble Enhance | Wandelt eine vorhandene Sprachaufnahme anhand einer Referenzstimme in Audio zu Audio um; gibt das native 16-kHz-Ergebnis und optional eine neural restaurierte 44,1-kHz-Fassung aus. |
|
||||||
| `mike-ai-yue2-playground` | Image `mike-ai/yue2:3b-0.1.6` | Vorhandener, gestoppter Playground; in diesem Abgleich nicht funktional getestet. |
|
| `mike-ai-yue2-playground` | Image `mike-ai/yue2:3b-0.1.6` | Vorhandener, gestoppter Playground; in diesem Abgleich nicht funktional getestet. |
|
||||||
@@ -45,7 +48,8 @@ nicht automatisch ein ungenutzter Rest.
|
|||||||
|
|
||||||
Alle fünf llama-Container verwenden llama.cpp 0.4.1 (`b29c606`). Medium lief
|
Alle fünf llama-Container verwenden llama.cpp 0.4.1 (`b29c606`). Medium lief
|
||||||
beim Abgleich gesund mit zwei Slots; die anderen vier Textprofile waren
|
beim Abgleich gesund mit zwei Slots; die anderen vier Textprofile waren
|
||||||
planmäßig nicht aktiv. Der Image-Worker war ebenfalls gestoppt. Die
|
planmäßig nicht aktiv. Der Image-Worker und beide Prompt-Enhancer waren
|
||||||
|
ebenfalls gestoppt. Die
|
||||||
Infrastruktur und die Musik-/Applio-Oberflächen liefen. Diese Zustände ändern
|
Infrastruktur und die Musik-/Applio-Oberflächen liefen. Diese Zustände ändern
|
||||||
sich mit der Nutzung.
|
sich mit der Nutzung.
|
||||||
Die Detailbeschreibungen der Spezialmodelle stammen aus der bestehenden
|
Die Detailbeschreibungen der Spezialmodelle stammen aus der bestehenden
|
||||||
|
|||||||
+31
-7
@@ -1,9 +1,24 @@
|
|||||||
# Geprüfter Live-Stand auf Athena
|
# Geprüfter Live-Stand auf Athena
|
||||||
|
|
||||||
|
Nachtrag vom 25. September 2026: Die produktive Spracherkennung läuft über
|
||||||
|
Qwen3-ASR 0.6B Q8 auf der CPU (`mike-ai-qwen-asr` und
|
||||||
|
`mike-ai-qwen-asr-worker`). Der Router bietet nur `qwen3-asr` als STT-Modell
|
||||||
|
an; OpenClaw und die Voice-Brücke verwenden denselben Namen. Eine M4A-Aufnahme
|
||||||
|
wurde über `/v1/audio/transcriptions` geprüft. Whisper-Container,
|
||||||
|
Images und Modellvolume wurden entfernt. Das aktive Ultra-Profil wurde bei
|
||||||
|
dieser Umstellung nicht gewechselt. Die Angaben zu Whisper weiter unten
|
||||||
|
beschreiben den historischen Stand vom 21. September.
|
||||||
|
|
||||||
|
Nachtrag vom 24. September 2026: Ultra verarbeitet nun Bilder mit einem
|
||||||
|
CPU-seitigen BF16-Vision-Projektor. Der Stand und der Funktionstest sind in
|
||||||
|
[Ultra-Vision mit CPU-Projektor](ULTRA_CPU_VISION_20260924.md) dokumentiert.
|
||||||
|
Die folgende Tabelle bildet weiterhin den historischen Stand vom 21. September ab.
|
||||||
|
|
||||||
Stand: **21. September 2026**. Quelle: lesender SSH-Abgleich von Docker,
|
Stand: **21. September 2026**. Quelle: lesender SSH-Abgleich von Docker,
|
||||||
Compose-Dateien, Git-Inhalten, Images, Health-Endpunkten und Backup-Timern.
|
Compose-Dateien, Git-Inhalten, Images, Health-Endpunkten und Backup-Timern.
|
||||||
Containerzustände sind Momentaufnahmen; der Controller darf Profile danach
|
Containerzustände sind Momentaufnahmen; der Controller darf Profile danach
|
||||||
umschalten. Es wurde für diesen Abgleich kein Dienst neu gestartet.
|
umschalten. Für die Prompt-Enhancer-Integration wurden Router und Controller
|
||||||
|
gezielt neu gebaut und beide Bildpfade transaktional geprüft.
|
||||||
|
|
||||||
Aktueller Abschluss: [Test- und Änderungsübersicht](ATHENA_SESSION_20260920.md).
|
Aktueller Abschluss: [Test- und Änderungsübersicht](ATHENA_SESSION_20260920.md).
|
||||||
Die folgenden MTP-/Microbatch-Werte enthalten die abschließenden Übernahmen.
|
Die folgenden MTP-/Microbatch-Werte enthalten die abschließenden Übernahmen.
|
||||||
@@ -41,7 +56,11 @@ nicht mehr den aktuellen Containerzustand.
|
|||||||
|
|
||||||
- Qwen-Image-2.1 INT8 läuft produktiv über gepinntes ComfyUI mit Low-VRAM auf
|
- Qwen-Image-2.1 INT8 läuft produktiv über gepinntes ComfyUI mit Low-VRAM auf
|
||||||
der RTX 5080. Der Router verwendet 25 Schritte, Guidance 1 und bis zu vier
|
der RTX 5080. Der Router verwendet 25 Schritte, Guidance 1 und bis zu vier
|
||||||
Referenzbilder. FLUX.2 Klein 9B FP8 bleibt mit seinen Gewichten als
|
Referenzbilder. Vor jedem Bild startet er automatisch den passenden
|
||||||
|
offiziellen Q5-Prompt-Enhancer: PE-T2I für reine Textaufträge oder PE-I2I
|
||||||
|
für Referenzbilder. Der Enhancer läuft kurzzeitig auf der RTX 3060, wird vor
|
||||||
|
dem Rendern wieder gestoppt und speichert Rohprompt sowie Rewrite im
|
||||||
|
Bild-Sidecar. FLUX.2 Klein 9B FP8 bleibt mit seinen Gewichten als
|
||||||
gestoppter, explizit allowlist-beschränkter Rückfallcontainer erhalten.
|
gestoppter, explizit allowlist-beschränkter Rückfallcontainer erhalten.
|
||||||
- Qwen3-TTS 1.7B auf 3060 hinter dem TTS-Gateway; kein Piper-Fallback.
|
- Qwen3-TTS 1.7B auf 3060 hinter dem TTS-Gateway; kein Piper-Fallback.
|
||||||
- Whisper.cpp `ggml-small` auf CPU über `/v1/audio/transcriptions`.
|
- Whisper.cpp `ggml-small` auf CPU über `/v1/audio/transcriptions`.
|
||||||
@@ -59,7 +78,12 @@ nicht mehr den aktuellen Containerzustand.
|
|||||||
Er braucht keine GPU und keinen Profilwechsel. Das Plugin liegt in
|
Er braucht keine GPU und keinen Profilwechsel. Das Plugin liegt in
|
||||||
`integrations/openclaw-athena-talk` und läuft auf Unraid, nicht auf Athena.
|
`integrations/openclaw-athena-talk` und läuft auf Unraid, nicht auf Athena.
|
||||||
Diktat und hochgeladene M4A-Sprachnachrichten nutzen den separaten
|
Diktat und hochgeladene M4A-Sprachnachrichten nutzen den separaten
|
||||||
Transkriptionspfad des Plugins. OpenClaw liefert den fertigen Agententext
|
Transkriptionspfad des Plugins. Plugin 1.3.0 zerlegt längere Browser-Diktate
|
||||||
|
während der Aufnahme in überlappende Sechs-Sekunden-Abschnitte und hält so
|
||||||
|
den Abschluss innerhalb von OpenClaws festem Fünf-Sekunden-Fenster. OpenClaw
|
||||||
|
selbst wurde dafür nicht gepatcht. Die lokale Sicherheitsgrenze
|
||||||
|
`maxSpeechSeconds` steht produktiv auf 180 Sekunden. OpenClaw liefert den
|
||||||
|
fertigen Agententext
|
||||||
an die Brücke; das Sprechen beginnt daher erst nach Abschluss der
|
an die Brücke; das Sprechen beginnt daher erst nach Abschluss der
|
||||||
Agentenantwort. Details und Grenzen stehen in der [Voice-Doku](../services/athena-realtime-voice/README.md).
|
Agentenantwort. Details und Grenzen stehen in der [Voice-Doku](../services/athena-realtime-voice/README.md).
|
||||||
- `mike-ai-mikes-applio-ui` läuft gesund und ohne GPU. Die Quelle liegt im
|
- `mike-ai-mikes-applio-ui` läuft gesund und ohne GPU. Die Quelle liegt im
|
||||||
@@ -68,7 +92,7 @@ nicht mehr den aktuellen Containerzustand.
|
|||||||
trotzdem Status, Modelle und vorbereitende Aufgaben anzeigen; eine echte
|
trotzdem Status, Modelle und vorbereitende Aufgaben anzeigen; eine echte
|
||||||
Konvertierung verlangt den Applio-Modus.
|
Konvertierung verlangt den Applio-Modus.
|
||||||
|
|
||||||
Der vollständige Bestand der **31** angelegten Docker-Container steht im
|
Der vollständige Bestand der **33** angelegten Docker-Container steht im
|
||||||
[Container-Inventar](CONTAINER_INVENTORY.md). `Created` oder `Exited` bei
|
[Container-Inventar](CONTAINER_INVENTORY.md). `Created` oder `Exited` bei
|
||||||
Spezialworkern bedeutet nicht automatisch einen zu entfernenden Testrest.
|
Spezialworkern bedeutet nicht automatisch einen zu entfernenden Testrest.
|
||||||
|
|
||||||
@@ -104,7 +128,7 @@ externe Restic-Timer zwar ebenfalls aktiviert, aber ohne seine erforderliche
|
|||||||
`/etc/mike-ai/disaster-backup.env` **nicht** als funktionierende externe
|
`/etc/mike-ai/disaster-backup.env` **nicht** als funktionierende externe
|
||||||
Sicherung zu werten. Details und Restore-Grenzen: [Backup](RECOVERY.md).
|
Sicherung zu werten. Details und Restore-Grenzen: [Backup](RECOVERY.md).
|
||||||
|
|
||||||
Geprüft wurden aktuelle Container- und Quellenstände sowie ausgewählte
|
Geprüft wurden aktuelle Container- und Quellenstände, Health-Endpunkte sowie
|
||||||
Health-Endpunkte. Keine erneuten Bild-, Langkontext-, Trainings- oder
|
Text-zu-Bild und Referenzbildbearbeitung mit den beiden Prompt-Enhancern. Keine
|
||||||
Restore-Tests; ältere Ergebnisse sind im [Update-Audit](UPDATE_AUDIT_20260915.md)
|
erneuten Langkontext-, Trainings- oder Restore-Tests; ältere Ergebnisse sind im [Update-Audit](UPDATE_AUDIT_20260915.md)
|
||||||
und [FLUX-Auflösungstest](FLUX_RESOLUTION_TEST_20260912.md) dokumentiert.
|
und [FLUX-Auflösungstest](FLUX_RESOLUTION_TEST_20260912.md) dokumentiert.
|
||||||
+86
-7
@@ -31,10 +31,14 @@ Ein Bildauftrag läuft transaktional:
|
|||||||
|
|
||||||
1. Der Router merkt sich das aktive Textprofil.
|
1. Der Router merkt sich das aktive Textprofil.
|
||||||
2. Der Profile Controller stoppt LLM, TTS und andere GPU-Spezialworker.
|
2. Der Profile Controller stoppt LLM, TTS und andere GPU-Spezialworker.
|
||||||
3. `mike-ai-image-worker` startet Qwen Image auf der RTX 5080.
|
3. Bei einem Textauftrag startet `mike-ai-image-prompt-enhancer-t2i`, bei
|
||||||
4. Der Adapter erzeugt das Bild oder bearbeitet bis zu vier Referenzbilder.
|
Referenzbildern `mike-ai-image-prompt-enhancer-i2i` auf der RTX 3060.
|
||||||
5. Der Qwen-Worker wird vollständig gestoppt.
|
4. Der offizielle Qwen-Enhancer schreibt den knappen Benutzertext um und wird
|
||||||
6. Das vorherige Textprofil und Qwen3-TTS werden wiederhergestellt.
|
anschließend vollständig gestoppt.
|
||||||
|
5. `mike-ai-image-worker` startet Qwen Image auf der RTX 5080 und erzeugt das
|
||||||
|
Bild oder bearbeitet bis zu vier Referenzbilder.
|
||||||
|
6. Der Qwen-Worker wird vollständig gestoppt.
|
||||||
|
7. Das vorherige Textprofil und Qwen3-TTS werden wiederhergestellt.
|
||||||
|
|
||||||
Der Container ist außerhalb eines Auftrags gestoppt. Das ist der Sollzustand.
|
Der Container ist außerhalb eines Auftrags gestoppt. Das ist der Sollzustand.
|
||||||
|
|
||||||
@@ -46,13 +50,14 @@ sudo ./scripts/prepare-qwen-image-21.sh
|
|||||||
```
|
```
|
||||||
|
|
||||||
Das Skript lädt fehlende Dateien mit festen SHA-256-Prüfsummen, baut den
|
Das Skript lädt fehlende Dateien mit festen SHA-256-Prüfsummen, baut den
|
||||||
produktiven Qwen-Worker und legt Qwen sowie FLUX gestoppt an. Es lädt dabei
|
produktiven Qwen-Worker und legt Qwen, beide Prompt-Enhancer sowie FLUX
|
||||||
kein Modell in den VRAM.
|
gestoppt an. Es lädt dabei kein Modell in den VRAM.
|
||||||
|
|
||||||
## Funktionsprobe über den echten Routerweg
|
## Funktionsprobe über den echten Routerweg
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
curl -sS http://127.0.0.1:8081/v1/images/generations \
|
curl -sS http://127.0.0.1:8081/v1/images/generations \
|
||||||
|
-H "Authorization: Bearer $ROUTER_API_KEY" \
|
||||||
-H 'Content-Type: application/json' \
|
-H 'Content-Type: application/json' \
|
||||||
-d '{
|
-d '{
|
||||||
"model":"Qwen-Image-2.1-int8",
|
"model":"Qwen-Image-2.1-int8",
|
||||||
@@ -64,7 +69,7 @@ curl -sS http://127.0.0.1:8081/v1/images/generations \
|
|||||||
}'
|
}'
|
||||||
```
|
```
|
||||||
|
|
||||||
Nach dem Lauf müssen `mike-ai-image-worker` und
|
Nach dem Lauf müssen `mike-ai-image-worker`, beide Prompt-Enhancer und
|
||||||
`mike-ai-flux-image-worker` gestoppt sein und das vorherige LLM-Profil wieder
|
`mike-ai-flux-image-worker` gestoppt sein und das vorherige LLM-Profil wieder
|
||||||
gesund laufen.
|
gesund laufen.
|
||||||
|
|
||||||
@@ -150,3 +155,77 @@ und anschließend Router und Controller neu ausgerollt werden. Keine dieser
|
|||||||
Historische FLUX-Messwerte und Grenzen stehen in
|
Historische FLUX-Messwerte und Grenzen stehen in
|
||||||
[FLUX_9B_BETA.md](FLUX_9B_BETA.md) und
|
[FLUX_9B_BETA.md](FLUX_9B_BETA.md) und
|
||||||
[FLUX_RESOLUTION_TEST_20260912.md](FLUX_RESOLUTION_TEST_20260912.md).
|
[FLUX_RESOLUTION_TEST_20260912.md](FLUX_RESOLUTION_TEST_20260912.md).
|
||||||
|
|
||||||
|
## Offizielle Prompt-Enhancer im produktiven Router
|
||||||
|
|
||||||
|
Qwen veröffentlicht zwei getrennte, auf Qwen3.5-VL 9B nachtrainierte
|
||||||
|
Prompt-Enhancer. Sie sind kein Bestandteil der Qwen-Image-Gewichte. Athena
|
||||||
|
startet automatisch die zum Auftrag passende Variante; OpenClaw ruft weiterhin
|
||||||
|
nur sein normales Werkzeug `image_generate` mit dem unveränderten Modellnamen
|
||||||
|
`openai/Qwen-Image-2.1-int8` auf.
|
||||||
|
|
||||||
|
- `mike-ai-image-prompt-enhancer-t2i` bereitet reine Textaufträge auf.
|
||||||
|
- `mike-ai-image-prompt-enhancer-i2i` wertet ein bis vier Referenzbilder
|
||||||
|
gemeinsam aus und trennt Identität, gewünschte Änderungen und neue Szene.
|
||||||
|
- Beide verwenden getrennte offizielle Systemprompts, `Q5_K_M`, llama.cpp
|
||||||
|
Build 10930, 16.384 Token Kontext und höchstens 2.048 Denktokens.
|
||||||
|
- Sie laufen ausschließlich und nacheinander auf der RTX 3060. Nach dem Rewrite
|
||||||
|
werden sie gestoppt; Qwen-Image rendert anschließend auf der RTX 5080.
|
||||||
|
- Falls der Client keine Größe vorgibt, übernimmt der Router das vom Enhancer
|
||||||
|
vorgeschlagene Seitenverhältnis und bildet es auf eine erlaubte Auflösung ab.
|
||||||
|
Eine ausdrücklich angegebene gültige Größe bleibt unverändert.
|
||||||
|
- Das Sidecar jedes PNG enthält `original_prompt`, den tatsächlich verwendeten
|
||||||
|
`prompt` und Modell, Quantisierung, Laufzeit und Verhältnis unter
|
||||||
|
`prompt_enhancer`. Verdeckte Denktokens werden nicht gespeichert.
|
||||||
|
- Ein fehlgeschlagener Rewrite bricht den Bildauftrag sichtbar ab. Der Router
|
||||||
|
sendet in diesem Fall keinen unaufbereiteten Ersatzprompt an Qwen-Image.
|
||||||
|
|
||||||
|
### Gepinnte Dateien
|
||||||
|
|
||||||
|
PE-I2I stammt aus `Qwen/Qwen-Image-2.1-PE-I2I`, Revision
|
||||||
|
`72927bc08afc99b7888ceb7d7d51a12db3700bbd`, und dem GGUF-Repository
|
||||||
|
`prithivMLmods/Qwen-Image-2.1-PE-I2I-GGUF`, Revision
|
||||||
|
`55b9c1a326599e142d59bcad8715d5601ccf8daa`. Die Dateien liegen unter
|
||||||
|
`/data/models/qwen-image-2.1-pe-i2i-q5`:
|
||||||
|
|
||||||
|
| Datei | SHA-256 |
|
||||||
|
|---|---|
|
||||||
|
| `Qwen-Image-2.1-PE-I2I.Q5_K_M.gguf` | `cb71f71fe5fe5938570d65a7c403fbcb1a9065dff9dd9a021f58aec42785b84e` |
|
||||||
|
| `Qwen-Image-2.1-PE-I2I.mmproj-bf16.gguf` | `8dedb71dbc3092dc47de9108ad373d68a12854e2527d59bd9399601738f3bce1` |
|
||||||
|
| `system_prompt.txt` | `e378fea686a1431581ba4c654d332ae96adad633f144ae738ec8ce9c4fd66439` |
|
||||||
|
|
||||||
|
PE-T2I stammt aus `Qwen/Qwen-Image-2.1-PE-T2I`, Revision
|
||||||
|
`f3ed7985c788ad75b3ab7223e0c4c51e2a43545b`, und dem GGUF-Repository
|
||||||
|
`prithivMLmods/Qwen-Image-2.1-PE-T2I-GGUF`, Revision
|
||||||
|
`e18d4a3e0830ab157770738b16830e6fcf5f57d4`. Die Dateien liegen unter
|
||||||
|
`/data/models/qwen-image-2.1-pe-t2i-q5`:
|
||||||
|
|
||||||
|
| Datei | SHA-256 |
|
||||||
|
|---|---|
|
||||||
|
| `Qwen-Image-2.1-PE-T2I.Q5_K_M.gguf` | `749f5652fd6e8b760ca860091f7a5bffa254a7954203ae5c9bcdf5f93197702f` |
|
||||||
|
| `system_prompt.txt` | `a77c9a06c59b120741141d9514b95682bb8761d02bec49ca61def7b2b3d9fb99` |
|
||||||
|
|
||||||
|
Der zuvor geprüfte PE-I2I-FP8-Checkpoint ist 13,54 GB groß und passt mit
|
||||||
|
Aktivierungs- und Laufzeitspeicher nicht in die 12-GB-RTX-3060. CPU-Offload war
|
||||||
|
unpraktisch langsam. Er wurde verworfen und entfernt. Der Q5-I2I-Worker belegt
|
||||||
|
etwa 7.167 MiB VRAM und ist deshalb der produktive Stand.
|
||||||
|
|
||||||
|
### Abnahme vom 21. September 2026
|
||||||
|
|
||||||
|
Der vollständige öffentliche Routerweg wurde nach der Integration zweimal
|
||||||
|
geprüft:
|
||||||
|
|
||||||
|
- Ein kurzer deutscher Textprompt wurde von PE-T2I in 36,6 Sekunden in einen
|
||||||
|
strukturierten englischen Produktionsprompt umgeschrieben. Qwen-Image
|
||||||
|
erzeugte das 1024×1024-PNG anschließend in 49,0 Sekunden.
|
||||||
|
- Ein hochgeladenes Referenzbild wurde von PE-I2I in 65,9 Sekunden aufbereitet.
|
||||||
|
Der Rewrite verlangte ausdrücklich eine neue Szene, neue Pose und Kleidung
|
||||||
|
sowie das Entfernen der alten Gegenstände. Qwen-Image erzeugte in 111,3
|
||||||
|
Sekunden ein neues 1024×1536-PNG: genau eine Person in einer Blumenwiese an
|
||||||
|
einer Autobahn, ohne Badewanne, Gitarre oder Gummiente.
|
||||||
|
|
||||||
|
Nach beiden Aufträgen waren der Bildworker und beide Enhancer gestoppt. Medium
|
||||||
|
und Qwen3-TTS liefen wieder gesund. Der erste I2I-Versuch deckte eine relative
|
||||||
|
Pfadauflösung bei temporären Multipart-Uploads auf; der korrigierte Pfad ist mit
|
||||||
|
einem eigenen Regressionstest abgedeckt. Insgesamt laufen 96 schnelle
|
||||||
|
GPU-freie Release-Tests.
|
||||||
+21
-2
@@ -16,13 +16,25 @@ eingeschlossen: Im Archiv `athena-2026-09-16T13-02-56.tar.gz` wurden sowohl
|
|||||||
`compose.yaml` als auch `services/athena-realtime-voice/server.py` geprüft.
|
`compose.yaml` als auch `services/athena-realtime-voice/server.py` geprüft.
|
||||||
Das OpenClaw-Plugin auf Unraid liegt außerhalb dieses Athena-Backups; seine
|
Das OpenClaw-Plugin auf Unraid liegt außerhalb dieses Athena-Backups; seine
|
||||||
Quelle ist im Git-Repository unter `integrations/openclaw-athena-talk` erfasst.
|
Quelle ist im Git-Repository unter `integrations/openclaw-athena-talk` erfasst.
|
||||||
|
Die produktiv installierte Version 1.3.0 liegt zusätzlich im persistenten
|
||||||
|
OpenClaw-Appdata. Vor ihrer Installation wurde die bisherige Version als
|
||||||
|
`/mnt/nvme-storage/appdata/OpenClaw/config/plugin-backups/athena-talk-1.2.1-before-streaming.tar.gz`
|
||||||
|
gesichert. Für eine Neuinstallation ist der im Git dokumentierte Build mit
|
||||||
|
`openclaw plugins install <paket.tgz> --force --accept-capabilities` zu
|
||||||
|
installieren; eine Änderung an OpenClaw-Core-Dateien ist nicht erforderlich.
|
||||||
|
|
||||||
Piper-Daten sind kein aktueller Sicherungsbestand. Modellgewichte unter
|
Piper-Daten sind kein aktueller Sicherungsbestand. Modellgewichte unter
|
||||||
`/data/models` und das reproduzierbare Whisper-Volume gehören nicht zu diesen
|
`/data/models`, einschließlich Qwen3-ASR unter
|
||||||
Backup-Mounts. Ein Backup ausschließlich auf `/data` schützt nicht vor deren Ausfall.
|
`/data/models/qwen3-asr-0.6b-q8`, gehören nicht zu diesen Backup-Mounts.
|
||||||
|
Das frühere Whisper-Volume wurde am 25. September 2026 entfernt.
|
||||||
|
Ein Backup ausschließlich auf `/data` schützt nicht vor einem Ausfall der Datenplatte.
|
||||||
Das gilt auch für das reproduzierbare EmbeddingGemma-Gewicht unter
|
Das gilt auch für das reproduzierbare EmbeddingGemma-Gewicht unter
|
||||||
`/data/models/embeddinggemma`; URL und SHA-256 stehen in
|
`/data/models/embeddinggemma`; URL und SHA-256 stehen in
|
||||||
`config/install.env.example`, sodass der Installer es erneut laden und prüfen kann.
|
`config/install.env.example`, sodass der Installer es erneut laden und prüfen kann.
|
||||||
|
Auch die Qwen-Image-Gewichte und beide Prompt-Enhancer liegen unter
|
||||||
|
`/data/models` und damit außerhalb des regulären Volume-Backups. Das Skript
|
||||||
|
`scripts/prepare-qwen-image-21.sh` lädt sämtliche gepinnten Dateien anhand
|
||||||
|
fester SHA-256-Prüfsummen erneut und legt Bildworker sowie Enhancer gestoppt an.
|
||||||
|
|
||||||
Die verschlüsselten Notfallpakete über `athena-export-backup.timer` laufen
|
Die verschlüsselten Notfallpakete über `athena-export-backup.timer` laufen
|
||||||
ebenfalls im Fünf-Stunden-Takt; beim Abgleich war der Timer aktiv und der
|
ebenfalls im Fünf-Stunden-Takt; beim Abgleich war der Timer aktiv und der
|
||||||
@@ -30,6 +42,13 @@ letzte Lauf am 16. September 2026 um 13:50 Uhr verzeichnet. Sie
|
|||||||
liegen unter `/data/emergency-backups`. Schutz vor Datenplattenausfall setzt
|
liegen unter `/data/emergency-backups`. Schutz vor Datenplattenausfall setzt
|
||||||
eine außerhalb Athenas aufbewahrte Kopie voraus.
|
eine außerhalb Athenas aufbewahrte Kopie voraus.
|
||||||
|
|
||||||
|
Die lokale Rotation hält fünf verschlüsselte Generationen. Da ein Paket rund
|
||||||
|
46 GB umfasst, gibt das Exportscript bei bereits erreichter Aufbewahrungszahl
|
||||||
|
vor dem Schreiben genau den ältesten Slot frei. Ohne diese Reihenfolge hätte
|
||||||
|
der Lauf am 21. September bei nur noch 17 GB freiem Speicher die neue Datei
|
||||||
|
nicht mehr vollständig schreiben können. Scheitert der neue Lauf danach,
|
||||||
|
bleiben weiterhin vier gültige Generationen erhalten.
|
||||||
|
|
||||||
Die externe Restic-Sicherung über `athena-disaster-backup.timer` ist derzeit
|
Die externe Restic-Sicherung über `athena-disaster-backup.timer` ist derzeit
|
||||||
nicht eingerichtet: Der Timer ist zwar aktiviert, aber
|
nicht eingerichtet: Der Timer ist zwar aktiviert, aber
|
||||||
`/etc/mike-ai/disaster-backup.env` fehlt. Bis ein externes
|
`/etc/mike-ai/disaster-backup.env` fehlt. Bis ein externes
|
||||||
|
|||||||
@@ -9,7 +9,7 @@ Standardprofil: **medium** · globales Ausgabelimit: **8192 Token**
|
|||||||
| fast | `qwen-fast` | 76,800 | 1 | Qwen3.8-27B IQ4 Mix | 5080 only | ja | 2 |
|
| fast | `qwen-fast` | 76,800 | 1 | Qwen3.8-27B IQ4 Mix | 5080 only | ja | 2 |
|
||||||
| medium | `qwen-medium` | 160,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 85:15 | ja | 2 |
|
| medium | `qwen-medium` | 160,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 85:15 | ja | 2 |
|
||||||
| large | `qwen-large` | 192,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 86:14 | ja | 2 |
|
| large | `qwen-large` | 192,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 86:14 | ja | 2 |
|
||||||
| ultra | `qwen-ultra` | 262,144 | 1 | Qwen3.8-27B IQ4 XS Pure | 80:20 | nein | 2 |
|
| ultra | `qwen-ultra` | 262,144 | 1 | Qwen3.8-27B IQ4 XS Pure | 80:20 | ja | 2 |
|
||||||
| uncensored | `qwen-uncensored` | 80,000 | 1 | Qwen3.8-27B Abliterated Q4_K_M | 90:10 | ja | 2 |
|
| uncensored | `qwen-uncensored` | 80,000 | 1 | Qwen3.8-27B Abliterated Q4_K_M | 90:10 | ja | 2 |
|
||||||
|
|
||||||
## Zweck
|
## Zweck
|
||||||
@@ -17,5 +17,5 @@ Standardprofil: **medium** · globales Ausgabelimit: **8192 Token**
|
|||||||
- **fast**: Schnelles Profil für kurze Chats und zügige Werkzeugaufgaben.
|
- **fast**: Schnelles Profil für kurze Chats und zügige Werkzeugaufgaben.
|
||||||
- **medium**: Ausgewogenes Standardprofil für Alltag und lange agentische Aufgaben.
|
- **medium**: Ausgewogenes Standardprofil für Alltag und lange agentische Aufgaben.
|
||||||
- **large**: Großes Profil für umfangreiche Dokumente und lange technische Arbeiten.
|
- **large**: Großes Profil für umfangreiche Dokumente und lange technische Arbeiten.
|
||||||
- **ultra**: Maximaler Textkontext; bewusst ohne Vision-Projektor.
|
- **ultra**: Maximaler Kontext mit Vision-Projektor auf der CPU.
|
||||||
- **uncensored**: Weniger restriktives Spezialprofil; Werkzeugrechte bleiben unverändert.
|
- **uncensored**: Weniger restriktives Spezialprofil; Werkzeugrechte bleiben unverändert.
|
||||||
@@ -1,6 +1,6 @@
|
|||||||
# Register getesteter Modelle
|
# Register getesteter Modelle
|
||||||
|
|
||||||
Stand: 20. September 2026
|
Stand: 21. September 2026
|
||||||
|
|
||||||
Dieses Dokument ist die zentrale Sperrliste gegen doppelte Modelltests. Vor
|
Dieses Dokument ist die zentrale Sperrliste gegen doppelte Modelltests. Vor
|
||||||
jedem Download müssen Repository, Dateiname, Basismodell, Fine-Tune und
|
jedem Download müssen Repository, Dateiname, Basismodell, Fine-Tune und
|
||||||
@@ -55,6 +55,8 @@ Titelgenerierung und Kontextkompression in Hermes.
|
|||||||
| bis 07.09.2026 | FLUX.2 Klein 4B | funktional, aber schwächere räumliche und motivische Konsistenz | **ersetzt** durch 9B FP8 | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
| bis 07.09.2026 | FLUX.2 Klein 4B | funktional, aber schwächere räumliche und motivische Konsistenz | **ersetzt** durch 9B FP8 | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
||||||
| 07.09.2026 | FLUX.2 Klein 9B FP8 | bessere Prompttreue als 4B; produktiver Zwei-GPU-Pfad, auf 1024 × 1024 begrenzt | **Standby seit 21.09.2026** | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
| 07.09.2026 | FLUX.2 Klein 9B FP8 | bessere Prompttreue als 4B; produktiver Zwei-GPU-Pfad, auf 1024 × 1024 begrenzt | **Standby seit 21.09.2026** | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
||||||
| 21.09.2026 | Qwen-Image-2.1 INT8 | im direkten Vergleich fotorealistischer, bessere Textdarstellung und Prompttreue; 1024 × 1024 bei 25 Schritten in rund 49 s Pipeline-Lauf | **produktiv** | [QWEN_IMAGE_21.md](QWEN_IMAGE_21.md) |
|
| 21.09.2026 | Qwen-Image-2.1 INT8 | im direkten Vergleich fotorealistischer, bessere Textdarstellung und Prompttreue; 1024 × 1024 bei 25 Schritten in rund 49 s Pipeline-Lauf | **produktiv** | [QWEN_IMAGE_21.md](QWEN_IMAGE_21.md) |
|
||||||
|
| 21.09.2026 | Qwen-Image-2.1 PE-I2I, FP8 und Q5_K_M | FP8 passt nicht vollständig in 12 GB und ist mit CPU-Offload unpraktisch langsam; Q5_K_M samt BF16-MMProj belegt 7.167 MiB auf der RTX 3060. Der produktive Router-Rewrite mit einem Referenzbild dauerte 65,9 s und führte zu einer vollständig neuen Szene | **Q5_K_M produktiv**, FP8 verworfen | [QWEN_IMAGE_21.md](QWEN_IMAGE_21.md#offizielle-prompt-enhancer-im-produktiven-router) |
|
||||||
|
| 21.09.2026 | Qwen-Image-2.1 PE-T2I Q5_K_M | automatischer Rewrite eines knappen deutschen Textprompts in 36,6 s; nachfolgende Bildgenerierung erfolgreich | **produktiv** für Text-zu-Bild | [QWEN_IMAGE_21.md](QWEN_IMAGE_21.md#offizielle-prompt-enhancer-im-produktiven-router) |
|
||||||
| 08.09.2026 | HYPIR-SD2 | glättete oder erfand Details und veränderte kleine Strukturen | **verworfen** | [IMAGE_RESTORATION.md](IMAGE_RESTORATION.md) |
|
| 08.09.2026 | HYPIR-SD2 | glättete oder erfand Details und veränderte kleine Strukturen | **verworfen** | [IMAGE_RESTORATION.md](IMAGE_RESTORATION.md) |
|
||||||
| 08.09.2026 | SeedVR2 7B FP8 | bewahrte Identität besser als HYPIR, brachte beim realen unscharfen Foto aber kaum nutzbare Details zurück | **verworfen** | [IMAGE_RESTORATION.md](IMAGE_RESTORATION.md) |
|
| 08.09.2026 | SeedVR2 7B FP8 | bewahrte Identität besser als HYPIR, brachte beim realen unscharfen Foto aber kaum nutzbare Details zurück | **verworfen** | [IMAGE_RESTORATION.md](IMAGE_RESTORATION.md) |
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,23 @@
|
|||||||
|
# Ultra-Vision mit CPU-Projektor
|
||||||
|
|
||||||
|
Stand: 24. September 2026. Ultra verwendet weiterhin Qwen3.8-27B IQ4_XS Pure
|
||||||
|
mit 262.144 Token Kontext. Der BF16-Vision-Projektor wird mit `--mmproj`
|
||||||
|
geladen und durch `--no-mmproj-offload` auf der CPU gehalten. So kann Ultra
|
||||||
|
Bilder verarbeiten, ohne den knapp bemessenen GPU-Speicher des 256K-Profils
|
||||||
|
zusätzlich mit dem Projektor zu belegen. Router und Profilmatrix melden Ultra
|
||||||
|
als vision-fähig.
|
||||||
|
|
||||||
|
Der frühere text-only-Modus war die Ursache dafür, dass der Router Bildanfragen
|
||||||
|
an `qwen-ultra` abwies. Ein Test über den Router nach dem Deployment lieferte
|
||||||
|
HTTP 200 für ein einzelnes synthetisches Farbbild (`Red`) und für fünf
|
||||||
|
synthetische Bilder (`5`). Der Ultra-Start lud den multimodalen Projektor.
|
||||||
|
Gemessen wurden 11,47 Sekunden bis zur Betriebsbereitschaft und 1,66 bzw.
|
||||||
|
0,9 Sekunden für die beiden kleinen Testanfragen. Große Screenshots und lange
|
||||||
|
Kontexte wurden damit nicht vermessen; deren Laufzeit kann deutlich höher
|
||||||
|
sein. Nach dem Test wurde Medium wieder aktiviert, Ultra ist gestoppt.
|
||||||
|
|
||||||
|
Vor der Änderung wurden die drei produktiven Dateien auf Athena unter
|
||||||
|
`/var/tmp/ultra-vision-cpu-20260924/` gesichert. Ein Rückbau muss Compose,
|
||||||
|
Router-Profilregister und Profilmatrix gemeinsam auf den vorigen Stand setzen
|
||||||
|
und den Router neu laden. Das Verzeichnis liegt nur temporär auf dem Host;
|
||||||
|
die dauerhafte Versionierung erfolgt im Git-Repository.
|
||||||
@@ -5,7 +5,7 @@ existing Athena speech stack. An experimental browser WebRTC path is available
|
|||||||
through the separate `services/athena-realtime-voice` service:
|
through the separate `services/athena-realtime-voice` service:
|
||||||
|
|
||||||
1. local VAD collects a spoken utterance,
|
1. local VAD collects a spoken utterance,
|
||||||
2. Athena Whisper transcribes it,
|
2. Athena Qwen3-ASR transcribes it,
|
||||||
3. OpenClaw's normal agent-consult path answers with its configured model and
|
3. OpenClaw's normal agent-consult path answers with its configured model and
|
||||||
tools,
|
tools,
|
||||||
4. Athena Qwen3-TTS returns PCM audio to the Talk client.
|
4. Athena Qwen3-TTS returns PCM audio to the Talk client.
|
||||||
@@ -43,7 +43,7 @@ Recommended `talk.realtime` configuration:
|
|||||||
"vadThreshold": 0.018,
|
"vadThreshold": 0.018,
|
||||||
"silenceDurationMs": 750,
|
"silenceDurationMs": 750,
|
||||||
"prefixPaddingMs": 300,
|
"prefixPaddingMs": 300,
|
||||||
"maxSpeechSeconds": 45
|
"maxSpeechSeconds": 180
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -60,19 +60,35 @@ openclaw plugins install . --force --accept-capabilities
|
|||||||
openclaw plugins inspect athena-talk --runtime --json
|
openclaw plugins inspect athena-talk --runtime --json
|
||||||
```
|
```
|
||||||
|
|
||||||
Version 1.2.1 also registers **Athena Whisper (Diktieren)** as a separate
|
Version 1.3.1 sends `qwen3-asr` throughout dictation and Talk. The plugin
|
||||||
|
registers **Athena Qwen3-ASR (Diktieren)** as a separate
|
||||||
realtime transcription provider through OpenClaw's official plugin API. In the
|
realtime transcription provider through OpenClaw's official plugin API. In the
|
||||||
browser composer, hold the microphone for dictation, then release it to send
|
browser composer, hold the microphone for dictation, then release it to send
|
||||||
the 8 kHz G.711 audio through the Gateway. The plugin converts it to PCM WAV
|
the 8 kHz G.711 audio through the Gateway. Short recordings are converted to
|
||||||
and calls the same Athena `/audio/transcriptions` endpoint used by Talk. The
|
PCM WAV and sent to Athena's existing `/audio/transcriptions` endpoint in one
|
||||||
transcribed text is returned to the composer; this path does not invoke the
|
request. Longer recordings are split while the user is still speaking into
|
||||||
agent or TTS. The transcription provider reuses `talk.realtime.providers.athena-talk`
|
six-second windows with 0.5 seconds of overlap. The plugin sends these windows
|
||||||
and the configured model provider for its URL/key. If that model provider has
|
sequentially to the persistent Qwen3-ASR service, removes duplicated overlap
|
||||||
no key, it reuses `tts.providers.openai.apiKey` only when the TTS and STT URLs
|
words, and caches finished
|
||||||
have the same origin. No second credential is needed. In `talk.catalog`, it
|
segments until recording stops. Only the short final tail then remains inside
|
||||||
appears under `transcription.providers`. OpenClaw
|
OpenClaw's fixed five-second final-drain window. Each STT request is capped
|
||||||
currently gives a transcription provider five seconds to return its final text
|
at 4.5 seconds.
|
||||||
after recording stops; the plugin caps its Whisper request at 4.5 seconds.
|
|
||||||
|
This is incremental pre-transcription over OpenClaw's official transcription
|
||||||
|
provider API. Qwen3-ASR still receives complete short WAV segments; it is
|
||||||
|
not a native token-streaming STT protocol. No OpenClaw core file was patched.
|
||||||
|
The transcribed text is
|
||||||
|
returned to the composer; this path does not invoke the agent or TTS. The
|
||||||
|
provider reuses `talk.realtime.providers.athena-talk` and the configured model
|
||||||
|
provider for its URL/key. If that model provider has no key, it reuses
|
||||||
|
`tts.providers.openai.apiKey` only when the TTS and STT URLs have the same
|
||||||
|
origin. No second credential is needed. In `talk.catalog`, it appears under
|
||||||
|
`transcription.providers`.
|
||||||
|
|
||||||
|
The production provider allows up to 180 seconds per recording. This limit is
|
||||||
|
a local safety cap shared by dictation and Talk, not an OpenClaw or Qwen3-ASR
|
||||||
|
restriction. Incremental segmentation keeps long dictation bounded while it
|
||||||
|
is being recorded.
|
||||||
|
|
||||||
### Voice-note file attachments
|
### Voice-note file attachments
|
||||||
|
|
||||||
@@ -80,7 +96,7 @@ An M4A voice note uploaded as a chat attachment does **not** use the realtime
|
|||||||
dictation provider above. OpenClaw processes it through its built-in
|
dictation provider above. OpenClaw processes it through its built-in
|
||||||
`tools.media.audio` path. On the Unraid installation, automatic provider
|
`tools.media.audio` path. On the Unraid installation, automatic provider
|
||||||
selection hit `SsrFBlockedError` for the private Athena address. Configure the
|
selection hit `SsrFBlockedError` for the private Athena address. Configure the
|
||||||
existing OpenAI-compatible provider and select Whisper explicitly:
|
existing OpenAI-compatible provider with the `qwen3-asr` model:
|
||||||
|
|
||||||
```json5
|
```json5
|
||||||
{
|
{
|
||||||
@@ -97,7 +113,7 @@ existing OpenAI-compatible provider and select Whisper explicitly:
|
|||||||
models: [
|
models: [
|
||||||
{
|
{
|
||||||
provider: "openai",
|
provider: "openai",
|
||||||
model: "whisper-1",
|
model: "qwen3-asr",
|
||||||
baseUrl: "http://192.168.1.212:8081/v1",
|
baseUrl: "http://192.168.1.212:8081/v1",
|
||||||
capabilities: ["audio"],
|
capabilities: ["audio"],
|
||||||
},
|
},
|
||||||
@@ -112,7 +128,7 @@ OpenClaw 2026.9.4 accepts `request.allowPrivateNetwork` under
|
|||||||
settings hot-reload without restarting the Gateway. The existing
|
settings hot-reload without restarting the Gateway. The existing
|
||||||
`ATHENA_ROUTER_API_KEY` SecretRef is reused; do not add another literal key.
|
`ATHENA_ROUTER_API_KEY` SecretRef is reused; do not add another literal key.
|
||||||
On 2026-09-16, OpenClaw's official audio module transcribed the affected M4A
|
On 2026-09-16, OpenClaw's official audio module transcribed the affected M4A
|
||||||
through Athena Whisper (127 characters returned). This confirms the endpoint
|
through Athena's then-active Whisper service (127 characters returned). This confirms the endpoint
|
||||||
and file format; a fresh attachment in the chat is still needed to verify the
|
and file format; a fresh attachment in the chat is still needed to verify the
|
||||||
full message-to-transcript flow.
|
full message-to-transcript flow.
|
||||||
|
|
||||||
|
|||||||
+100
-33
@@ -4,6 +4,10 @@ const AUDIO_FORMAT = { encoding: "pcm16", sampleRateHz: 24000, channels: 1 };
|
|||||||
const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls";
|
const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls";
|
||||||
const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key";
|
const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key";
|
||||||
const MAX_OFFER_BYTES = 64 * 1024;
|
const MAX_OFFER_BYTES = 64 * 1024;
|
||||||
|
const DICTATION_SAMPLE_RATE_HZ = 8000;
|
||||||
|
const DICTATION_SEGMENT_BYTES = DICTATION_SAMPLE_RATE_HZ * 6;
|
||||||
|
const DICTATION_OVERLAP_BYTES = DICTATION_SAMPLE_RATE_HZ / 2;
|
||||||
|
const DICTATION_REQUEST_TIMEOUT_MS = 4500;
|
||||||
const browserKeys = generateKeyPairSync("ed25519");
|
const browserKeys = generateKeyPairSync("ed25519");
|
||||||
const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString();
|
const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString();
|
||||||
function base64url(value) {
|
function base64url(value) {
|
||||||
@@ -134,7 +138,7 @@ function wavFromPcm16(pcm, sampleRate = 24000) {
|
|||||||
return Buffer.concat([header, pcm]);
|
return Buffer.concat([header, pcm]);
|
||||||
}
|
}
|
||||||
// The browser transcription relay sends 8 kHz G.711 mu-law, while Athena's
|
// The browser transcription relay sends 8 kHz G.711 mu-law, while Athena's
|
||||||
// existing Whisper endpoint accepts PCM WAV uploads.
|
// transcription endpoint accepts PCM WAV uploads.
|
||||||
function wavFromMulaw8k(audio) {
|
function wavFromMulaw8k(audio) {
|
||||||
const pcm = Buffer.allocUnsafe(audio.length * 2);
|
const pcm = Buffer.allocUnsafe(audio.length * 2);
|
||||||
for (let i = 0; i < audio.length; i += 1) {
|
for (let i = 0; i < audio.length; i += 1) {
|
||||||
@@ -149,13 +153,38 @@ function resolveTranscriptionConfig(cfg, rawConfig) {
|
|||||||
const talkProvider = record(record(talkConfig.providers)["athena-talk"]);
|
const talkProvider = record(record(talkConfig.providers)["athena-talk"]);
|
||||||
return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } });
|
return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } });
|
||||||
}
|
}
|
||||||
|
function comparableWord(value) {
|
||||||
|
return value.toLocaleLowerCase("de-DE").replace(/[^\p{L}\p{N}]+/gu, "");
|
||||||
|
}
|
||||||
|
function removeTranscriptOverlap(previous, current) {
|
||||||
|
const priorWords = previous.trim().split(/\s+/).filter(Boolean);
|
||||||
|
const currentWords = current.trim().split(/\s+/).filter(Boolean);
|
||||||
|
const maximum = Math.min(12, priorWords.length, currentWords.length);
|
||||||
|
for (let count = maximum; count >= 1; count -= 1) {
|
||||||
|
const left = priorWords.slice(-count).map(comparableWord);
|
||||||
|
const right = currentWords.slice(0, count).map(comparableWord);
|
||||||
|
if (!left.every((word, index) => word && word === right[index]))
|
||||||
|
continue;
|
||||||
|
// A single short word is too ambiguous to remove safely. Longer words are
|
||||||
|
// sufficient because the audio overlap is only half a second.
|
||||||
|
if (count === 1 && left[0].length < 5)
|
||||||
|
continue;
|
||||||
|
return currentWords.slice(count).join(" ");
|
||||||
|
}
|
||||||
|
return currentWords.join(" ");
|
||||||
|
}
|
||||||
class AthenaTranscriptionSession {
|
class AthenaTranscriptionSession {
|
||||||
req;
|
req;
|
||||||
config;
|
config;
|
||||||
connected = false;
|
connected = false;
|
||||||
closed = false;
|
closed = false;
|
||||||
audio = [];
|
audio = [];
|
||||||
bytes = 0;
|
bufferedBytes = 0;
|
||||||
|
totalBytes = 0;
|
||||||
|
processing = Promise.resolve();
|
||||||
|
completedTranscripts = [];
|
||||||
|
emittedTranscripts = 0;
|
||||||
|
processingError = null;
|
||||||
constructor(req, config) {
|
constructor(req, config) {
|
||||||
this.req = req;
|
this.req = req;
|
||||||
this.config = config;
|
this.config = config;
|
||||||
@@ -165,50 +194,88 @@ class AthenaTranscriptionSession {
|
|||||||
sendAudio(audio) {
|
sendAudio(audio) {
|
||||||
if (!this.isConnected() || audio.length === 0)
|
if (!this.isConnected() || audio.length === 0)
|
||||||
return;
|
return;
|
||||||
if (this.bytes === 0)
|
if (this.totalBytes === 0)
|
||||||
this.req.onSpeechStart?.();
|
this.req.onSpeechStart?.();
|
||||||
const maxBytes = this.config.maxSpeechSeconds * 8000;
|
const maxBytes = this.config.maxSpeechSeconds * 8000;
|
||||||
if (this.bytes + audio.length > maxBytes) {
|
if (this.totalBytes + audio.length > maxBytes) {
|
||||||
this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`));
|
this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`));
|
||||||
this.close();
|
this.close();
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
this.audio.push(Buffer.from(audio));
|
this.audio.push(Buffer.from(audio));
|
||||||
this.bytes += audio.length;
|
this.bufferedBytes += audio.length;
|
||||||
|
this.totalBytes += audio.length;
|
||||||
|
while (this.bufferedBytes >= DICTATION_SEGMENT_BYTES) {
|
||||||
|
this.queueFullSegment();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
close() {
|
close() {
|
||||||
if (this.closed)
|
if (this.closed)
|
||||||
return;
|
return;
|
||||||
this.closed = true;
|
this.closed = true;
|
||||||
this.connected = false;
|
this.connected = false;
|
||||||
if (!this.bytes)
|
if (!this.totalBytes)
|
||||||
return;
|
return;
|
||||||
const audio = Buffer.concat(this.audio);
|
this.emitCompletedTranscripts();
|
||||||
this.audio.length = 0;
|
const tail = Buffer.concat(this.audio);
|
||||||
void this.transcribe(audio);
|
this.audio = [];
|
||||||
|
this.bufferedBytes = 0;
|
||||||
|
if (tail.length > DICTATION_OVERLAP_BYTES || this.completedTranscripts.length === 0) {
|
||||||
|
this.queueTranscription(tail);
|
||||||
|
}
|
||||||
|
void this.processing.finally(() => {
|
||||||
|
this.emitCompletedTranscripts();
|
||||||
|
if (this.processingError && this.completedTranscripts.length === 0) {
|
||||||
|
this.req.onError?.(this.processingError);
|
||||||
|
}
|
||||||
|
});
|
||||||
}
|
}
|
||||||
async transcribe(audio) {
|
queueFullSegment() {
|
||||||
try {
|
const buffered = Buffer.concat(this.audio);
|
||||||
const form = new FormData();
|
const segment = Buffer.from(buffered.subarray(0, DICTATION_SEGMENT_BYTES));
|
||||||
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
const retained = Buffer.from(buffered.subarray(DICTATION_SEGMENT_BYTES - DICTATION_OVERLAP_BYTES));
|
||||||
form.append("model", "whisper-1");
|
this.audio = retained.length ? [retained] : [];
|
||||||
form.append("language", this.config.language);
|
this.bufferedBytes = retained.length;
|
||||||
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
|
this.queueTranscription(segment);
|
||||||
method: "POST",
|
}
|
||||||
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
|
queueTranscription(audio) {
|
||||||
body: form,
|
if (!audio.length)
|
||||||
signal: AbortSignal.timeout(4500),
|
return;
|
||||||
});
|
this.processing = this.processing.then(async () => {
|
||||||
if (!response.ok)
|
const previous = this.completedTranscripts.join(" ");
|
||||||
throw new Error(`Athena STT failed (HTTP ${response.status})`);
|
const text = await this.transcribe(audio, previous.slice(-240));
|
||||||
const text = String(record(await response.json()).text || "").trim();
|
const novel = removeTranscriptOverlap(previous, text);
|
||||||
if (text)
|
if (novel)
|
||||||
this.req.onTranscript?.(text);
|
this.completedTranscripts.push(novel);
|
||||||
}
|
if (this.closed)
|
||||||
catch (error) {
|
this.emitCompletedTranscripts();
|
||||||
this.req.onError?.(error instanceof Error ? error : new Error(String(error)));
|
}).catch((error) => {
|
||||||
|
this.processingError = error instanceof Error ? error : new Error(String(error));
|
||||||
|
});
|
||||||
|
}
|
||||||
|
emitCompletedTranscripts() {
|
||||||
|
while (this.emittedTranscripts < this.completedTranscripts.length) {
|
||||||
|
this.req.onTranscript?.(this.completedTranscripts[this.emittedTranscripts]);
|
||||||
|
this.emittedTranscripts += 1;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
async transcribe(audio, prompt) {
|
||||||
|
const form = new FormData();
|
||||||
|
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
||||||
|
form.append("model", "qwen3-asr");
|
||||||
|
form.append("language", this.config.language);
|
||||||
|
if (prompt)
|
||||||
|
form.append("prompt", prompt);
|
||||||
|
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
|
||||||
|
method: "POST",
|
||||||
|
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
|
||||||
|
body: form,
|
||||||
|
signal: AbortSignal.timeout(DICTATION_REQUEST_TIMEOUT_MS),
|
||||||
|
});
|
||||||
|
if (!response.ok)
|
||||||
|
throw new Error(`Athena STT failed (HTTP ${response.status})`);
|
||||||
|
return String(record(await response.json()).text || "").trim();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
function pcmRms(pcm) {
|
function pcmRms(pcm) {
|
||||||
if (pcm.length < 2)
|
if (pcm.length < 2)
|
||||||
@@ -404,7 +471,7 @@ class AthenaTalkBridge {
|
|||||||
const form = new FormData();
|
const form = new FormData();
|
||||||
const wavBytes = Uint8Array.from(wavFromPcm16(pcm));
|
const wavBytes = Uint8Array.from(wavFromPcm16(pcm));
|
||||||
form.append("file", new Blob([wavBytes], { type: "audio/wav" }), "talk.wav");
|
form.append("file", new Blob([wavBytes], { type: "audio/wav" }), "talk.wav");
|
||||||
form.append("model", "whisper-1");
|
form.append("model", "qwen3-asr");
|
||||||
form.append("language", this.cfg.language);
|
form.append("language", this.cfg.language);
|
||||||
const response = await fetch(`${this.cfg.baseUrl}/audio/transcriptions`, {
|
const response = await fetch(`${this.cfg.baseUrl}/audio/transcriptions`, {
|
||||||
method: "POST",
|
method: "POST",
|
||||||
@@ -500,9 +567,9 @@ export default definePluginEntry({
|
|||||||
register(api) {
|
register(api) {
|
||||||
api.registerRealtimeTranscriptionProvider({
|
api.registerRealtimeTranscriptionProvider({
|
||||||
id: "athena-talk",
|
id: "athena-talk",
|
||||||
label: "Athena Whisper (Diktieren)",
|
label: "Athena Qwen3-ASR (Diktieren)",
|
||||||
defaultModel: "whisper-1",
|
defaultModel: "qwen3-asr",
|
||||||
models: ["whisper-1"],
|
models: ["qwen3-asr"],
|
||||||
autoSelectOrder: 1,
|
autoSelectOrder: 1,
|
||||||
resolveConfig: ({ cfg, rawConfig }) => resolveTranscriptionConfig(cfg, rawConfig),
|
resolveConfig: ({ cfg, rawConfig }) => resolveTranscriptionConfig(cfg, rawConfig),
|
||||||
isConfigured: ({ providerConfig }) => Boolean(record(providerConfig).baseUrl),
|
isConfigured: ({ providerConfig }) => Boolean(record(providerConfig).baseUrl),
|
||||||
|
|||||||
@@ -20,6 +20,10 @@ type ProviderConfig = {
|
|||||||
const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls";
|
const BROWSER_OFFER_PATH = "/plugins/athena-talk/realtime/calls";
|
||||||
const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key";
|
const BROWSER_KEY_PATH = "/plugins/athena-talk/realtime/public-key";
|
||||||
const MAX_OFFER_BYTES = 64 * 1024;
|
const MAX_OFFER_BYTES = 64 * 1024;
|
||||||
|
const DICTATION_SAMPLE_RATE_HZ = 8000;
|
||||||
|
const DICTATION_SEGMENT_BYTES = DICTATION_SAMPLE_RATE_HZ * 6;
|
||||||
|
const DICTATION_OVERLAP_BYTES = DICTATION_SAMPLE_RATE_HZ / 2;
|
||||||
|
const DICTATION_REQUEST_TIMEOUT_MS = 4500;
|
||||||
const browserKeys = generateKeyPairSync("ed25519");
|
const browserKeys = generateKeyPairSync("ed25519");
|
||||||
const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString();
|
const publicKeyPem = browserKeys.publicKey.export({ format: "pem", type: "spki" }).toString();
|
||||||
|
|
||||||
@@ -154,7 +158,7 @@ function wavFromPcm16(pcm: Buffer, sampleRate = 24000): Buffer {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// The browser transcription relay sends 8 kHz G.711 mu-law, while Athena's
|
// The browser transcription relay sends 8 kHz G.711 mu-law, while Athena's
|
||||||
// existing Whisper endpoint accepts PCM WAV uploads.
|
// transcription endpoint accepts PCM WAV uploads.
|
||||||
function wavFromMulaw8k(audio: Buffer): Buffer {
|
function wavFromMulaw8k(audio: Buffer): Buffer {
|
||||||
const pcm = Buffer.allocUnsafe(audio.length * 2);
|
const pcm = Buffer.allocUnsafe(audio.length * 2);
|
||||||
for (let i = 0; i < audio.length; i += 1) {
|
for (let i = 0; i < audio.length; i += 1) {
|
||||||
@@ -171,11 +175,36 @@ function resolveTranscriptionConfig(cfg: unknown, rawConfig: unknown): Required<
|
|||||||
return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } });
|
return resolveConfig({ cfg, providerConfig: { ...talkProvider, ...record(rawConfig) } });
|
||||||
}
|
}
|
||||||
|
|
||||||
|
function comparableWord(value: string): string {
|
||||||
|
return value.toLocaleLowerCase("de-DE").replace(/[^\p{L}\p{N}]+/gu, "");
|
||||||
|
}
|
||||||
|
|
||||||
|
function removeTranscriptOverlap(previous: string, current: string): string {
|
||||||
|
const priorWords = previous.trim().split(/\s+/).filter(Boolean);
|
||||||
|
const currentWords = current.trim().split(/\s+/).filter(Boolean);
|
||||||
|
const maximum = Math.min(12, priorWords.length, currentWords.length);
|
||||||
|
for (let count = maximum; count >= 1; count -= 1) {
|
||||||
|
const left = priorWords.slice(-count).map(comparableWord);
|
||||||
|
const right = currentWords.slice(0, count).map(comparableWord);
|
||||||
|
if (!left.every((word, index) => word && word === right[index])) continue;
|
||||||
|
// A single short word is too ambiguous to remove safely. Longer words are
|
||||||
|
// sufficient because the audio overlap is only half a second.
|
||||||
|
if (count === 1 && left[0].length < 5) continue;
|
||||||
|
return currentWords.slice(count).join(" ");
|
||||||
|
}
|
||||||
|
return currentWords.join(" ");
|
||||||
|
}
|
||||||
|
|
||||||
class AthenaTranscriptionSession {
|
class AthenaTranscriptionSession {
|
||||||
private connected = false;
|
private connected = false;
|
||||||
private closed = false;
|
private closed = false;
|
||||||
private readonly audio: Buffer[] = [];
|
private audio: Buffer[] = [];
|
||||||
private bytes = 0;
|
private bufferedBytes = 0;
|
||||||
|
private totalBytes = 0;
|
||||||
|
private processing: Promise<void> = Promise.resolve();
|
||||||
|
private readonly completedTranscripts: string[] = [];
|
||||||
|
private emittedTranscripts = 0;
|
||||||
|
private processingError: Error | null = null;
|
||||||
|
|
||||||
constructor(
|
constructor(
|
||||||
private readonly req: {
|
private readonly req: {
|
||||||
@@ -193,46 +222,85 @@ class AthenaTranscriptionSession {
|
|||||||
|
|
||||||
sendAudio(audio: Buffer): void {
|
sendAudio(audio: Buffer): void {
|
||||||
if (!this.isConnected() || audio.length === 0) return;
|
if (!this.isConnected() || audio.length === 0) return;
|
||||||
if (this.bytes === 0) this.req.onSpeechStart?.();
|
if (this.totalBytes === 0) this.req.onSpeechStart?.();
|
||||||
const maxBytes = this.config.maxSpeechSeconds * 8000;
|
const maxBytes = this.config.maxSpeechSeconds * 8000;
|
||||||
if (this.bytes + audio.length > maxBytes) {
|
if (this.totalBytes + audio.length > maxBytes) {
|
||||||
this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`));
|
this.req.onError?.(new Error(`Athena dictation is limited to ${this.config.maxSpeechSeconds} seconds`));
|
||||||
this.close();
|
this.close();
|
||||||
return;
|
return;
|
||||||
}
|
}
|
||||||
this.audio.push(Buffer.from(audio));
|
this.audio.push(Buffer.from(audio));
|
||||||
this.bytes += audio.length;
|
this.bufferedBytes += audio.length;
|
||||||
|
this.totalBytes += audio.length;
|
||||||
|
while (this.bufferedBytes >= DICTATION_SEGMENT_BYTES) {
|
||||||
|
this.queueFullSegment();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
close(): void {
|
close(): void {
|
||||||
if (this.closed) return;
|
if (this.closed) return;
|
||||||
this.closed = true;
|
this.closed = true;
|
||||||
this.connected = false;
|
this.connected = false;
|
||||||
if (!this.bytes) return;
|
if (!this.totalBytes) return;
|
||||||
const audio = Buffer.concat(this.audio);
|
this.emitCompletedTranscripts();
|
||||||
this.audio.length = 0;
|
const tail = Buffer.concat(this.audio);
|
||||||
void this.transcribe(audio);
|
this.audio = [];
|
||||||
|
this.bufferedBytes = 0;
|
||||||
|
if (tail.length > DICTATION_OVERLAP_BYTES || this.completedTranscripts.length === 0) {
|
||||||
|
this.queueTranscription(tail);
|
||||||
|
}
|
||||||
|
void this.processing.finally(() => {
|
||||||
|
this.emitCompletedTranscripts();
|
||||||
|
if (this.processingError && this.completedTranscripts.length === 0) {
|
||||||
|
this.req.onError?.(this.processingError);
|
||||||
|
}
|
||||||
|
});
|
||||||
}
|
}
|
||||||
|
|
||||||
private async transcribe(audio: Buffer): Promise<void> {
|
private queueFullSegment(): void {
|
||||||
try {
|
const buffered = Buffer.concat(this.audio);
|
||||||
const form = new FormData();
|
const segment = Buffer.from(buffered.subarray(0, DICTATION_SEGMENT_BYTES));
|
||||||
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
const retained = Buffer.from(buffered.subarray(DICTATION_SEGMENT_BYTES - DICTATION_OVERLAP_BYTES));
|
||||||
form.append("model", "whisper-1");
|
this.audio = retained.length ? [retained] : [];
|
||||||
form.append("language", this.config.language);
|
this.bufferedBytes = retained.length;
|
||||||
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
|
this.queueTranscription(segment);
|
||||||
method: "POST",
|
}
|
||||||
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
|
|
||||||
body: form,
|
private queueTranscription(audio: Buffer): void {
|
||||||
signal: AbortSignal.timeout(4500),
|
if (!audio.length) return;
|
||||||
});
|
this.processing = this.processing.then(async () => {
|
||||||
if (!response.ok) throw new Error(`Athena STT failed (HTTP ${response.status})`);
|
const previous = this.completedTranscripts.join(" ");
|
||||||
const text = String(record(await response.json()).text || "").trim();
|
const text = await this.transcribe(audio, previous.slice(-240));
|
||||||
if (text) this.req.onTranscript?.(text);
|
const novel = removeTranscriptOverlap(previous, text);
|
||||||
} catch (error) {
|
if (novel) this.completedTranscripts.push(novel);
|
||||||
this.req.onError?.(error instanceof Error ? error : new Error(String(error)));
|
if (this.closed) this.emitCompletedTranscripts();
|
||||||
|
}).catch((error: unknown) => {
|
||||||
|
this.processingError = error instanceof Error ? error : new Error(String(error));
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
private emitCompletedTranscripts(): void {
|
||||||
|
while (this.emittedTranscripts < this.completedTranscripts.length) {
|
||||||
|
this.req.onTranscript?.(this.completedTranscripts[this.emittedTranscripts]);
|
||||||
|
this.emittedTranscripts += 1;
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
private async transcribe(audio: Buffer, prompt: string): Promise<string> {
|
||||||
|
const form = new FormData();
|
||||||
|
form.append("file", new Blob([Uint8Array.from(wavFromMulaw8k(audio))], { type: "audio/wav" }), "dictation.wav");
|
||||||
|
form.append("model", "qwen3-asr");
|
||||||
|
form.append("language", this.config.language);
|
||||||
|
if (prompt) form.append("prompt", prompt);
|
||||||
|
const response = await fetch(`${this.config.baseUrl}/audio/transcriptions`, {
|
||||||
|
method: "POST",
|
||||||
|
headers: this.config.apiKey ? { Authorization: `Bearer ${this.config.apiKey}` } : {},
|
||||||
|
body: form,
|
||||||
|
signal: AbortSignal.timeout(DICTATION_REQUEST_TIMEOUT_MS),
|
||||||
|
});
|
||||||
|
if (!response.ok) throw new Error(`Athena STT failed (HTTP ${response.status})`);
|
||||||
|
return String(record(await response.json()).text || "").trim();
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
function pcmRms(pcm: Buffer): number {
|
function pcmRms(pcm: Buffer): number {
|
||||||
@@ -422,7 +490,7 @@ class AthenaTalkBridge {
|
|||||||
const form = new FormData();
|
const form = new FormData();
|
||||||
const wavBytes = Uint8Array.from(wavFromPcm16(pcm));
|
const wavBytes = Uint8Array.from(wavFromPcm16(pcm));
|
||||||
form.append("file", new Blob([wavBytes], { type: "audio/wav" }), "talk.wav");
|
form.append("file", new Blob([wavBytes], { type: "audio/wav" }), "talk.wav");
|
||||||
form.append("model", "whisper-1");
|
form.append("model", "qwen3-asr");
|
||||||
form.append("language", this.cfg.language);
|
form.append("language", this.cfg.language);
|
||||||
const response = await fetch(`${this.cfg.baseUrl}/audio/transcriptions`, {
|
const response = await fetch(`${this.cfg.baseUrl}/audio/transcriptions`, {
|
||||||
method: "POST",
|
method: "POST",
|
||||||
@@ -510,9 +578,9 @@ export default definePluginEntry({
|
|||||||
register(api) {
|
register(api) {
|
||||||
api.registerRealtimeTranscriptionProvider({
|
api.registerRealtimeTranscriptionProvider({
|
||||||
id: "athena-talk",
|
id: "athena-talk",
|
||||||
label: "Athena Whisper (Diktieren)",
|
label: "Athena Qwen3-ASR (Diktieren)",
|
||||||
defaultModel: "whisper-1",
|
defaultModel: "qwen3-asr",
|
||||||
models: ["whisper-1"],
|
models: ["qwen3-asr"],
|
||||||
autoSelectOrder: 1,
|
autoSelectOrder: 1,
|
||||||
resolveConfig: ({ cfg, rawConfig }) => resolveTranscriptionConfig(cfg, rawConfig),
|
resolveConfig: ({ cfg, rawConfig }) => resolveTranscriptionConfig(cfg, rawConfig),
|
||||||
isConfigured: ({ providerConfig }) => Boolean(record(providerConfig).baseUrl),
|
isConfigured: ({ providerConfig }) => Boolean(record(providerConfig).baseUrl),
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
{
|
{
|
||||||
"id": "athena-talk",
|
"id": "athena-talk",
|
||||||
"name": "Athena Local Talk",
|
"name": "Athena Local Talk",
|
||||||
"description": "Private OpenClaw Talk provider using Athena Whisper and Qwen3-TTS.",
|
"description": "Private OpenClaw Talk provider using Athena Qwen3-ASR and Qwen3-TTS.",
|
||||||
"activation": {
|
"activation": {
|
||||||
"onStartup": true
|
"onStartup": true
|
||||||
},
|
},
|
||||||
|
|||||||
+2
-2
@@ -1,12 +1,12 @@
|
|||||||
{
|
{
|
||||||
"name": "@casaderoll/openclaw-athena-talk",
|
"name": "@casaderoll/openclaw-athena-talk",
|
||||||
"version": "1.2.1",
|
"version": "1.3.1",
|
||||||
"lockfileVersion": 3,
|
"lockfileVersion": 3,
|
||||||
"requires": true,
|
"requires": true,
|
||||||
"packages": {
|
"packages": {
|
||||||
"": {
|
"": {
|
||||||
"name": "@casaderoll/openclaw-athena-talk",
|
"name": "@casaderoll/openclaw-athena-talk",
|
||||||
"version": "1.2.1",
|
"version": "1.3.1",
|
||||||
"devDependencies": {
|
"devDependencies": {
|
||||||
"@types/node": "^24.0.0",
|
"@types/node": "^24.0.0",
|
||||||
"openclaw": "2026.9.4",
|
"openclaw": "2026.9.4",
|
||||||
|
|||||||
@@ -1,8 +1,8 @@
|
|||||||
{
|
{
|
||||||
"name": "@casaderoll/openclaw-athena-talk",
|
"name": "@casaderoll/openclaw-athena-talk",
|
||||||
"version": "1.2.1",
|
"version": "1.3.1",
|
||||||
"private": true,
|
"private": true,
|
||||||
"description": "Local OpenClaw Talk provider backed by Athena Whisper and Qwen3-TTS",
|
"description": "Local OpenClaw Talk provider backed by Athena Qwen3-ASR and Qwen3-TTS",
|
||||||
"type": "module",
|
"type": "module",
|
||||||
"files": [
|
"files": [
|
||||||
"dist",
|
"dist",
|
||||||
|
|||||||
@@ -3,7 +3,15 @@ import { createServer } from "node:http";
|
|||||||
import { test } from "node:test";
|
import { test } from "node:test";
|
||||||
import plugin from "./dist/index.js";
|
import plugin from "./dist/index.js";
|
||||||
|
|
||||||
test("dictation registers separately and sends G.711 audio to Athena Whisper", async () => {
|
async function waitFor(predicate, timeoutMs = 2000) {
|
||||||
|
const deadline = Date.now() + timeoutMs;
|
||||||
|
while (!predicate()) {
|
||||||
|
if (Date.now() >= deadline) throw new Error("timed out waiting for condition");
|
||||||
|
await new Promise((resolve) => setTimeout(resolve, 10));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
test("dictation registers separately and sends G.711 audio to Athena Qwen3-ASR", async () => {
|
||||||
let transcription;
|
let transcription;
|
||||||
plugin.register({
|
plugin.register({
|
||||||
registerRealtimeTranscriptionProvider: (value) => { transcription = value; },
|
registerRealtimeTranscriptionProvider: (value) => { transcription = value; },
|
||||||
@@ -29,7 +37,7 @@ test("dictation registers separately and sends G.711 audio to Athena Whisper", a
|
|||||||
assert.equal(wav.readUInt16LE(34), 16);
|
assert.equal(wav.readUInt16LE(34), 16);
|
||||||
assert.equal(wav.length, 48);
|
assert.equal(wav.length, 48);
|
||||||
assert.equal(form.get("language"), "de");
|
assert.equal(form.get("language"), "de");
|
||||||
assert.equal(form.get("model"), "whisper-1");
|
assert.equal(form.get("model"), "qwen3-asr");
|
||||||
res.writeHead(200, { "Content-Type": "application/json" }).end(JSON.stringify({ text: "Hallo Athena" }));
|
res.writeHead(200, { "Content-Type": "application/json" }).end(JSON.stringify({ text: "Hallo Athena" }));
|
||||||
});
|
});
|
||||||
await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve));
|
await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve));
|
||||||
@@ -79,3 +87,73 @@ test("dictation reuses only a TTS key for the same Athena origin", () => {
|
|||||||
cfg: withTts("http://other:8081/v1"), rawConfig: {},
|
cfg: withTts("http://other:8081/v1"), rawConfig: {},
|
||||||
}).apiKey, "");
|
}).apiKey, "");
|
||||||
});
|
});
|
||||||
|
|
||||||
|
test("long dictation is transcribed incrementally before the recording closes", async () => {
|
||||||
|
let transcription;
|
||||||
|
plugin.register({
|
||||||
|
registerRealtimeTranscriptionProvider: (value) => { transcription = value; },
|
||||||
|
registerRealtimeVoiceProvider: () => {},
|
||||||
|
registerHttpRoute: () => {},
|
||||||
|
});
|
||||||
|
|
||||||
|
const answers = [
|
||||||
|
"Dies ist ein langer Abschnitt",
|
||||||
|
"langer Abschnitt mit einer Fortsetzung",
|
||||||
|
"einer Fortsetzung und einem Ende.",
|
||||||
|
];
|
||||||
|
let uploads = 0;
|
||||||
|
const server = createServer(async (req, res) => {
|
||||||
|
const index = uploads++;
|
||||||
|
const form = await new Request("http://localhost", {
|
||||||
|
method: "POST",
|
||||||
|
headers: { "Content-Type": req.headers["content-type"] },
|
||||||
|
body: req,
|
||||||
|
duplex: "half",
|
||||||
|
}).formData();
|
||||||
|
const wav = Buffer.from(await form.get("file").arrayBuffer());
|
||||||
|
assert.equal(wav.readUInt32LE(24), 8000);
|
||||||
|
if (index > 0) assert.ok(String(form.get("prompt") || "").length > 0);
|
||||||
|
res.writeHead(200, { "Content-Type": "application/json" })
|
||||||
|
.end(JSON.stringify({ text: answers[index] }));
|
||||||
|
});
|
||||||
|
await new Promise((resolve) => server.listen(0, "127.0.0.1", resolve));
|
||||||
|
const cfg = { talk: { realtime: { providers: { "athena-talk": {
|
||||||
|
baseUrl: `http://127.0.0.1:${server.address().port}/v1`, language: "de",
|
||||||
|
} } } } };
|
||||||
|
const providerConfig = transcription.resolveConfig({ cfg, rawConfig: {} });
|
||||||
|
const transcripts = [];
|
||||||
|
const errors = [];
|
||||||
|
try {
|
||||||
|
const session = transcription.createSession({
|
||||||
|
cfg,
|
||||||
|
providerConfig,
|
||||||
|
onTranscript: (text) => transcripts.push(text),
|
||||||
|
onError: (error) => errors.push(error),
|
||||||
|
});
|
||||||
|
await session.connect();
|
||||||
|
|
||||||
|
// Six seconds start the first request while dictation is still active.
|
||||||
|
session.sendAudio(Buffer.alloc(48_000, 0xff));
|
||||||
|
await waitFor(() => uploads === 1);
|
||||||
|
assert.deepEqual(transcripts, []);
|
||||||
|
|
||||||
|
// Another 5.5 seconds form the next overlapping segment. The remaining
|
||||||
|
// 2 seconds are finalized only when the user stops dictation.
|
||||||
|
session.sendAudio(Buffer.alloc(44_000, 0xff));
|
||||||
|
await waitFor(() => uploads === 2);
|
||||||
|
session.sendAudio(Buffer.alloc(16_000, 0xff));
|
||||||
|
session.close();
|
||||||
|
|
||||||
|
await waitFor(() => transcripts.length === 3);
|
||||||
|
assert.deepEqual(transcripts, [
|
||||||
|
"Dies ist ein langer Abschnitt",
|
||||||
|
"mit einer Fortsetzung",
|
||||||
|
"und einem Ende.",
|
||||||
|
]);
|
||||||
|
assert.equal(uploads, 3);
|
||||||
|
assert.deepEqual(errors, []);
|
||||||
|
} finally {
|
||||||
|
server.closeAllConnections();
|
||||||
|
await new Promise((resolve) => server.close(resolve));
|
||||||
|
}
|
||||||
|
});
|
||||||
@@ -52,7 +52,8 @@ case "$command" in
|
|||||||
[[ $# -eq 0 ]] || { echo "core akzeptiert keine weiteren Services" >&2; exit 2; }
|
[[ $# -eq 0 ]] || { echo "core akzeptiert keine weiteren Services" >&2; exit 2; }
|
||||||
run "$ROOT_DIR/platform/mcp/install-tools.sh"
|
run "$ROOT_DIR/platform/mcp/install-tools.sh"
|
||||||
run "${compose[@]}" up -d --build \
|
run "${compose[@]}" up -d --build \
|
||||||
wireguard-gateway embedding qwen3-tts tts-gateway profile-controller router llama-dashboard portainer backup
|
wireguard-gateway embedding qwen3-tts tts-gateway profile-controller \
|
||||||
|
qwen-asr qwen-asr-worker router llama-dashboard portainer backup
|
||||||
else
|
else
|
||||||
run "${compose[@]}" up -d --build --no-deps "$@"
|
run "${compose[@]}" up -d --build --no-deps "$@"
|
||||||
fi
|
fi
|
||||||
|
|||||||
@@ -37,6 +37,22 @@ target="$OUTPUT_DIR/$name"
|
|||||||
list=$(mktemp /tmp/athena-export-list.XXXXXX)
|
list=$(mktemp /tmp/athena-export-list.XXXXXX)
|
||||||
trap 'rm -f "$list" "$partial"' EXIT
|
trap 'rm -f "$list" "$partial"' EXIT
|
||||||
|
|
||||||
|
# The encrypted export is roughly the size of its predecessor. Keeping all
|
||||||
|
# generations until after writing the next one requires one extra archive of
|
||||||
|
# free space and can deadlock a full backup disk. Release exactly one slot
|
||||||
|
# before writing when the configured retention is already reached. If the new
|
||||||
|
# export fails, KEEP-1 valid generations remain available.
|
||||||
|
mapfile -t existing < <(find "$OUTPUT_DIR" -maxdepth 1 -type f \
|
||||||
|
-name 'athena-portable-*.tar.zst.age' -printf '%T@ %p\n' \
|
||||||
|
| sort -rn | cut -d' ' -f2-)
|
||||||
|
while ((${#existing[@]} >= KEEP)); do
|
||||||
|
last_index=$((${#existing[@]} - 1))
|
||||||
|
oldest=${existing[$last_index]}
|
||||||
|
rm -f -- "$oldest" "$oldest.sha256"
|
||||||
|
unset 'existing[last_index]'
|
||||||
|
existing=("${existing[@]}")
|
||||||
|
done
|
||||||
|
|
||||||
add_path() {
|
add_path() {
|
||||||
local path=${1#/}
|
local path=${1#/}
|
||||||
[[ ! -e /$path ]] || printf '%s\0' "$path" >>"$list"
|
[[ ! -e /$path ]] || printf '%s\0' "$path" >>"$list"
|
||||||
|
|||||||
@@ -27,6 +27,10 @@ LABEL_KEY = "com.mike-ai.llama-profile"
|
|||||||
IMAGE_LABEL_KEY = "com.mike-ai.image-worker"
|
IMAGE_LABEL_KEY = "com.mike-ai.image-worker"
|
||||||
IMAGE_WORKER = os.environ.get("IMAGE_WORKER", "image")
|
IMAGE_WORKER = os.environ.get("IMAGE_WORKER", "image")
|
||||||
RESTORE_WORKER = os.environ.get("RESTORE_WORKER", "restore")
|
RESTORE_WORKER = os.environ.get("RESTORE_WORKER", "restore")
|
||||||
|
IMAGE_PROMPT_I2I_WORKER = os.environ.get(
|
||||||
|
"IMAGE_PROMPT_I2I_WORKER", "image-prompt-i2i").strip()
|
||||||
|
IMAGE_PROMPT_T2I_WORKER = os.environ.get(
|
||||||
|
"IMAGE_PROMPT_T2I_WORKER", "image-prompt-t2i").strip()
|
||||||
FLUX_STANDBY_WORKER = os.environ.get(
|
FLUX_STANDBY_WORKER = os.environ.get(
|
||||||
"FLUX_STANDBY_WORKER", "flux-standby").strip()
|
"FLUX_STANDBY_WORKER", "flux-standby").strip()
|
||||||
TTS_LABEL_KEY = "com.mike-ai.tts-worker"
|
TTS_LABEL_KEY = "com.mike-ai.tts-worker"
|
||||||
@@ -102,7 +106,12 @@ def image_container(kind: str = IMAGE_WORKER) -> dict:
|
|||||||
|
|
||||||
def image_containers() -> list[dict]:
|
def image_containers() -> list[dict]:
|
||||||
"""All allowlisted GPU workers that must never overlap an LLM."""
|
"""All allowlisted GPU workers that must never overlap an LLM."""
|
||||||
allowed = {IMAGE_WORKER, RESTORE_WORKER}
|
allowed = {
|
||||||
|
IMAGE_WORKER,
|
||||||
|
RESTORE_WORKER,
|
||||||
|
IMAGE_PROMPT_I2I_WORKER,
|
||||||
|
IMAGE_PROMPT_T2I_WORKER,
|
||||||
|
}
|
||||||
if FLUX_STANDBY_WORKER:
|
if FLUX_STANDBY_WORKER:
|
||||||
allowed.add(FLUX_STANDBY_WORKER)
|
allowed.add(FLUX_STANDBY_WORKER)
|
||||||
return [item for item in labelled_containers(IMAGE_LABEL_KEY)
|
return [item for item in labelled_containers(IMAGE_LABEL_KEY)
|
||||||
@@ -363,7 +372,13 @@ def stop_inference() -> dict:
|
|||||||
|
|
||||||
|
|
||||||
def set_image_worker(running: bool, kind: str = IMAGE_WORKER) -> dict:
|
def set_image_worker(running: bool, kind: str = IMAGE_WORKER) -> dict:
|
||||||
if kind not in {IMAGE_WORKER, RESTORE_WORKER, FLUX_STANDBY_WORKER}:
|
if kind not in {
|
||||||
|
IMAGE_WORKER,
|
||||||
|
RESTORE_WORKER,
|
||||||
|
FLUX_STANDBY_WORKER,
|
||||||
|
IMAGE_PROMPT_I2I_WORKER,
|
||||||
|
IMAGE_PROMPT_T2I_WORKER,
|
||||||
|
}:
|
||||||
raise ValueError("worker is not allowlisted")
|
raise ValueError("worker is not allowlisted")
|
||||||
with LOCK:
|
with LOCK:
|
||||||
item = image_container(kind)
|
item = image_container(kind)
|
||||||
@@ -799,6 +814,10 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
worker_paths = {
|
worker_paths = {
|
||||||
"/workers/image/start": (IMAGE_WORKER, True),
|
"/workers/image/start": (IMAGE_WORKER, True),
|
||||||
"/workers/image/stop": (IMAGE_WORKER, False),
|
"/workers/image/stop": (IMAGE_WORKER, False),
|
||||||
|
"/workers/image-prompt-i2i/start": (IMAGE_PROMPT_I2I_WORKER, True),
|
||||||
|
"/workers/image-prompt-i2i/stop": (IMAGE_PROMPT_I2I_WORKER, False),
|
||||||
|
"/workers/image-prompt-t2i/start": (IMAGE_PROMPT_T2I_WORKER, True),
|
||||||
|
"/workers/image-prompt-t2i/stop": (IMAGE_PROMPT_T2I_WORKER, False),
|
||||||
"/workers/restore/start": (RESTORE_WORKER, True),
|
"/workers/restore/start": (RESTORE_WORKER, True),
|
||||||
"/workers/restore/stop": (RESTORE_WORKER, False),
|
"/workers/restore/stop": (RESTORE_WORKER, False),
|
||||||
"/workers/flux-standby/start": (FLUX_STANDBY_WORKER, True),
|
"/workers/flux-standby/start": (FLUX_STANDBY_WORKER, True),
|
||||||
|
|||||||
@@ -0,0 +1,17 @@
|
|||||||
|
FROM python:3.13.7-slim-bookworm
|
||||||
|
|
||||||
|
RUN apt-get update \
|
||||||
|
&& apt-get install -y --no-install-recommends ffmpeg \
|
||||||
|
&& rm -rf /var/lib/apt/lists/* \
|
||||||
|
&& useradd --system --uid 10005 --home-dir /nonexistent --shell /usr/sbin/nologin stt
|
||||||
|
|
||||||
|
WORKDIR /app
|
||||||
|
COPY router/qwen_asr_worker.py /app/qwen_asr_worker.py
|
||||||
|
|
||||||
|
ENV QWEN_ASR_HOST=0.0.0.0 \
|
||||||
|
QWEN_ASR_PORT=8084 \
|
||||||
|
QWEN_ASR_LANGUAGE=de
|
||||||
|
|
||||||
|
USER 10005:10005
|
||||||
|
EXPOSE 8084
|
||||||
|
ENTRYPOINT ["python", "/app/qwen_asr_worker.py"]
|
||||||
@@ -133,8 +133,8 @@ display:
|
|||||||
language: "de"
|
language: "de"
|
||||||
show_reasoning: false
|
show_reasoning: false
|
||||||
|
|
||||||
# Speech input stays local and private. A fixed German hint avoids Whisper
|
# Speech input stays local and private. The fixed German hint helps with
|
||||||
# interpreting short utterances as English while retaining the fast base model.
|
# short utterances and product names.
|
||||||
stt:
|
stt:
|
||||||
enabled: true
|
enabled: true
|
||||||
provider: "local"
|
provider: "local"
|
||||||
|
|||||||
@@ -32,7 +32,7 @@ gateway, UI, CPU-STT, backup and operator containers may remain active.
|
|||||||
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
|
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
|
||||||
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
|
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
|
||||||
- Large: the same Q4 model, 192,000 tokens, vision.
|
- Large: the same Q4 model, 192,000 tokens, vision.
|
||||||
- Ultra: the same Q4 model, 262,144 tokens, no vision projector.
|
- Ultra: the same Q4 model, 262,144 tokens, vision projector on CPU.
|
||||||
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
|
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
|
||||||
|
|
||||||
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
|
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
|
||||||
|
|||||||
@@ -16,10 +16,10 @@ nach Standardbenchmark, Tool-Calling-Test und Kontexttest übernommen.
|
|||||||
|
|
||||||
- Fast: Qwen3.8-27B IQ4-MIX mit MTP2
|
- Fast: Qwen3.8-27B IQ4-MIX mit MTP2
|
||||||
- Medium und Large: Qwen3.8-27B IQ4_XS Pure mit MTP3
|
- Medium und Large: Qwen3.8-27B IQ4_XS Pure mit MTP3
|
||||||
- Ultra: Qwen3.8-27B IQ4_XS Pure mit MTP2 und maximalem Textkontext
|
- Ultra: Qwen3.8-27B IQ4_XS Pure mit MTP2, maximalem Kontext und Vision-Projektor auf CPU
|
||||||
- Uncensored: Blackfrost Qwen3.8-27B Abliterated Q4_K_M mit MTP2
|
- Uncensored: Blackfrost Qwen3.8-27B Abliterated Q4_K_M mit MTP2
|
||||||
- Fast, Medium, Large und Uncensored: integrierte Vision; der jeweilige
|
- Alle fünf Profile: integrierte Vision. Bei Fast, Medium, Large und
|
||||||
Projektor liegt vollständig auf der RTX 3060
|
Uncensored liegt der Projektor auf der RTX 3060; bei Ultra auf der CPU.
|
||||||
|
|
||||||
Die produktive Runtime ist seit dem 12. September 2026 auf llama.cpp
|
Die produktive Runtime ist seit dem 12. September 2026 auf llama.cpp
|
||||||
**Build 10930**, Commit `56381e407c0ccfb3a6f71e668a27a901001d22ce`,
|
**Build 10930**, Commit `56381e407c0ccfb3a6f71e668a27a901001d22ce`,
|
||||||
|
|||||||
@@ -3,4 +3,4 @@ Description=Legacy native Qwen Ultra 256K profile (Docker is the production path
|
|||||||
|
|
||||||
[Service]
|
[Service]
|
||||||
ExecStart=
|
ExecStart=
|
||||||
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --alias qwen-ultra --ctx-size 262144 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --load-mode none --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
|
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-ultra --ctx-size 262144 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --load-mode none --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
|
||||||
+193
-12
@@ -29,12 +29,11 @@ Sprachausgabe (Qwen3-TTS auf RTX 3060):
|
|||||||
POST /v1/audio/speech (OpenAI-kompatibel)
|
POST /v1/audio/speech (OpenAI-kompatibel)
|
||||||
GET /v1/audio/voices (verfügbare Stimmen)
|
GET /v1/audio/voices (verfügbare Stimmen)
|
||||||
|
|
||||||
Spracherkennung (whisper.cpp, deutsch, CPU-only):
|
Spracherkennung (Qwen3-ASR, deutsch, CPU-only):
|
||||||
POST /v1/audio/transcriptions (OpenAI-kompatibel)
|
POST /v1/audio/transcriptions (OpenAI-kompatibel)
|
||||||
GET /v1/audio/models (verfügbare Audio-Modelle)
|
GET /v1/audio/models (verfügbare Audio-Modelle)
|
||||||
|
|
||||||
Der TTS-Worker (mike-ai-xtts.service) und der STT-Worker
|
TTS und Qwen3-ASR laufen als separate, langlebige Dienste.
|
||||||
(mike-ai-whisper.service) laufen als separate, langlebige Prozesse.
|
|
||||||
Der Router leitet /v1/audio/speech und /v1/audio/transcriptions
|
Der Router leitet /v1/audio/speech und /v1/audio/transcriptions
|
||||||
per HTTP an die Worker weiter.
|
per HTTP an die Worker weiter.
|
||||||
|
|
||||||
@@ -146,6 +145,18 @@ IMAGE_PYTHON = os.environ.get(
|
|||||||
"IMAGE_PYTHON", "/opt/mike-ai/ai-profile-router/venv/bin/python")
|
"IMAGE_PYTHON", "/opt/mike-ai/ai-profile-router/venv/bin/python")
|
||||||
IMAGE_WORKER_URL = os.environ.get("IMAGE_WORKER_URL", "").rstrip("/")
|
IMAGE_WORKER_URL = os.environ.get("IMAGE_WORKER_URL", "").rstrip("/")
|
||||||
IMAGE_WORKER_TOKEN = os.environ.get("IMAGE_WORKER_TOKEN", "").strip()
|
IMAGE_WORKER_TOKEN = os.environ.get("IMAGE_WORKER_TOKEN", "").strip()
|
||||||
|
IMAGE_PROMPT_I2I_URL = os.environ.get(
|
||||||
|
"IMAGE_PROMPT_I2I_URL", "").rstrip("/")
|
||||||
|
IMAGE_PROMPT_T2I_URL = os.environ.get(
|
||||||
|
"IMAGE_PROMPT_T2I_URL", "").rstrip("/")
|
||||||
|
IMAGE_PROMPT_I2I_SYSTEM_FILE = os.environ.get(
|
||||||
|
"IMAGE_PROMPT_I2I_SYSTEM_FILE",
|
||||||
|
"/etc/mike-ai/qwen-image-pe-i2i-system-prompt.txt")
|
||||||
|
IMAGE_PROMPT_T2I_SYSTEM_FILE = os.environ.get(
|
||||||
|
"IMAGE_PROMPT_T2I_SYSTEM_FILE",
|
||||||
|
"/etc/mike-ai/qwen-image-pe-t2i-system-prompt.txt")
|
||||||
|
IMAGE_PROMPT_ENHANCER_TIMEOUT = float(os.environ.get(
|
||||||
|
"IMAGE_PROMPT_ENHANCER_TIMEOUT", "180"))
|
||||||
IMAGE_MODEL_NAME = os.environ.get(
|
IMAGE_MODEL_NAME = os.environ.get(
|
||||||
"IMAGE_MODEL_NAME", "Qwen-Image-2.1-int8")
|
"IMAGE_MODEL_NAME", "Qwen-Image-2.1-int8")
|
||||||
IMAGE_INFERENCE_STEPS = int(os.environ.get("IMAGE_INFERENCE_STEPS", "25"))
|
IMAGE_INFERENCE_STEPS = int(os.environ.get("IMAGE_INFERENCE_STEPS", "25"))
|
||||||
@@ -182,6 +193,16 @@ IMAGE_QUALITY = {"standard": IMAGE_INFERENCE_STEPS,
|
|||||||
IMAGE_DEFAULT_QUALITY = "standard"
|
IMAGE_DEFAULT_QUALITY = "standard"
|
||||||
IMAGE_MAX_N = 4
|
IMAGE_MAX_N = 4
|
||||||
|
|
||||||
|
IMAGE_RATIO_SIZES = {
|
||||||
|
"1:1": (1024, 1024),
|
||||||
|
"3:2": (1536, 1024),
|
||||||
|
"2:3": (1024, 1536),
|
||||||
|
"4:3": (1536, 1024),
|
||||||
|
"3:4": (1024, 1536),
|
||||||
|
"16:9": (1920, 1088),
|
||||||
|
"9:16": (1088, 1920),
|
||||||
|
}
|
||||||
|
|
||||||
# --- Sprachausgabe (Qwen3-TTS über das interne Normalisierungs-Gateway) ---
|
# --- Sprachausgabe (Qwen3-TTS über das interne Normalisierungs-Gateway) ---
|
||||||
TTS_WORKER_URL = os.environ.get("TTS_WORKER_URL", "http://127.0.0.1:8085")
|
TTS_WORKER_URL = os.environ.get("TTS_WORKER_URL", "http://127.0.0.1:8085")
|
||||||
TTS_TIMEOUT = float(os.environ.get("TTS_TIMEOUT", "300")) # s, pro Synthese
|
TTS_TIMEOUT = float(os.environ.get("TTS_TIMEOUT", "300")) # s, pro Synthese
|
||||||
@@ -194,11 +215,11 @@ TTS_DEFAULT_VOICE = os.environ.get(
|
|||||||
TTS_FORMATS = ("mp3", "wav", "pcm")
|
TTS_FORMATS = ("mp3", "wav", "pcm")
|
||||||
TTS_DEFAULT_FORMAT = "mp3"
|
TTS_DEFAULT_FORMAT = "mp3"
|
||||||
|
|
||||||
# --- Spracherkennung (whisper.cpp, deutsch, CPU-only) ---
|
# --- Spracherkennung (Qwen3-ASR, deutsch, CPU-only) ---
|
||||||
STT_WORKER_URL = os.environ.get("STT_WORKER_URL", "http://127.0.0.1:8084")
|
STT_WORKER_URL = os.environ.get("STT_WORKER_URL", "http://127.0.0.1:8084")
|
||||||
STT_TIMEOUT = float(os.environ.get("STT_TIMEOUT", "120")) # s, pro Transkription
|
STT_TIMEOUT = float(os.environ.get("STT_TIMEOUT", "120")) # s, pro Transkription
|
||||||
STT_CONNECT_TIMEOUT = float(os.environ.get("STT_CONNECT_TIMEOUT", "5"))
|
STT_CONNECT_TIMEOUT = float(os.environ.get("STT_CONNECT_TIMEOUT", "5"))
|
||||||
STT_MODEL = "whisper-1" # virtuelles Modell für /v1/audio/transcriptions
|
STT_MODEL = "qwen3-asr"
|
||||||
|
|
||||||
# Maximale Upload-Größe (Bytes) – verhindert unbegrenzten RAM-Verbrauch.
|
# Maximale Upload-Größe (Bytes) – verhindert unbegrenzten RAM-Verbrauch.
|
||||||
# 50 MB ist für Audio-Dateien (WebM/Opus, WAV, MP3) mehr als ausreichend.
|
# 50 MB ist für Audio-Dateien (WebM/Opus, WAV, MP3) mehr als ausreichend.
|
||||||
@@ -1386,11 +1407,151 @@ def _restore_qwen(profile: str) -> None:
|
|||||||
RUNTIME.save(last_profile=profile, phase="idle")
|
RUNTIME.save(last_profile=profile, phase="idle")
|
||||||
|
|
||||||
|
|
||||||
|
def _image_data_url(path: str) -> str:
|
||||||
|
"""Read one already validated local reference image as a data URL."""
|
||||||
|
image_root = os.path.abspath(IMAGE_DIR)
|
||||||
|
resolved = (os.path.abspath(path) if os.path.isabs(path)
|
||||||
|
else os.path.abspath(os.path.join(image_root, path)))
|
||||||
|
if os.path.commonpath((image_root, resolved)) != image_root:
|
||||||
|
raise RuntimeError("Referenzbild liegt außerhalb des Bildverzeichnisses")
|
||||||
|
with open(resolved, "rb") as handle:
|
||||||
|
data = handle.read()
|
||||||
|
if data.startswith(b"\x89PNG\r\n\x1a\n"):
|
||||||
|
mime = "image/png"
|
||||||
|
elif data.startswith(b"\xff\xd8\xff"):
|
||||||
|
mime = "image/jpeg"
|
||||||
|
elif data.startswith(b"RIFF") and data[8:12] == b"WEBP":
|
||||||
|
mime = "image/webp"
|
||||||
|
else:
|
||||||
|
raise RuntimeError(f"Referenzbild hat ein unbekanntes Format: {path}")
|
||||||
|
return f"data:{mime};base64,{base64.b64encode(data).decode()}"
|
||||||
|
|
||||||
|
|
||||||
|
def _parse_prompt_enhancer_result(content: object) -> dict:
|
||||||
|
if not isinstance(content, str) or not content.strip():
|
||||||
|
raise RuntimeError("Prompt-Enhancer lieferte keine Antwort")
|
||||||
|
text = content.strip()
|
||||||
|
if text.startswith("```"):
|
||||||
|
text = re.sub(r"^```(?:json)?\s*", "", text, flags=re.IGNORECASE)
|
||||||
|
text = re.sub(r"\s*```$", "", text)
|
||||||
|
try:
|
||||||
|
result = json.loads(text)
|
||||||
|
except ValueError:
|
||||||
|
start, end = text.find("{"), text.rfind("}")
|
||||||
|
if start < 0 or end <= start:
|
||||||
|
raise RuntimeError("Prompt-Enhancer lieferte kein JSON")
|
||||||
|
try:
|
||||||
|
result = json.loads(text[start:end + 1])
|
||||||
|
except ValueError as exc:
|
||||||
|
raise RuntimeError("Prompt-Enhancer lieferte ungültiges JSON") from exc
|
||||||
|
if not isinstance(result, dict):
|
||||||
|
raise RuntimeError("Prompt-Enhancer lieferte kein JSON-Objekt")
|
||||||
|
rewritten = result.get("rewritten_prompt")
|
||||||
|
if not isinstance(rewritten, str) or not rewritten.strip():
|
||||||
|
raise RuntimeError("Prompt-Enhancer lieferte keinen rewritten_prompt")
|
||||||
|
if len(rewritten) > 8000:
|
||||||
|
raise RuntimeError("Aufbereiteter Bildprompt ist länger als 8000 Zeichen")
|
||||||
|
result["rewritten_prompt"] = rewritten.strip()
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
def _enhance_image_prompt(prompt: str, source_files: list[str]) -> tuple[str, dict]:
|
||||||
|
"""Run the official Qwen Image 2.1 prompt enhancer on the RTX 3060."""
|
||||||
|
editing = bool(source_files)
|
||||||
|
kind = "image-prompt-i2i" if editing else "image-prompt-t2i"
|
||||||
|
url = IMAGE_PROMPT_I2I_URL if editing else IMAGE_PROMPT_T2I_URL
|
||||||
|
system_file = (IMAGE_PROMPT_I2I_SYSTEM_FILE if editing
|
||||||
|
else IMAGE_PROMPT_T2I_SYSTEM_FILE)
|
||||||
|
model = "qwen-image-pe-i2i" if editing else "qwen-image-pe-t2i"
|
||||||
|
if not PROFILE_CONTROL_URL or not url:
|
||||||
|
raise RuntimeError("Qwen-Image-Prompt-Enhancer ist nicht konfiguriert")
|
||||||
|
try:
|
||||||
|
with open(system_file, encoding="utf-8") as handle:
|
||||||
|
system_prompt = handle.read().strip()
|
||||||
|
except OSError as exc:
|
||||||
|
raise RuntimeError(f"Systemprompt des Prompt-Enhancers fehlt: {exc}") from exc
|
||||||
|
|
||||||
|
started = time.monotonic()
|
||||||
|
_profile_controller_request("POST", f"/workers/{kind}/start")
|
||||||
|
try:
|
||||||
|
deadline = time.monotonic() + IMAGE_START_TIMEOUT
|
||||||
|
while True:
|
||||||
|
try:
|
||||||
|
with urllib.request.urlopen(url + "/health", timeout=3) as response:
|
||||||
|
health = json.load(response)
|
||||||
|
if health.get("status") == "ok":
|
||||||
|
break
|
||||||
|
except (OSError, urllib.error.URLError, TimeoutError, ValueError):
|
||||||
|
pass
|
||||||
|
if time.monotonic() >= deadline:
|
||||||
|
raise RuntimeError("Prompt-Enhancer hat nicht gestartet")
|
||||||
|
time.sleep(1)
|
||||||
|
|
||||||
|
if editing:
|
||||||
|
user_content: object = [
|
||||||
|
{"type": "image_url", "image_url": {"url": _image_data_url(path)}}
|
||||||
|
for path in source_files
|
||||||
|
]
|
||||||
|
user_content.append({"type": "text", "text": prompt})
|
||||||
|
else:
|
||||||
|
user_content = prompt
|
||||||
|
request_body = {
|
||||||
|
"model": model,
|
||||||
|
"messages": [
|
||||||
|
{"role": "system", "content": system_prompt},
|
||||||
|
{"role": "user", "content": user_content},
|
||||||
|
],
|
||||||
|
"temperature": 1.0,
|
||||||
|
"top_p": 0.95,
|
||||||
|
"top_k": 20,
|
||||||
|
"presence_penalty": 0.0 if editing else 1.5,
|
||||||
|
"max_tokens": 4096,
|
||||||
|
"thinking_budget_tokens": 2048,
|
||||||
|
"chat_template_kwargs": {
|
||||||
|
"enable_thinking": True,
|
||||||
|
"reasoning_effort": "low",
|
||||||
|
},
|
||||||
|
}
|
||||||
|
req = urllib.request.Request(
|
||||||
|
url + "/v1/chat/completions",
|
||||||
|
data=json.dumps(request_body).encode(),
|
||||||
|
method="POST",
|
||||||
|
headers={"Content-Type": "application/json"},
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
with urllib.request.urlopen(
|
||||||
|
req, timeout=IMAGE_PROMPT_ENHANCER_TIMEOUT) as response:
|
||||||
|
payload = json.load(response)
|
||||||
|
except urllib.error.HTTPError as exc:
|
||||||
|
detail = exc.read(4096).decode(errors="replace")
|
||||||
|
raise RuntimeError(
|
||||||
|
f"Prompt-Enhancer HTTP {exc.code}: {detail[-500:]}") from exc
|
||||||
|
try:
|
||||||
|
content = payload["choices"][0]["message"]["content"]
|
||||||
|
except (KeyError, IndexError, TypeError) as exc:
|
||||||
|
raise RuntimeError("Prompt-Enhancer-Antwort ist unvollständig") from exc
|
||||||
|
result = _parse_prompt_enhancer_result(content)
|
||||||
|
metadata = {
|
||||||
|
"model": model,
|
||||||
|
"quantization": "Q5_K_M",
|
||||||
|
"seconds": round(time.monotonic() - started, 3),
|
||||||
|
"wh_ratio": result.get("wh_ratio"),
|
||||||
|
"ratio_follow": result.get("ratio_follow"),
|
||||||
|
}
|
||||||
|
return result["rewritten_prompt"], metadata
|
||||||
|
finally:
|
||||||
|
try:
|
||||||
|
_profile_controller_request("POST", f"/workers/{kind}/stop")
|
||||||
|
except Exception as exc:
|
||||||
|
log.warning("Prompt-Enhancer ließ sich nicht stoppen: %s", exc)
|
||||||
|
|
||||||
|
|
||||||
def generate_image(prompt: str, width: int, height: int, steps: int,
|
def generate_image(prompt: str, width: int, height: int, steps: int,
|
||||||
guidance: float, seed: int | None, n: int,
|
guidance: float, seed: int | None, n: int,
|
||||||
quality: str = "standard",
|
quality: str = "standard",
|
||||||
source_files: list[str] | None = None,
|
source_files: list[str] | None = None,
|
||||||
model: str = IMAGE_MODEL_NAME,
|
model: str = IMAGE_MODEL_NAME,
|
||||||
|
size_explicit: bool = False,
|
||||||
) -> tuple[list[str], str | None]:
|
) -> tuple[list[str], str | None]:
|
||||||
"""Orchestriert die Bildgenerierung inkl. Qwen-Hotswap.
|
"""Orchestriert die Bildgenerierung inkl. Qwen-Hotswap.
|
||||||
|
|
||||||
@@ -1411,6 +1572,8 @@ def generate_image(prompt: str, width: int, height: int, steps: int,
|
|||||||
warning: str | None = None
|
warning: str | None = None
|
||||||
img.last_error = None
|
img.last_error = None
|
||||||
img.current_model = model
|
img.current_model = model
|
||||||
|
original_prompt = prompt
|
||||||
|
prompt_enhancer: dict | None = None
|
||||||
# Qwen wird gestoppt → für Chats nicht verfügbar (die warten).
|
# Qwen wird gestoppt → für Chats nicht verfügbar (die warten).
|
||||||
_set_qwen_unavailable(True)
|
_set_qwen_unavailable(True)
|
||||||
try:
|
try:
|
||||||
@@ -1433,11 +1596,26 @@ def generate_image(prompt: str, width: int, height: int, steps: int,
|
|||||||
f"(Exit {proc.returncode}): {out[-500:]}")
|
f"(Exit {proc.returncode}): {out[-500:]}")
|
||||||
_wait_upstream_down(time.monotonic() + 60)
|
_wait_upstream_down(time.monotonic() + 60)
|
||||||
|
|
||||||
# 2) Worker starten (Modell wird beim ersten generate geladen).
|
# 2) Den knappen Benutzerprompt mit dem offiziellen Qwen-
|
||||||
|
# Prompt-Enhancer auf der RTX 3060 in einen belastbaren
|
||||||
|
# Produktionsprompt umschreiben. Danach wird der Enhancer wieder
|
||||||
|
# beendet, bevor der eigentliche Bildworker startet.
|
||||||
|
img.phase = "enhancing-prompt"
|
||||||
|
prompt, prompt_enhancer = _enhance_image_prompt(
|
||||||
|
original_prompt, source_files or [])
|
||||||
|
ratio = prompt_enhancer.get("wh_ratio")
|
||||||
|
if not size_explicit and isinstance(ratio, str):
|
||||||
|
width, height = IMAGE_RATIO_SIZES.get(
|
||||||
|
ratio.strip(), (width, height))
|
||||||
|
log.info("Bildprompt aufbereitet (%s, %.1f s, Verhältnis %s)",
|
||||||
|
prompt_enhancer["model"],
|
||||||
|
prompt_enhancer["seconds"], ratio)
|
||||||
|
|
||||||
|
# 3) Worker starten (Modell wird beim ersten generate geladen).
|
||||||
img.phase = "loading-image"
|
img.phase = "loading-image"
|
||||||
worker = _worker()
|
worker = _worker()
|
||||||
|
|
||||||
# 3) Generieren.
|
# 4) Generieren.
|
||||||
for i in range(n):
|
for i in range(n):
|
||||||
img.phase = "generating"
|
img.phase = "generating"
|
||||||
filename = time.strftime("%Y%m%d-%H%M%S") + \
|
filename = time.strftime("%Y%m%d-%H%M%S") + \
|
||||||
@@ -1465,6 +1643,8 @@ def generate_image(prompt: str, width: int, height: int, steps: int,
|
|||||||
# Metadaten speichern (Sidecar-JSON).
|
# Metadaten speichern (Sidecar-JSON).
|
||||||
meta = {
|
meta = {
|
||||||
"prompt": prompt,
|
"prompt": prompt,
|
||||||
|
"original_prompt": original_prompt,
|
||||||
|
"prompt_enhancer": prompt_enhancer,
|
||||||
"seed": seed,
|
"seed": seed,
|
||||||
"width": width,
|
"width": width,
|
||||||
"height": height,
|
"height": height,
|
||||||
@@ -1493,7 +1673,7 @@ def generate_image(prompt: str, width: int, height: int, steps: int,
|
|||||||
if removed:
|
if removed:
|
||||||
log.info("Bild-Retention: %d alte Bilder entfernt", len(removed))
|
log.info("Bild-Retention: %d alte Bilder entfernt", len(removed))
|
||||||
|
|
||||||
# 4) Worker vollständig beenden (VRAM + CUDA-Kontext freigeben).
|
# 5) Worker vollständig beenden (VRAM + CUDA-Kontext freigeben).
|
||||||
img.phase = "unloading-image"
|
img.phase = "unloading-image"
|
||||||
worker.stop()
|
worker.stop()
|
||||||
img.worker = None
|
img.worker = None
|
||||||
@@ -1510,7 +1690,7 @@ def generate_image(prompt: str, width: int, height: int, steps: int,
|
|||||||
img.worker = None
|
img.worker = None
|
||||||
raise
|
raise
|
||||||
finally:
|
finally:
|
||||||
# 5) Qwen immer wiederherstellen.
|
# 6) Qwen immer wiederherstellen.
|
||||||
img.phase = "restoring-qwen"
|
img.phase = "restoring-qwen"
|
||||||
try:
|
try:
|
||||||
_restore_qwen(profile)
|
_restore_qwen(profile)
|
||||||
@@ -2358,6 +2538,7 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
return
|
return
|
||||||
|
|
||||||
# Größe
|
# Größe
|
||||||
|
size_explicit = "size" in data
|
||||||
size = data.get("size", "1024x1024")
|
size = data.get("size", "1024x1024")
|
||||||
if size not in IMAGE_SIZES:
|
if size not in IMAGE_SIZES:
|
||||||
self._send_error(
|
self._send_error(
|
||||||
@@ -2425,7 +2606,7 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
try:
|
try:
|
||||||
results, warning = generate_image(
|
results, warning = generate_image(
|
||||||
prompt.strip(), width, height, steps, guidance, seed, n,
|
prompt.strip(), width, height, steps, guidance, seed, n,
|
||||||
quality, source_files, model)
|
quality, source_files, model, size_explicit)
|
||||||
except (ValueError, RuntimeError) as e:
|
except (ValueError, RuntimeError) as e:
|
||||||
self._send_error(503, str(e), "server_error", "image_generation_failed")
|
self._send_error(503, str(e), "server_error", "image_generation_failed")
|
||||||
return
|
return
|
||||||
@@ -2664,7 +2845,7 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
models.append({
|
models.append({
|
||||||
"id": STT_MODEL,
|
"id": STT_MODEL,
|
||||||
"object": "model",
|
"object": "model",
|
||||||
"owned_by": "whisper.cpp",
|
"owned_by": "qwen3-asr",
|
||||||
"type": "transcription",
|
"type": "transcription",
|
||||||
})
|
})
|
||||||
if tts.get("ready"):
|
if tts.get("ready"):
|
||||||
@@ -2787,7 +2968,7 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
|
|
||||||
# Modell-Validierung
|
# Modell-Validierung
|
||||||
model = fields.get("model", STT_MODEL)
|
model = fields.get("model", STT_MODEL)
|
||||||
if model not in (STT_MODEL, "whisper"):
|
if model != STT_MODEL:
|
||||||
self._send_error(400, f"unbekanntes Modell: {model!r} "
|
self._send_error(400, f"unbekanntes Modell: {model!r} "
|
||||||
f"(erwartet: {STT_MODEL})",
|
f"(erwartet: {STT_MODEL})",
|
||||||
"invalid_request_error", "unknown_model")
|
"invalid_request_error", "unknown_model")
|
||||||
|
|||||||
@@ -0,0 +1,147 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""OpenAI-router STT adapter for the persistent, CPU-only Qwen3-ASR server."""
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import subprocess
|
||||||
|
import tempfile
|
||||||
|
import time
|
||||||
|
import urllib.request
|
||||||
|
import uuid
|
||||||
|
from email import policy
|
||||||
|
from email.parser import BytesParser
|
||||||
|
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||||
|
|
||||||
|
|
||||||
|
HOST = os.environ.get("QWEN_ASR_HOST", "0.0.0.0")
|
||||||
|
PORT = int(os.environ.get("QWEN_ASR_PORT", "8084"))
|
||||||
|
SERVER_URL = os.environ.get("QWEN_ASR_SERVER_URL", "http://qwen-asr:8080").rstrip("/")
|
||||||
|
LANGUAGE = os.environ.get("QWEN_ASR_LANGUAGE", "de")
|
||||||
|
MAX_BODY_BYTES = 25 * 1024 * 1024
|
||||||
|
|
||||||
|
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
|
||||||
|
log = logging.getLogger("qwen-asr-worker")
|
||||||
|
|
||||||
|
|
||||||
|
def clean_transcript(value: str) -> str:
|
||||||
|
"""llama.cpp may include a Qwen task marker before the spoken words."""
|
||||||
|
if "<asr_text>" in value:
|
||||||
|
value = value.split("<asr_text>", 1)[1]
|
||||||
|
return value.replace("<|endoftext|>", "").strip()
|
||||||
|
|
||||||
|
|
||||||
|
def transcribe(audio: bytes, filename: str, language: str) -> dict:
|
||||||
|
suffix = os.path.splitext(filename)[1].lower() or ".wav"
|
||||||
|
with tempfile.TemporaryDirectory(prefix="qwen_asr_") as directory:
|
||||||
|
source = os.path.join(directory, "input" + suffix)
|
||||||
|
wav = os.path.join(directory, "audio.wav")
|
||||||
|
with open(source, "wb") as handle:
|
||||||
|
handle.write(audio)
|
||||||
|
result = subprocess.run(
|
||||||
|
["ffmpeg", "-nostdin", "-hide_banner", "-loglevel", "error", "-y",
|
||||||
|
"-i", source, "-ac", "1", "-ar", "16000", "-c:a", "pcm_s16le", wav],
|
||||||
|
capture_output=True, text=True, timeout=30,
|
||||||
|
)
|
||||||
|
if result.returncode:
|
||||||
|
raise ValueError("Audio konnte nicht gelesen werden: " + result.stderr[-300:])
|
||||||
|
with open(wav, "rb") as handle:
|
||||||
|
pcm = handle.read()
|
||||||
|
|
||||||
|
boundary = "athena-qwen-asr-" + uuid.uuid4().hex
|
||||||
|
body = b"".join([
|
||||||
|
f"--{boundary}\r\n".encode(),
|
||||||
|
b'Content-Disposition: form-data; name="file"; filename="audio.wav"\r\n',
|
||||||
|
b"Content-Type: audio/wav\r\n\r\n", pcm, b"\r\n",
|
||||||
|
f"--{boundary}\r\n".encode(),
|
||||||
|
b'Content-Disposition: form-data; name="model"\r\n\r\n',
|
||||||
|
b"qwen3-asr-0.6b\r\n",
|
||||||
|
f"--{boundary}\r\n".encode(),
|
||||||
|
b'Content-Disposition: form-data; name="language"\r\n\r\n',
|
||||||
|
language.encode(), b"\r\n",
|
||||||
|
f"--{boundary}--\r\n".encode(),
|
||||||
|
])
|
||||||
|
request = urllib.request.Request(
|
||||||
|
SERVER_URL + "/v1/audio/transcriptions", data=body,
|
||||||
|
headers={"Content-Type": f"multipart/form-data; boundary={boundary}"},
|
||||||
|
method="POST",
|
||||||
|
)
|
||||||
|
started = time.monotonic()
|
||||||
|
with urllib.request.urlopen(request, timeout=60) as response:
|
||||||
|
payload = json.loads(response.read())
|
||||||
|
if not isinstance(payload, dict) or not isinstance(payload.get("text"), str):
|
||||||
|
raise RuntimeError("Qwen3-ASR returned no transcription")
|
||||||
|
text = clean_transcript(payload["text"])
|
||||||
|
elapsed = int((time.monotonic() - started) * 1000)
|
||||||
|
log.info("Qwen3-ASR transcribed %d characters in %d ms", len(text), elapsed)
|
||||||
|
return {"text": text, "language": language, "duration_ms": elapsed,
|
||||||
|
"engine": "qwen3-asr-0.6b"}
|
||||||
|
|
||||||
|
|
||||||
|
def parse_audio(body: bytes, content_type: str) -> tuple[bytes, str, str]:
|
||||||
|
if "multipart/form-data" not in content_type.lower():
|
||||||
|
return body, "audio.wav", LANGUAGE
|
||||||
|
message = BytesParser(policy=policy.default).parsebytes(
|
||||||
|
b"MIME-Version: 1.0\r\nContent-Type: " + content_type.encode() +
|
||||||
|
b"\r\n\r\n" + body
|
||||||
|
)
|
||||||
|
if not message.is_multipart():
|
||||||
|
raise ValueError("Invalid multipart upload")
|
||||||
|
audio = b""
|
||||||
|
filename = "audio.wav"
|
||||||
|
language = LANGUAGE
|
||||||
|
for part in message.iter_parts():
|
||||||
|
name = part.get_param("name", header="content-disposition")
|
||||||
|
if name == "file":
|
||||||
|
audio = part.get_payload(decode=True) or b""
|
||||||
|
filename = os.path.basename(part.get_filename() or filename)
|
||||||
|
elif name == "language":
|
||||||
|
language = (part.get_payload(decode=True) or b"").decode("utf-8").strip()
|
||||||
|
return audio, filename, language if language and language != "auto" else LANGUAGE
|
||||||
|
|
||||||
|
|
||||||
|
class Handler(BaseHTTPRequestHandler):
|
||||||
|
def send_json(self, status: int, data: dict) -> None:
|
||||||
|
body = json.dumps(data, ensure_ascii=False).encode()
|
||||||
|
self.send_response(status)
|
||||||
|
self.send_header("Content-Type", "application/json; charset=utf-8")
|
||||||
|
self.send_header("Content-Length", str(len(body)))
|
||||||
|
self.end_headers()
|
||||||
|
self.wfile.write(body)
|
||||||
|
|
||||||
|
def do_GET(self) -> None:
|
||||||
|
if self.path != "/status":
|
||||||
|
self.send_json(404, {"error": "not found"})
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
with urllib.request.urlopen(SERVER_URL + "/health", timeout=2) as response:
|
||||||
|
ready = response.status == 200
|
||||||
|
except Exception:
|
||||||
|
ready = False
|
||||||
|
self.send_json(200, {"ready": ready, "model": "qwen3-asr-0.6b",
|
||||||
|
"engine": "qwen3-asr", "language": LANGUAGE})
|
||||||
|
|
||||||
|
def do_POST(self) -> None:
|
||||||
|
if self.path != "/transcribe":
|
||||||
|
self.send_json(404, {"error": "not found"})
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
size = int(self.headers.get("Content-Length", "0"))
|
||||||
|
if not 0 < size <= MAX_BODY_BYTES:
|
||||||
|
self.send_json(413, {"error": "Invalid audio size"})
|
||||||
|
return
|
||||||
|
audio, filename, language = parse_audio(
|
||||||
|
self.rfile.read(size), self.headers.get("Content-Type", "")
|
||||||
|
)
|
||||||
|
if not audio:
|
||||||
|
raise ValueError("Missing audio file")
|
||||||
|
self.send_json(200, transcribe(audio, filename, language))
|
||||||
|
except ValueError as exc:
|
||||||
|
self.send_json(400, {"error": str(exc)})
|
||||||
|
except Exception:
|
||||||
|
log.exception("Transcription failed")
|
||||||
|
self.send_json(503, {"error": "Qwen3-ASR unavailable"})
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
ThreadingHTTPServer((HOST, PORT), Handler).serve_forever()
|
||||||
@@ -18,7 +18,7 @@
|
|||||||
"ultra": {
|
"ultra": {
|
||||||
"context": 262144,
|
"context": 262144,
|
||||||
"model_alias": "qwen-ultra",
|
"model_alias": "qwen-ultra",
|
||||||
"vision": false
|
"vision": true
|
||||||
},
|
},
|
||||||
"uncensored": {
|
"uncensored": {
|
||||||
"context": 80000,
|
"context": 80000,
|
||||||
|
|||||||
@@ -5,6 +5,8 @@ set -Eeuo pipefail
|
|||||||
ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||||
ENV_FILE=${STACK_ENV:-/etc/mike-ai/stack.env}
|
ENV_FILE=${STACK_ENV:-/etc/mike-ai/stack.env}
|
||||||
MODEL_DIR=${QWEN_IMAGE_21_MODEL_DIR:-/data/models/Qwen-Image-2.1-ComfyUI}
|
MODEL_DIR=${QWEN_IMAGE_21_MODEL_DIR:-/data/models/Qwen-Image-2.1-ComfyUI}
|
||||||
|
PE_I2I_DIR=${QWEN_IMAGE_PE_I2I_MODEL_DIR:-/data/models/qwen-image-2.1-pe-i2i-q5}
|
||||||
|
PE_T2I_DIR=${QWEN_IMAGE_PE_T2I_MODEL_DIR:-/data/models/qwen-image-2.1-pe-t2i-q5}
|
||||||
BASE=https://huggingface.co/Comfy-Org/Qwen-Image-2.1/resolve/ace0edeb3791a594ddfa36ed5f41a178a394e921
|
BASE=https://huggingface.co/Comfy-Org/Qwen-Image-2.1/resolve/ace0edeb3791a594ddfa36ed5f41a178a394e921
|
||||||
|
|
||||||
download() {
|
download() {
|
||||||
@@ -19,15 +21,52 @@ download() {
|
|||||||
mv "$target.part" "$target"
|
mv "$target.part" "$target"
|
||||||
}
|
}
|
||||||
|
|
||||||
|
download_url() {
|
||||||
|
local url=$1 target=$2 expected=$3
|
||||||
|
install -d -m 0755 "$(dirname "$target")"
|
||||||
|
if [[ -f $target ]] && echo "$expected $target" | sha256sum -c --status; then
|
||||||
|
echo "OK: $target"
|
||||||
|
return
|
||||||
|
fi
|
||||||
|
curl -fL --retry 5 --retry-all-errors -C - -o "$target.part" "$url"
|
||||||
|
echo "$expected $target.part" | sha256sum -c --status
|
||||||
|
mv "$target.part" "$target"
|
||||||
|
}
|
||||||
|
|
||||||
download diffusion_models/qwen_image_2.1_int8_convrot.safetensors cb74113cb03faecd79611b01fd7fd642f0aa60d6f0b95086abee214d75eaa57d
|
download diffusion_models/qwen_image_2.1_int8_convrot.safetensors cb74113cb03faecd79611b01fd7fd642f0aa60d6f0b95086abee214d75eaa57d
|
||||||
download text_encoders/qwen3vl_8b_int8_convrot.safetensors 8bfd0f6e12abf2d2d697ecc888e5e90b0d6741d6708f05799f53afa560452e8f
|
download text_encoders/qwen3vl_8b_int8_convrot.safetensors 8bfd0f6e12abf2d2d697ecc888e5e90b0d6741d6708f05799f53afa560452e8f
|
||||||
download vae/qwen_image_2.1_vae_bf16.safetensors bb21f7473051e1ac368515dd3f2e15cd44d7a11748ee8823e1ddca3e4876b7c9
|
download vae/qwen_image_2.1_vae_bf16.safetensors bb21f7473051e1ac368515dd3f2e15cd44d7a11748ee8823e1ddca3e4876b7c9
|
||||||
|
|
||||||
|
# Official Qwen Image 2.1 Prompt Enhancers, quantized to Q5_K_M for the
|
||||||
|
# 12-GiB RTX 3060. I2I uses the separate multimodal projector; T2I does not.
|
||||||
|
download_url \
|
||||||
|
https://huggingface.co/prithivMLmods/Qwen-Image-2.1-PE-I2I-GGUF/resolve/55b9c1a326599e142d59bcad8715d5601ccf8daa/Qwen-Image-2.1-PE-I2I.Q5_K_M.gguf \
|
||||||
|
"$PE_I2I_DIR/Qwen-Image-2.1-PE-I2I.Q5_K_M.gguf" \
|
||||||
|
cb71f71fe5fe5938570d65a7c403fbcb1a9065dff9dd9a021f58aec42785b84e
|
||||||
|
download_url \
|
||||||
|
https://huggingface.co/prithivMLmods/Qwen-Image-2.1-PE-I2I-GGUF/resolve/55b9c1a326599e142d59bcad8715d5601ccf8daa/Qwen-Image-2.1-PE-I2I.mmproj-bf16.gguf \
|
||||||
|
"$PE_I2I_DIR/Qwen-Image-2.1-PE-I2I.mmproj-bf16.gguf" \
|
||||||
|
8dedb71dbc3092dc47de9108ad373d68a12854e2527d59bd9399601738f3bce1
|
||||||
|
download_url \
|
||||||
|
https://huggingface.co/Qwen/Qwen-Image-2.1-PE-I2I/resolve/72927bc08afc99b7888ceb7d7d51a12db3700bbd/system_prompt.txt \
|
||||||
|
"$PE_I2I_DIR/system_prompt.txt" \
|
||||||
|
e378fea686a1431581ba4c654d332ae96adad633f144ae738ec8ce9c4fd66439
|
||||||
|
download_url \
|
||||||
|
https://huggingface.co/prithivMLmods/Qwen-Image-2.1-PE-T2I-GGUF/resolve/e18d4a3e0830ab157770738b16830e6fcf5f57d4/Qwen-Image-2.1-PE-T2I.Q5_K_M.gguf \
|
||||||
|
"$PE_T2I_DIR/Qwen-Image-2.1-PE-T2I.Q5_K_M.gguf" \
|
||||||
|
749f5652fd6e8b760ca860091f7a5bffa254a7954203ae5c9bcdf5f93197702f
|
||||||
|
download_url \
|
||||||
|
https://huggingface.co/Qwen/Qwen-Image-2.1-PE-T2I/resolve/f3ed7985c788ad75b3ab7223e0c4c51e2a43545b/system_prompt.txt \
|
||||||
|
"$PE_T2I_DIR/system_prompt.txt" \
|
||||||
|
a77c9a06c59b120741141d9514b95682bb8761d02bec49ca61def7b2b3d9fb99
|
||||||
|
|
||||||
compose=(docker compose --env-file "$ENV_FILE" -f "$ROOT_DIR/compose.yaml" --profile image --profile flux-standby)
|
compose=(docker compose --env-file "$ENV_FILE" -f "$ROOT_DIR/compose.yaml" --profile image --profile flux-standby)
|
||||||
"${compose[@]}" build image-worker
|
"${compose[@]}" build image-worker
|
||||||
"${compose[@]}" up -d --build --no-deps profile-controller
|
"${compose[@]}" up -d --build --no-deps profile-controller
|
||||||
"${compose[@]}" create image-worker flux-image-worker
|
"${compose[@]}" create image-worker image-prompt-enhancer-i2i \
|
||||||
"${compose[@]}" stop image-worker flux-image-worker
|
image-prompt-enhancer-t2i flux-image-worker
|
||||||
|
"${compose[@]}" stop image-worker image-prompt-enhancer-i2i \
|
||||||
|
image-prompt-enhancer-t2i flux-image-worker
|
||||||
state=$(docker inspect -f '{{.State.Status}}' mike-ai-image-worker)
|
state=$(docker inspect -f '{{.State.Status}}' mike-ai-image-worker)
|
||||||
[[ $state == exited || $state == created ]]
|
[[ $state == exited || $state == created ]]
|
||||||
echo "Qwen-Image-2.1 production worker and FLUX standby are prepared and stopped."
|
echo "Qwen-Image-2.1, both prompt enhancers and the FLUX standby are prepared and stopped."
|
||||||
@@ -47,7 +47,7 @@ def main() -> int:
|
|||||||
body = (
|
body = (
|
||||||
f"--{boundary}\r\n"
|
f"--{boundary}\r\n"
|
||||||
'Content-Disposition: form-data; name="model"\r\n\r\n'
|
'Content-Disposition: form-data; name="model"\r\n\r\n'
|
||||||
"whisper-1\r\n"
|
"qwen3-asr\r\n"
|
||||||
f"--{boundary}\r\n"
|
f"--{boundary}\r\n"
|
||||||
'Content-Disposition: form-data; name="language"\r\n\r\n'
|
'Content-Disposition: form-data; name="language"\r\n\r\n'
|
||||||
"de\r\n"
|
"de\r\n"
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
# Athena realtime voice bridge
|
# Athena realtime voice bridge
|
||||||
|
|
||||||
This independent service lets OpenClaw's existing browser Talk UI use Athena
|
This independent service lets OpenClaw's existing browser Talk UI use Athena
|
||||||
Whisper, OpenClaw's agent, and Athena Qwen3-TTS through the browser's supported
|
Qwen3-ASR, OpenClaw's agent, and Athena Qwen3-TTS through the browser's supported
|
||||||
OpenAI-style WebRTC transport. It does not modify OpenClaw or switch an Athena
|
OpenAI-style WebRTC transport. It does not modify OpenClaw or switch an Athena
|
||||||
profile. The existing `gateway-relay` path remains available.
|
profile. The existing `gateway-relay` path remains available.
|
||||||
|
|
||||||
@@ -69,7 +69,7 @@ cd /opt/mike-ai/stack
|
|||||||
docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice
|
docker compose --env-file /etc/mike-ai/stack.env up -d --no-deps --build realtime-voice
|
||||||
```
|
```
|
||||||
|
|
||||||
OpenClaw 2026.9.4 uses the `athena-talk` plugin version 1.2.1 with
|
OpenClaw uses the `athena-talk` plugin version 1.3.0 with
|
||||||
`talk.realtime.transport` set to `webrtc` and
|
`talk.realtime.transport` set to `webrtc` and
|
||||||
`talk.realtime.providers.athena-talk.realtimeUpstreamUrl` set to
|
`talk.realtime.providers.athena-talk.realtimeUpstreamUrl` set to
|
||||||
`http://192.168.1.212:8090/v1/realtime/calls`. On the Mac, turn off
|
`http://192.168.1.212:8090/v1/realtime/calls`. On the Mac, turn off
|
||||||
@@ -90,10 +90,14 @@ python3.11 -m venv .venv
|
|||||||
```
|
```
|
||||||
|
|
||||||
The synthetic microphone test passed over the actual OpenClaw HTTPS offer
|
The synthetic microphone test passed over the actual OpenClaw HTTPS offer
|
||||||
route and WireGuard media path with production Whisper and Qwen3-TTS on
|
route and WireGuard media path with production Qwen3-ASR and Qwen3-TTS on
|
||||||
2026-09-16. Subsequent real browser Talk sessions successfully transcribed
|
2026-09-16. Subsequent real browser Talk sessions successfully transcribed
|
||||||
and answered multiple user turns. Browser dictation through the same plugin
|
and answered multiple user turns. Browser dictation uses a separate plugin
|
||||||
also produced text after a correction to its separate transcription path.
|
path. Since version 1.3.0, recordings longer than six seconds are
|
||||||
|
pre-transcribed incrementally in overlapping short windows while the
|
||||||
|
microphone remains active. This keeps the final tail inside OpenClaw's
|
||||||
|
five-second completion window. Talk through this WebRTC service still ends
|
||||||
|
and transcribes one utterance at a time.
|
||||||
The first user report of a second turn becoming stuck was addressed by
|
The first user report of a second turn becoming stuck was addressed by
|
||||||
matching conversation item and predecessor IDs; the later two-turn test
|
matching conversation item and predecessor IDs; the later two-turn test
|
||||||
passed. This is still half-duplex and needs a private route for WebRTC media.
|
passed. This is still half-duplex and needs a private route for WebRTC media.
|
||||||
|
|||||||
@@ -281,7 +281,7 @@ class RealtimeSession:
|
|||||||
try:
|
try:
|
||||||
form = FormData()
|
form = FormData()
|
||||||
form.add_field("file", make_wav(pcm), filename="talk.wav", content_type="audio/wav")
|
form.add_field("file", make_wav(pcm), filename="talk.wav", content_type="audio/wav")
|
||||||
form.add_field("model", "whisper-1")
|
form.add_field("model", "qwen3-asr")
|
||||||
form.add_field("language", os.environ.get("STT_LANGUAGE", "de"))
|
form.add_field("language", os.environ.get("STT_LANGUAGE", "de"))
|
||||||
async with self.http.post(
|
async with self.http.post(
|
||||||
os.environ["ATHENA_API_BASE_URL"].rstrip("/") + "/audio/transcriptions",
|
os.environ["ATHENA_API_BASE_URL"].rstrip("/") + "/audio/transcriptions",
|
||||||
|
|||||||
+4
-1
@@ -27,6 +27,8 @@ for name in \
|
|||||||
mike-ai-wireguard-gateway \
|
mike-ai-wireguard-gateway \
|
||||||
mike-ai-profile-controller \
|
mike-ai-profile-controller \
|
||||||
mike-ai-embedding \
|
mike-ai-embedding \
|
||||||
|
mike-ai-qwen-asr \
|
||||||
|
mike-ai-qwen-asr-worker \
|
||||||
mike-ai-router \
|
mike-ai-router \
|
||||||
mike-ai-qwen3-tts \
|
mike-ai-qwen3-tts \
|
||||||
mike-ai-tts-gateway \
|
mike-ai-tts-gateway \
|
||||||
@@ -70,7 +72,8 @@ for legacy in \
|
|||||||
mike-ai-mcp-deemix \
|
mike-ai-mcp-deemix \
|
||||||
mike-ai-mcp-github \
|
mike-ai-mcp-github \
|
||||||
mike-ai-mcp-homeassistant \
|
mike-ai-mcp-homeassistant \
|
||||||
mike-ai-mcp-navidrome; do
|
mike-ai-mcp-navidrome \
|
||||||
|
mike-ai-whisper; do
|
||||||
[[ -z $(docker inspect -f '{{.Name}}' "$legacy" 2>/dev/null || true) ]] || \
|
[[ -z $(docker inspect -f '{{.Name}}' "$legacy" 2>/dev/null || true) ]] || \
|
||||||
fail "Altlast existiert noch: $legacy"
|
fail "Altlast existiert noch: $legacy"
|
||||||
done
|
done
|
||||||
|
|||||||
Reference in new issue
Block a user