Make Qwen Image 2.1 the production image worker
This commit is contained in:
@@ -20,7 +20,7 @@ Sie betreibt:
|
|||||||
|
|
||||||
- llama.cpp mit genau einem aktiven Qwen-Profil,
|
- llama.cpp mit genau einem aktiven Qwen-Profil,
|
||||||
- den OpenAI-kompatiblen Profile Router,
|
- den OpenAI-kompatiblen Profile Router,
|
||||||
- FLUX.2 Klein 9B FP8 Beta für Textbilder und Referenzbild-Bearbeitung,
|
- Qwen-Image-2.1 INT8 für Textbilder und Referenzbild-Bearbeitung,
|
||||||
- Qwen3-TTS und TTS-Gateway für Sprache,
|
- Qwen3-TTS und TTS-Gateway für Sprache,
|
||||||
- Whisper.cpp und die WebRTC-Brücke für OpenClaw Talk,
|
- Whisper.cpp und die WebRTC-Brücke für OpenClaw Talk,
|
||||||
- EmbeddingGemma auf der CPU für OpenClaws semantische Memory-Suche,
|
- EmbeddingGemma auf der CPU für OpenClaws semantische Memory-Suche,
|
||||||
@@ -90,7 +90,7 @@ nicht direkt. Kein automatischer Host-Neustart ist vorgesehen.
|
|||||||
- Fast: kurze, interaktive Aufgaben
|
- Fast: kurze, interaktive Aufgaben
|
||||||
- Medium/Large/Ultra: steigende Kontextgrößen desselben lokalen Qwen-Modells
|
- Medium/Large/Ultra: steigende Kontextgrößen desselben lokalen Qwen-Modells
|
||||||
- Uncensored: separates lokales Profil
|
- Uncensored: separates lokales Profil
|
||||||
- FLUX.2 Klein 9B FP8 Beta: Bildgenerierung und Editing; Qwen wird dafür kurz entladen und danach
|
- Qwen-Image-2.1 INT8: Bildgenerierung und Editing auf der RTX 5080; das Textmodell wird dafür kurz entladen und danach
|
||||||
automatisch wiederhergestellt
|
automatisch wiederhergestellt
|
||||||
- Qwen3-TTS 1.7B: RTX 3060; kein Piper-Fallback
|
- Qwen3-TTS 1.7B: RTX 3060; kein Piper-Fallback
|
||||||
- EmbeddingGemma 300M Q8: CPU, OpenAI-kompatibel auf Port 8082; kein GPU-Zugriff
|
- EmbeddingGemma 300M Q8: CPU, OpenAI-kompatibel auf Port 8082; kein GPU-Zugriff
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# Athena AI
|
# Athena AI
|
||||||
|
|
||||||
Stand: **20. September 2026**, auf Athena geprüft. Die Textprofil-Images verwenden
|
Stand: **21. September 2026**, auf Athena geprüft. Die Textprofil-Images verwenden
|
||||||
**llama.cpp 0.4.1** (`b29c606`).
|
**llama.cpp 0.4.1** (`b29c606`).
|
||||||
Der [geprüfte Live-Stand](docs/LIVE_STATE.md) beschreibt Profile, GPUs und die
|
Der [geprüfte Live-Stand](docs/LIVE_STATE.md) beschreibt Profile, GPUs und die
|
||||||
Abweichung zwischen dem bereitgestellten Stack und dem Git-Checkout.
|
Abweichung zwischen dem bereitgestellten Stack und dem Git-Checkout.
|
||||||
@@ -21,9 +21,8 @@ dokumentieren den ausgerollten Stand, reduzierte Statuslatenzen und Tests.
|
|||||||
|
|
||||||
- genau ein aktives llama.cpp-Profil: Fast, Medium, Large, Ultra oder Uncensored
|
- genau ein aktives llama.cpp-Profil: Fast, Medium, Large, Ultra oder Uncensored
|
||||||
- Profile Router auf Port 8081
|
- Profile Router auf Port 8081
|
||||||
- FLUX.2 Klein 9B FP8 Beta: Transformer/VAE auf RTX 5080, Textencoder auf RTX 3060
|
- Qwen-Image-2.1 INT8 als produktiver Bildworker auf der RTX 5080
|
||||||
- Qwen-Image-2.1 als isolierter, standardmäßig gestoppter Vergleichsworker
|
- FLUX.2 Klein 9B FP8 Beta als gestoppter Rückfallcontainer
|
||||||
([Testablauf](docs/QWEN_IMAGE_21_TEST.md)); nicht produktiv an den Router angebunden
|
|
||||||
- Qwen3-TTS 1.7B auf der RTX 3060 hinter dem TTS-Gateway; kein Piper-Fallback
|
- Qwen3-TTS 1.7B auf der RTX 3060 hinter dem TTS-Gateway; kein Piper-Fallback
|
||||||
- Whisper.cpp `small` auf der CPU für lokale deutsche Spracherkennung
|
- Whisper.cpp `small` auf der CPU für lokale deutsche Spracherkennung
|
||||||
- EmbeddingGemma 300M Q8 auf der CPU für OpenClaws hybride Memory-Suche
|
- EmbeddingGemma 300M Q8 auf der CPU für OpenClaws hybride Memory-Suche
|
||||||
@@ -44,15 +43,16 @@ dokumentieren den ausgerollten Stand, reduzierte Statuslatenzen und Tests.
|
|||||||
Hermes nutzt Athenas Router unter `http://192.168.1.212:8081/v1`. Ein MCPHub
|
Hermes nutzt Athenas Router unter `http://192.168.1.212:8081/v1`. Ein MCPHub
|
||||||
ist nicht mehr Bestandteil der produktiven Architektur.
|
ist nicht mehr Bestandteil der produktiven Architektur.
|
||||||
|
|
||||||
## Verifizierte Bildauflösung auf dem Live-System
|
## Produktiver Bildpfad
|
||||||
|
|
||||||
Der am 12. September geprüfte Live-Worker verwendet **FLUX.2 Klein 9B FP8
|
Bildanfragen an den Router verwenden seit dem 21. September
|
||||||
Beta**. Die frühere 4B-Angabe war veraltet. Im Auflösungstest
|
**Qwen-Image-2.1 INT8** über gepinntes ComfyUI. Der Worker läuft mit Low-VRAM
|
||||||
war **1024 × 1024 erfolgreich**, während **1280 × 1280 einen CUDA-OOM** auf
|
auf der RTX 5080, 25 Schritten, CFG 1 und kann bis zu vier Referenzbilder
|
||||||
der RTX 5080 auslöste. Höhere Auflösungen sind damit nicht freigegeben;
|
verarbeiten. Das aktive Textprofil und TTS werden für den Auftrag angehalten
|
||||||
Zwischenwerte wurden nicht getestet. Die produktive 1024-Begrenzung bleibt
|
und anschließend wiederhergestellt. Der bisherige FLUX.2-Worker und seine
|
||||||
bestehen. Messwerte, Testbedingungen und die Korrektur früherer optimistischer
|
Gewichte bleiben als gestoppter, explizit allowlist-beschränkter Rückfallpfad
|
||||||
Schätzungen stehen in [Auflösungstest vom 12. September](docs/FLUX_RESOLUTION_TEST_20260912.md).
|
erhalten. Reproduzierbarer Stand und Rückschaltung:
|
||||||
|
[Qwen-Image-2.1](docs/QWEN_IMAGE_21.md).
|
||||||
|
|
||||||
## Installation des Repository-Stands
|
## Installation des Repository-Stands
|
||||||
|
|
||||||
|
|||||||
+51
-75
@@ -572,7 +572,7 @@ services:
|
|||||||
ALLOWED_PROFILES: fast,medium,large,ultra,uncensored
|
ALLOWED_PROFILES: fast,medium,large,ultra,uncensored
|
||||||
IMAGE_WORKER: image
|
IMAGE_WORKER: image
|
||||||
RESTORE_WORKER: restore
|
RESTORE_WORKER: restore
|
||||||
QWEN_IMAGE_TEST_WORKER: qwen-image-2.1-test
|
FLUX_STANDBY_WORKER: flux-standby
|
||||||
TTS_WORKER: qwen3
|
TTS_WORKER: qwen3
|
||||||
MUSIC_WORKER: acestep
|
MUSIC_WORKER: acestep
|
||||||
YUE2_WORKER: yue2
|
YUE2_WORKER: yue2
|
||||||
@@ -631,7 +631,8 @@ services:
|
|||||||
IMAGE_DIR: /data/images
|
IMAGE_DIR: /data/images
|
||||||
IMAGE_WORKER_URL: http://image-worker:8086
|
IMAGE_WORKER_URL: http://image-worker:8086
|
||||||
IMAGE_WORKER_TOKEN: "${CONTROLLER_TOKEN:?CONTROLLER_TOKEN is required}"
|
IMAGE_WORKER_TOKEN: "${CONTROLLER_TOKEN:?CONTROLLER_TOKEN is required}"
|
||||||
IMAGE_MODEL_NAME: FLUX.2-klein-9B-fp8-beta
|
IMAGE_MODEL_NAME: Qwen-Image-2.1-int8
|
||||||
|
IMAGE_INFERENCE_STEPS: "25"
|
||||||
CHAT_IMAGE_ALLOW_REMOTE_URLS: "false"
|
CHAT_IMAGE_ALLOW_REMOTE_URLS: "false"
|
||||||
ENABLE_IMAGE_GENERATION: "true"
|
ENABLE_IMAGE_GENERATION: "true"
|
||||||
ENABLE_TTS: "true"
|
ENABLE_TTS: "true"
|
||||||
@@ -673,6 +674,51 @@ services:
|
|||||||
condition: service_healthy
|
condition: service_healthy
|
||||||
|
|
||||||
image-worker:
|
image-worker:
|
||||||
|
build:
|
||||||
|
context: platform/docker/qwen-image-worker
|
||||||
|
args:
|
||||||
|
COMFYUI_COMMIT: b0f4b7b294ce482a2e071d9d762c133d38c7aa07
|
||||||
|
image: mike-ai/qwen-image-worker:local
|
||||||
|
container_name: mike-ai-image-worker
|
||||||
|
restart: "no"
|
||||||
|
profiles: [image]
|
||||||
|
labels:
|
||||||
|
com.mike-ai.image-worker: image
|
||||||
|
deploy:
|
||||||
|
resources:
|
||||||
|
reservations:
|
||||||
|
devices:
|
||||||
|
- driver: nvidia
|
||||||
|
device_ids: ["${QWEN_IMAGE_21_GPU:-1}"]
|
||||||
|
capabilities: [gpu]
|
||||||
|
read_only: true
|
||||||
|
tmpfs:
|
||||||
|
- /tmp:size=4g,mode=1777
|
||||||
|
- /opt/ComfyUI/user:size=64m,mode=0755
|
||||||
|
- /opt/ComfyUI/temp:size=4g,mode=1777
|
||||||
|
- /opt/ComfyUI/input:size=512m,mode=0755
|
||||||
|
volumes:
|
||||||
|
- "${QWEN_IMAGE_21_MODEL_DIR:-/data/models/Qwen-Image-2.1-ComfyUI}:/opt/ComfyUI/models:ro"
|
||||||
|
- router-images:/data/images
|
||||||
|
environment:
|
||||||
|
NVIDIA_DRIVER_CAPABILITIES: compute,utility
|
||||||
|
PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True
|
||||||
|
WORKER_TOKEN: "${CONTROLLER_TOKEN:?CONTROLLER_TOKEN is required}"
|
||||||
|
IMAGE_DIR: /data/images
|
||||||
|
networks: [inference]
|
||||||
|
security_opt: ["no-new-privileges:true"]
|
||||||
|
cap_drop: [ALL]
|
||||||
|
healthcheck:
|
||||||
|
test: [CMD, python, -c, "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8086/health', timeout=2)"]
|
||||||
|
interval: 5s
|
||||||
|
timeout: 3s
|
||||||
|
retries: 90
|
||||||
|
start_period: 10s
|
||||||
|
|
||||||
|
# Previous production image model, retained as a stopped rollback target.
|
||||||
|
# It is outside the normal router path and can only be started through the
|
||||||
|
# controller's allowlisted flux-standby endpoint.
|
||||||
|
flux-image-worker:
|
||||||
build:
|
build:
|
||||||
context: platform/docker/image-worker
|
context: platform/docker/image-worker
|
||||||
args:
|
args:
|
||||||
@@ -681,11 +727,11 @@ services:
|
|||||||
ACCELERATE_VERSION: ${ACCELERATE_VERSION:-1.14.0}
|
ACCELERATE_VERSION: ${ACCELERATE_VERSION:-1.14.0}
|
||||||
HF_HUB_VERSION: ${HF_HUB_VERSION:-1.28.0}
|
HF_HUB_VERSION: ${HF_HUB_VERSION:-1.28.0}
|
||||||
image: mike-ai/image-worker:local
|
image: mike-ai/image-worker:local
|
||||||
container_name: mike-ai-image-worker
|
container_name: mike-ai-flux-image-worker
|
||||||
restart: "no"
|
restart: "no"
|
||||||
profiles: [image]
|
profiles: [flux-standby]
|
||||||
labels:
|
labels:
|
||||||
com.mike-ai.image-worker: image
|
com.mike-ai.image-worker: flux-standby
|
||||||
gpus: all
|
gpus: all
|
||||||
read_only: true
|
read_only: true
|
||||||
tmpfs: ["/tmp:size=1g,mode=1777"]
|
tmpfs: ["/tmp:size=1g,mode=1777"]
|
||||||
@@ -709,76 +755,6 @@ services:
|
|||||||
timeout: 3s
|
timeout: 3s
|
||||||
retries: 12
|
retries: 12
|
||||||
|
|
||||||
# Isolated evaluation target. It is created in a stopped state by the
|
|
||||||
# preparation script and can only be started through the controller's
|
|
||||||
# allowlisted, GPU-exclusive test endpoint.
|
|
||||||
qwen-image-21-test:
|
|
||||||
build:
|
|
||||||
context: platform/docker/qwen-image-21-test
|
|
||||||
args:
|
|
||||||
COMFYUI_COMMIT: b0f4b7b294ce482a2e071d9d762c133d38c7aa07
|
|
||||||
image: mike-ai/qwen-image-2.1-test:local
|
|
||||||
container_name: mike-ai-qwen-image-2.1-test
|
|
||||||
restart: "no"
|
|
||||||
profiles: [qwen-image-test]
|
|
||||||
labels:
|
|
||||||
com.mike-ai.image-worker: qwen-image-2.1-test
|
|
||||||
deploy:
|
|
||||||
resources:
|
|
||||||
reservations:
|
|
||||||
devices:
|
|
||||||
- driver: nvidia
|
|
||||||
device_ids: ["${QWEN_IMAGE_21_GPU:-1}"]
|
|
||||||
capabilities: [gpu]
|
|
||||||
read_only: true
|
|
||||||
tmpfs:
|
|
||||||
- /tmp:size=4g,mode=1777
|
|
||||||
- /opt/ComfyUI/user:size=64m,mode=0755
|
|
||||||
- /opt/ComfyUI/temp:size=4g,mode=1777
|
|
||||||
volumes:
|
|
||||||
- "${QWEN_IMAGE_21_MODEL_DIR:-/data/models/Qwen-Image-2.1-ComfyUI}:/opt/ComfyUI/models:ro"
|
|
||||||
- "${QWEN_IMAGE_21_OUTPUT_DIR:-/data/qwen-image-2.1-test-output}:/opt/ComfyUI/output"
|
|
||||||
environment:
|
|
||||||
NVIDIA_DRIVER_CAPABILITIES: compute,utility
|
|
||||||
PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True
|
|
||||||
command:
|
|
||||||
- python
|
|
||||||
- main.py
|
|
||||||
- --listen
|
|
||||||
- 0.0.0.0
|
|
||||||
- --port
|
|
||||||
- "8188"
|
|
||||||
- --lowvram
|
|
||||||
- --preview-method
|
|
||||||
- none
|
|
||||||
networks: [inference]
|
|
||||||
security_opt: ["no-new-privileges:true"]
|
|
||||||
cap_drop: [ALL]
|
|
||||||
healthcheck:
|
|
||||||
test: [CMD, python, -c, "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8188/system_stats', timeout=2)"]
|
|
||||||
interval: 5s
|
|
||||||
timeout: 3s
|
|
||||||
retries: 60
|
|
||||||
start_period: 10s
|
|
||||||
|
|
||||||
qwen-image-21-test-runner:
|
|
||||||
image: python:3.12-slim
|
|
||||||
profiles: [qwen-image-test]
|
|
||||||
read_only: true
|
|
||||||
tmpfs: ["/tmp:size=32m,mode=1777"]
|
|
||||||
volumes:
|
|
||||||
- ./scripts/qwen-image-21-test.py:/opt/test/run.py:ro
|
|
||||||
- "${QWEN_IMAGE_21_OUTPUT_DIR:-/data/qwen-image-2.1-test-output}:/output"
|
|
||||||
environment:
|
|
||||||
CONTROLLER_URL: http://profile-controller:8090
|
|
||||||
CONTROLLER_TOKEN: "${CONTROLLER_TOKEN:?CONTROLLER_TOKEN is required}"
|
|
||||||
COMFY_URL: http://qwen-image-21-test:8188
|
|
||||||
OUTPUT_DIR: /output
|
|
||||||
command: [python, /opt/test/run.py]
|
|
||||||
networks: [control, inference]
|
|
||||||
security_opt: ["no-new-privileges:true"]
|
|
||||||
cap_drop: [ALL]
|
|
||||||
|
|
||||||
qwen3-tts:
|
qwen3-tts:
|
||||||
image: ${QWEN3_TTS_IMAGE:-ghcr.io/malaiwah/qwen3-tts-server:latest@sha256:b363a01d08b1bbecbfc3ca6f585368fae2cfdc591f9ecca6643738369f9a9d98}
|
image: ${QWEN3_TTS_IMAGE:-ghcr.io/malaiwah/qwen3-tts-server:latest@sha256:b363a01d08b1bbecbfc3ca6f585368fae2cfdc591f9ecca6643738369f9a9d98}
|
||||||
container_name: mike-ai-qwen3-tts
|
container_name: mike-ai-qwen3-tts
|
||||||
|
|||||||
+1
-1
@@ -364,7 +364,7 @@ png = open('/tmp/test_dl.png', 'rb').read()
|
|||||||
print(json.dumps({
|
print(json.dumps({
|
||||||
'prompt': 'Behalte die Person bei und ändere nur den Hintergrund',
|
'prompt': 'Behalte die Person bei und ändere nur den Hintergrund',
|
||||||
'size': '1024x1024',
|
'size': '1024x1024',
|
||||||
'steps': 4,
|
'steps': 25,
|
||||||
'guidance': 1.0,
|
'guidance': 1.0,
|
||||||
'response_format': 'b64_json',
|
'response_format': 'b64_json',
|
||||||
'image_b64': base64.b64encode(png).decode(),
|
'image_b64': base64.b64encode(png).decode(),
|
||||||
|
|||||||
@@ -23,7 +23,7 @@ def item(profile, state="exited"):
|
|||||||
|
|
||||||
|
|
||||||
def image_item(state="exited"):
|
def image_item(state="exited"):
|
||||||
return {"Id": "id-flux", "State": state,
|
return {"Id": "id-qwen-image", "State": state,
|
||||||
"Labels": {controller.IMAGE_LABEL_KEY: controller.IMAGE_WORKER}}
|
"Labels": {controller.IMAGE_LABEL_KEY: controller.IMAGE_WORKER}}
|
||||||
|
|
||||||
|
|
||||||
@@ -32,9 +32,9 @@ def restore_item(state="exited"):
|
|||||||
"Labels": {controller.IMAGE_LABEL_KEY: controller.RESTORE_WORKER}}
|
"Labels": {controller.IMAGE_LABEL_KEY: controller.RESTORE_WORKER}}
|
||||||
|
|
||||||
|
|
||||||
def qwen_image_test_item(state="exited"):
|
def flux_standby_item(state="exited"):
|
||||||
return {"Id": "id-qwen-image-test", "State": state,
|
return {"Id": "id-flux-standby", "State": state,
|
||||||
"Labels": {controller.IMAGE_LABEL_KEY: controller.QWEN_IMAGE_TEST_WORKER}}
|
"Labels": {controller.IMAGE_LABEL_KEY: controller.FLUX_STANDBY_WORKER}}
|
||||||
|
|
||||||
|
|
||||||
def tts_item(state="running"):
|
def tts_item(state="running"):
|
||||||
@@ -97,7 +97,7 @@ class ProfileControllerTests(unittest.TestCase):
|
|||||||
self.assertEqual(result, {"music_worker": "acestep", "state": "running"})
|
self.assertEqual(result, {"music_worker": "acestep", "state": "running"})
|
||||||
self.assertEqual(calls, [
|
self.assertEqual(calls, [
|
||||||
("POST", "/containers/id-ultra/stop?t=120"),
|
("POST", "/containers/id-ultra/stop?t=120"),
|
||||||
("POST", "/containers/id-flux/stop?t=20"),
|
("POST", "/containers/id-qwen-image/stop?t=20"),
|
||||||
("POST", "/containers/id-tts/stop?t=30"),
|
("POST", "/containers/id-tts/stop?t=30"),
|
||||||
("POST", "/containers/id-music/start"),
|
("POST", "/containers/id-music/start"),
|
||||||
])
|
])
|
||||||
@@ -158,10 +158,10 @@ class ProfileControllerTests(unittest.TestCase):
|
|||||||
self.assertEqual(calls, [
|
self.assertEqual(calls, [
|
||||||
("POST", "/containers/id-medium/stop?t=120"),
|
("POST", "/containers/id-medium/stop?t=120"),
|
||||||
("POST", "/containers/id-tts/stop?t=30"),
|
("POST", "/containers/id-tts/stop?t=30"),
|
||||||
("POST", "/containers/id-flux/start"),
|
("POST", "/containers/id-qwen-image/start"),
|
||||||
])
|
])
|
||||||
|
|
||||||
def test_restore_start_stops_flux_and_starts_restore(self):
|
def test_restore_start_stops_qwen_image_and_starts_restore(self):
|
||||||
profiles = {name: item(name) for name in controller.ALLOWED}
|
profiles = {name: item(name) for name in controller.ALLOWED}
|
||||||
calls = []
|
calls = []
|
||||||
|
|
||||||
@@ -178,11 +178,11 @@ class ProfileControllerTests(unittest.TestCase):
|
|||||||
controller.set_image_worker(True, controller.RESTORE_WORKER)
|
controller.set_image_worker(True, controller.RESTORE_WORKER)
|
||||||
self.assertEqual(calls, [
|
self.assertEqual(calls, [
|
||||||
("POST", "/containers/id-tts/stop?t=30"),
|
("POST", "/containers/id-tts/stop?t=30"),
|
||||||
("POST", "/containers/id-flux/stop?t=20"),
|
("POST", "/containers/id-qwen-image/stop?t=20"),
|
||||||
("POST", "/containers/id-restore/start"),
|
("POST", "/containers/id-restore/start"),
|
||||||
])
|
])
|
||||||
|
|
||||||
def test_qwen_image_test_is_allowlisted_and_exclusive(self):
|
def test_flux_standby_is_allowlisted_and_exclusive(self):
|
||||||
profiles = {name: item(name) for name in controller.ALLOWED}
|
profiles = {name: item(name) for name in controller.ALLOWED}
|
||||||
profiles["medium"] = item("medium", "running")
|
profiles["medium"] = item("medium", "running")
|
||||||
calls = []
|
calls = []
|
||||||
@@ -192,21 +192,21 @@ class ProfileControllerTests(unittest.TestCase):
|
|||||||
return 204, b""
|
return 204, b""
|
||||||
|
|
||||||
with patch.object(controller, "containers", return_value=profiles), \
|
with patch.object(controller, "containers", return_value=profiles), \
|
||||||
patch.object(controller, "image_container", return_value=qwen_image_test_item()), \
|
patch.object(controller, "image_container", return_value=flux_standby_item()), \
|
||||||
patch.object(controller, "image_containers", return_value=[
|
patch.object(controller, "image_containers", return_value=[
|
||||||
image_item("running"), restore_item(), qwen_image_test_item()]), \
|
image_item("running"), restore_item(), flux_standby_item()]), \
|
||||||
patch.object(controller, "tts_container", return_value=tts_item()), \
|
patch.object(controller, "tts_container", return_value=tts_item()), \
|
||||||
patch.object(controller, "docker_request", side_effect=request):
|
patch.object(controller, "docker_request", side_effect=request):
|
||||||
result = controller.set_image_worker(
|
result = controller.set_image_worker(
|
||||||
True, controller.QWEN_IMAGE_TEST_WORKER)
|
True, controller.FLUX_STANDBY_WORKER)
|
||||||
|
|
||||||
self.assertEqual(result, {
|
self.assertEqual(result, {
|
||||||
"image_worker": "qwen-image-2.1-test", "state": "running"})
|
"image_worker": "flux-standby", "state": "running"})
|
||||||
self.assertEqual(calls, [
|
self.assertEqual(calls, [
|
||||||
("POST", "/containers/id-medium/stop?t=120"),
|
("POST", "/containers/id-medium/stop?t=120"),
|
||||||
("POST", "/containers/id-tts/stop?t=30"),
|
("POST", "/containers/id-tts/stop?t=30"),
|
||||||
("POST", "/containers/id-flux/stop?t=20"),
|
("POST", "/containers/id-qwen-image/stop?t=20"),
|
||||||
("POST", "/containers/id-qwen-image-test/start"),
|
("POST", "/containers/id-flux-standby/start"),
|
||||||
])
|
])
|
||||||
|
|
||||||
def test_profile_activation_stops_image_worker_first(self):
|
def test_profile_activation_stops_image_worker_first(self):
|
||||||
@@ -224,7 +224,7 @@ class ProfileControllerTests(unittest.TestCase):
|
|||||||
patch.object(controller, "docker_request", side_effect=request):
|
patch.object(controller, "docker_request", side_effect=request):
|
||||||
controller.activate("fast")
|
controller.activate("fast")
|
||||||
self.assertEqual(calls, [
|
self.assertEqual(calls, [
|
||||||
("POST", "/containers/id-flux/stop?t=120"),
|
("POST", "/containers/id-qwen-image/stop?t=120"),
|
||||||
("POST", "/containers/id-fast/start"),
|
("POST", "/containers/id-fast/start"),
|
||||||
])
|
])
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,47 @@
|
|||||||
|
import importlib.util
|
||||||
|
import os
|
||||||
|
import pathlib
|
||||||
|
import unittest
|
||||||
|
|
||||||
|
|
||||||
|
os.environ.setdefault("WORKER_TOKEN", "x" * 48)
|
||||||
|
PATH = (pathlib.Path(__file__).parents[1]
|
||||||
|
/ "platform/docker/qwen-image-worker/qwen_image_worker.py")
|
||||||
|
SPEC = importlib.util.spec_from_file_location("qwen_image_worker", PATH)
|
||||||
|
worker = importlib.util.module_from_spec(SPEC)
|
||||||
|
assert SPEC and SPEC.loader
|
||||||
|
SPEC.loader.exec_module(worker)
|
||||||
|
|
||||||
|
|
||||||
|
class QwenImageWorkerTests(unittest.TestCase):
|
||||||
|
def test_text_to_image_uses_empty_latent(self):
|
||||||
|
workflow = worker.make_workflow("test", 7, 25, 1024, 1024, [], "job")
|
||||||
|
self.assertEqual(workflow["6"]["inputs"]["latent_image"], ["5", 0])
|
||||||
|
self.assertIn("5", workflow)
|
||||||
|
self.assertNotIn("vae", workflow["4"]["inputs"])
|
||||||
|
|
||||||
|
def test_edit_uses_reference_latent_and_qwen_dynamic_inputs(self):
|
||||||
|
workflow = worker.make_workflow(
|
||||||
|
"edit <image1>", 7, 25, 1024, 1024,
|
||||||
|
["job-ref-1.png", "job-ref-2.jpg"], "job")
|
||||||
|
encode = workflow["4"]["inputs"]
|
||||||
|
self.assertEqual(workflow["6"]["inputs"]["latent_image"], ["4", 2])
|
||||||
|
self.assertEqual(encode["vae"], ["3", 0])
|
||||||
|
self.assertEqual(encode["image_1"], ["9", 0])
|
||||||
|
self.assertEqual(encode["image_2"], ["10", 0])
|
||||||
|
self.assertNotIn("5", workflow)
|
||||||
|
|
||||||
|
def test_router_contract_is_fixed_to_production_parameters(self):
|
||||||
|
request = {
|
||||||
|
"prompt": "test", "filename": "result.png", "width": 1024,
|
||||||
|
"height": 1024, "steps": 25, "guidance": 1.0, "seed": 42,
|
||||||
|
}
|
||||||
|
self.assertEqual(worker.validate_request(request)[1:],
|
||||||
|
("result.png", 1024, 1024, 25, 1.0, 42))
|
||||||
|
request["steps"] = 4
|
||||||
|
with self.assertRaisesRegex(ValueError, "steps=25"):
|
||||||
|
worker.validate_request(request)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
unittest.main()
|
||||||
@@ -12,7 +12,7 @@ flowchart LR
|
|||||||
H -->|OpenAI API| R[Profile Router<br/>Athena :8081]
|
H -->|OpenAI API| R[Profile Router<br/>Athena :8081]
|
||||||
R --> P[Profile Controller]
|
R --> P[Profile Controller]
|
||||||
P --> Q[genau ein llama.cpp-Profil<br/>Qwen Fast / Medium / Large / Ultra / Uncensored]
|
P --> Q[genau ein llama.cpp-Profil<br/>Qwen Fast / Medium / Large / Ultra / Uncensored]
|
||||||
R --> I[FLUX.2 Klein 9B FP8 Beta<br/>RTX 5080, Text + Editing]
|
R --> I[Qwen-Image-2.1 INT8<br/>RTX 5080, Text + Editing]
|
||||||
R --> T[Qwen3-TTS 1.7B RTX 3060<br/>TTS-Gateway]
|
R --> T[Qwen3-TTS 1.7B RTX 3060<br/>TTS-Gateway]
|
||||||
R --> STT[Whisper.cpp small<br/>CPU, lokale Spracherkennung]
|
R --> STT[Whisper.cpp small<br/>CPU, lokale Spracherkennung]
|
||||||
|
|
||||||
@@ -32,8 +32,9 @@ flowchart LR
|
|||||||
K[Backup alle 5 Stunden] --> DATA[/data und /etc/mike-ai]
|
K[Backup alle 5 Stunden] --> DATA[/data und /etc/mike-ai]
|
||||||
```
|
```
|
||||||
|
|
||||||
Stand: 20. September 2026. Versionen und Abweichungen: [Live-Stand](LIVE_STATE.md).
|
Stand: 21. September 2026. Versionen und Abweichungen: [Live-Stand](LIVE_STATE.md).
|
||||||
FLUX nutzt beide GPUs und pausiert dafür LLM und Qwen3-TTS. Dashboard und
|
Qwen Image nutzt die RTX 5080 und pausiert dafür LLM und Qwen3-TTS. FLUX
|
||||||
|
bleibt als gestoppter Rückfallworker vorhanden. Dashboard und
|
||||||
Portainer besitzen eigene Netzwerk-Namespaces in `mike-ai_frontend`; der Router
|
Portainer besitzen eigene Netzwerk-Namespaces in `mike-ai_frontend`; der Router
|
||||||
läuft in `mike-ai_control`. Der Gateway vermittelt den privaten Zugriff.
|
läuft in `mike-ai_control`. Der Gateway vermittelt den privaten Zugriff.
|
||||||
Zusätzliche Musik-, Audio-, Voice- und 3D-Container stehen im
|
Zusätzliche Musik-, Audio-, Voice- und 3D-Container stehen im
|
||||||
|
|||||||
@@ -14,7 +14,8 @@ nicht automatisch ein ungenutzter Rest.
|
|||||||
| `mike-ai-embedding` | EmbeddingGemma 300M QAT Q8 | Dauerhafter CPU-only-Endpunkt für OpenClaws semantische und hybride Memory-Suche; teilt ausschließlich den privaten Netzwerk-Namespace des WireGuard-Gateways. |
|
| `mike-ai-embedding` | EmbeddingGemma 300M QAT Q8 | Dauerhafter CPU-only-Endpunkt für OpenClaws semantische und hybride Memory-Suche; teilt ausschließlich den privaten Netzwerk-Namespace des WireGuard-Gateways. |
|
||||||
| `mike-ai-bonsai2-ab` | Bonsai-2-Vergleichsmodell | Gestoppter, reproduzierbar dokumentierter A/B-Testcontainer; kein Produktivprofil. |
|
| `mike-ai-bonsai2-ab` | Bonsai-2-Vergleichsmodell | Gestoppter, reproduzierbar dokumentierter A/B-Testcontainer; kein Produktivprofil. |
|
||||||
| `mike-ai-applio-studio` | Applio/RVC; Stimmenmodelle werden nutzerseitig ergänzt | Vollständige RVC-Oberfläche für Inferenz, Modellverwaltung und Training auf der RTX 5080. Für eine Konvertierung ist ein importiertes oder trainiertes `.pth`-Modell nötig; eine Referenzaufnahme allein reicht nicht. |
|
| `mike-ai-applio-studio` | Applio/RVC; Stimmenmodelle werden nutzerseitig ergänzt | Vollständige RVC-Oberfläche für Inferenz, Modellverwaltung und Training auf der RTX 5080. Für eine Konvertierung ist ein importiertes oder trainiertes `.pth`-Modell nötig; eine Referenzaufnahme allein reicht nicht. |
|
||||||
| `mike-ai-image-worker` | FLUX.2 Klein 9B FP8, Qwen3-8B NF4 Textencoder und VAE | Erzeugt und bearbeitet Bilder transaktional; nutzt während eines Auftrags RTX 5080 und RTX 3060. |
|
| `mike-ai-image-worker` | Qwen-Image-2.1 INT8, Qwen3-VL-8B INT8 und VAE | Produktiver Bildworker; erzeugt und bearbeitet Bilder transaktional auf der RTX 5080. |
|
||||||
|
| `mike-ai-flux-image-worker` | FLUX.2 Klein 9B FP8, Qwen3-8B NF4 Textencoder und VAE | Gestoppter Rückfallworker; wird vom normalen Router-Bildpfad nicht gestartet. |
|
||||||
| `mike-ai-llama-dashboard` | kein Modell | Zeigt Telemetrie, Profile, GPU-Nutzung, Slot-Kontextbelegung mit Verlauf und Betriebsarten an und bietet die Modusumschaltung. |
|
| `mike-ai-llama-dashboard` | kein Modell | Zeigt Telemetrie, Profile, GPU-Nutzung, Slot-Kontextbelegung mit Verlauf und Betriebsarten an und bietet die Modusumschaltung. |
|
||||||
| `mike-ai-llama-fast` | Qwen3.8-27B `IQ4-MIX`, Qwen-MMProj BF16 | Schnelles Q4-Text-/Vision-Profil mit 76.800 Token Kontext. |
|
| `mike-ai-llama-fast` | Qwen3.8-27B `IQ4-MIX`, Qwen-MMProj BF16 | Schnelles Q4-Text-/Vision-Profil mit 76.800 Token Kontext. |
|
||||||
| `mike-ai-llama-large` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Q4-Text-/Vision-Profil mit 192.000 Token Kontext und Verteilung auf beide GPUs. |
|
| `mike-ai-llama-large` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Q4-Text-/Vision-Profil mit 192.000 Token Kontext und Verteilung auf beide GPUs. |
|
||||||
@@ -55,4 +56,4 @@ sämtliche Modellgewichte oder Funktionen neu.
|
|||||||
|
|
||||||
Gestoppte Container sind nicht automatisch Testreste. Vor einer Entfernung
|
Gestoppte Container sind nicht automatisch Testreste. Vor einer Entfernung
|
||||||
Compose-Zuordnung, Mounts und Controller-Verweise prüfen. Die vorhandenen
|
Compose-Zuordnung, Mounts und Controller-Verweise prüfen. Die vorhandenen
|
||||||
Spezialcontainer wurden durch den FLUX-Auflösungstest weder angelegt noch entfernt.
|
Die historischen FLUX-Tests sind keine Aussage über den produktiven Qwen-Pfad.
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# Aktuelle Laufzeitnotizen
|
# Aktuelle Laufzeitnotizen
|
||||||
|
|
||||||
Stand: 20. September 2026.
|
Stand: 21. September 2026.
|
||||||
|
|
||||||
Der verbindliche Abgleich steht im [geprüften Live-Stand](LIVE_STATE.md).
|
Der verbindliche Abgleich steht im [geprüften Live-Stand](LIVE_STATE.md).
|
||||||
Produktiv läuft **llama.cpp 0.4.1** auf Commit `b29c606`, zuvor b10930.
|
Produktiv läuft **llama.cpp 0.4.1** auf Commit `b29c606`, zuvor b10930.
|
||||||
@@ -18,9 +18,10 @@ Funktionsproben, Vergleichsmessungen und gesicherter Rückfallstand.
|
|||||||
[Update-Audit vom 15. September](UPDATE_AUDIT_20260915.md): aktualisierte
|
[Update-Audit vom 15. September](UPDATE_AUDIT_20260915.md): aktualisierte
|
||||||
Komponenten, unveränderte aktuelle Komponenten, Aufräumarbeiten und Rollback.
|
Komponenten, unveränderte aktuelle Komponenten, Aufräumarbeiten und Rollback.
|
||||||
|
|
||||||
FLUX verwendet **9B FP8 Beta** mit Textencoder auf der RTX 3060.
|
Der produktive Bildpfad verwendet **Qwen-Image-2.1 INT8** mit ComfyUI auf der
|
||||||
[Auflösungstest](FLUX_RESOLUTION_TEST_20260912.md): 1024 erfolgreich,
|
RTX 5080. FLUX 9B FP8 und seine Gewichte bleiben als gestoppter Rückfallpfad.
|
||||||
1280 CUDA-OOM; produktive Begrenzung weiterhin 1024.
|
[Qwen-Betriebsdoku](QWEN_IMAGE_21.md); der historische
|
||||||
|
[FLUX-Auflösungstest](FLUX_RESOLUTION_TEST_20260912.md) gilt nur für FLUX.
|
||||||
Sprachausgabe: Qwen3-TTS 1.7B ohne Piper. Spracherkennung: Whisper.cpp 1.9.4
|
Sprachausgabe: Qwen3-TTS 1.7B ohne Piper. Spracherkennung: Whisper.cpp 1.9.4
|
||||||
mit dem Modell `small` auf der CPU. Applio basiert auf dem stabilen Release 3.6.4.
|
mit dem Modell `small` auf der CPU. Applio basiert auf dem stabilen Release 3.6.4.
|
||||||
|
|
||||||
|
|||||||
+4
-4
@@ -39,10 +39,10 @@ nicht mehr den aktuellen Containerzustand.
|
|||||||
|
|
||||||
## Weitere Dienste
|
## Weitere Dienste
|
||||||
|
|
||||||
- FLUX.2 Klein 9B FP8 Beta: Transformer/VAE auf 5080, Qwen3-8B-NF4-
|
- Qwen-Image-2.1 INT8 läuft produktiv über gepinntes ComfyUI mit Low-VRAM auf
|
||||||
Textencoder auf 3060. Produktiv freigegeben sind höchstens 1024 × 1024,
|
der RTX 5080. Der Router verwendet 25 Schritte, Guidance 1 und bis zu vier
|
||||||
vier Schritte, Guidance 1, vier Referenzbilder und 128 Encoder-Tokens.
|
Referenzbilder. FLUX.2 Klein 9B FP8 bleibt mit seinen Gewichten als
|
||||||
1280 × 1280 lief im Test in CUDA-OOM.
|
gestoppter, explizit allowlist-beschränkter Rückfallcontainer erhalten.
|
||||||
- Qwen3-TTS 1.7B auf 3060 hinter dem TTS-Gateway; kein Piper-Fallback.
|
- Qwen3-TTS 1.7B auf 3060 hinter dem TTS-Gateway; kein Piper-Fallback.
|
||||||
- Whisper.cpp `ggml-small` auf CPU über `/v1/audio/transcriptions`.
|
- Whisper.cpp `ggml-small` auf CPU über `/v1/audio/transcriptions`.
|
||||||
- `mike-ai-embedding` stellt EmbeddingGemma 300M Q8 CPU-only über den privaten
|
- `mike-ai-embedding` stellt EmbeddingGemma 300M Q8 CPU-only über den privaten
|
||||||
|
|||||||
@@ -0,0 +1,83 @@
|
|||||||
|
# Qwen-Image-2.1 auf Athena
|
||||||
|
|
||||||
|
Stand: 21. September 2026. Qwen-Image-2.1 ist der produktive Bildworker des
|
||||||
|
Routers. FLUX.2 Klein 9B FP8 bleibt als gestoppter Rückfallcontainer erhalten.
|
||||||
|
|
||||||
|
## Gepinnter Stand
|
||||||
|
|
||||||
|
- ComfyUI: Commit `b0f4b7b294ce482a2e071d9d762c133d38c7aa07`
|
||||||
|
- Modell-Repository: `Comfy-Org/Qwen-Image-2.1`, Revision
|
||||||
|
`ace0edeb3791a594ddfa36ed5f41a178a394e921`
|
||||||
|
- Transformer: `qwen_image_2.1_int8_convrot.safetensors`
|
||||||
|
(`cb74113cb03faecd79611b01fd7fd642f0aa60d6f0b95086abee214d75eaa57d`)
|
||||||
|
- Textencoder: `qwen3vl_8b_int8_convrot.safetensors`
|
||||||
|
(`8bfd0f6e12abf2d2d697ecc888e5e90b0d6741d6708f05799f53afa560452e8f`)
|
||||||
|
- VAE: `qwen_image_2.1_vae_bf16.safetensors`
|
||||||
|
(`bb21f7473051e1ac368515dd3f2e15cd44d7a11748ee8823e1ddca3e4876b7c9`)
|
||||||
|
- Pipeline: 25 Schritte, CFG 1, Euler/Simple, `--lowvram`
|
||||||
|
- GPU: ausschließlich Athena-GPU 1, die RTX 5080 mit 16 GB
|
||||||
|
|
||||||
|
Die Gewichte liegen außerhalb von Git unter
|
||||||
|
`/data/models/Qwen-Image-2.1-ComfyUI`. Der Adapter
|
||||||
|
`platform/docker/qwen-image-worker/qwen_image_worker.py` stellt vor ComfyUI
|
||||||
|
die bestehende private Worker-API auf Port 8086 bereit. Der öffentliche
|
||||||
|
Routervertrag bleibt damit `/v1/images/generations` und `/v1/images/edits`.
|
||||||
|
|
||||||
|
## Lebenszyklus
|
||||||
|
|
||||||
|
Ein Bildauftrag läuft transaktional:
|
||||||
|
|
||||||
|
1. Der Router merkt sich das aktive Textprofil.
|
||||||
|
2. Der Profile Controller stoppt LLM, TTS und andere GPU-Spezialworker.
|
||||||
|
3. `mike-ai-image-worker` startet Qwen Image auf der RTX 5080.
|
||||||
|
4. Der Adapter erzeugt das Bild oder bearbeitet bis zu vier Referenzbilder.
|
||||||
|
5. Der Qwen-Worker wird vollständig gestoppt.
|
||||||
|
6. Das vorherige Textprofil und Qwen3-TTS werden wiederhergestellt.
|
||||||
|
|
||||||
|
Der Container ist außerhalb eines Auftrags gestoppt. Das ist der Sollzustand.
|
||||||
|
|
||||||
|
## Reproduzierbare Vorbereitung
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd /opt/mike-ai/stack
|
||||||
|
sudo ./scripts/prepare-qwen-image-21.sh
|
||||||
|
```
|
||||||
|
|
||||||
|
Das Skript lädt fehlende Dateien mit festen SHA-256-Prüfsummen, baut den
|
||||||
|
produktiven Qwen-Worker und legt Qwen sowie FLUX gestoppt an. Es lädt dabei
|
||||||
|
kein Modell in den VRAM.
|
||||||
|
|
||||||
|
## Funktionsprobe über den echten Routerweg
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -sS http://127.0.0.1:8081/v1/images/generations \
|
||||||
|
-H 'Content-Type: application/json' \
|
||||||
|
-d '{
|
||||||
|
"model":"Qwen-Image-2.1-int8",
|
||||||
|
"prompt":"Fotorealistisches rotes Haus bei Tageslicht",
|
||||||
|
"size":"1024x1024",
|
||||||
|
"steps":25,
|
||||||
|
"guidance":1.0,
|
||||||
|
"response_format":"url"
|
||||||
|
}'
|
||||||
|
```
|
||||||
|
|
||||||
|
Nach dem Lauf müssen `mike-ai-image-worker` und
|
||||||
|
`mike-ai-flux-image-worker` gestoppt sein und das vorherige LLM-Profil wieder
|
||||||
|
gesund laufen.
|
||||||
|
|
||||||
|
## FLUX-Rückfallpfad
|
||||||
|
|
||||||
|
FLUX verwendet weiterhin das Image `mike-ai/image-worker:local`, die
|
||||||
|
bestehenden Modellverzeichnisse und den Container
|
||||||
|
`mike-ai-flux-image-worker`. Er hat das Controller-Label `flux-standby` und
|
||||||
|
kann deshalb nicht durch einen normalen Router-Bildauftrag gestartet werden.
|
||||||
|
Eine manuelle Diagnose ist nur über den allowlist-beschränkten Controllerpfad
|
||||||
|
`/workers/flux-standby/start` möglich. Für eine dauerhafte Rückschaltung müssen
|
||||||
|
Service-DNS, Modellname und Schrittzahl gemeinsam in Compose auf FLUX gesetzt
|
||||||
|
und anschließend Router und Controller neu ausgerollt werden. Keine dieser
|
||||||
|
Änderungen erfolgt automatisch.
|
||||||
|
|
||||||
|
Historische FLUX-Messwerte und Grenzen stehen in
|
||||||
|
[FLUX_9B_BETA.md](FLUX_9B_BETA.md) und
|
||||||
|
[FLUX_RESOLUTION_TEST_20260912.md](FLUX_RESOLUTION_TEST_20260912.md).
|
||||||
@@ -1,60 +0,0 @@
|
|||||||
# Qwen-Image-2.1 – isolierter Test
|
|
||||||
|
|
||||||
Stand: 21. September 2026. Dieser Pfad dient nur dem Vergleich mit dem
|
|
||||||
produktiven FLUX.2-Worker. Er ersetzt FLUX nicht und ist nicht an den Router
|
|
||||||
angebunden.
|
|
||||||
|
|
||||||
## Reproduzierbarer Stand
|
|
||||||
|
|
||||||
- ComfyUI: Commit `b0f4b7b294ce482a2e071d9d762c133d38c7aa07`
|
|
||||||
- Modell-Repository: `Comfy-Org/Qwen-Image-2.1`, Revision
|
|
||||||
`ace0edeb3791a594ddfa36ed5f41a178a394e921`
|
|
||||||
- Transformer: `qwen_image_2.1_int8_convrot.safetensors` (7.256.783.064 Byte)
|
|
||||||
- Textencoder: `qwen3vl_8b_int8_convrot.safetensors` (9.350.798.360 Byte)
|
|
||||||
- VAE: `qwen_image_2.1_vae_bf16.safetensors` (675.509.688 Byte)
|
|
||||||
- erster Vergleich: 1024 × 1024, 25 Schritte, CFG 1, Euler/Simple
|
|
||||||
|
|
||||||
Die Gewichte sind die von Qwen verlinkte offizielle ComfyUI-Aufbereitung. Der
|
|
||||||
Worker verwendet `--lowvram` auf der RTX 5080. So bleibt der Test mit 16 GB
|
|
||||||
VRAM möglich; der Preis ist CPU-Offload und damit eine längere Laufzeit.
|
|
||||||
|
|
||||||
## Vorbereitung ohne GPU-Last
|
|
||||||
|
|
||||||
```bash
|
|
||||||
cd /opt/mike-ai/stack
|
|
||||||
sudo ./scripts/prepare-qwen-image-21-test.sh
|
|
||||||
```
|
|
||||||
|
|
||||||
Das Skript prüft feste SHA-256-Summen, baut das Image, aktualisiert den
|
|
||||||
Profile Controller und erstellt den Testcontainer **gestoppt**. Es lädt kein
|
|
||||||
Modell in die GPU.
|
|
||||||
|
|
||||||
## Einmaliger Test nach ausdrücklichem „Go“
|
|
||||||
|
|
||||||
```bash
|
|
||||||
cd /opt/mike-ai/stack
|
|
||||||
docker compose --env-file /etc/mike-ai/stack.env \
|
|
||||||
--profile qwen-image-test run --rm qwen-image-21-test-runner \
|
|
||||||
python /opt/test/run.py 'HIER DEN VEREINBARTEN PROMPT EINSETZEN'
|
|
||||||
```
|
|
||||||
|
|
||||||
Der Runner merkt sich das aktive LLM-Profil, startet den Qwen-Worker über den
|
|
||||||
allowlist-beschränkten Controller, erzeugt genau ein Bild und stellt danach
|
|
||||||
auch bei einem Fehler das vorherige Profil wieder her. Das Ergebnis liegt in
|
|
||||||
`/data/qwen-image-2.1-test-output/`. FLUX-Konfiguration und FLUX-Gewichte
|
|
||||||
werden nicht verändert.
|
|
||||||
|
|
||||||
## Vollständig entfernen
|
|
||||||
|
|
||||||
```bash
|
|
||||||
cd /opt/mike-ai/stack
|
|
||||||
docker compose --env-file /etc/mike-ai/stack.env \
|
|
||||||
--profile qwen-image-test rm -sf qwen-image-21-test
|
|
||||||
docker image rm mike-ai/qwen-image-2.1-test:local
|
|
||||||
rm -rf /data/models/Qwen-Image-2.1-ComfyUI \
|
|
||||||
/data/qwen-image-2.1-test-output
|
|
||||||
```
|
|
||||||
|
|
||||||
Anschließend kann die Controller-Erweiterung bei Bedarf aus Git zurückgenommen
|
|
||||||
und nur `profile-controller` neu gebaut werden. Die Entfernung ist nicht Teil
|
|
||||||
der Vorbereitung und wird niemals automatisch ausgeführt.
|
|
||||||
@@ -53,7 +53,8 @@ Titelgenerierung und Kontextkompression in Hermes.
|
|||||||
| Datum | Modell | Ergebnis | Status / Entscheidung | Beleg |
|
| Datum | Modell | Ergebnis | Status / Entscheidung | Beleg |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| bis 07.09.2026 | FLUX.2 Klein 4B | funktional, aber schwächere räumliche und motivische Konsistenz | **ersetzt** durch 9B FP8 | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
| bis 07.09.2026 | FLUX.2 Klein 4B | funktional, aber schwächere räumliche und motivische Konsistenz | **ersetzt** durch 9B FP8 | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
||||||
| 07.09.2026 | FLUX.2 Klein 9B FP8 | bessere Prompttreue; produktiver Zwei-GPU-Pfad, derzeit auf 1024 × 1024 begrenzt | **produktiv als Beta** | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
| 07.09.2026 | FLUX.2 Klein 9B FP8 | bessere Prompttreue als 4B; produktiver Zwei-GPU-Pfad, auf 1024 × 1024 begrenzt | **Standby seit 21.09.2026** | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
||||||
|
| 21.09.2026 | Qwen-Image-2.1 INT8 | im direkten Vergleich fotorealistischer, bessere Textdarstellung und Prompttreue; 1024 × 1024 bei 25 Schritten in rund 49 s Pipeline-Lauf | **produktiv** | [QWEN_IMAGE_21.md](QWEN_IMAGE_21.md) |
|
||||||
| 08.09.2026 | HYPIR-SD2 | glättete oder erfand Details und veränderte kleine Strukturen | **verworfen** | [IMAGE_RESTORATION.md](IMAGE_RESTORATION.md) |
|
| 08.09.2026 | HYPIR-SD2 | glättete oder erfand Details und veränderte kleine Strukturen | **verworfen** | [IMAGE_RESTORATION.md](IMAGE_RESTORATION.md) |
|
||||||
| 08.09.2026 | SeedVR2 7B FP8 | bewahrte Identität besser als HYPIR, brachte beim realen unscharfen Foto aber kaum nutzbare Details zurück | **verworfen** | [IMAGE_RESTORATION.md](IMAGE_RESTORATION.md) |
|
| 08.09.2026 | SeedVR2 7B FP8 | bewahrte Identität besser als HYPIR, brachte beim realen unscharfen Foto aber kaum nutzbare Details zurück | **verworfen** | [IMAGE_RESTORATION.md](IMAGE_RESTORATION.md) |
|
||||||
|
|
||||||
|
|||||||
@@ -27,8 +27,8 @@ LABEL_KEY = "com.mike-ai.llama-profile"
|
|||||||
IMAGE_LABEL_KEY = "com.mike-ai.image-worker"
|
IMAGE_LABEL_KEY = "com.mike-ai.image-worker"
|
||||||
IMAGE_WORKER = os.environ.get("IMAGE_WORKER", "image")
|
IMAGE_WORKER = os.environ.get("IMAGE_WORKER", "image")
|
||||||
RESTORE_WORKER = os.environ.get("RESTORE_WORKER", "restore")
|
RESTORE_WORKER = os.environ.get("RESTORE_WORKER", "restore")
|
||||||
QWEN_IMAGE_TEST_WORKER = os.environ.get(
|
FLUX_STANDBY_WORKER = os.environ.get(
|
||||||
"QWEN_IMAGE_TEST_WORKER", "qwen-image-2.1-test").strip()
|
"FLUX_STANDBY_WORKER", "flux-standby").strip()
|
||||||
TTS_LABEL_KEY = "com.mike-ai.tts-worker"
|
TTS_LABEL_KEY = "com.mike-ai.tts-worker"
|
||||||
TTS_WORKER = os.environ.get("TTS_WORKER", "qwen3")
|
TTS_WORKER = os.environ.get("TTS_WORKER", "qwen3")
|
||||||
MUSIC_LABEL_KEY = "com.mike-ai.music-worker"
|
MUSIC_LABEL_KEY = "com.mike-ai.music-worker"
|
||||||
@@ -103,8 +103,8 @@ def image_container(kind: str = IMAGE_WORKER) -> dict:
|
|||||||
def image_containers() -> list[dict]:
|
def image_containers() -> list[dict]:
|
||||||
"""All allowlisted GPU workers that must never overlap an LLM."""
|
"""All allowlisted GPU workers that must never overlap an LLM."""
|
||||||
allowed = {IMAGE_WORKER, RESTORE_WORKER}
|
allowed = {IMAGE_WORKER, RESTORE_WORKER}
|
||||||
if QWEN_IMAGE_TEST_WORKER:
|
if FLUX_STANDBY_WORKER:
|
||||||
allowed.add(QWEN_IMAGE_TEST_WORKER)
|
allowed.add(FLUX_STANDBY_WORKER)
|
||||||
return [item for item in labelled_containers(IMAGE_LABEL_KEY)
|
return [item for item in labelled_containers(IMAGE_LABEL_KEY)
|
||||||
if item.get("Labels", {}).get(IMAGE_LABEL_KEY) in allowed]
|
if item.get("Labels", {}).get(IMAGE_LABEL_KEY) in allowed]
|
||||||
|
|
||||||
@@ -363,7 +363,7 @@ def stop_inference() -> dict:
|
|||||||
|
|
||||||
|
|
||||||
def set_image_worker(running: bool, kind: str = IMAGE_WORKER) -> dict:
|
def set_image_worker(running: bool, kind: str = IMAGE_WORKER) -> dict:
|
||||||
if kind not in {IMAGE_WORKER, RESTORE_WORKER, QWEN_IMAGE_TEST_WORKER}:
|
if kind not in {IMAGE_WORKER, RESTORE_WORKER, FLUX_STANDBY_WORKER}:
|
||||||
raise ValueError("worker is not allowlisted")
|
raise ValueError("worker is not allowlisted")
|
||||||
with LOCK:
|
with LOCK:
|
||||||
item = image_container(kind)
|
item = image_container(kind)
|
||||||
@@ -371,8 +371,8 @@ def set_image_worker(running: bool, kind: str = IMAGE_WORKER) -> dict:
|
|||||||
# The image worker may never overlap a llama profile on the 5080.
|
# The image worker may never overlap a llama profile on the 5080.
|
||||||
for profile_item in containers().values():
|
for profile_item in containers().values():
|
||||||
stop_container(profile_item)
|
stop_container(profile_item)
|
||||||
# The 9B beta text encoder temporarily borrows the RTX 3060 from
|
# Image workers are GPU-exclusive. The FLUX standby also borrows
|
||||||
# Qwen3-TTS. TTS is unavailable during this exclusive GPU phase.
|
# the RTX 3060, while Qwen Image runs only on the RTX 5080.
|
||||||
stop_container(tts_container(), timeout=30)
|
stop_container(tts_container(), timeout=30)
|
||||||
stop_music_if_configured()
|
stop_music_if_configured()
|
||||||
stop_yue2_if_configured()
|
stop_yue2_if_configured()
|
||||||
@@ -801,8 +801,8 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
"/workers/image/stop": (IMAGE_WORKER, False),
|
"/workers/image/stop": (IMAGE_WORKER, False),
|
||||||
"/workers/restore/start": (RESTORE_WORKER, True),
|
"/workers/restore/start": (RESTORE_WORKER, True),
|
||||||
"/workers/restore/stop": (RESTORE_WORKER, False),
|
"/workers/restore/stop": (RESTORE_WORKER, False),
|
||||||
"/workers/qwen-image-test/start": (QWEN_IMAGE_TEST_WORKER, True),
|
"/workers/flux-standby/start": (FLUX_STANDBY_WORKER, True),
|
||||||
"/workers/qwen-image-test/stop": (QWEN_IMAGE_TEST_WORKER, False),
|
"/workers/flux-standby/stop": (FLUX_STANDBY_WORKER, False),
|
||||||
}
|
}
|
||||||
if self.path in worker_paths:
|
if self.path in worker_paths:
|
||||||
try:
|
try:
|
||||||
|
|||||||
+3
-1
@@ -13,4 +13,6 @@ RUN test -n "$COMFYUI_COMMIT" \
|
|||||||
&& rm -rf /var/lib/apt/lists/* /root/.cache
|
&& rm -rf /var/lib/apt/lists/* /root/.cache
|
||||||
|
|
||||||
WORKDIR /opt/ComfyUI
|
WORKDIR /opt/ComfyUI
|
||||||
EXPOSE 8188
|
COPY qwen_image_worker.py /opt/qwen-image-worker.py
|
||||||
|
EXPOSE 8086
|
||||||
|
CMD ["python", "/opt/qwen-image-worker.py"]
|
||||||
@@ -0,0 +1,268 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Router-compatible Qwen-Image-2.1 worker backed by pinned ComfyUI."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
import random
|
||||||
|
import shutil
|
||||||
|
import signal
|
||||||
|
import subprocess
|
||||||
|
import threading
|
||||||
|
import time
|
||||||
|
import urllib.request
|
||||||
|
import uuid
|
||||||
|
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
|
||||||
|
HOST = os.environ.get("WORKER_HOST", "0.0.0.0")
|
||||||
|
PORT = int(os.environ.get("WORKER_PORT", "8086"))
|
||||||
|
TOKEN = os.environ.get("WORKER_TOKEN", "").strip()
|
||||||
|
COMFY = "http://127.0.0.1:8188"
|
||||||
|
COMFY_DIR = Path("/opt/ComfyUI")
|
||||||
|
INPUT_DIR = COMFY_DIR / "input"
|
||||||
|
OUTPUT_DIR = Path(os.environ.get("IMAGE_DIR", "/data/images")).resolve()
|
||||||
|
MODEL = "Qwen-Image-2.1-int8"
|
||||||
|
GENERATION_LOCK = threading.Lock()
|
||||||
|
READY = threading.Event()
|
||||||
|
ACTIVE = False
|
||||||
|
COMFY_PROCESS: subprocess.Popen | None = None
|
||||||
|
|
||||||
|
if len(TOKEN) < 32:
|
||||||
|
raise RuntimeError("WORKER_TOKEN is missing or too short")
|
||||||
|
|
||||||
|
|
||||||
|
def request_json(path: str, payload: dict | None = None,
|
||||||
|
timeout: float = 30) -> dict:
|
||||||
|
body = None if payload is None else json.dumps(payload).encode()
|
||||||
|
headers = {"Content-Type": "application/json"} if body is not None else {}
|
||||||
|
method = "POST" if body is not None else "GET"
|
||||||
|
request = urllib.request.Request(COMFY + path, data=body,
|
||||||
|
headers=headers, method=method)
|
||||||
|
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||||
|
return json.load(response)
|
||||||
|
|
||||||
|
|
||||||
|
def wait_for_comfy() -> None:
|
||||||
|
for _ in range(300):
|
||||||
|
if COMFY_PROCESS is not None and COMFY_PROCESS.poll() is not None:
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
request_json("/system_stats", timeout=2)
|
||||||
|
READY.set()
|
||||||
|
return
|
||||||
|
except Exception:
|
||||||
|
time.sleep(1)
|
||||||
|
|
||||||
|
|
||||||
|
def start_comfy() -> None:
|
||||||
|
global COMFY_PROCESS
|
||||||
|
INPUT_DIR.mkdir(parents=True, exist_ok=True)
|
||||||
|
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
|
||||||
|
command = [
|
||||||
|
"python", "main.py", "--listen", "127.0.0.1", "--port", "8188",
|
||||||
|
"--lowvram", "--preview-method", "none",
|
||||||
|
"--input-directory", str(INPUT_DIR),
|
||||||
|
"--output-directory", str(OUTPUT_DIR),
|
||||||
|
]
|
||||||
|
COMFY_PROCESS = subprocess.Popen(command, cwd=COMFY_DIR)
|
||||||
|
threading.Thread(target=wait_for_comfy, daemon=True).start()
|
||||||
|
|
||||||
|
|
||||||
|
def stop(*_: object) -> None:
|
||||||
|
if COMFY_PROCESS is not None and COMFY_PROCESS.poll() is None:
|
||||||
|
COMFY_PROCESS.terminate()
|
||||||
|
try:
|
||||||
|
COMFY_PROCESS.wait(timeout=15)
|
||||||
|
except subprocess.TimeoutExpired:
|
||||||
|
COMFY_PROCESS.kill()
|
||||||
|
raise SystemExit(0)
|
||||||
|
|
||||||
|
|
||||||
|
def validate_request(data: dict) -> tuple[str, str, int, int, int, float, int]:
|
||||||
|
prompt = data.get("prompt")
|
||||||
|
filename = data.get("filename")
|
||||||
|
if not isinstance(prompt, str) or not prompt.strip() or len(prompt) > 8000:
|
||||||
|
raise ValueError("invalid prompt")
|
||||||
|
if (not isinstance(filename, str) or Path(filename).name != filename
|
||||||
|
or not filename.endswith(".png")):
|
||||||
|
raise ValueError("invalid filename")
|
||||||
|
width = int(data.get("width", 1024))
|
||||||
|
height = int(data.get("height", 1024))
|
||||||
|
if width % 32 or height % 32 or not 512 <= width <= 2048 or not 512 <= height <= 2048:
|
||||||
|
raise ValueError("width and height must be multiples of 32 between 512 and 2048")
|
||||||
|
steps = int(data.get("steps", 25))
|
||||||
|
guidance = float(data.get("guidance", 1.0))
|
||||||
|
if steps != 25 or guidance != 1.0:
|
||||||
|
raise ValueError("Qwen-Image-2.1 requires steps=25 and guidance=1.0")
|
||||||
|
seed = data.get("seed")
|
||||||
|
seed = random.randrange(2**32) if seed is None else int(seed)
|
||||||
|
if not 0 <= seed <= 2**32 - 1:
|
||||||
|
raise ValueError("invalid seed")
|
||||||
|
return prompt.strip(), filename, width, height, steps, guidance, seed
|
||||||
|
|
||||||
|
|
||||||
|
def make_workflow(prompt: str, seed: int, steps: int, width: int, height: int,
|
||||||
|
references: list[str], prefix: str) -> dict:
|
||||||
|
encode_inputs: dict = {
|
||||||
|
"clip": ["2", 0], "prompt": prompt, "negative_prompt": "",
|
||||||
|
"resolution": max(width, height),
|
||||||
|
}
|
||||||
|
workflow: dict = {
|
||||||
|
"1": {"class_type": "UNETLoader", "inputs": {
|
||||||
|
"unet_name": "qwen_image_2.1_int8_convrot.safetensors",
|
||||||
|
"weight_dtype": "default"}},
|
||||||
|
"2": {"class_type": "CLIPLoader", "inputs": {
|
||||||
|
"clip_name": "qwen3vl_8b_int8_convrot.safetensors",
|
||||||
|
"type": "qwen_image", "device": "default"}},
|
||||||
|
"3": {"class_type": "VAELoader", "inputs": {
|
||||||
|
"vae_name": "qwen_image_2.1_vae_bf16.safetensors"}},
|
||||||
|
"4": {"class_type": "TextEncodeQwenImage21", "inputs": encode_inputs},
|
||||||
|
"6": {"class_type": "KSampler", "inputs": {
|
||||||
|
"model": ["1", 0], "positive": ["4", 0], "negative": ["4", 1],
|
||||||
|
"latent_image": ["4", 2] if references else ["5", 0],
|
||||||
|
"seed": seed, "steps": steps, "cfg": 1.0,
|
||||||
|
"sampler_name": "euler", "scheduler": "simple", "denoise": 1.0}},
|
||||||
|
"7": {"class_type": "VAEDecode", "inputs": {
|
||||||
|
"samples": ["6", 0], "vae": ["3", 0]}},
|
||||||
|
"8": {"class_type": "SaveImage", "inputs": {
|
||||||
|
"filename_prefix": prefix, "images": ["7", 0]}},
|
||||||
|
}
|
||||||
|
if references:
|
||||||
|
encode_inputs["vae"] = ["3", 0]
|
||||||
|
for index, name in enumerate(references, 1):
|
||||||
|
node = str(8 + index)
|
||||||
|
workflow[node] = {"class_type": "LoadImage", "inputs": {"image": name}}
|
||||||
|
encode_inputs[f"image_{index}"] = [node, 0]
|
||||||
|
else:
|
||||||
|
workflow["5"] = {"class_type": "EmptyLatentImage", "inputs": {
|
||||||
|
"width": width, "height": height, "batch_size": 1}}
|
||||||
|
return workflow
|
||||||
|
|
||||||
|
|
||||||
|
def wait_result(prompt_id: str, timeout: int = 1800) -> dict:
|
||||||
|
deadline = time.monotonic() + timeout
|
||||||
|
while time.monotonic() < deadline:
|
||||||
|
history = request_json(f"/history/{prompt_id}", timeout=10)
|
||||||
|
if prompt_id in history:
|
||||||
|
result = history[prompt_id]
|
||||||
|
status = result.get("status", {})
|
||||||
|
if status.get("status_str") == "error" or not status.get("completed", True):
|
||||||
|
raise RuntimeError("ComfyUI generation failed: " + json.dumps(status))
|
||||||
|
return result
|
||||||
|
time.sleep(1)
|
||||||
|
raise TimeoutError("Qwen image generation timed out")
|
||||||
|
|
||||||
|
|
||||||
|
def prepare_references(data: dict, job: str) -> list[str]:
|
||||||
|
source_files = data.get("source_files") or []
|
||||||
|
if not isinstance(source_files, list) or len(source_files) > 4:
|
||||||
|
raise ValueError("invalid source image list")
|
||||||
|
copied: list[str] = []
|
||||||
|
for index, name in enumerate(source_files, 1):
|
||||||
|
if not isinstance(name, str) or Path(name).name != name:
|
||||||
|
raise ValueError("invalid source image filename")
|
||||||
|
source = (OUTPUT_DIR / name).resolve()
|
||||||
|
if source.parent != OUTPUT_DIR or not source.is_file():
|
||||||
|
raise ValueError("source image not found")
|
||||||
|
with source.open("rb") as stream:
|
||||||
|
header = stream.read(16)
|
||||||
|
if header.startswith(b"\x89PNG\r\n\x1a\n"):
|
||||||
|
suffix = ".png"
|
||||||
|
elif header.startswith(b"\xff\xd8\xff"):
|
||||||
|
suffix = ".jpg"
|
||||||
|
elif header.startswith((b"RIFF",)) and header[8:12] == b"WEBP":
|
||||||
|
suffix = ".webp"
|
||||||
|
else:
|
||||||
|
raise ValueError("unsupported source image format")
|
||||||
|
target_name = f"{job}-ref-{index}{suffix}"
|
||||||
|
shutil.copyfile(source, INPUT_DIR / target_name)
|
||||||
|
copied.append(target_name)
|
||||||
|
return copied
|
||||||
|
|
||||||
|
|
||||||
|
def generate(data: dict) -> dict:
|
||||||
|
global ACTIVE
|
||||||
|
with GENERATION_LOCK:
|
||||||
|
ACTIVE = True
|
||||||
|
started = time.monotonic()
|
||||||
|
copied: list[str] = []
|
||||||
|
try:
|
||||||
|
prompt, filename, width, height, steps, _, seed = validate_request(data)
|
||||||
|
job = "router-" + uuid.uuid4().hex
|
||||||
|
copied = prepare_references(data, job)
|
||||||
|
queued = request_json("/prompt", {"prompt": make_workflow(
|
||||||
|
prompt, seed, steps, width, height, copied, job),
|
||||||
|
"client_id": "athena-image-router"})
|
||||||
|
result = wait_result(queued["prompt_id"])
|
||||||
|
images = [image for output in result.get("outputs", {}).values()
|
||||||
|
for image in output.get("images", [])]
|
||||||
|
if not images:
|
||||||
|
raise RuntimeError("generation completed without an image")
|
||||||
|
image = images[0]
|
||||||
|
generated = (OUTPUT_DIR / image.get("subfolder", "") /
|
||||||
|
image["filename"]).resolve()
|
||||||
|
if OUTPUT_DIR not in generated.parents or not generated.is_file():
|
||||||
|
raise RuntimeError("ComfyUI returned an invalid output path")
|
||||||
|
target = (OUTPUT_DIR / filename).resolve()
|
||||||
|
if target.parent != OUTPUT_DIR:
|
||||||
|
raise RuntimeError("invalid output path")
|
||||||
|
generated.replace(target)
|
||||||
|
return {"status": "ok", "filename": filename, "seed": seed,
|
||||||
|
"seconds": round(time.monotonic() - started, 3),
|
||||||
|
"model": MODEL}
|
||||||
|
finally:
|
||||||
|
for name in copied:
|
||||||
|
try:
|
||||||
|
(INPUT_DIR / name).unlink()
|
||||||
|
except FileNotFoundError:
|
||||||
|
pass
|
||||||
|
ACTIVE = False
|
||||||
|
|
||||||
|
|
||||||
|
class Handler(BaseHTTPRequestHandler):
|
||||||
|
def log_message(self, fmt: str, *args: object) -> None:
|
||||||
|
print(f"[qwen-image-2.1] {self.client_address[0]} {fmt % args}", flush=True)
|
||||||
|
|
||||||
|
def reply(self, status: int, payload: dict) -> None:
|
||||||
|
body = json.dumps(payload, separators=(",", ":")).encode()
|
||||||
|
self.send_response(status)
|
||||||
|
self.send_header("Content-Type", "application/json")
|
||||||
|
self.send_header("Content-Length", str(len(body)))
|
||||||
|
self.end_headers()
|
||||||
|
self.wfile.write(body)
|
||||||
|
|
||||||
|
def do_GET(self) -> None: # noqa: N802
|
||||||
|
if self.path != "/health":
|
||||||
|
self.reply(404, {"error": "not found"})
|
||||||
|
elif not READY.is_set():
|
||||||
|
self.reply(503, {"status": "starting", "model": MODEL})
|
||||||
|
else:
|
||||||
|
self.reply(200, {"status": "ok", "model_loaded": ACTIVE,
|
||||||
|
"model": MODEL})
|
||||||
|
|
||||||
|
def do_POST(self) -> None: # noqa: N802
|
||||||
|
if self.headers.get("Authorization", "") != f"Bearer {TOKEN}":
|
||||||
|
self.reply(401, {"error": "unauthorized"})
|
||||||
|
return
|
||||||
|
if self.path != "/generate":
|
||||||
|
self.reply(404, {"error": "not found"})
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
length = int(self.headers.get("Content-Length", "0"))
|
||||||
|
if length < 2 or length > 32768:
|
||||||
|
raise ValueError("invalid request size")
|
||||||
|
self.reply(200, generate(json.loads(self.rfile.read(length))))
|
||||||
|
except Exception as exc:
|
||||||
|
print(f"[qwen-image-2.1] generation failed: {type(exc).__name__}: "
|
||||||
|
f"{str(exc)[:1000]}", flush=True)
|
||||||
|
self.reply(500, {"status": "error", "message": str(exc)})
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
signal.signal(signal.SIGTERM, stop)
|
||||||
|
signal.signal(signal.SIGINT, stop)
|
||||||
|
start_comfy()
|
||||||
|
ThreadingHTTPServer((HOST, PORT), Handler).serve_forever()
|
||||||
+14
-10
@@ -19,7 +19,7 @@ Kommandos: POST /fast, /medium, /large, /ultra,
|
|||||||
/uncensored
|
/uncensored
|
||||||
GET /status (Zustand)
|
GET /status (Zustand)
|
||||||
|
|
||||||
Bildgenerierung und Editing (FLUX.2 Klein 9B FP8 beta):
|
Bildgenerierung und Editing (Qwen-Image-2.1 INT8):
|
||||||
POST /v1/images/generations (OpenAI-kompatibel)
|
POST /v1/images/generations (OpenAI-kompatibel)
|
||||||
POST /v1/images/edits (lokal, Referenzbilder)
|
POST /v1/images/edits (lokal, Referenzbilder)
|
||||||
GET /images (Liste)
|
GET /images (Liste)
|
||||||
@@ -39,8 +39,8 @@ Der Router leitet /v1/audio/speech und /v1/audio/transcriptions
|
|||||||
per HTTP an die Worker weiter.
|
per HTTP an die Worker weiter.
|
||||||
|
|
||||||
Der Router agiert als Modell-Orchestrator: vor der Generierung wird
|
Der Router agiert als Modell-Orchestrator: vor der Generierung wird
|
||||||
llama.cpp und Qwen3-TTS gestoppt, der Bild-Worker lädt FLUX.2 und den
|
llama.cpp und Qwen3-TTS gestoppt, der Bild-Worker lädt Qwen-Image und den
|
||||||
Text-Encoder auf getrennte GPUs, generiert/bearbeitet und entlädt
|
Textencoder auf die RTX 5080, generiert/bearbeitet und entlädt
|
||||||
das Modell wieder; danach wird das vorherige Qwen-Profil wiederher-
|
das Modell wieder; danach wird das vorherige Qwen-Profil wiederher-
|
||||||
gestellt und erst dann geantwortet (try/finally – Qwen wird auch bei
|
gestellt und erst dann geantwortet (try/finally – Qwen wird auch bei
|
||||||
Fehlgeschlagener Generierung wiederhergestellt).
|
Fehlgeschlagener Generierung wiederhergestellt).
|
||||||
@@ -137,7 +137,7 @@ DEFAULT_REASONING_EFFORT = os.environ.get(
|
|||||||
GLOBAL_SYSTEM_POLICY_FILE = os.environ.get(
|
GLOBAL_SYSTEM_POLICY_FILE = os.environ.get(
|
||||||
"GLOBAL_SYSTEM_POLICY_FILE", "").strip()
|
"GLOBAL_SYSTEM_POLICY_FILE", "").strip()
|
||||||
|
|
||||||
# --- Bildgenerierung und Referenzbild-Bearbeitung (FLUX.2 Klein 9B FP8) ---
|
# --- Bildgenerierung und Referenzbild-Bearbeitung (Qwen-Image-2.1 INT8) ---
|
||||||
LLAMA_SERVICE = os.environ.get("LLAMA_SERVICE", "mike-ai-llama-ui.service")
|
LLAMA_SERVICE = os.environ.get("LLAMA_SERVICE", "mike-ai-llama-ui.service")
|
||||||
SYSTEMCTL_BIN = os.environ.get("SYSTEMCTL_BIN", "systemctl")
|
SYSTEMCTL_BIN = os.environ.get("SYSTEMCTL_BIN", "systemctl")
|
||||||
IMAGE_WORKER = os.environ.get(
|
IMAGE_WORKER = os.environ.get(
|
||||||
@@ -147,7 +147,8 @@ IMAGE_PYTHON = os.environ.get(
|
|||||||
IMAGE_WORKER_URL = os.environ.get("IMAGE_WORKER_URL", "").rstrip("/")
|
IMAGE_WORKER_URL = os.environ.get("IMAGE_WORKER_URL", "").rstrip("/")
|
||||||
IMAGE_WORKER_TOKEN = os.environ.get("IMAGE_WORKER_TOKEN", "").strip()
|
IMAGE_WORKER_TOKEN = os.environ.get("IMAGE_WORKER_TOKEN", "").strip()
|
||||||
IMAGE_MODEL_NAME = os.environ.get(
|
IMAGE_MODEL_NAME = os.environ.get(
|
||||||
"IMAGE_MODEL_NAME", "FLUX.2-klein-9B-fp8-beta")
|
"IMAGE_MODEL_NAME", "Qwen-Image-2.1-int8")
|
||||||
|
IMAGE_INFERENCE_STEPS = int(os.environ.get("IMAGE_INFERENCE_STEPS", "25"))
|
||||||
IMAGE_DIR = os.environ.get(
|
IMAGE_DIR = os.environ.get(
|
||||||
"IMAGE_DIR", "/opt/mike-ai/ai-profile-router/images")
|
"IMAGE_DIR", "/opt/mike-ai/ai-profile-router/images")
|
||||||
IMAGE_WORKER_LOG = os.environ.get(
|
IMAGE_WORKER_LOG = os.environ.get(
|
||||||
@@ -175,8 +176,9 @@ IMAGE_SIZES = {
|
|||||||
"1920x1088": (1920, 1088),
|
"1920x1088": (1920, 1088),
|
||||||
"1088x1920": (1088, 1920),
|
"1088x1920": (1088, 1920),
|
||||||
}
|
}
|
||||||
# Das destillierte FLUX.2 Klein 9B ist auf vier Schritte ausgelegt.
|
# Der produktive Qwen-Image-2.1-Workflow ist auf 25 Schritte festgelegt.
|
||||||
IMAGE_QUALITY = {"standard": 4, "high": 4}
|
IMAGE_QUALITY = {"standard": IMAGE_INFERENCE_STEPS,
|
||||||
|
"high": IMAGE_INFERENCE_STEPS}
|
||||||
IMAGE_DEFAULT_QUALITY = "standard"
|
IMAGE_DEFAULT_QUALITY = "standard"
|
||||||
IMAGE_MAX_N = 4
|
IMAGE_MAX_N = 4
|
||||||
|
|
||||||
@@ -1134,7 +1136,7 @@ def switch_profile(profile: str, implicit: bool = False) -> dict:
|
|||||||
|
|
||||||
|
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
# Bildgenerierung und Editing (FLUX.2 Klein 9B FP8)
|
# Bildgenerierung und Editing (Qwen-Image-2.1 INT8)
|
||||||
# ---------------------------------------------------------------------------
|
# ---------------------------------------------------------------------------
|
||||||
|
|
||||||
class _Worker:
|
class _Worker:
|
||||||
@@ -2342,8 +2344,10 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
return
|
return
|
||||||
steps = data.get("steps", IMAGE_QUALITY[quality])
|
steps = data.get("steps", IMAGE_QUALITY[quality])
|
||||||
if (not isinstance(steps, int) or isinstance(steps, bool)
|
if (not isinstance(steps, int) or isinstance(steps, bool)
|
||||||
or steps != 4):
|
or steps != IMAGE_INFERENCE_STEPS):
|
||||||
self._send_error(400, f"{IMAGE_MODEL_NAME} erfordert 'steps'=4",
|
self._send_error(
|
||||||
|
400, f"{IMAGE_MODEL_NAME} erfordert "
|
||||||
|
f"'steps'={IMAGE_INFERENCE_STEPS}",
|
||||||
"invalid_request_error", "invalid_steps")
|
"invalid_request_error", "invalid_steps")
|
||||||
return
|
return
|
||||||
guidance = data.get("guidance", 1.0)
|
guidance = data.get("guidance", 1.0)
|
||||||
|
|||||||
@@ -1,11 +1,10 @@
|
|||||||
#!/usr/bin/env bash
|
#!/usr/bin/env bash
|
||||||
# Download pinned official weights, build the isolated worker and leave it stopped.
|
# Download pinned official weights and prepare the production worker stopped.
|
||||||
set -Eeuo pipefail
|
set -Eeuo pipefail
|
||||||
|
|
||||||
ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||||
ENV_FILE=${STACK_ENV:-/etc/mike-ai/stack.env}
|
ENV_FILE=${STACK_ENV:-/etc/mike-ai/stack.env}
|
||||||
MODEL_DIR=${QWEN_IMAGE_21_MODEL_DIR:-/data/models/Qwen-Image-2.1-ComfyUI}
|
MODEL_DIR=${QWEN_IMAGE_21_MODEL_DIR:-/data/models/Qwen-Image-2.1-ComfyUI}
|
||||||
OUTPUT_DIR=${QWEN_IMAGE_21_OUTPUT_DIR:-/data/qwen-image-2.1-test-output}
|
|
||||||
BASE=https://huggingface.co/Comfy-Org/Qwen-Image-2.1/resolve/ace0edeb3791a594ddfa36ed5f41a178a394e921
|
BASE=https://huggingface.co/Comfy-Org/Qwen-Image-2.1/resolve/ace0edeb3791a594ddfa36ed5f41a178a394e921
|
||||||
|
|
||||||
download() {
|
download() {
|
||||||
@@ -23,14 +22,12 @@ download() {
|
|||||||
download diffusion_models/qwen_image_2.1_int8_convrot.safetensors cb74113cb03faecd79611b01fd7fd642f0aa60d6f0b95086abee214d75eaa57d
|
download diffusion_models/qwen_image_2.1_int8_convrot.safetensors cb74113cb03faecd79611b01fd7fd642f0aa60d6f0b95086abee214d75eaa57d
|
||||||
download text_encoders/qwen3vl_8b_int8_convrot.safetensors 8bfd0f6e12abf2d2d697ecc888e5e90b0d6741d6708f05799f53afa560452e8f
|
download text_encoders/qwen3vl_8b_int8_convrot.safetensors 8bfd0f6e12abf2d2d697ecc888e5e90b0d6741d6708f05799f53afa560452e8f
|
||||||
download vae/qwen_image_2.1_vae_bf16.safetensors bb21f7473051e1ac368515dd3f2e15cd44d7a11748ee8823e1ddca3e4876b7c9
|
download vae/qwen_image_2.1_vae_bf16.safetensors bb21f7473051e1ac368515dd3f2e15cd44d7a11748ee8823e1ddca3e4876b7c9
|
||||||
install -d -m 0755 "$OUTPUT_DIR"
|
|
||||||
|
|
||||||
compose=(docker compose --env-file "$ENV_FILE" -f "$ROOT_DIR/compose.yaml" --profile qwen-image-test)
|
compose=(docker compose --env-file "$ENV_FILE" -f "$ROOT_DIR/compose.yaml" --profile image --profile flux-standby)
|
||||||
"${compose[@]}" pull qwen-image-21-test-runner
|
"${compose[@]}" build image-worker
|
||||||
"${compose[@]}" build qwen-image-21-test
|
|
||||||
"${compose[@]}" up -d --build --no-deps profile-controller
|
"${compose[@]}" up -d --build --no-deps profile-controller
|
||||||
"${compose[@]}" create qwen-image-21-test
|
"${compose[@]}" create image-worker flux-image-worker
|
||||||
"${compose[@]}" stop qwen-image-21-test
|
"${compose[@]}" stop image-worker flux-image-worker
|
||||||
state=$(docker inspect -f '{{.State.Status}}' mike-ai-qwen-image-2.1-test)
|
state=$(docker inspect -f '{{.State.Status}}' mike-ai-image-worker)
|
||||||
[[ $state == exited || $state == created ]]
|
[[ $state == exited || $state == created ]]
|
||||||
echo "Qwen-Image-2.1 test is prepared and stopped. No GPU model was loaded."
|
echo "Qwen-Image-2.1 production worker and FLUX standby are prepared and stopped."
|
||||||
@@ -1,136 +0,0 @@
|
|||||||
#!/usr/bin/env python3
|
|
||||||
"""One-shot Qwen-Image-2.1 evaluation with guaranteed profile restoration."""
|
|
||||||
|
|
||||||
from __future__ import annotations
|
|
||||||
|
|
||||||
import argparse
|
|
||||||
import json
|
|
||||||
import os
|
|
||||||
import pathlib
|
|
||||||
import random
|
|
||||||
import time
|
|
||||||
import urllib.parse
|
|
||||||
import urllib.request
|
|
||||||
|
|
||||||
|
|
||||||
CONTROLLER = os.environ.get("CONTROLLER_URL", "http://profile-controller:8090")
|
|
||||||
TOKEN = os.environ["CONTROLLER_TOKEN"]
|
|
||||||
COMFY = os.environ.get("COMFY_URL", "http://qwen-image-21-test:8188")
|
|
||||||
OUTPUT = pathlib.Path(os.environ.get("OUTPUT_DIR", "/output"))
|
|
||||||
|
|
||||||
|
|
||||||
def request_json(url: str, *, payload: dict | None = None,
|
|
||||||
authenticated: bool = False, timeout: float = 30) -> dict:
|
|
||||||
body = None if payload is None else json.dumps(payload).encode()
|
|
||||||
headers = {"Content-Type": "application/json"}
|
|
||||||
if authenticated:
|
|
||||||
headers["Authorization"] = f"Bearer {TOKEN}"
|
|
||||||
method = "POST" if payload is not None else "GET"
|
|
||||||
req = urllib.request.Request(url, data=body, headers=headers, method=method)
|
|
||||||
with urllib.request.urlopen(req, timeout=timeout) as response:
|
|
||||||
return json.load(response)
|
|
||||||
|
|
||||||
|
|
||||||
def controller(path: str, *, post: bool = False) -> dict:
|
|
||||||
return request_json(CONTROLLER + path, payload={} if post else None,
|
|
||||||
authenticated=True, timeout=900)
|
|
||||||
|
|
||||||
|
|
||||||
def wait_comfy(timeout: int = 900) -> None:
|
|
||||||
deadline = time.monotonic() + timeout
|
|
||||||
while time.monotonic() < deadline:
|
|
||||||
try:
|
|
||||||
request_json(COMFY + "/system_stats", timeout=3)
|
|
||||||
return
|
|
||||||
except Exception:
|
|
||||||
time.sleep(2)
|
|
||||||
raise TimeoutError("Qwen-Image worker did not become ready")
|
|
||||||
|
|
||||||
|
|
||||||
def workflow(prompt: str, seed: int, steps: int, width: int, height: int) -> dict:
|
|
||||||
return {
|
|
||||||
"1": {"class_type": "UNETLoader", "inputs": {
|
|
||||||
"unet_name": "qwen_image_2.1_int8_convrot.safetensors",
|
|
||||||
"weight_dtype": "default"}},
|
|
||||||
"2": {"class_type": "CLIPLoader", "inputs": {
|
|
||||||
"clip_name": "qwen3vl_8b_int8_convrot.safetensors",
|
|
||||||
"type": "qwen_image", "device": "default"}},
|
|
||||||
"3": {"class_type": "VAELoader", "inputs": {
|
|
||||||
"vae_name": "qwen_image_2.1_vae_bf16.safetensors"}},
|
|
||||||
"4": {"class_type": "TextEncodeQwenImage21", "inputs": {
|
|
||||||
"clip": ["2", 0], "prompt": prompt, "negative_prompt": "",
|
|
||||||
"resolution": max(width, height)}},
|
|
||||||
"5": {"class_type": "EmptyLatentImage", "inputs": {
|
|
||||||
"width": width, "height": height, "batch_size": 1}},
|
|
||||||
"6": {"class_type": "KSampler", "inputs": {
|
|
||||||
"model": ["1", 0], "positive": ["4", 0], "negative": ["4", 1],
|
|
||||||
"latent_image": ["5", 0], "seed": seed, "steps": steps,
|
|
||||||
"cfg": 1.0, "sampler_name": "euler", "scheduler": "simple",
|
|
||||||
"denoise": 1.0}},
|
|
||||||
"7": {"class_type": "VAEDecode", "inputs": {
|
|
||||||
"samples": ["6", 0], "vae": ["3", 0]}},
|
|
||||||
"8": {"class_type": "SaveImage", "inputs": {
|
|
||||||
"filename_prefix": "qwen-image-2.1-test", "images": ["7", 0]}},
|
|
||||||
}
|
|
||||||
|
|
||||||
|
|
||||||
def wait_result(prompt_id: str, timeout: int = 3600) -> dict:
|
|
||||||
deadline = time.monotonic() + timeout
|
|
||||||
while time.monotonic() < deadline:
|
|
||||||
history = request_json(f"{COMFY}/history/{prompt_id}", timeout=10)
|
|
||||||
if prompt_id in history:
|
|
||||||
result = history[prompt_id]
|
|
||||||
status = result.get("status", {})
|
|
||||||
if status.get("status_str") == "error" or not status.get("completed", True):
|
|
||||||
raise RuntimeError("generation failed: " + json.dumps(status))
|
|
||||||
return result
|
|
||||||
time.sleep(2)
|
|
||||||
raise TimeoutError("Qwen-Image generation timed out")
|
|
||||||
|
|
||||||
|
|
||||||
def main() -> None:
|
|
||||||
parser = argparse.ArgumentParser()
|
|
||||||
parser.add_argument("prompt")
|
|
||||||
parser.add_argument("--seed", type=int, default=None)
|
|
||||||
parser.add_argument("--steps", type=int, default=25)
|
|
||||||
parser.add_argument("--width", type=int, default=1024)
|
|
||||||
parser.add_argument("--height", type=int, default=1024)
|
|
||||||
args = parser.parse_args()
|
|
||||||
if args.width % 32 or args.height % 32:
|
|
||||||
parser.error("width and height must be multiples of 32")
|
|
||||||
seed = args.seed if args.seed is not None else random.randrange(2**53)
|
|
||||||
previous = controller("/profiles/status").get("active_profile")
|
|
||||||
started = time.monotonic()
|
|
||||||
try:
|
|
||||||
controller("/workers/qwen-image-test/start", post=True)
|
|
||||||
wait_comfy()
|
|
||||||
queued = request_json(COMFY + "/prompt", payload={
|
|
||||||
"prompt": workflow(args.prompt, seed, args.steps, args.width, args.height),
|
|
||||||
"client_id": "athena-qwen-image-21-test"}, timeout=30)
|
|
||||||
prompt_id = queued["prompt_id"]
|
|
||||||
result = wait_result(prompt_id)
|
|
||||||
images = []
|
|
||||||
for node in result.get("outputs", {}).values():
|
|
||||||
images.extend(node.get("images", []))
|
|
||||||
if not images:
|
|
||||||
raise RuntimeError("generation completed without an image")
|
|
||||||
image = images[0]
|
|
||||||
query = urllib.parse.urlencode({
|
|
||||||
"filename": image["filename"], "subfolder": image.get("subfolder", ""),
|
|
||||||
"type": image.get("type", "output")})
|
|
||||||
target = OUTPUT / image["filename"]
|
|
||||||
target.parent.mkdir(parents=True, exist_ok=True)
|
|
||||||
with urllib.request.urlopen(COMFY + "/view?" + query, timeout=120) as src:
|
|
||||||
target.write_bytes(src.read())
|
|
||||||
print(json.dumps({"status": "ok", "file": str(target), "seed": seed,
|
|
||||||
"seconds": round(time.monotonic() - started, 1)}, indent=2))
|
|
||||||
finally:
|
|
||||||
try:
|
|
||||||
controller("/workers/qwen-image-test/stop", post=True)
|
|
||||||
finally:
|
|
||||||
if previous:
|
|
||||||
controller(f"/profiles/{previous}/activate", post=True)
|
|
||||||
|
|
||||||
|
|
||||||
if __name__ == "__main__":
|
|
||||||
main()
|
|
||||||
Reference in New Issue
Block a user