Make Qwen Image 2.1 the production image worker
This commit is contained in:
@@ -20,7 +20,7 @@ Sie betreibt:
|
||||
|
||||
- llama.cpp mit genau einem aktiven Qwen-Profil,
|
||||
- den OpenAI-kompatiblen Profile Router,
|
||||
- FLUX.2 Klein 9B FP8 Beta für Textbilder und Referenzbild-Bearbeitung,
|
||||
- Qwen-Image-2.1 INT8 für Textbilder und Referenzbild-Bearbeitung,
|
||||
- Qwen3-TTS und TTS-Gateway für Sprache,
|
||||
- Whisper.cpp und die WebRTC-Brücke für OpenClaw Talk,
|
||||
- EmbeddingGemma auf der CPU für OpenClaws semantische Memory-Suche,
|
||||
@@ -90,7 +90,7 @@ nicht direkt. Kein automatischer Host-Neustart ist vorgesehen.
|
||||
- Fast: kurze, interaktive Aufgaben
|
||||
- Medium/Large/Ultra: steigende Kontextgrößen desselben lokalen Qwen-Modells
|
||||
- Uncensored: separates lokales Profil
|
||||
- FLUX.2 Klein 9B FP8 Beta: Bildgenerierung und Editing; Qwen wird dafür kurz entladen und danach
|
||||
- Qwen-Image-2.1 INT8: Bildgenerierung und Editing auf der RTX 5080; das Textmodell wird dafür kurz entladen und danach
|
||||
automatisch wiederhergestellt
|
||||
- Qwen3-TTS 1.7B: RTX 3060; kein Piper-Fallback
|
||||
- EmbeddingGemma 300M Q8: CPU, OpenAI-kompatibel auf Port 8082; kein GPU-Zugriff
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Athena AI
|
||||
|
||||
Stand: **20. September 2026**, auf Athena geprüft. Die Textprofil-Images verwenden
|
||||
Stand: **21. September 2026**, auf Athena geprüft. Die Textprofil-Images verwenden
|
||||
**llama.cpp 0.4.1** (`b29c606`).
|
||||
Der [geprüfte Live-Stand](docs/LIVE_STATE.md) beschreibt Profile, GPUs und die
|
||||
Abweichung zwischen dem bereitgestellten Stack und dem Git-Checkout.
|
||||
@@ -21,9 +21,8 @@ dokumentieren den ausgerollten Stand, reduzierte Statuslatenzen und Tests.
|
||||
|
||||
- genau ein aktives llama.cpp-Profil: Fast, Medium, Large, Ultra oder Uncensored
|
||||
- Profile Router auf Port 8081
|
||||
- FLUX.2 Klein 9B FP8 Beta: Transformer/VAE auf RTX 5080, Textencoder auf RTX 3060
|
||||
- Qwen-Image-2.1 als isolierter, standardmäßig gestoppter Vergleichsworker
|
||||
([Testablauf](docs/QWEN_IMAGE_21_TEST.md)); nicht produktiv an den Router angebunden
|
||||
- Qwen-Image-2.1 INT8 als produktiver Bildworker auf der RTX 5080
|
||||
- FLUX.2 Klein 9B FP8 Beta als gestoppter Rückfallcontainer
|
||||
- Qwen3-TTS 1.7B auf der RTX 3060 hinter dem TTS-Gateway; kein Piper-Fallback
|
||||
- Whisper.cpp `small` auf der CPU für lokale deutsche Spracherkennung
|
||||
- EmbeddingGemma 300M Q8 auf der CPU für OpenClaws hybride Memory-Suche
|
||||
@@ -44,15 +43,16 @@ dokumentieren den ausgerollten Stand, reduzierte Statuslatenzen und Tests.
|
||||
Hermes nutzt Athenas Router unter `http://192.168.1.212:8081/v1`. Ein MCPHub
|
||||
ist nicht mehr Bestandteil der produktiven Architektur.
|
||||
|
||||
## Verifizierte Bildauflösung auf dem Live-System
|
||||
## Produktiver Bildpfad
|
||||
|
||||
Der am 12. September geprüfte Live-Worker verwendet **FLUX.2 Klein 9B FP8
|
||||
Beta**. Die frühere 4B-Angabe war veraltet. Im Auflösungstest
|
||||
war **1024 × 1024 erfolgreich**, während **1280 × 1280 einen CUDA-OOM** auf
|
||||
der RTX 5080 auslöste. Höhere Auflösungen sind damit nicht freigegeben;
|
||||
Zwischenwerte wurden nicht getestet. Die produktive 1024-Begrenzung bleibt
|
||||
bestehen. Messwerte, Testbedingungen und die Korrektur früherer optimistischer
|
||||
Schätzungen stehen in [Auflösungstest vom 12. September](docs/FLUX_RESOLUTION_TEST_20260912.md).
|
||||
Bildanfragen an den Router verwenden seit dem 21. September
|
||||
**Qwen-Image-2.1 INT8** über gepinntes ComfyUI. Der Worker läuft mit Low-VRAM
|
||||
auf der RTX 5080, 25 Schritten, CFG 1 und kann bis zu vier Referenzbilder
|
||||
verarbeiten. Das aktive Textprofil und TTS werden für den Auftrag angehalten
|
||||
und anschließend wiederhergestellt. Der bisherige FLUX.2-Worker und seine
|
||||
Gewichte bleiben als gestoppter, explizit allowlist-beschränkter Rückfallpfad
|
||||
erhalten. Reproduzierbarer Stand und Rückschaltung:
|
||||
[Qwen-Image-2.1](docs/QWEN_IMAGE_21.md).
|
||||
|
||||
## Installation des Repository-Stands
|
||||
|
||||
|
||||
+51
-75
@@ -572,7 +572,7 @@ services:
|
||||
ALLOWED_PROFILES: fast,medium,large,ultra,uncensored
|
||||
IMAGE_WORKER: image
|
||||
RESTORE_WORKER: restore
|
||||
QWEN_IMAGE_TEST_WORKER: qwen-image-2.1-test
|
||||
FLUX_STANDBY_WORKER: flux-standby
|
||||
TTS_WORKER: qwen3
|
||||
MUSIC_WORKER: acestep
|
||||
YUE2_WORKER: yue2
|
||||
@@ -631,7 +631,8 @@ services:
|
||||
IMAGE_DIR: /data/images
|
||||
IMAGE_WORKER_URL: http://image-worker:8086
|
||||
IMAGE_WORKER_TOKEN: "${CONTROLLER_TOKEN:?CONTROLLER_TOKEN is required}"
|
||||
IMAGE_MODEL_NAME: FLUX.2-klein-9B-fp8-beta
|
||||
IMAGE_MODEL_NAME: Qwen-Image-2.1-int8
|
||||
IMAGE_INFERENCE_STEPS: "25"
|
||||
CHAT_IMAGE_ALLOW_REMOTE_URLS: "false"
|
||||
ENABLE_IMAGE_GENERATION: "true"
|
||||
ENABLE_TTS: "true"
|
||||
@@ -673,6 +674,51 @@ services:
|
||||
condition: service_healthy
|
||||
|
||||
image-worker:
|
||||
build:
|
||||
context: platform/docker/qwen-image-worker
|
||||
args:
|
||||
COMFYUI_COMMIT: b0f4b7b294ce482a2e071d9d762c133d38c7aa07
|
||||
image: mike-ai/qwen-image-worker:local
|
||||
container_name: mike-ai-image-worker
|
||||
restart: "no"
|
||||
profiles: [image]
|
||||
labels:
|
||||
com.mike-ai.image-worker: image
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids: ["${QWEN_IMAGE_21_GPU:-1}"]
|
||||
capabilities: [gpu]
|
||||
read_only: true
|
||||
tmpfs:
|
||||
- /tmp:size=4g,mode=1777
|
||||
- /opt/ComfyUI/user:size=64m,mode=0755
|
||||
- /opt/ComfyUI/temp:size=4g,mode=1777
|
||||
- /opt/ComfyUI/input:size=512m,mode=0755
|
||||
volumes:
|
||||
- "${QWEN_IMAGE_21_MODEL_DIR:-/data/models/Qwen-Image-2.1-ComfyUI}:/opt/ComfyUI/models:ro"
|
||||
- router-images:/data/images
|
||||
environment:
|
||||
NVIDIA_DRIVER_CAPABILITIES: compute,utility
|
||||
PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True
|
||||
WORKER_TOKEN: "${CONTROLLER_TOKEN:?CONTROLLER_TOKEN is required}"
|
||||
IMAGE_DIR: /data/images
|
||||
networks: [inference]
|
||||
security_opt: ["no-new-privileges:true"]
|
||||
cap_drop: [ALL]
|
||||
healthcheck:
|
||||
test: [CMD, python, -c, "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8086/health', timeout=2)"]
|
||||
interval: 5s
|
||||
timeout: 3s
|
||||
retries: 90
|
||||
start_period: 10s
|
||||
|
||||
# Previous production image model, retained as a stopped rollback target.
|
||||
# It is outside the normal router path and can only be started through the
|
||||
# controller's allowlisted flux-standby endpoint.
|
||||
flux-image-worker:
|
||||
build:
|
||||
context: platform/docker/image-worker
|
||||
args:
|
||||
@@ -681,11 +727,11 @@ services:
|
||||
ACCELERATE_VERSION: ${ACCELERATE_VERSION:-1.14.0}
|
||||
HF_HUB_VERSION: ${HF_HUB_VERSION:-1.28.0}
|
||||
image: mike-ai/image-worker:local
|
||||
container_name: mike-ai-image-worker
|
||||
container_name: mike-ai-flux-image-worker
|
||||
restart: "no"
|
||||
profiles: [image]
|
||||
profiles: [flux-standby]
|
||||
labels:
|
||||
com.mike-ai.image-worker: image
|
||||
com.mike-ai.image-worker: flux-standby
|
||||
gpus: all
|
||||
read_only: true
|
||||
tmpfs: ["/tmp:size=1g,mode=1777"]
|
||||
@@ -709,76 +755,6 @@ services:
|
||||
timeout: 3s
|
||||
retries: 12
|
||||
|
||||
# Isolated evaluation target. It is created in a stopped state by the
|
||||
# preparation script and can only be started through the controller's
|
||||
# allowlisted, GPU-exclusive test endpoint.
|
||||
qwen-image-21-test:
|
||||
build:
|
||||
context: platform/docker/qwen-image-21-test
|
||||
args:
|
||||
COMFYUI_COMMIT: b0f4b7b294ce482a2e071d9d762c133d38c7aa07
|
||||
image: mike-ai/qwen-image-2.1-test:local
|
||||
container_name: mike-ai-qwen-image-2.1-test
|
||||
restart: "no"
|
||||
profiles: [qwen-image-test]
|
||||
labels:
|
||||
com.mike-ai.image-worker: qwen-image-2.1-test
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids: ["${QWEN_IMAGE_21_GPU:-1}"]
|
||||
capabilities: [gpu]
|
||||
read_only: true
|
||||
tmpfs:
|
||||
- /tmp:size=4g,mode=1777
|
||||
- /opt/ComfyUI/user:size=64m,mode=0755
|
||||
- /opt/ComfyUI/temp:size=4g,mode=1777
|
||||
volumes:
|
||||
- "${QWEN_IMAGE_21_MODEL_DIR:-/data/models/Qwen-Image-2.1-ComfyUI}:/opt/ComfyUI/models:ro"
|
||||
- "${QWEN_IMAGE_21_OUTPUT_DIR:-/data/qwen-image-2.1-test-output}:/opt/ComfyUI/output"
|
||||
environment:
|
||||
NVIDIA_DRIVER_CAPABILITIES: compute,utility
|
||||
PYTORCH_CUDA_ALLOC_CONF: expandable_segments:True
|
||||
command:
|
||||
- python
|
||||
- main.py
|
||||
- --listen
|
||||
- 0.0.0.0
|
||||
- --port
|
||||
- "8188"
|
||||
- --lowvram
|
||||
- --preview-method
|
||||
- none
|
||||
networks: [inference]
|
||||
security_opt: ["no-new-privileges:true"]
|
||||
cap_drop: [ALL]
|
||||
healthcheck:
|
||||
test: [CMD, python, -c, "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8188/system_stats', timeout=2)"]
|
||||
interval: 5s
|
||||
timeout: 3s
|
||||
retries: 60
|
||||
start_period: 10s
|
||||
|
||||
qwen-image-21-test-runner:
|
||||
image: python:3.12-slim
|
||||
profiles: [qwen-image-test]
|
||||
read_only: true
|
||||
tmpfs: ["/tmp:size=32m,mode=1777"]
|
||||
volumes:
|
||||
- ./scripts/qwen-image-21-test.py:/opt/test/run.py:ro
|
||||
- "${QWEN_IMAGE_21_OUTPUT_DIR:-/data/qwen-image-2.1-test-output}:/output"
|
||||
environment:
|
||||
CONTROLLER_URL: http://profile-controller:8090
|
||||
CONTROLLER_TOKEN: "${CONTROLLER_TOKEN:?CONTROLLER_TOKEN is required}"
|
||||
COMFY_URL: http://qwen-image-21-test:8188
|
||||
OUTPUT_DIR: /output
|
||||
command: [python, /opt/test/run.py]
|
||||
networks: [control, inference]
|
||||
security_opt: ["no-new-privileges:true"]
|
||||
cap_drop: [ALL]
|
||||
|
||||
qwen3-tts:
|
||||
image: ${QWEN3_TTS_IMAGE:-ghcr.io/malaiwah/qwen3-tts-server:latest@sha256:b363a01d08b1bbecbfc3ca6f585368fae2cfdc591f9ecca6643738369f9a9d98}
|
||||
container_name: mike-ai-qwen3-tts
|
||||
|
||||
+1
-1
@@ -364,7 +364,7 @@ png = open('/tmp/test_dl.png', 'rb').read()
|
||||
print(json.dumps({
|
||||
'prompt': 'Behalte die Person bei und ändere nur den Hintergrund',
|
||||
'size': '1024x1024',
|
||||
'steps': 4,
|
||||
'steps': 25,
|
||||
'guidance': 1.0,
|
||||
'response_format': 'b64_json',
|
||||
'image_b64': base64.b64encode(png).decode(),
|
||||
|
||||
@@ -23,7 +23,7 @@ def item(profile, state="exited"):
|
||||
|
||||
|
||||
def image_item(state="exited"):
|
||||
return {"Id": "id-flux", "State": state,
|
||||
return {"Id": "id-qwen-image", "State": state,
|
||||
"Labels": {controller.IMAGE_LABEL_KEY: controller.IMAGE_WORKER}}
|
||||
|
||||
|
||||
@@ -32,9 +32,9 @@ def restore_item(state="exited"):
|
||||
"Labels": {controller.IMAGE_LABEL_KEY: controller.RESTORE_WORKER}}
|
||||
|
||||
|
||||
def qwen_image_test_item(state="exited"):
|
||||
return {"Id": "id-qwen-image-test", "State": state,
|
||||
"Labels": {controller.IMAGE_LABEL_KEY: controller.QWEN_IMAGE_TEST_WORKER}}
|
||||
def flux_standby_item(state="exited"):
|
||||
return {"Id": "id-flux-standby", "State": state,
|
||||
"Labels": {controller.IMAGE_LABEL_KEY: controller.FLUX_STANDBY_WORKER}}
|
||||
|
||||
|
||||
def tts_item(state="running"):
|
||||
@@ -97,7 +97,7 @@ class ProfileControllerTests(unittest.TestCase):
|
||||
self.assertEqual(result, {"music_worker": "acestep", "state": "running"})
|
||||
self.assertEqual(calls, [
|
||||
("POST", "/containers/id-ultra/stop?t=120"),
|
||||
("POST", "/containers/id-flux/stop?t=20"),
|
||||
("POST", "/containers/id-qwen-image/stop?t=20"),
|
||||
("POST", "/containers/id-tts/stop?t=30"),
|
||||
("POST", "/containers/id-music/start"),
|
||||
])
|
||||
@@ -158,10 +158,10 @@ class ProfileControllerTests(unittest.TestCase):
|
||||
self.assertEqual(calls, [
|
||||
("POST", "/containers/id-medium/stop?t=120"),
|
||||
("POST", "/containers/id-tts/stop?t=30"),
|
||||
("POST", "/containers/id-flux/start"),
|
||||
("POST", "/containers/id-qwen-image/start"),
|
||||
])
|
||||
|
||||
def test_restore_start_stops_flux_and_starts_restore(self):
|
||||
def test_restore_start_stops_qwen_image_and_starts_restore(self):
|
||||
profiles = {name: item(name) for name in controller.ALLOWED}
|
||||
calls = []
|
||||
|
||||
@@ -178,11 +178,11 @@ class ProfileControllerTests(unittest.TestCase):
|
||||
controller.set_image_worker(True, controller.RESTORE_WORKER)
|
||||
self.assertEqual(calls, [
|
||||
("POST", "/containers/id-tts/stop?t=30"),
|
||||
("POST", "/containers/id-flux/stop?t=20"),
|
||||
("POST", "/containers/id-qwen-image/stop?t=20"),
|
||||
("POST", "/containers/id-restore/start"),
|
||||
])
|
||||
|
||||
def test_qwen_image_test_is_allowlisted_and_exclusive(self):
|
||||
def test_flux_standby_is_allowlisted_and_exclusive(self):
|
||||
profiles = {name: item(name) for name in controller.ALLOWED}
|
||||
profiles["medium"] = item("medium", "running")
|
||||
calls = []
|
||||
@@ -192,21 +192,21 @@ class ProfileControllerTests(unittest.TestCase):
|
||||
return 204, b""
|
||||
|
||||
with patch.object(controller, "containers", return_value=profiles), \
|
||||
patch.object(controller, "image_container", return_value=qwen_image_test_item()), \
|
||||
patch.object(controller, "image_container", return_value=flux_standby_item()), \
|
||||
patch.object(controller, "image_containers", return_value=[
|
||||
image_item("running"), restore_item(), qwen_image_test_item()]), \
|
||||
image_item("running"), restore_item(), flux_standby_item()]), \
|
||||
patch.object(controller, "tts_container", return_value=tts_item()), \
|
||||
patch.object(controller, "docker_request", side_effect=request):
|
||||
result = controller.set_image_worker(
|
||||
True, controller.QWEN_IMAGE_TEST_WORKER)
|
||||
True, controller.FLUX_STANDBY_WORKER)
|
||||
|
||||
self.assertEqual(result, {
|
||||
"image_worker": "qwen-image-2.1-test", "state": "running"})
|
||||
"image_worker": "flux-standby", "state": "running"})
|
||||
self.assertEqual(calls, [
|
||||
("POST", "/containers/id-medium/stop?t=120"),
|
||||
("POST", "/containers/id-tts/stop?t=30"),
|
||||
("POST", "/containers/id-flux/stop?t=20"),
|
||||
("POST", "/containers/id-qwen-image-test/start"),
|
||||
("POST", "/containers/id-qwen-image/stop?t=20"),
|
||||
("POST", "/containers/id-flux-standby/start"),
|
||||
])
|
||||
|
||||
def test_profile_activation_stops_image_worker_first(self):
|
||||
@@ -224,7 +224,7 @@ class ProfileControllerTests(unittest.TestCase):
|
||||
patch.object(controller, "docker_request", side_effect=request):
|
||||
controller.activate("fast")
|
||||
self.assertEqual(calls, [
|
||||
("POST", "/containers/id-flux/stop?t=120"),
|
||||
("POST", "/containers/id-qwen-image/stop?t=120"),
|
||||
("POST", "/containers/id-fast/start"),
|
||||
])
|
||||
|
||||
|
||||
@@ -0,0 +1,47 @@
|
||||
import importlib.util
|
||||
import os
|
||||
import pathlib
|
||||
import unittest
|
||||
|
||||
|
||||
os.environ.setdefault("WORKER_TOKEN", "x" * 48)
|
||||
PATH = (pathlib.Path(__file__).parents[1]
|
||||
/ "platform/docker/qwen-image-worker/qwen_image_worker.py")
|
||||
SPEC = importlib.util.spec_from_file_location("qwen_image_worker", PATH)
|
||||
worker = importlib.util.module_from_spec(SPEC)
|
||||
assert SPEC and SPEC.loader
|
||||
SPEC.loader.exec_module(worker)
|
||||
|
||||
|
||||
class QwenImageWorkerTests(unittest.TestCase):
|
||||
def test_text_to_image_uses_empty_latent(self):
|
||||
workflow = worker.make_workflow("test", 7, 25, 1024, 1024, [], "job")
|
||||
self.assertEqual(workflow["6"]["inputs"]["latent_image"], ["5", 0])
|
||||
self.assertIn("5", workflow)
|
||||
self.assertNotIn("vae", workflow["4"]["inputs"])
|
||||
|
||||
def test_edit_uses_reference_latent_and_qwen_dynamic_inputs(self):
|
||||
workflow = worker.make_workflow(
|
||||
"edit <image1>", 7, 25, 1024, 1024,
|
||||
["job-ref-1.png", "job-ref-2.jpg"], "job")
|
||||
encode = workflow["4"]["inputs"]
|
||||
self.assertEqual(workflow["6"]["inputs"]["latent_image"], ["4", 2])
|
||||
self.assertEqual(encode["vae"], ["3", 0])
|
||||
self.assertEqual(encode["image_1"], ["9", 0])
|
||||
self.assertEqual(encode["image_2"], ["10", 0])
|
||||
self.assertNotIn("5", workflow)
|
||||
|
||||
def test_router_contract_is_fixed_to_production_parameters(self):
|
||||
request = {
|
||||
"prompt": "test", "filename": "result.png", "width": 1024,
|
||||
"height": 1024, "steps": 25, "guidance": 1.0, "seed": 42,
|
||||
}
|
||||
self.assertEqual(worker.validate_request(request)[1:],
|
||||
("result.png", 1024, 1024, 25, 1.0, 42))
|
||||
request["steps"] = 4
|
||||
with self.assertRaisesRegex(ValueError, "steps=25"):
|
||||
worker.validate_request(request)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -12,7 +12,7 @@ flowchart LR
|
||||
H -->|OpenAI API| R[Profile Router<br/>Athena :8081]
|
||||
R --> P[Profile Controller]
|
||||
P --> Q[genau ein llama.cpp-Profil<br/>Qwen Fast / Medium / Large / Ultra / Uncensored]
|
||||
R --> I[FLUX.2 Klein 9B FP8 Beta<br/>RTX 5080, Text + Editing]
|
||||
R --> I[Qwen-Image-2.1 INT8<br/>RTX 5080, Text + Editing]
|
||||
R --> T[Qwen3-TTS 1.7B RTX 3060<br/>TTS-Gateway]
|
||||
R --> STT[Whisper.cpp small<br/>CPU, lokale Spracherkennung]
|
||||
|
||||
@@ -32,8 +32,9 @@ flowchart LR
|
||||
K[Backup alle 5 Stunden] --> DATA[/data und /etc/mike-ai]
|
||||
```
|
||||
|
||||
Stand: 20. September 2026. Versionen und Abweichungen: [Live-Stand](LIVE_STATE.md).
|
||||
FLUX nutzt beide GPUs und pausiert dafür LLM und Qwen3-TTS. Dashboard und
|
||||
Stand: 21. September 2026. Versionen und Abweichungen: [Live-Stand](LIVE_STATE.md).
|
||||
Qwen Image nutzt die RTX 5080 und pausiert dafür LLM und Qwen3-TTS. FLUX
|
||||
bleibt als gestoppter Rückfallworker vorhanden. Dashboard und
|
||||
Portainer besitzen eigene Netzwerk-Namespaces in `mike-ai_frontend`; der Router
|
||||
läuft in `mike-ai_control`. Der Gateway vermittelt den privaten Zugriff.
|
||||
Zusätzliche Musik-, Audio-, Voice- und 3D-Container stehen im
|
||||
|
||||
@@ -14,7 +14,8 @@ nicht automatisch ein ungenutzter Rest.
|
||||
| `mike-ai-embedding` | EmbeddingGemma 300M QAT Q8 | Dauerhafter CPU-only-Endpunkt für OpenClaws semantische und hybride Memory-Suche; teilt ausschließlich den privaten Netzwerk-Namespace des WireGuard-Gateways. |
|
||||
| `mike-ai-bonsai2-ab` | Bonsai-2-Vergleichsmodell | Gestoppter, reproduzierbar dokumentierter A/B-Testcontainer; kein Produktivprofil. |
|
||||
| `mike-ai-applio-studio` | Applio/RVC; Stimmenmodelle werden nutzerseitig ergänzt | Vollständige RVC-Oberfläche für Inferenz, Modellverwaltung und Training auf der RTX 5080. Für eine Konvertierung ist ein importiertes oder trainiertes `.pth`-Modell nötig; eine Referenzaufnahme allein reicht nicht. |
|
||||
| `mike-ai-image-worker` | FLUX.2 Klein 9B FP8, Qwen3-8B NF4 Textencoder und VAE | Erzeugt und bearbeitet Bilder transaktional; nutzt während eines Auftrags RTX 5080 und RTX 3060. |
|
||||
| `mike-ai-image-worker` | Qwen-Image-2.1 INT8, Qwen3-VL-8B INT8 und VAE | Produktiver Bildworker; erzeugt und bearbeitet Bilder transaktional auf der RTX 5080. |
|
||||
| `mike-ai-flux-image-worker` | FLUX.2 Klein 9B FP8, Qwen3-8B NF4 Textencoder und VAE | Gestoppter Rückfallworker; wird vom normalen Router-Bildpfad nicht gestartet. |
|
||||
| `mike-ai-llama-dashboard` | kein Modell | Zeigt Telemetrie, Profile, GPU-Nutzung, Slot-Kontextbelegung mit Verlauf und Betriebsarten an und bietet die Modusumschaltung. |
|
||||
| `mike-ai-llama-fast` | Qwen3.8-27B `IQ4-MIX`, Qwen-MMProj BF16 | Schnelles Q4-Text-/Vision-Profil mit 76.800 Token Kontext. |
|
||||
| `mike-ai-llama-large` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Q4-Text-/Vision-Profil mit 192.000 Token Kontext und Verteilung auf beide GPUs. |
|
||||
@@ -55,4 +56,4 @@ sämtliche Modellgewichte oder Funktionen neu.
|
||||
|
||||
Gestoppte Container sind nicht automatisch Testreste. Vor einer Entfernung
|
||||
Compose-Zuordnung, Mounts und Controller-Verweise prüfen. Die vorhandenen
|
||||
Spezialcontainer wurden durch den FLUX-Auflösungstest weder angelegt noch entfernt.
|
||||
Die historischen FLUX-Tests sind keine Aussage über den produktiven Qwen-Pfad.
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Aktuelle Laufzeitnotizen
|
||||
|
||||
Stand: 20. September 2026.
|
||||
Stand: 21. September 2026.
|
||||
|
||||
Der verbindliche Abgleich steht im [geprüften Live-Stand](LIVE_STATE.md).
|
||||
Produktiv läuft **llama.cpp 0.4.1** auf Commit `b29c606`, zuvor b10930.
|
||||
@@ -18,9 +18,10 @@ Funktionsproben, Vergleichsmessungen und gesicherter Rückfallstand.
|
||||
[Update-Audit vom 15. September](UPDATE_AUDIT_20260915.md): aktualisierte
|
||||
Komponenten, unveränderte aktuelle Komponenten, Aufräumarbeiten und Rollback.
|
||||
|
||||
FLUX verwendet **9B FP8 Beta** mit Textencoder auf der RTX 3060.
|
||||
[Auflösungstest](FLUX_RESOLUTION_TEST_20260912.md): 1024 erfolgreich,
|
||||
1280 CUDA-OOM; produktive Begrenzung weiterhin 1024.
|
||||
Der produktive Bildpfad verwendet **Qwen-Image-2.1 INT8** mit ComfyUI auf der
|
||||
RTX 5080. FLUX 9B FP8 und seine Gewichte bleiben als gestoppter Rückfallpfad.
|
||||
[Qwen-Betriebsdoku](QWEN_IMAGE_21.md); der historische
|
||||
[FLUX-Auflösungstest](FLUX_RESOLUTION_TEST_20260912.md) gilt nur für FLUX.
|
||||
Sprachausgabe: Qwen3-TTS 1.7B ohne Piper. Spracherkennung: Whisper.cpp 1.9.4
|
||||
mit dem Modell `small` auf der CPU. Applio basiert auf dem stabilen Release 3.6.4.
|
||||
|
||||
|
||||
+4
-4
@@ -39,10 +39,10 @@ nicht mehr den aktuellen Containerzustand.
|
||||
|
||||
## Weitere Dienste
|
||||
|
||||
- FLUX.2 Klein 9B FP8 Beta: Transformer/VAE auf 5080, Qwen3-8B-NF4-
|
||||
Textencoder auf 3060. Produktiv freigegeben sind höchstens 1024 × 1024,
|
||||
vier Schritte, Guidance 1, vier Referenzbilder und 128 Encoder-Tokens.
|
||||
1280 × 1280 lief im Test in CUDA-OOM.
|
||||
- Qwen-Image-2.1 INT8 läuft produktiv über gepinntes ComfyUI mit Low-VRAM auf
|
||||
der RTX 5080. Der Router verwendet 25 Schritte, Guidance 1 und bis zu vier
|
||||
Referenzbilder. FLUX.2 Klein 9B FP8 bleibt mit seinen Gewichten als
|
||||
gestoppter, explizit allowlist-beschränkter Rückfallcontainer erhalten.
|
||||
- Qwen3-TTS 1.7B auf 3060 hinter dem TTS-Gateway; kein Piper-Fallback.
|
||||
- Whisper.cpp `ggml-small` auf CPU über `/v1/audio/transcriptions`.
|
||||
- `mike-ai-embedding` stellt EmbeddingGemma 300M Q8 CPU-only über den privaten
|
||||
|
||||
@@ -0,0 +1,83 @@
|
||||
# Qwen-Image-2.1 auf Athena
|
||||
|
||||
Stand: 21. September 2026. Qwen-Image-2.1 ist der produktive Bildworker des
|
||||
Routers. FLUX.2 Klein 9B FP8 bleibt als gestoppter Rückfallcontainer erhalten.
|
||||
|
||||
## Gepinnter Stand
|
||||
|
||||
- ComfyUI: Commit `b0f4b7b294ce482a2e071d9d762c133d38c7aa07`
|
||||
- Modell-Repository: `Comfy-Org/Qwen-Image-2.1`, Revision
|
||||
`ace0edeb3791a594ddfa36ed5f41a178a394e921`
|
||||
- Transformer: `qwen_image_2.1_int8_convrot.safetensors`
|
||||
(`cb74113cb03faecd79611b01fd7fd642f0aa60d6f0b95086abee214d75eaa57d`)
|
||||
- Textencoder: `qwen3vl_8b_int8_convrot.safetensors`
|
||||
(`8bfd0f6e12abf2d2d697ecc888e5e90b0d6741d6708f05799f53afa560452e8f`)
|
||||
- VAE: `qwen_image_2.1_vae_bf16.safetensors`
|
||||
(`bb21f7473051e1ac368515dd3f2e15cd44d7a11748ee8823e1ddca3e4876b7c9`)
|
||||
- Pipeline: 25 Schritte, CFG 1, Euler/Simple, `--lowvram`
|
||||
- GPU: ausschließlich Athena-GPU 1, die RTX 5080 mit 16 GB
|
||||
|
||||
Die Gewichte liegen außerhalb von Git unter
|
||||
`/data/models/Qwen-Image-2.1-ComfyUI`. Der Adapter
|
||||
`platform/docker/qwen-image-worker/qwen_image_worker.py` stellt vor ComfyUI
|
||||
die bestehende private Worker-API auf Port 8086 bereit. Der öffentliche
|
||||
Routervertrag bleibt damit `/v1/images/generations` und `/v1/images/edits`.
|
||||
|
||||
## Lebenszyklus
|
||||
|
||||
Ein Bildauftrag läuft transaktional:
|
||||
|
||||
1. Der Router merkt sich das aktive Textprofil.
|
||||
2. Der Profile Controller stoppt LLM, TTS und andere GPU-Spezialworker.
|
||||
3. `mike-ai-image-worker` startet Qwen Image auf der RTX 5080.
|
||||
4. Der Adapter erzeugt das Bild oder bearbeitet bis zu vier Referenzbilder.
|
||||
5. Der Qwen-Worker wird vollständig gestoppt.
|
||||
6. Das vorherige Textprofil und Qwen3-TTS werden wiederhergestellt.
|
||||
|
||||
Der Container ist außerhalb eines Auftrags gestoppt. Das ist der Sollzustand.
|
||||
|
||||
## Reproduzierbare Vorbereitung
|
||||
|
||||
```bash
|
||||
cd /opt/mike-ai/stack
|
||||
sudo ./scripts/prepare-qwen-image-21.sh
|
||||
```
|
||||
|
||||
Das Skript lädt fehlende Dateien mit festen SHA-256-Prüfsummen, baut den
|
||||
produktiven Qwen-Worker und legt Qwen sowie FLUX gestoppt an. Es lädt dabei
|
||||
kein Modell in den VRAM.
|
||||
|
||||
## Funktionsprobe über den echten Routerweg
|
||||
|
||||
```bash
|
||||
curl -sS http://127.0.0.1:8081/v1/images/generations \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{
|
||||
"model":"Qwen-Image-2.1-int8",
|
||||
"prompt":"Fotorealistisches rotes Haus bei Tageslicht",
|
||||
"size":"1024x1024",
|
||||
"steps":25,
|
||||
"guidance":1.0,
|
||||
"response_format":"url"
|
||||
}'
|
||||
```
|
||||
|
||||
Nach dem Lauf müssen `mike-ai-image-worker` und
|
||||
`mike-ai-flux-image-worker` gestoppt sein und das vorherige LLM-Profil wieder
|
||||
gesund laufen.
|
||||
|
||||
## FLUX-Rückfallpfad
|
||||
|
||||
FLUX verwendet weiterhin das Image `mike-ai/image-worker:local`, die
|
||||
bestehenden Modellverzeichnisse und den Container
|
||||
`mike-ai-flux-image-worker`. Er hat das Controller-Label `flux-standby` und
|
||||
kann deshalb nicht durch einen normalen Router-Bildauftrag gestartet werden.
|
||||
Eine manuelle Diagnose ist nur über den allowlist-beschränkten Controllerpfad
|
||||
`/workers/flux-standby/start` möglich. Für eine dauerhafte Rückschaltung müssen
|
||||
Service-DNS, Modellname und Schrittzahl gemeinsam in Compose auf FLUX gesetzt
|
||||
und anschließend Router und Controller neu ausgerollt werden. Keine dieser
|
||||
Änderungen erfolgt automatisch.
|
||||
|
||||
Historische FLUX-Messwerte und Grenzen stehen in
|
||||
[FLUX_9B_BETA.md](FLUX_9B_BETA.md) und
|
||||
[FLUX_RESOLUTION_TEST_20260912.md](FLUX_RESOLUTION_TEST_20260912.md).
|
||||
@@ -1,60 +0,0 @@
|
||||
# Qwen-Image-2.1 – isolierter Test
|
||||
|
||||
Stand: 21. September 2026. Dieser Pfad dient nur dem Vergleich mit dem
|
||||
produktiven FLUX.2-Worker. Er ersetzt FLUX nicht und ist nicht an den Router
|
||||
angebunden.
|
||||
|
||||
## Reproduzierbarer Stand
|
||||
|
||||
- ComfyUI: Commit `b0f4b7b294ce482a2e071d9d762c133d38c7aa07`
|
||||
- Modell-Repository: `Comfy-Org/Qwen-Image-2.1`, Revision
|
||||
`ace0edeb3791a594ddfa36ed5f41a178a394e921`
|
||||
- Transformer: `qwen_image_2.1_int8_convrot.safetensors` (7.256.783.064 Byte)
|
||||
- Textencoder: `qwen3vl_8b_int8_convrot.safetensors` (9.350.798.360 Byte)
|
||||
- VAE: `qwen_image_2.1_vae_bf16.safetensors` (675.509.688 Byte)
|
||||
- erster Vergleich: 1024 × 1024, 25 Schritte, CFG 1, Euler/Simple
|
||||
|
||||
Die Gewichte sind die von Qwen verlinkte offizielle ComfyUI-Aufbereitung. Der
|
||||
Worker verwendet `--lowvram` auf der RTX 5080. So bleibt der Test mit 16 GB
|
||||
VRAM möglich; der Preis ist CPU-Offload und damit eine längere Laufzeit.
|
||||
|
||||
## Vorbereitung ohne GPU-Last
|
||||
|
||||
```bash
|
||||
cd /opt/mike-ai/stack
|
||||
sudo ./scripts/prepare-qwen-image-21-test.sh
|
||||
```
|
||||
|
||||
Das Skript prüft feste SHA-256-Summen, baut das Image, aktualisiert den
|
||||
Profile Controller und erstellt den Testcontainer **gestoppt**. Es lädt kein
|
||||
Modell in die GPU.
|
||||
|
||||
## Einmaliger Test nach ausdrücklichem „Go“
|
||||
|
||||
```bash
|
||||
cd /opt/mike-ai/stack
|
||||
docker compose --env-file /etc/mike-ai/stack.env \
|
||||
--profile qwen-image-test run --rm qwen-image-21-test-runner \
|
||||
python /opt/test/run.py 'HIER DEN VEREINBARTEN PROMPT EINSETZEN'
|
||||
```
|
||||
|
||||
Der Runner merkt sich das aktive LLM-Profil, startet den Qwen-Worker über den
|
||||
allowlist-beschränkten Controller, erzeugt genau ein Bild und stellt danach
|
||||
auch bei einem Fehler das vorherige Profil wieder her. Das Ergebnis liegt in
|
||||
`/data/qwen-image-2.1-test-output/`. FLUX-Konfiguration und FLUX-Gewichte
|
||||
werden nicht verändert.
|
||||
|
||||
## Vollständig entfernen
|
||||
|
||||
```bash
|
||||
cd /opt/mike-ai/stack
|
||||
docker compose --env-file /etc/mike-ai/stack.env \
|
||||
--profile qwen-image-test rm -sf qwen-image-21-test
|
||||
docker image rm mike-ai/qwen-image-2.1-test:local
|
||||
rm -rf /data/models/Qwen-Image-2.1-ComfyUI \
|
||||
/data/qwen-image-2.1-test-output
|
||||
```
|
||||
|
||||
Anschließend kann die Controller-Erweiterung bei Bedarf aus Git zurückgenommen
|
||||
und nur `profile-controller` neu gebaut werden. Die Entfernung ist nicht Teil
|
||||
der Vorbereitung und wird niemals automatisch ausgeführt.
|
||||
@@ -53,7 +53,8 @@ Titelgenerierung und Kontextkompression in Hermes.
|
||||
| Datum | Modell | Ergebnis | Status / Entscheidung | Beleg |
|
||||
|---|---|---|---|---|
|
||||
| bis 07.09.2026 | FLUX.2 Klein 4B | funktional, aber schwächere räumliche und motivische Konsistenz | **ersetzt** durch 9B FP8 | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
||||
| 07.09.2026 | FLUX.2 Klein 9B FP8 | bessere Prompttreue; produktiver Zwei-GPU-Pfad, derzeit auf 1024 × 1024 begrenzt | **produktiv als Beta** | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
||||
| 07.09.2026 | FLUX.2 Klein 9B FP8 | bessere Prompttreue als 4B; produktiver Zwei-GPU-Pfad, auf 1024 × 1024 begrenzt | **Standby seit 21.09.2026** | [FLUX_9B_BETA.md](FLUX_9B_BETA.md) |
|
||||
| 21.09.2026 | Qwen-Image-2.1 INT8 | im direkten Vergleich fotorealistischer, bessere Textdarstellung und Prompttreue; 1024 × 1024 bei 25 Schritten in rund 49 s Pipeline-Lauf | **produktiv** | [QWEN_IMAGE_21.md](QWEN_IMAGE_21.md) |
|
||||
| 08.09.2026 | HYPIR-SD2 | glättete oder erfand Details und veränderte kleine Strukturen | **verworfen** | [IMAGE_RESTORATION.md](IMAGE_RESTORATION.md) |
|
||||
| 08.09.2026 | SeedVR2 7B FP8 | bewahrte Identität besser als HYPIR, brachte beim realen unscharfen Foto aber kaum nutzbare Details zurück | **verworfen** | [IMAGE_RESTORATION.md](IMAGE_RESTORATION.md) |
|
||||
|
||||
|
||||
@@ -27,8 +27,8 @@ LABEL_KEY = "com.mike-ai.llama-profile"
|
||||
IMAGE_LABEL_KEY = "com.mike-ai.image-worker"
|
||||
IMAGE_WORKER = os.environ.get("IMAGE_WORKER", "image")
|
||||
RESTORE_WORKER = os.environ.get("RESTORE_WORKER", "restore")
|
||||
QWEN_IMAGE_TEST_WORKER = os.environ.get(
|
||||
"QWEN_IMAGE_TEST_WORKER", "qwen-image-2.1-test").strip()
|
||||
FLUX_STANDBY_WORKER = os.environ.get(
|
||||
"FLUX_STANDBY_WORKER", "flux-standby").strip()
|
||||
TTS_LABEL_KEY = "com.mike-ai.tts-worker"
|
||||
TTS_WORKER = os.environ.get("TTS_WORKER", "qwen3")
|
||||
MUSIC_LABEL_KEY = "com.mike-ai.music-worker"
|
||||
@@ -103,8 +103,8 @@ def image_container(kind: str = IMAGE_WORKER) -> dict:
|
||||
def image_containers() -> list[dict]:
|
||||
"""All allowlisted GPU workers that must never overlap an LLM."""
|
||||
allowed = {IMAGE_WORKER, RESTORE_WORKER}
|
||||
if QWEN_IMAGE_TEST_WORKER:
|
||||
allowed.add(QWEN_IMAGE_TEST_WORKER)
|
||||
if FLUX_STANDBY_WORKER:
|
||||
allowed.add(FLUX_STANDBY_WORKER)
|
||||
return [item for item in labelled_containers(IMAGE_LABEL_KEY)
|
||||
if item.get("Labels", {}).get(IMAGE_LABEL_KEY) in allowed]
|
||||
|
||||
@@ -363,7 +363,7 @@ def stop_inference() -> dict:
|
||||
|
||||
|
||||
def set_image_worker(running: bool, kind: str = IMAGE_WORKER) -> dict:
|
||||
if kind not in {IMAGE_WORKER, RESTORE_WORKER, QWEN_IMAGE_TEST_WORKER}:
|
||||
if kind not in {IMAGE_WORKER, RESTORE_WORKER, FLUX_STANDBY_WORKER}:
|
||||
raise ValueError("worker is not allowlisted")
|
||||
with LOCK:
|
||||
item = image_container(kind)
|
||||
@@ -371,8 +371,8 @@ def set_image_worker(running: bool, kind: str = IMAGE_WORKER) -> dict:
|
||||
# The image worker may never overlap a llama profile on the 5080.
|
||||
for profile_item in containers().values():
|
||||
stop_container(profile_item)
|
||||
# The 9B beta text encoder temporarily borrows the RTX 3060 from
|
||||
# Qwen3-TTS. TTS is unavailable during this exclusive GPU phase.
|
||||
# Image workers are GPU-exclusive. The FLUX standby also borrows
|
||||
# the RTX 3060, while Qwen Image runs only on the RTX 5080.
|
||||
stop_container(tts_container(), timeout=30)
|
||||
stop_music_if_configured()
|
||||
stop_yue2_if_configured()
|
||||
@@ -801,8 +801,8 @@ class Handler(BaseHTTPRequestHandler):
|
||||
"/workers/image/stop": (IMAGE_WORKER, False),
|
||||
"/workers/restore/start": (RESTORE_WORKER, True),
|
||||
"/workers/restore/stop": (RESTORE_WORKER, False),
|
||||
"/workers/qwen-image-test/start": (QWEN_IMAGE_TEST_WORKER, True),
|
||||
"/workers/qwen-image-test/stop": (QWEN_IMAGE_TEST_WORKER, False),
|
||||
"/workers/flux-standby/start": (FLUX_STANDBY_WORKER, True),
|
||||
"/workers/flux-standby/stop": (FLUX_STANDBY_WORKER, False),
|
||||
}
|
||||
if self.path in worker_paths:
|
||||
try:
|
||||
|
||||
+3
-1
@@ -13,4 +13,6 @@ RUN test -n "$COMFYUI_COMMIT" \
|
||||
&& rm -rf /var/lib/apt/lists/* /root/.cache
|
||||
|
||||
WORKDIR /opt/ComfyUI
|
||||
EXPOSE 8188
|
||||
COPY qwen_image_worker.py /opt/qwen-image-worker.py
|
||||
EXPOSE 8086
|
||||
CMD ["python", "/opt/qwen-image-worker.py"]
|
||||
@@ -0,0 +1,268 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Router-compatible Qwen-Image-2.1 worker backed by pinned ComfyUI."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
import random
|
||||
import shutil
|
||||
import signal
|
||||
import subprocess
|
||||
import threading
|
||||
import time
|
||||
import urllib.request
|
||||
import uuid
|
||||
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
HOST = os.environ.get("WORKER_HOST", "0.0.0.0")
|
||||
PORT = int(os.environ.get("WORKER_PORT", "8086"))
|
||||
TOKEN = os.environ.get("WORKER_TOKEN", "").strip()
|
||||
COMFY = "http://127.0.0.1:8188"
|
||||
COMFY_DIR = Path("/opt/ComfyUI")
|
||||
INPUT_DIR = COMFY_DIR / "input"
|
||||
OUTPUT_DIR = Path(os.environ.get("IMAGE_DIR", "/data/images")).resolve()
|
||||
MODEL = "Qwen-Image-2.1-int8"
|
||||
GENERATION_LOCK = threading.Lock()
|
||||
READY = threading.Event()
|
||||
ACTIVE = False
|
||||
COMFY_PROCESS: subprocess.Popen | None = None
|
||||
|
||||
if len(TOKEN) < 32:
|
||||
raise RuntimeError("WORKER_TOKEN is missing or too short")
|
||||
|
||||
|
||||
def request_json(path: str, payload: dict | None = None,
|
||||
timeout: float = 30) -> dict:
|
||||
body = None if payload is None else json.dumps(payload).encode()
|
||||
headers = {"Content-Type": "application/json"} if body is not None else {}
|
||||
method = "POST" if body is not None else "GET"
|
||||
request = urllib.request.Request(COMFY + path, data=body,
|
||||
headers=headers, method=method)
|
||||
with urllib.request.urlopen(request, timeout=timeout) as response:
|
||||
return json.load(response)
|
||||
|
||||
|
||||
def wait_for_comfy() -> None:
|
||||
for _ in range(300):
|
||||
if COMFY_PROCESS is not None and COMFY_PROCESS.poll() is not None:
|
||||
return
|
||||
try:
|
||||
request_json("/system_stats", timeout=2)
|
||||
READY.set()
|
||||
return
|
||||
except Exception:
|
||||
time.sleep(1)
|
||||
|
||||
|
||||
def start_comfy() -> None:
|
||||
global COMFY_PROCESS
|
||||
INPUT_DIR.mkdir(parents=True, exist_ok=True)
|
||||
OUTPUT_DIR.mkdir(parents=True, exist_ok=True)
|
||||
command = [
|
||||
"python", "main.py", "--listen", "127.0.0.1", "--port", "8188",
|
||||
"--lowvram", "--preview-method", "none",
|
||||
"--input-directory", str(INPUT_DIR),
|
||||
"--output-directory", str(OUTPUT_DIR),
|
||||
]
|
||||
COMFY_PROCESS = subprocess.Popen(command, cwd=COMFY_DIR)
|
||||
threading.Thread(target=wait_for_comfy, daemon=True).start()
|
||||
|
||||
|
||||
def stop(*_: object) -> None:
|
||||
if COMFY_PROCESS is not None and COMFY_PROCESS.poll() is None:
|
||||
COMFY_PROCESS.terminate()
|
||||
try:
|
||||
COMFY_PROCESS.wait(timeout=15)
|
||||
except subprocess.TimeoutExpired:
|
||||
COMFY_PROCESS.kill()
|
||||
raise SystemExit(0)
|
||||
|
||||
|
||||
def validate_request(data: dict) -> tuple[str, str, int, int, int, float, int]:
|
||||
prompt = data.get("prompt")
|
||||
filename = data.get("filename")
|
||||
if not isinstance(prompt, str) or not prompt.strip() or len(prompt) > 8000:
|
||||
raise ValueError("invalid prompt")
|
||||
if (not isinstance(filename, str) or Path(filename).name != filename
|
||||
or not filename.endswith(".png")):
|
||||
raise ValueError("invalid filename")
|
||||
width = int(data.get("width", 1024))
|
||||
height = int(data.get("height", 1024))
|
||||
if width % 32 or height % 32 or not 512 <= width <= 2048 or not 512 <= height <= 2048:
|
||||
raise ValueError("width and height must be multiples of 32 between 512 and 2048")
|
||||
steps = int(data.get("steps", 25))
|
||||
guidance = float(data.get("guidance", 1.0))
|
||||
if steps != 25 or guidance != 1.0:
|
||||
raise ValueError("Qwen-Image-2.1 requires steps=25 and guidance=1.0")
|
||||
seed = data.get("seed")
|
||||
seed = random.randrange(2**32) if seed is None else int(seed)
|
||||
if not 0 <= seed <= 2**32 - 1:
|
||||
raise ValueError("invalid seed")
|
||||
return prompt.strip(), filename, width, height, steps, guidance, seed
|
||||
|
||||
|
||||
def make_workflow(prompt: str, seed: int, steps: int, width: int, height: int,
|
||||
references: list[str], prefix: str) -> dict:
|
||||
encode_inputs: dict = {
|
||||
"clip": ["2", 0], "prompt": prompt, "negative_prompt": "",
|
||||
"resolution": max(width, height),
|
||||
}
|
||||
workflow: dict = {
|
||||
"1": {"class_type": "UNETLoader", "inputs": {
|
||||
"unet_name": "qwen_image_2.1_int8_convrot.safetensors",
|
||||
"weight_dtype": "default"}},
|
||||
"2": {"class_type": "CLIPLoader", "inputs": {
|
||||
"clip_name": "qwen3vl_8b_int8_convrot.safetensors",
|
||||
"type": "qwen_image", "device": "default"}},
|
||||
"3": {"class_type": "VAELoader", "inputs": {
|
||||
"vae_name": "qwen_image_2.1_vae_bf16.safetensors"}},
|
||||
"4": {"class_type": "TextEncodeQwenImage21", "inputs": encode_inputs},
|
||||
"6": {"class_type": "KSampler", "inputs": {
|
||||
"model": ["1", 0], "positive": ["4", 0], "negative": ["4", 1],
|
||||
"latent_image": ["4", 2] if references else ["5", 0],
|
||||
"seed": seed, "steps": steps, "cfg": 1.0,
|
||||
"sampler_name": "euler", "scheduler": "simple", "denoise": 1.0}},
|
||||
"7": {"class_type": "VAEDecode", "inputs": {
|
||||
"samples": ["6", 0], "vae": ["3", 0]}},
|
||||
"8": {"class_type": "SaveImage", "inputs": {
|
||||
"filename_prefix": prefix, "images": ["7", 0]}},
|
||||
}
|
||||
if references:
|
||||
encode_inputs["vae"] = ["3", 0]
|
||||
for index, name in enumerate(references, 1):
|
||||
node = str(8 + index)
|
||||
workflow[node] = {"class_type": "LoadImage", "inputs": {"image": name}}
|
||||
encode_inputs[f"image_{index}"] = [node, 0]
|
||||
else:
|
||||
workflow["5"] = {"class_type": "EmptyLatentImage", "inputs": {
|
||||
"width": width, "height": height, "batch_size": 1}}
|
||||
return workflow
|
||||
|
||||
|
||||
def wait_result(prompt_id: str, timeout: int = 1800) -> dict:
|
||||
deadline = time.monotonic() + timeout
|
||||
while time.monotonic() < deadline:
|
||||
history = request_json(f"/history/{prompt_id}", timeout=10)
|
||||
if prompt_id in history:
|
||||
result = history[prompt_id]
|
||||
status = result.get("status", {})
|
||||
if status.get("status_str") == "error" or not status.get("completed", True):
|
||||
raise RuntimeError("ComfyUI generation failed: " + json.dumps(status))
|
||||
return result
|
||||
time.sleep(1)
|
||||
raise TimeoutError("Qwen image generation timed out")
|
||||
|
||||
|
||||
def prepare_references(data: dict, job: str) -> list[str]:
|
||||
source_files = data.get("source_files") or []
|
||||
if not isinstance(source_files, list) or len(source_files) > 4:
|
||||
raise ValueError("invalid source image list")
|
||||
copied: list[str] = []
|
||||
for index, name in enumerate(source_files, 1):
|
||||
if not isinstance(name, str) or Path(name).name != name:
|
||||
raise ValueError("invalid source image filename")
|
||||
source = (OUTPUT_DIR / name).resolve()
|
||||
if source.parent != OUTPUT_DIR or not source.is_file():
|
||||
raise ValueError("source image not found")
|
||||
with source.open("rb") as stream:
|
||||
header = stream.read(16)
|
||||
if header.startswith(b"\x89PNG\r\n\x1a\n"):
|
||||
suffix = ".png"
|
||||
elif header.startswith(b"\xff\xd8\xff"):
|
||||
suffix = ".jpg"
|
||||
elif header.startswith((b"RIFF",)) and header[8:12] == b"WEBP":
|
||||
suffix = ".webp"
|
||||
else:
|
||||
raise ValueError("unsupported source image format")
|
||||
target_name = f"{job}-ref-{index}{suffix}"
|
||||
shutil.copyfile(source, INPUT_DIR / target_name)
|
||||
copied.append(target_name)
|
||||
return copied
|
||||
|
||||
|
||||
def generate(data: dict) -> dict:
|
||||
global ACTIVE
|
||||
with GENERATION_LOCK:
|
||||
ACTIVE = True
|
||||
started = time.monotonic()
|
||||
copied: list[str] = []
|
||||
try:
|
||||
prompt, filename, width, height, steps, _, seed = validate_request(data)
|
||||
job = "router-" + uuid.uuid4().hex
|
||||
copied = prepare_references(data, job)
|
||||
queued = request_json("/prompt", {"prompt": make_workflow(
|
||||
prompt, seed, steps, width, height, copied, job),
|
||||
"client_id": "athena-image-router"})
|
||||
result = wait_result(queued["prompt_id"])
|
||||
images = [image for output in result.get("outputs", {}).values()
|
||||
for image in output.get("images", [])]
|
||||
if not images:
|
||||
raise RuntimeError("generation completed without an image")
|
||||
image = images[0]
|
||||
generated = (OUTPUT_DIR / image.get("subfolder", "") /
|
||||
image["filename"]).resolve()
|
||||
if OUTPUT_DIR not in generated.parents or not generated.is_file():
|
||||
raise RuntimeError("ComfyUI returned an invalid output path")
|
||||
target = (OUTPUT_DIR / filename).resolve()
|
||||
if target.parent != OUTPUT_DIR:
|
||||
raise RuntimeError("invalid output path")
|
||||
generated.replace(target)
|
||||
return {"status": "ok", "filename": filename, "seed": seed,
|
||||
"seconds": round(time.monotonic() - started, 3),
|
||||
"model": MODEL}
|
||||
finally:
|
||||
for name in copied:
|
||||
try:
|
||||
(INPUT_DIR / name).unlink()
|
||||
except FileNotFoundError:
|
||||
pass
|
||||
ACTIVE = False
|
||||
|
||||
|
||||
class Handler(BaseHTTPRequestHandler):
|
||||
def log_message(self, fmt: str, *args: object) -> None:
|
||||
print(f"[qwen-image-2.1] {self.client_address[0]} {fmt % args}", flush=True)
|
||||
|
||||
def reply(self, status: int, payload: dict) -> None:
|
||||
body = json.dumps(payload, separators=(",", ":")).encode()
|
||||
self.send_response(status)
|
||||
self.send_header("Content-Type", "application/json")
|
||||
self.send_header("Content-Length", str(len(body)))
|
||||
self.end_headers()
|
||||
self.wfile.write(body)
|
||||
|
||||
def do_GET(self) -> None: # noqa: N802
|
||||
if self.path != "/health":
|
||||
self.reply(404, {"error": "not found"})
|
||||
elif not READY.is_set():
|
||||
self.reply(503, {"status": "starting", "model": MODEL})
|
||||
else:
|
||||
self.reply(200, {"status": "ok", "model_loaded": ACTIVE,
|
||||
"model": MODEL})
|
||||
|
||||
def do_POST(self) -> None: # noqa: N802
|
||||
if self.headers.get("Authorization", "") != f"Bearer {TOKEN}":
|
||||
self.reply(401, {"error": "unauthorized"})
|
||||
return
|
||||
if self.path != "/generate":
|
||||
self.reply(404, {"error": "not found"})
|
||||
return
|
||||
try:
|
||||
length = int(self.headers.get("Content-Length", "0"))
|
||||
if length < 2 or length > 32768:
|
||||
raise ValueError("invalid request size")
|
||||
self.reply(200, generate(json.loads(self.rfile.read(length))))
|
||||
except Exception as exc:
|
||||
print(f"[qwen-image-2.1] generation failed: {type(exc).__name__}: "
|
||||
f"{str(exc)[:1000]}", flush=True)
|
||||
self.reply(500, {"status": "error", "message": str(exc)})
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
signal.signal(signal.SIGTERM, stop)
|
||||
signal.signal(signal.SIGINT, stop)
|
||||
start_comfy()
|
||||
ThreadingHTTPServer((HOST, PORT), Handler).serve_forever()
|
||||
+14
-10
@@ -19,7 +19,7 @@ Kommandos: POST /fast, /medium, /large, /ultra,
|
||||
/uncensored
|
||||
GET /status (Zustand)
|
||||
|
||||
Bildgenerierung und Editing (FLUX.2 Klein 9B FP8 beta):
|
||||
Bildgenerierung und Editing (Qwen-Image-2.1 INT8):
|
||||
POST /v1/images/generations (OpenAI-kompatibel)
|
||||
POST /v1/images/edits (lokal, Referenzbilder)
|
||||
GET /images (Liste)
|
||||
@@ -39,8 +39,8 @@ Der Router leitet /v1/audio/speech und /v1/audio/transcriptions
|
||||
per HTTP an die Worker weiter.
|
||||
|
||||
Der Router agiert als Modell-Orchestrator: vor der Generierung wird
|
||||
llama.cpp und Qwen3-TTS gestoppt, der Bild-Worker lädt FLUX.2 und den
|
||||
Text-Encoder auf getrennte GPUs, generiert/bearbeitet und entlädt
|
||||
llama.cpp und Qwen3-TTS gestoppt, der Bild-Worker lädt Qwen-Image und den
|
||||
Textencoder auf die RTX 5080, generiert/bearbeitet und entlädt
|
||||
das Modell wieder; danach wird das vorherige Qwen-Profil wiederher-
|
||||
gestellt und erst dann geantwortet (try/finally – Qwen wird auch bei
|
||||
Fehlgeschlagener Generierung wiederhergestellt).
|
||||
@@ -137,7 +137,7 @@ DEFAULT_REASONING_EFFORT = os.environ.get(
|
||||
GLOBAL_SYSTEM_POLICY_FILE = os.environ.get(
|
||||
"GLOBAL_SYSTEM_POLICY_FILE", "").strip()
|
||||
|
||||
# --- Bildgenerierung und Referenzbild-Bearbeitung (FLUX.2 Klein 9B FP8) ---
|
||||
# --- Bildgenerierung und Referenzbild-Bearbeitung (Qwen-Image-2.1 INT8) ---
|
||||
LLAMA_SERVICE = os.environ.get("LLAMA_SERVICE", "mike-ai-llama-ui.service")
|
||||
SYSTEMCTL_BIN = os.environ.get("SYSTEMCTL_BIN", "systemctl")
|
||||
IMAGE_WORKER = os.environ.get(
|
||||
@@ -147,7 +147,8 @@ IMAGE_PYTHON = os.environ.get(
|
||||
IMAGE_WORKER_URL = os.environ.get("IMAGE_WORKER_URL", "").rstrip("/")
|
||||
IMAGE_WORKER_TOKEN = os.environ.get("IMAGE_WORKER_TOKEN", "").strip()
|
||||
IMAGE_MODEL_NAME = os.environ.get(
|
||||
"IMAGE_MODEL_NAME", "FLUX.2-klein-9B-fp8-beta")
|
||||
"IMAGE_MODEL_NAME", "Qwen-Image-2.1-int8")
|
||||
IMAGE_INFERENCE_STEPS = int(os.environ.get("IMAGE_INFERENCE_STEPS", "25"))
|
||||
IMAGE_DIR = os.environ.get(
|
||||
"IMAGE_DIR", "/opt/mike-ai/ai-profile-router/images")
|
||||
IMAGE_WORKER_LOG = os.environ.get(
|
||||
@@ -175,8 +176,9 @@ IMAGE_SIZES = {
|
||||
"1920x1088": (1920, 1088),
|
||||
"1088x1920": (1088, 1920),
|
||||
}
|
||||
# Das destillierte FLUX.2 Klein 9B ist auf vier Schritte ausgelegt.
|
||||
IMAGE_QUALITY = {"standard": 4, "high": 4}
|
||||
# Der produktive Qwen-Image-2.1-Workflow ist auf 25 Schritte festgelegt.
|
||||
IMAGE_QUALITY = {"standard": IMAGE_INFERENCE_STEPS,
|
||||
"high": IMAGE_INFERENCE_STEPS}
|
||||
IMAGE_DEFAULT_QUALITY = "standard"
|
||||
IMAGE_MAX_N = 4
|
||||
|
||||
@@ -1134,7 +1136,7 @@ def switch_profile(profile: str, implicit: bool = False) -> dict:
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Bildgenerierung und Editing (FLUX.2 Klein 9B FP8)
|
||||
# Bildgenerierung und Editing (Qwen-Image-2.1 INT8)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
class _Worker:
|
||||
@@ -2342,8 +2344,10 @@ class Handler(BaseHTTPRequestHandler):
|
||||
return
|
||||
steps = data.get("steps", IMAGE_QUALITY[quality])
|
||||
if (not isinstance(steps, int) or isinstance(steps, bool)
|
||||
or steps != 4):
|
||||
self._send_error(400, f"{IMAGE_MODEL_NAME} erfordert 'steps'=4",
|
||||
or steps != IMAGE_INFERENCE_STEPS):
|
||||
self._send_error(
|
||||
400, f"{IMAGE_MODEL_NAME} erfordert "
|
||||
f"'steps'={IMAGE_INFERENCE_STEPS}",
|
||||
"invalid_request_error", "invalid_steps")
|
||||
return
|
||||
guidance = data.get("guidance", 1.0)
|
||||
|
||||
@@ -1,11 +1,10 @@
|
||||
#!/usr/bin/env bash
|
||||
# Download pinned official weights, build the isolated worker and leave it stopped.
|
||||
# Download pinned official weights and prepare the production worker stopped.
|
||||
set -Eeuo pipefail
|
||||
|
||||
ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
ENV_FILE=${STACK_ENV:-/etc/mike-ai/stack.env}
|
||||
MODEL_DIR=${QWEN_IMAGE_21_MODEL_DIR:-/data/models/Qwen-Image-2.1-ComfyUI}
|
||||
OUTPUT_DIR=${QWEN_IMAGE_21_OUTPUT_DIR:-/data/qwen-image-2.1-test-output}
|
||||
BASE=https://huggingface.co/Comfy-Org/Qwen-Image-2.1/resolve/ace0edeb3791a594ddfa36ed5f41a178a394e921
|
||||
|
||||
download() {
|
||||
@@ -23,14 +22,12 @@ download() {
|
||||
download diffusion_models/qwen_image_2.1_int8_convrot.safetensors cb74113cb03faecd79611b01fd7fd642f0aa60d6f0b95086abee214d75eaa57d
|
||||
download text_encoders/qwen3vl_8b_int8_convrot.safetensors 8bfd0f6e12abf2d2d697ecc888e5e90b0d6741d6708f05799f53afa560452e8f
|
||||
download vae/qwen_image_2.1_vae_bf16.safetensors bb21f7473051e1ac368515dd3f2e15cd44d7a11748ee8823e1ddca3e4876b7c9
|
||||
install -d -m 0755 "$OUTPUT_DIR"
|
||||
|
||||
compose=(docker compose --env-file "$ENV_FILE" -f "$ROOT_DIR/compose.yaml" --profile qwen-image-test)
|
||||
"${compose[@]}" pull qwen-image-21-test-runner
|
||||
"${compose[@]}" build qwen-image-21-test
|
||||
compose=(docker compose --env-file "$ENV_FILE" -f "$ROOT_DIR/compose.yaml" --profile image --profile flux-standby)
|
||||
"${compose[@]}" build image-worker
|
||||
"${compose[@]}" up -d --build --no-deps profile-controller
|
||||
"${compose[@]}" create qwen-image-21-test
|
||||
"${compose[@]}" stop qwen-image-21-test
|
||||
state=$(docker inspect -f '{{.State.Status}}' mike-ai-qwen-image-2.1-test)
|
||||
"${compose[@]}" create image-worker flux-image-worker
|
||||
"${compose[@]}" stop image-worker flux-image-worker
|
||||
state=$(docker inspect -f '{{.State.Status}}' mike-ai-image-worker)
|
||||
[[ $state == exited || $state == created ]]
|
||||
echo "Qwen-Image-2.1 test is prepared and stopped. No GPU model was loaded."
|
||||
echo "Qwen-Image-2.1 production worker and FLUX standby are prepared and stopped."
|
||||
@@ -1,136 +0,0 @@
|
||||
#!/usr/bin/env python3
|
||||
"""One-shot Qwen-Image-2.1 evaluation with guaranteed profile restoration."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import pathlib
|
||||
import random
|
||||
import time
|
||||
import urllib.parse
|
||||
import urllib.request
|
||||
|
||||
|
||||
CONTROLLER = os.environ.get("CONTROLLER_URL", "http://profile-controller:8090")
|
||||
TOKEN = os.environ["CONTROLLER_TOKEN"]
|
||||
COMFY = os.environ.get("COMFY_URL", "http://qwen-image-21-test:8188")
|
||||
OUTPUT = pathlib.Path(os.environ.get("OUTPUT_DIR", "/output"))
|
||||
|
||||
|
||||
def request_json(url: str, *, payload: dict | None = None,
|
||||
authenticated: bool = False, timeout: float = 30) -> dict:
|
||||
body = None if payload is None else json.dumps(payload).encode()
|
||||
headers = {"Content-Type": "application/json"}
|
||||
if authenticated:
|
||||
headers["Authorization"] = f"Bearer {TOKEN}"
|
||||
method = "POST" if payload is not None else "GET"
|
||||
req = urllib.request.Request(url, data=body, headers=headers, method=method)
|
||||
with urllib.request.urlopen(req, timeout=timeout) as response:
|
||||
return json.load(response)
|
||||
|
||||
|
||||
def controller(path: str, *, post: bool = False) -> dict:
|
||||
return request_json(CONTROLLER + path, payload={} if post else None,
|
||||
authenticated=True, timeout=900)
|
||||
|
||||
|
||||
def wait_comfy(timeout: int = 900) -> None:
|
||||
deadline = time.monotonic() + timeout
|
||||
while time.monotonic() < deadline:
|
||||
try:
|
||||
request_json(COMFY + "/system_stats", timeout=3)
|
||||
return
|
||||
except Exception:
|
||||
time.sleep(2)
|
||||
raise TimeoutError("Qwen-Image worker did not become ready")
|
||||
|
||||
|
||||
def workflow(prompt: str, seed: int, steps: int, width: int, height: int) -> dict:
|
||||
return {
|
||||
"1": {"class_type": "UNETLoader", "inputs": {
|
||||
"unet_name": "qwen_image_2.1_int8_convrot.safetensors",
|
||||
"weight_dtype": "default"}},
|
||||
"2": {"class_type": "CLIPLoader", "inputs": {
|
||||
"clip_name": "qwen3vl_8b_int8_convrot.safetensors",
|
||||
"type": "qwen_image", "device": "default"}},
|
||||
"3": {"class_type": "VAELoader", "inputs": {
|
||||
"vae_name": "qwen_image_2.1_vae_bf16.safetensors"}},
|
||||
"4": {"class_type": "TextEncodeQwenImage21", "inputs": {
|
||||
"clip": ["2", 0], "prompt": prompt, "negative_prompt": "",
|
||||
"resolution": max(width, height)}},
|
||||
"5": {"class_type": "EmptyLatentImage", "inputs": {
|
||||
"width": width, "height": height, "batch_size": 1}},
|
||||
"6": {"class_type": "KSampler", "inputs": {
|
||||
"model": ["1", 0], "positive": ["4", 0], "negative": ["4", 1],
|
||||
"latent_image": ["5", 0], "seed": seed, "steps": steps,
|
||||
"cfg": 1.0, "sampler_name": "euler", "scheduler": "simple",
|
||||
"denoise": 1.0}},
|
||||
"7": {"class_type": "VAEDecode", "inputs": {
|
||||
"samples": ["6", 0], "vae": ["3", 0]}},
|
||||
"8": {"class_type": "SaveImage", "inputs": {
|
||||
"filename_prefix": "qwen-image-2.1-test", "images": ["7", 0]}},
|
||||
}
|
||||
|
||||
|
||||
def wait_result(prompt_id: str, timeout: int = 3600) -> dict:
|
||||
deadline = time.monotonic() + timeout
|
||||
while time.monotonic() < deadline:
|
||||
history = request_json(f"{COMFY}/history/{prompt_id}", timeout=10)
|
||||
if prompt_id in history:
|
||||
result = history[prompt_id]
|
||||
status = result.get("status", {})
|
||||
if status.get("status_str") == "error" or not status.get("completed", True):
|
||||
raise RuntimeError("generation failed: " + json.dumps(status))
|
||||
return result
|
||||
time.sleep(2)
|
||||
raise TimeoutError("Qwen-Image generation timed out")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("prompt")
|
||||
parser.add_argument("--seed", type=int, default=None)
|
||||
parser.add_argument("--steps", type=int, default=25)
|
||||
parser.add_argument("--width", type=int, default=1024)
|
||||
parser.add_argument("--height", type=int, default=1024)
|
||||
args = parser.parse_args()
|
||||
if args.width % 32 or args.height % 32:
|
||||
parser.error("width and height must be multiples of 32")
|
||||
seed = args.seed if args.seed is not None else random.randrange(2**53)
|
||||
previous = controller("/profiles/status").get("active_profile")
|
||||
started = time.monotonic()
|
||||
try:
|
||||
controller("/workers/qwen-image-test/start", post=True)
|
||||
wait_comfy()
|
||||
queued = request_json(COMFY + "/prompt", payload={
|
||||
"prompt": workflow(args.prompt, seed, args.steps, args.width, args.height),
|
||||
"client_id": "athena-qwen-image-21-test"}, timeout=30)
|
||||
prompt_id = queued["prompt_id"]
|
||||
result = wait_result(prompt_id)
|
||||
images = []
|
||||
for node in result.get("outputs", {}).values():
|
||||
images.extend(node.get("images", []))
|
||||
if not images:
|
||||
raise RuntimeError("generation completed without an image")
|
||||
image = images[0]
|
||||
query = urllib.parse.urlencode({
|
||||
"filename": image["filename"], "subfolder": image.get("subfolder", ""),
|
||||
"type": image.get("type", "output")})
|
||||
target = OUTPUT / image["filename"]
|
||||
target.parent.mkdir(parents=True, exist_ok=True)
|
||||
with urllib.request.urlopen(COMFY + "/view?" + query, timeout=120) as src:
|
||||
target.write_bytes(src.read())
|
||||
print(json.dumps({"status": "ok", "file": str(target), "seed": seed,
|
||||
"seconds": round(time.monotonic() - started, 1)}, indent=2))
|
||||
finally:
|
||||
try:
|
||||
controller("/workers/qwen-image-test/stop", post=True)
|
||||
finally:
|
||||
if previous:
|
||||
controller(f"/profiles/{previous}/activate", post=True)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user