router: Bildgenerierung mit FLUX.2 [klein] 4B Base (GPU-Hotswap)

- POST /v1/images/generations (OpenAI-kompatibel, prompt/size/n/seed/quality)
- quality: standard=30 Steps (Default), high=50 Steps
- Größen: 1024x1024, 1536x1024, 1024x1536, 1920x1088, 1088x1920
- GPU-Hotswap: Qwen stoppen -> FLUX laden -> Bild -> FLUX entladen -> Qwen
  wiederherstellen (exakt vorheriges Profil)
- Zentrales GPU/Modell-Lock (Profilwechsel und Bild teilen sich das Lock)
- Chat-Requests warten während Bild-Job (kein 502), Timeout CHAT_WAIT_TIMEOUT
- Robuste Recovery: try/finally, Worker-Beendigung, VRAM-Check, Qwen-Readiness
- /status: image.phase, image.worker, image.model_loaded, qwen.available,
  qwen.active_chats
- GET /images, GET /images/<datei> (validiert, nur images/-Verzeichnis)
- image_worker.py: FLUX-Worker (eigener Prozess, JSON-Protokoll, bf16 +
  enable_model_cpu_offload)
- deploy: venv (torch/diffusers/transformers/accelerate), Modell-Download,
  Image-Dir, systemd-Unit mit Image-Umgebungsvariablen
- dev: Mock-Worker, fake-systemctl, Benchmarks (GPU-Resident, Offload, Steps,
  Quality-Compare), 32 lokale Tests
- README: Bildgenerierung, Hotswap, Recovery, Benchmarks (RTX 5080),
  Python-Pakete

Benchmarks (RTX 5080, 16 GB, CPU-Offload):
- 512x512 / 10 Steps: ~9.3 s
- 1024x1024 / 30 Steps: ~31.3 s
- 1024x1024 / 50 Steps: ~45.3 s
- 1920x1088 / 50 Steps: ~91 s
- Peak-VRAM: ~8.4-8.9 GB
- Hotswap-Gesamtzeit: ~41-42 s (1024x1024 / 30 Steps)
This commit is contained in:
Mikei386
2026-08-19 08:51:37 +02:00
parent c5d92acd93
commit 7c5bbe2ffb
15 changed files with 1796 additions and 68 deletions
+66
View File
@@ -0,0 +1,66 @@
#!/usr/bin/env python3
"""FLUX.2 [klein] 4B Base – Steps-Vergleich (Qualität vs. Latenz).
1024x1024, fester Seed, cpu_offload. Vergleicht 20/30/40/50 Steps.
"""
import os
import subprocess
import time
os.environ["HF_HUB_DISABLE_PROGRESS_BARS"] = "1"
os.environ["TRANSFORMERS_VERBOSITY"] = "error"
os.environ["TOKENIZERS_PARALLELISM"] = "false"
MODEL = "/opt/mike-ai/models/FLUX.2-klein-base-4B"
PROMPT = ("A detailed photograph of a red cube on a white marble table, "
"soft studio lighting, shallow depth of field")
SEED = 0
WIDTH = HEIGHT = 1024
STEPS_LIST = [20, 30, 40, 50]
def nvidia_vram() -> int:
out = subprocess.check_output(
["nvidia-smi", "--query-gpu=memory.used",
"--format=csv,noheader,nounits"]).decode().strip()
return int(out.split()[0])
def main() -> None:
import torch
print(f"torch {torch.__version__} | {torch.cuda.get_device_name(0)}",
flush=True)
from diffusers import Flux2KleinPipeline
t0 = time.monotonic()
pipe = Flux2KleinPipeline.from_pretrained(MODEL, torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload()
print(f"[load] ready in {time.monotonic() - t0:.1f} s", flush=True)
for steps in STEPS_LIST:
out = f"/tmp/flux-steps-{steps}.png"
torch.cuda.synchronize()
torch.cuda.reset_peak_memory_stats()
t = time.monotonic()
try:
img = pipe(
prompt=PROMPT,
height=HEIGHT,
width=WIDTH,
guidance_scale=4.0,
num_inference_steps=steps,
generator=torch.Generator(device="cuda").manual_seed(SEED),
).images[0]
dt = time.monotonic() - t
img.save(out)
peak = torch.cuda.max_memory_allocated() / 1e9
print(f"[gen] {steps} steps: {dt:.1f} s | peak {peak:.2f} GB | "
f"nvidia-smi {nvidia_vram()} MiB | {out}", flush=True)
except Exception as e: # noqa: BLE001
print(f"[gen] {steps} steps: FEHLER: {e!r}", flush=True)
print("DONE", flush=True)
if __name__ == "__main__":
main()