Add RTX 5080 FLUX hot-swap worker

This commit is contained in:
Mikei386
2026-08-22 17:40:32 +02:00
parent 8267a85a96
commit 7ac93befc4
12 changed files with 398 additions and 41 deletions
+10 -10
View File
@@ -54,9 +54,9 @@ Zielplattform.
| Profil | Virtuelles Modell | Kontext | Besonderheit |
|---|---|---:|---|
| Fast | `qwen-fast` | 76.800 | IQ4-MIX, MTP2, vollständig GPU, CPU-mmproj |
| Medium **(Standard)** | `qwen-medium` | 160.000 | IQ4_XS Pure, MTP3, beide GPUs 90:10, CPU-mmproj |
| Large | `qwen-large` | 192.000 | IQ4_XS Pure, MTP3, beide GPUs 86:14, CPU-mmproj |
| Fast | `qwen-fast` | 76.800 | IQ4-MIX, MTP2, Text auf RTX 5080, mmproj auf RTX 3060 |
| Medium **(Standard)** | `qwen-medium` | 160.000 | IQ4_XS Pure, MTP3, beide GPUs 90:10, mmproj auf RTX 3060 |
| Large | `qwen-large` | 192.000 | IQ4_XS Pure, MTP3, beide GPUs 86:14, mmproj auf RTX 3060 |
| Ultra | `qwen-ultra` | 262.144 | IQ4_XS Pure, MTP2, beide GPUs 80:20, text-only; 68,2 Tok/s und 220K-Fülltest bestanden |
## Router
@@ -85,7 +85,7 @@ Der Router übernimmt:
|---|---|
| Text-/Visionmodell | jeweils aktives Qwen3.8-27B-Profil |
| Projektor | BF16-mmproj |
| Speicherort des Projektors | System-RAM (`--no-mmproj-offload`) |
| Speicherort des Projektors | RTX 3060 (`MTMD_BACKEND_DEVICE=CUDA1`) |
| Kontext | entspricht Fast/Medium/Large; Ultra ist bewusst text-only |
Vision ist Bestandteil von Fast, Medium und Large. Der Router prüft Bildgröße und URL,
@@ -93,13 +93,13 @@ leitet das Bild dann direkt weiter und führt keinen Modellwechsel mehr aus.
## Bildgenerierung
- Modell: FLUX.2 klein Base 4B
- Runtime: PyTorch/Diffusers
- CPU-Offload aktiviert
- Standard: 30 Schritte
- High: 50 Schritte
- Modell: FLUX.2 Klein 4B Distilled, Apache-2.0
- Runtime: eigener PyTorch-2.11/CUDA-12.8-/Diffusers-0.40-Container
- fest auf vier Schritte und Guidance 1,0 destilliert
- Worker läuft ausschließlich auf der RTX 5080 und ist im Normalbetrieb gestoppt
- der Controller beendet Qwen vor dem Job; Bildprompts bleiben im internen Netz
- Worker wird nach jedem Job vollständig beendet
- Qwen wird anschließend mit dem vorherigen Profil wiederhergestellt
- Qwen wird anschließend mit exakt dem vorherigen Profil wiederhergestellt
## Sprache
+3 -3
View File
@@ -6,9 +6,9 @@ und einer bewussten Aktualisierung dieser Datei.
| Profil | Virtuelles Modell | GGUF | Kontext | GPUs / Split | MTP | Vision | gemessene kurze Ausgabe |
|---|---|---|---:|---|---:|---|---:|
| Fast | `qwen-fast` | IQ4-MIX | 76.800 | RTX 5080 | 2 | ja, Projektor auf CPU | 85,5 Tok/s |
| **Medium (Default)** | `qwen-medium` | IQ4_XS Pure | 160.000 | RTX 5080 + RTX 3060, 90:10 | 3 | ja, Projektor auf CPU | 77,2 Tok/s |
| Large | `qwen-large` | IQ4_XS Pure | 192.000 | RTX 5080 + RTX 3060, 86:14 | 3 | ja, Projektor auf CPU | 75,3 Tok/s |
| Fast | `qwen-fast` | IQ4-MIX | 76.800 | RTX 5080 | 2 | ja, Projektor auf RTX 3060 | 85,5 Tok/s |
| **Medium (Default)** | `qwen-medium` | IQ4_XS Pure | 160.000 | RTX 5080 + RTX 3060, 90:10 | 3 | ja, Projektor auf RTX 3060 | 77,2 Tok/s |
| Large | `qwen-large` | IQ4_XS Pure | 192.000 | RTX 5080 + RTX 3060, 86:14 | 3 | ja, Projektor auf RTX 3060 | 75,3 Tok/s |
| Ultra | `qwen-ultra` | IQ4_XS Pure | 262.144 | RTX 5080 + RTX 3060, 80:20 | 2 | nein, text-only | 68,2 Tok/s |
## Standardverhalten