Replace vision hotswap with native multimodal profiles

This commit is contained in:
Mikei386
2026-08-20 14:24:56 +02:00
parent 8a45a0d851
commit cb07779f5a
19 changed files with 169 additions and 719 deletions
+10 -12
View File
@@ -36,7 +36,7 @@ Zielplattform.
### Aktives Fast-Profil
- Qwen3.8-27B IQ4-MIX
- Kontext 73.728
- Kontext 76.800
- vollständig auf CUDA0
- Flash Attention
- KV-Cache Q4_0 für K und V
@@ -51,9 +51,9 @@ Zielplattform.
| Profil | Virtuelles Modell | Kontext | Besonderheit |
|---|---|---:|---|
| Fast | `qwen-fast` | 73.728 | IQ4-MIX, MTP2, vollständig GPU |
| Medium | `qwen-medium` | 94.208 | IQ4_XS Pure, ohne MTP |
| Long | `qwen-long` | 131.072 | IQ4-MIX, MTP2, FFN-Blöcke 0–11 auf CPU |
| Fast | `qwen-fast` | 76.800 | IQ4-MIX, MTP2, vollständig GPU, CPU-mmproj |
| Medium | `qwen-medium` | 94.208 | IQ4_XS Pure, ohne MTP, CPU-mmproj |
| Long | `qwen-long` | 131.072 | IQ4-MIX, MTP2, FFN 0–11 auf CPU, CPU-mmproj |
## Router
@@ -69,7 +69,7 @@ Der Router übernimmt:
- OpenAI-kompatibles Chat-Proxying und Streaming
- virtuelle Modelle und automatische Profilumschaltung
- Tool Calls
- Vision-Hotswap mit Cache für Folgefragen
- direkte integrierte Vision in allen Qwen-Profilen
- FLUX-Hotswap zur Bildgenerierung
- Whisper Speech-to-Text
- XTTS Text-to-Speech
@@ -79,15 +79,13 @@ Der Router übernimmt:
| Bereich | Referenz |
|---|---|
| Text-/Visionmodell | Qwen3.8-27B Q3_K_M |
| Text-/Visionmodell | jeweils aktives Qwen3.8-27B-Profil |
| Projektor | BF16-mmproj |
| Kontext | 32.768 |
| temporärer Port | 8086 |
| maximale Ausgabe | 4.096 Tokens |
| Speicherort des Projektors | System-RAM (`--no-mmproj-offload`) |
| Kontext | entspricht Fast/Medium/Long |
Vision wird nicht dauerhaft parallel geladen. Der Router entlädt das
Textprofil, analysiert das Bild und stellt danach das ursprüngliche Profil
wieder her.
Vision ist Bestandteil jedes Profils. Der Router prüft Bildgröße und URL,
leitet das Bild dann direkt weiter und führt keinen Modellwechsel mehr aus.
## Bildgenerierung