Add final Qwen3.8 runtime and uncensored profile

This commit is contained in:
Mikei386
2026-08-23 00:10:53 +02:00
parent 1a9d94d22f
commit 2ef894160a
57 changed files with 10089 additions and 54 deletions
+5 -3
View File
@@ -26,7 +26,7 @@ Zielplattform.
|---|---|
| Runtime | llama.cpp |
| Repository | `https://github.com/ggml-org/llama.cpp.git` |
| Commit | `4df29be4f4c3673f428170fda944a5b19f743bb8` |
| Commit | `3f545beccee69d9975f466ec7e45fd9aacd8ba90` |
| Compiler | GCC 14.2 |
| Hauptdienst | jeweils ein Container `mike-ai-llama-<profil>` |
| llama.cpp-Port | 8080, ausschließlich im internen Docker-Netz |
@@ -44,6 +44,7 @@ Zielplattform.
- Flash Attention
- KV-Cache Q4_0 für K und V
- MTP Draft, maximal drei Tokens
- MTP-Akzeptanzschwelle 0,05; im Referenzlauf 77,26 statt 73,88 Tok/s
- sechs Threads und sechs Batch-Threads
- Batch 64, Micro-Batch 32
- ein paralleler Slot
@@ -58,6 +59,7 @@ Zielplattform.
| Medium **(Standard)** | `qwen-medium` | 160.000 | IQ4_XS Pure, MTP3, beide GPUs 90:10, mmproj auf RTX 3060 |
| Large | `qwen-large` | 192.000 | IQ4_XS Pure, MTP3, beide GPUs 86:14, mmproj auf RTX 3060 |
| Ultra | `qwen-ultra` | 262.144 | IQ4_XS Pure, MTP2, beide GPUs 80:20, text-only; 68,2 Tok/s und 220K-Fülltest bestanden |
| Uncensored | `qwen-uncensored` | 80.000 | Abliterated Q4_K_M, MTP2, beide GPUs 90:10, eigener mmproj auf RTX 3060; etwa 52,2 Tok/s |
## Router
@@ -73,7 +75,7 @@ Der Router übernimmt:
- OpenAI-kompatibles Chat-Proxying und Streaming
- virtuelle Modelle und automatische Profilumschaltung
- Tool Calls
- direkte integrierte Vision in Fast, Medium und Large
- direkte integrierte Vision in Fast, Medium, Large und Uncensored
- FLUX-Hotswap zur Bildgenerierung
- Whisper Speech-to-Text
- XTTS Text-to-Speech (historische Referenz; Zielsystem verwendet Piper)
@@ -88,7 +90,7 @@ Der Router übernimmt:
| Speicherort des Projektors | RTX 3060 (`MTMD_BACKEND_DEVICE=CUDA1`) |
| Kontext | entspricht Fast/Medium/Large; Ultra ist bewusst text-only |
Vision ist Bestandteil von Fast, Medium und Large. Der Router prüft Bildgröße und URL,
Vision ist Bestandteil von Fast, Medium, Large und Uncensored. Der Router prüft Bildgröße und URL,
leitet das Bild dann direkt weiter und führt keinen Modellwechsel mehr aus.
## Bildgenerierung