Use validated Medium MTP2 and microbatch256 combination
This commit is contained in:
+2
-2
@@ -206,7 +206,7 @@ services:
|
||||
- --batch-size
|
||||
- "${MEDIUM_BATCH_SIZE:-2048}"
|
||||
- --ubatch-size
|
||||
- "${MEDIUM_UBATCH_SIZE:-512}"
|
||||
- "${MEDIUM_UBATCH_SIZE:-256}"
|
||||
- --parallel
|
||||
- "${MEDIUM_PARALLEL_SLOTS:-1}"
|
||||
- --kv-unified
|
||||
@@ -247,7 +247,7 @@ services:
|
||||
- --spec-type
|
||||
- draft-mtp
|
||||
- --spec-draft-n-max
|
||||
- "3"
|
||||
- "2"
|
||||
- --spec-draft-type-k
|
||||
- f16
|
||||
- --spec-draft-type-v
|
||||
|
||||
@@ -24,7 +24,7 @@
|
||||
"model_family": "Qwen3.8-27B IQ4 XS Pure",
|
||||
"gpu_split": "85:15",
|
||||
"vision": true,
|
||||
"mtp": 3,
|
||||
"mtp": 2,
|
||||
"description": "Ausgewogenes Standardprofil für Alltag und lange agentische Aufgaben."
|
||||
},
|
||||
{
|
||||
|
||||
@@ -7,7 +7,7 @@ Standardprofil: **medium** · globales Ausgabelimit: **8192 Token**
|
||||
| Profil | API-Alias | Gesamtkontext | Slots | Modell | GPU-Verteilung | Vision | MTP |
|
||||
|---|---|---:|---:|---|---|---|---:|
|
||||
| fast | `qwen-fast` | 76,800 | 1 | Qwen3.8-27B IQ4 Mix | 5080 only | ja | 2 |
|
||||
| medium | `qwen-medium` | 160,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 85:15 | ja | 3 |
|
||||
| medium | `qwen-medium` | 160,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 85:15 | ja | 2 |
|
||||
| large | `qwen-large` | 192,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 86:14 | ja | 2 |
|
||||
| ultra | `qwen-ultra` | 262,144 | 1 | Qwen3.8-27B IQ4 XS Pure | 80:20 | nein | 2 |
|
||||
| uncensored | `qwen-uncensored` | 80,000 | 1 | Qwen3.8-27B Abliterated Q4_K_M | 90:10 | ja | 2 |
|
||||
|
||||
Reference in New Issue
Block a user