Enable benchmarked two-slot medium profile

This commit is contained in:
Mikei386
2026-08-30 20:54:22 +02:00
parent db3719af7a
commit a41d6dc28c
7 changed files with 60 additions and 21 deletions
+14 -1
View File
@@ -19,4 +19,17 @@ Full-context results with the selected settings:
| ultra | 257,998 | 698.2 tok/s | 21.1 tok/s | no | not configured for this profile |
| uncensored | 77,998 | 1,050.5 tok/s | 40.0 tok/s | no | passed |
The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. Moving Medium from a 90:10 to an 89:11 GPU split freed memory but reduced both prefill and output speed. The production setting therefore remains 2048 / 128 at 90:10.
The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. The original one-slot benchmark selected 2048 / 128 at 90:10.
## Medium two-slot benchmark
Medium now uses two parallel slots with unified KV, so both chats dynamically share one total 160K-token pool. The model weights remain loaded only once. To fit the additional scheduler buffers, the production GPU split is 85:15 while batch / ubatch remains 2048 / 128.
Identical fresh 100,297-token prompt with a deterministic 256-token completion:
| Mode | Slots | GPU split | Prefill | Output | Total time |
|---|---:|---:|---:|---:|---:|
| Previous | 1 | 90:10 | 1,072.12 tok/s | 48.51 tok/s | 98.81 s |
| Selected | 2 | 85:15 | 993.20 tok/s | 47.65 tok/s | 106.33 s |
Single-request cost: **7.4% lower prefill**, **1.8% lower output**, and **7.6% longer total time** for this near-full prompt. Two simultaneous fresh 10K prompts completed in 20 seconds; per-slot output measured 43.32 and 27.95 tok/s. Two very large simultaneous prefills can temporarily throttle an already-generating slot, so the total 160K pool should not be treated as two independent 160K contexts.
+7 -7
View File
@@ -4,13 +4,13 @@ Diese Datei wird aus `config/profile-matrix.json` erzeugt. Änderungen gehören
Standardprofil: **medium** · globales Ausgabelimit: **8192 Token**
| Profil | API-Alias | Kontext | Modell | GPU-Verteilung | Vision | MTP |
|---|---|---:|---|---|---|---:|
| fast | `qwen-fast` | 76,800 | Qwen3.8-27B IQ4 Mix | 5080 only | ja | 2 |
| medium | `qwen-medium` | 160,000 | Qwen3.8-27B IQ4 XS Pure | 90:10 | ja | 3 |
| large | `qwen-large` | 192,000 | Qwen3.8-27B IQ4 XS Pure | 86:14 | ja | 3 |
| ultra | `qwen-ultra` | 262,144 | Qwen3.8-27B IQ4 XS Pure | 80:20 | nein | 2 |
| uncensored | `qwen-uncensored` | 80,000 | Qwen3.8-27B Abliterated Q4_K_M | 90:10 | ja | 2 |
| Profil | API-Alias | Gesamtkontext | Slots | Modell | GPU-Verteilung | Vision | MTP |
|---|---|---:|---:|---|---|---|---:|
| fast | `qwen-fast` | 76,800 | 1 | Qwen3.8-27B IQ4 Mix | 5080 only | ja | 2 |
| medium | `qwen-medium` | 160,000 | 2 | Qwen3.8-27B IQ4 XS Pure | 85:15 | ja | 3 |
| large | `qwen-large` | 192,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 86:14 | ja | 3 |
| ultra | `qwen-ultra` | 262,144 | 1 | Qwen3.8-27B IQ4 XS Pure | 80:20 | nein | 2 |
| uncensored | `qwen-uncensored` | 80,000 | 1 | Qwen3.8-27B Abliterated Q4_K_M | 90:10 | ja | 2 |
## Zweck