Enable benchmarked two-slot medium profile
This commit is contained in:
1 parent
5b78d12112
commit
014583e7f6
7 files changed
+60
-21
No files matched your search
@@ -19,4 +19,17 @@ Full-context results with the selected settings:
|
||||
| ultra | 257,998 | 698.2 tok/s | 21.1 tok/s | no | not configured for this profile |
|
||||
| uncensored | 77,998 | 1,050.5 tok/s | 40.0 tok/s | no | passed |
|
||||
|
||||
The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. Moving Medium from a 90:10 to an 89:11 GPU split freed memory but reduced both prefill and output speed. The production setting therefore remains 2048 / 128 at 90:10.
|
||||
The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. The original one-slot benchmark selected 2048 / 128 at 90:10.
|
||||
|
||||
## Medium two-slot benchmark
|
||||
|
||||
Medium now uses two parallel slots with unified KV, so both chats dynamically share one total 160K-token pool. The model weights remain loaded only once. To fit the additional scheduler buffers, the production GPU split is 85:15 while batch / ubatch remains 2048 / 128.
|
||||
|
||||
Identical fresh 100,297-token prompt with a deterministic 256-token completion:
|
||||
|
||||
| Mode | Slots | GPU split | Prefill | Output | Total time |
|
||||
|---|---:|---:|---:|---:|---:|
|
||||
| Previous | 1 | 90:10 | 1,072.12 tok/s | 48.51 tok/s | 98.81 s |
|
||||
| Selected | 2 | 85:15 | 993.20 tok/s | 47.65 tok/s | 106.33 s |
|
||||
|
||||
Single-request cost: **7.4% lower prefill**, **1.8% lower output**, and **7.6% longer total time** for this near-full prompt. Two simultaneous fresh 10K prompts completed in 20 seconds; per-slot output measured 43.32 and 27.95 tok/s. Two very large simultaneous prefills can temporarily throttle an already-generating slot, so the total 160K pool should not be treated as two independent 160K contexts.
|
||||
Reference in new issue
Block a user