2.7 KiB
Qwen prefill batch benchmark (2026-08-30)
Identical 24K-token prompt, one slot, Q4 KV cache and the production GPU split of each profile. The output measurement uses the same deterministic 64-token completion. Full-context checks were run for every changed profile.
| Profile | Context | Previous batch / ubatch | Previous prefill | Previous output | Selected batch / ubatch | Selected prefill | Selected output | Result |
|---|---|---|---|---|---|---|---|---|
| fast | 76,800 | 64 / 32 | 907.2 tok/s | 75.4 tok/s | 64 / 32 | unchanged | unchanged | 2048 / 128 failed with CUDA OOM |
| medium | 160,000 | 64 / 32 | about 650–710 tok/s | about 65–70 tok/s | 2048 / 128 | 1,585.3 tok/s | 69.6 tok/s | selected; 155K plus vision passed |
| large | 192,000 | 64 / 32 | 730.6 tok/s | 78.4 tok/s | 2048 / 128 | 1,399.7 tok/s | 76.0 tok/s | selected; 190K plus vision passed |
| ultra | 262,144 | 64 / 32 | 672.0 tok/s | 57.1 tok/s | 2048 / 128 | 1,649.6 tok/s | 52.7 tok/s | selected; 258K passed |
| uncensored | 80,000 | 64 / 32 | 729.7 tok/s | 50.8 tok/s | 2048 / 128 | 1,331.8 tok/s | 49.9 tok/s | selected; 78K plus vision passed |
Full-context results with the selected settings:
| Profile | Tested prompt | Prefill | Output | Truncated | Vision after test |
|---|---|---|---|---|---|
| medium | 154,998 | 881.3 tok/s | 37.8 tok/s | no | passed |
| large | 189,998 | 738.0 tok/s | 27.6 tok/s | no | passed |
| ultra | 257,998 | 698.2 tok/s | 21.1 tok/s | no | not configured for this profile |
| uncensored | 77,998 | 1,050.5 tok/s | 40.0 tok/s | no | passed |
The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. The original one-slot benchmark selected 2048 / 128 at 90:10.
Medium two-slot benchmark
Medium now uses two parallel slots with unified KV, so both chats dynamically share one total 160K-token pool. The model weights remain loaded only once. To fit the additional scheduler buffers, the production GPU split is 85:15 while batch / ubatch remains 2048 / 128.
Identical fresh 100,297-token prompt with a deterministic 256-token completion:
| Mode | Slots | GPU split | Prefill | Output | Total time |
|---|---|---|---|---|---|
| Previous | 1 | 90:10 | 1,072.12 tok/s | 48.51 tok/s | 98.81 s |
| Selected | 2 | 85:15 | 993.20 tok/s | 47.65 tok/s | 106.33 s |
Single-request cost: 7.4% lower prefill, 1.8% lower output, and 7.6% longer total time for this near-full prompt. Two simultaneous fresh 10K prompts completed in 20 seconds; per-slot output measured 43.32 and 27.95 tok/s. Two very large simultaneous prefills can temporarily throttle an already-generating slot, so the total 160K pool should not be treated as two independent 160K contexts.