23 lines
1.8 KiB
Markdown
23 lines
1.8 KiB
Markdown
# Qwen prefill batch benchmark (2026-08-30)
|
||
|
||
Identical 24K-token prompt, one slot, Q4 KV cache and the production GPU split of each profile. The output measurement uses the same deterministic 64-token completion. Full-context checks were run for every changed profile.
|
||
|
||
| Profile | Context | Previous batch / ubatch | Previous prefill | Previous output | Selected batch / ubatch | Selected prefill | Selected output | Result |
|
||
|---|---:|---:|---:|---:|---:|---:|---:|---|
|
||
| fast | 76,800 | 64 / 32 | 907.2 tok/s | 75.4 tok/s | 64 / 32 | unchanged | unchanged | 2048 / 128 failed with CUDA OOM |
|
||
| medium | 160,000 | 64 / 32 | about 650–710 tok/s | about 65–70 tok/s | 2048 / 128 | 1,585.3 tok/s | 69.6 tok/s | selected; 155K plus vision passed |
|
||
| large | 192,000 | 64 / 32 | 730.6 tok/s | 78.4 tok/s | 2048 / 128 | 1,399.7 tok/s | 76.0 tok/s | selected; 190K plus vision passed |
|
||
| ultra | 262,144 | 64 / 32 | 672.0 tok/s | 57.1 tok/s | 2048 / 128 | 1,649.6 tok/s | 52.7 tok/s | selected; 258K passed |
|
||
| uncensored | 80,000 | 64 / 32 | 729.7 tok/s | 50.8 tok/s | 2048 / 128 | 1,331.8 tok/s | 49.9 tok/s | selected; 78K plus vision passed |
|
||
|
||
Full-context results with the selected settings:
|
||
|
||
| Profile | Tested prompt | Prefill | Output | Truncated | Vision after test |
|
||
|---|---:|---:|---:|---|---|
|
||
| medium | 154,998 | 881.3 tok/s | 37.8 tok/s | no | passed |
|
||
| large | 189,998 | 738.0 tok/s | 27.6 tok/s | no | passed |
|
||
| ultra | 257,998 | 698.2 tok/s | 21.1 tok/s | no | not configured for this profile |
|
||
| uncensored | 77,998 | 1,050.5 tok/s | 40.0 tok/s | no | passed |
|
||
|
||
The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. Moving Medium from a 90:10 to an 89:11 GPU split freed memory but reduced both prefill and output speed. The production setting therefore remains 2048 / 128 at 90:10.
|