Files
AI-Profile-Router/docs/PREFILL_BATCH_BENCHMARK_20260830.md
T

39 lines
2.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Qwen prefill batch benchmark (2026-08-30)
Identical 24K-token prompt, one slot, Q4 KV cache and the production GPU split of each profile. The output measurement uses the same deterministic 64-token completion. Full-context checks were run for every changed profile.
| Profile | Context | Previous batch / ubatch | Previous prefill | Previous output | Selected batch / ubatch | Selected prefill | Selected output | Result |
|---|---:|---:|---:|---:|---:|---:|---:|---|
| fast | 76,800 | 64 / 32 | 907.2 tok/s | 75.4 tok/s | 64 / 32 | unchanged | unchanged | 2048 / 128 failed with CUDA OOM |
| medium | 160,000 | 64 / 32 | about 650–710 tok/s | about 65–70 tok/s | 2048 / 128 | 1,585.3 tok/s | 69.6 tok/s | selected; 155K plus vision passed |
| large | 192,000 | 64 / 32 | 730.6 tok/s | 78.4 tok/s | 2048 / 128 | 1,399.7 tok/s | 76.0 tok/s | selected; 190K plus vision passed |
| ultra | 262,144 | 64 / 32 | 672.0 tok/s | 57.1 tok/s | 2048 / 128 | 1,649.6 tok/s | 52.7 tok/s | selected; 258K passed |
| uncensored | 80,000 | 64 / 32 | 729.7 tok/s | 50.8 tok/s | 2048 / 128 | 1,331.8 tok/s | 49.9 tok/s | selected; 78K plus vision passed |
Full-context results with the selected settings:
| Profile | Tested prompt | Prefill | Output | Truncated | Vision after test |
|---|---:|---:|---:|---|---|
| medium | 154,998 | 881.3 tok/s | 37.8 tok/s | no | passed |
| large | 189,998 | 738.0 tok/s | 27.6 tok/s | no | passed |
| ultra | 257,998 | 698.2 tok/s | 21.1 tok/s | no | not configured for this profile |
| uncensored | 77,998 | 1,050.5 tok/s | 40.0 tok/s | no | passed |
The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. The original one-slot benchmark selected 2048 / 128 at 90:10.
## Historischer Medium-Zwei-Slot-Test
This was an A/B candidate, not the current production configuration. Production
was returned to **one slot** because concurrent Hermes requests did not behave
reliably enough. The 85:15 GPU split and batch / ubatch 2048 / 128 remain in
production because they also work with the single-slot profile.
Identical fresh 100,297-token prompt with a deterministic 256-token completion:
| Mode | Slots | GPU split | Prefill | Output | Total time |
|---|---:|---:|---:|---:|---:|
| Previous | 1 | 90:10 | 1,072.12 tok/s | 48.51 tok/s | 98.81 s |
| Selected | 2 | 85:15 | 993.20 tok/s | 47.65 tok/s | 106.33 s |
Single-request cost: **7.4% lower prefill**, **1.8% lower output**, and **7.6% longer total time** for this near-full prompt. Two simultaneous fresh 10K prompts completed in 20 seconds; per-slot output measured 43.32 and 27.95 tok/s. Two very large simultaneous prefills can temporarily throttle an already-generating slot, so the total 160K pool should not be treated as two independent 160K contexts.