Files
AI-Profile-Router/docs/PREFILL_BATCH_BENCHMARK_20260830.md
T

41 lines
2.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Qwen prefill batch benchmark (2026-08-30)
Identical 24K-token prompt, one slot, Q4 KV cache and the production GPU split of each profile. The output measurement uses the same deterministic 64-token completion. Full-context checks were run for every changed profile.
| Profile | Context | Previous batch / ubatch | Previous prefill | Previous output | Selected batch / ubatch | Selected prefill | Selected output | Result |
|---|---:|---:|---:|---:|---:|---:|---:|---|
| fast | 76,800 | 64 / 32 | 907.2 tok/s | 75.4 tok/s | 64 / 32 | unchanged | unchanged | 2048 / 128 failed with CUDA OOM |
| medium | 160,000 | 64 / 32 | about 650–710 tok/s | about 65–70 tok/s | 2048 / 128 | 1,585.3 tok/s | 69.6 tok/s | selected; 155K plus vision passed |
| large | 192,000 | 64 / 32 | 730.6 tok/s | 78.4 tok/s | 2048 / 128 | 1,399.7 tok/s | 76.0 tok/s | selected; 190K plus vision passed |
| ultra | 262,144 | 64 / 32 | 672.0 tok/s | 57.1 tok/s | 2048 / 128 | 1,649.6 tok/s | 52.7 tok/s | selected; 258K passed |
| uncensored | 80,000 | 64 / 32 | 729.7 tok/s | 50.8 tok/s | 2048 / 128 | 1,331.8 tok/s | 49.9 tok/s | selected; 78K plus vision passed |
Full-context results with the selected settings:
| Profile | Tested prompt | Prefill | Output | Truncated | Vision after test |
|---|---:|---:|---:|---|---|
| medium | 154,998 | 881.3 tok/s | 37.8 tok/s | no | passed |
| large | 189,998 | 738.0 tok/s | 27.6 tok/s | no | passed |
| ultra | 257,998 | 698.2 tok/s | 21.1 tok/s | no | not configured for this profile |
| uncensored | 77,998 | 1,050.5 tok/s | 40.0 tok/s | no | passed |
The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. The original one-slot benchmark selected 2048 / 128 at 90:10.
## Historischer Medium-Zwei-Slot-Test
This began as an A/B candidate and was initially returned to **one slot**
because concurrent Hermes requests did not behave reliably enough. On
19 September 2026, Medium was enabled again with two unified-KV slots for a
controlled OpenClaw live trial. The figures below remain the historical
benchmark rather than a new performance claim. The current Medium ubatch is
512; the 85:15 GPU split is unchanged.
Identical fresh 100,297-token prompt with a deterministic 256-token completion:
| Mode | Slots | GPU split | Prefill | Output | Total time |
|---|---:|---:|---:|---:|---:|
| Previous | 1 | 90:10 | 1,072.12 tok/s | 48.51 tok/s | 98.81 s |
| Selected | 2 | 85:15 | 993.20 tok/s | 47.65 tok/s | 106.33 s |
Single-request cost: **7.4% lower prefill**, **1.8% lower output**, and **7.6% longer total time** for this near-full prompt. Two simultaneous fresh 10K prompts completed in 20 seconds; per-slot output measured 43.32 and 27.95 tok/s. Two very large simultaneous prefills can temporarily throttle an already-generating slot, so the total 160K pool should not be treated as two independent 160K contexts.