Files
AI-Profile-Router/docs/PREFILL_BATCH_BENCHMARK_20260830.md
T

2.7 KiB
Raw Blame History

Qwen prefill batch benchmark (2026-08-30)

Identical 24K-token prompt, one slot, Q4 KV cache and the production GPU split of each profile. The output measurement uses the same deterministic 64-token completion. Full-context checks were run for every changed profile.

Profile Context Previous batch / ubatch Previous prefill Previous output Selected batch / ubatch Selected prefill Selected output Result
fast 76,800 64 / 32 907.2 tok/s 75.4 tok/s 64 / 32 unchanged unchanged 2048 / 128 failed with CUDA OOM
medium 160,000 64 / 32 about 650–710 tok/s about 65–70 tok/s 2048 / 128 1,585.3 tok/s 69.6 tok/s selected; 155K plus vision passed
large 192,000 64 / 32 730.6 tok/s 78.4 tok/s 2048 / 128 1,399.7 tok/s 76.0 tok/s selected; 190K plus vision passed
ultra 262,144 64 / 32 672.0 tok/s 57.1 tok/s 2048 / 128 1,649.6 tok/s 52.7 tok/s selected; 258K passed
uncensored 80,000 64 / 32 729.7 tok/s 50.8 tok/s 2048 / 128 1,331.8 tok/s 49.9 tok/s selected; 78K plus vision passed

Full-context results with the selected settings:

Profile Tested prompt Prefill Output Truncated Vision after test
medium 154,998 881.3 tok/s 37.8 tok/s no passed
large 189,998 738.0 tok/s 27.6 tok/s no passed
ultra 257,998 698.2 tok/s 21.1 tok/s no not configured for this profile
uncensored 77,998 1,050.5 tok/s 40.0 tok/s no passed

The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. The original one-slot benchmark selected 2048 / 128 at 90:10.

Medium two-slot benchmark

Medium now uses two parallel slots with unified KV, so both chats dynamically share one total 160K-token pool. The model weights remain loaded only once. To fit the additional scheduler buffers, the production GPU split is 85:15 while batch / ubatch remains 2048 / 128.

Identical fresh 100,297-token prompt with a deterministic 256-token completion:

Mode Slots GPU split Prefill Output Total time
Previous 1 90:10 1,072.12 tok/s 48.51 tok/s 98.81 s
Selected 2 85:15 993.20 tok/s 47.65 tok/s 106.33 s

Single-request cost: 7.4% lower prefill, 1.8% lower output, and 7.6% longer total time for this near-full prompt. Two simultaneous fresh 10K prompts completed in 20 seconds; per-slot output measured 43.32 and 27.95 tok/s. Two very large simultaneous prefills can temporarily throttle an already-generating slot, so the total 160K pool should not be treated as two independent 160K contexts.