Files
AI-Profile-Router/docs/PREFILL_BATCH_BENCHMARK_20260830.md
T

1.8 KiB
Raw Blame History

Qwen prefill batch benchmark (2026-08-30)

Identical 24K-token prompt, one slot, Q4 KV cache and the production GPU split of each profile. The output measurement uses the same deterministic 64-token completion. Full-context checks were run for every changed profile.

Profile Context Previous batch / ubatch Previous prefill Previous output Selected batch / ubatch Selected prefill Selected output Result
fast 76,800 64 / 32 907.2 tok/s 75.4 tok/s 64 / 32 unchanged unchanged 2048 / 128 failed with CUDA OOM
medium 160,000 64 / 32 about 650–710 tok/s about 65–70 tok/s 2048 / 128 1,585.3 tok/s 69.6 tok/s selected; 155K plus vision passed
large 192,000 64 / 32 730.6 tok/s 78.4 tok/s 2048 / 128 1,399.7 tok/s 76.0 tok/s selected; 190K plus vision passed
ultra 262,144 64 / 32 672.0 tok/s 57.1 tok/s 2048 / 128 1,649.6 tok/s 52.7 tok/s selected; 258K passed
uncensored 80,000 64 / 32 729.7 tok/s 50.8 tok/s 2048 / 128 1,331.8 tok/s 49.9 tok/s selected; 78K plus vision passed

Full-context results with the selected settings:

Profile Tested prompt Prefill Output Truncated Vision after test
medium 154,998 881.3 tok/s 37.8 tok/s no passed
large 189,998 738.0 tok/s 27.6 tok/s no passed
ultra 257,998 698.2 tok/s 21.1 tok/s no not configured for this profile
uncensored 77,998 1,050.5 tok/s 40.0 tok/s no passed

The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. Moving Medium from a 90:10 to an 89:11 GPU split freed memory but reduced both prefill and output speed. The production setting therefore remains 2048 / 128 at 90:10.