Files
AI-Profile-Router/docs/DIRK_QWEN38_AB_20260901.md
T

3.7 KiB

Dirk Qwen3.8-27B A/B benchmark (2026-09-01)

Candidate: peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF, pinned revision 12362f2b3d7dc11044e99c9e7e99fb9f530528c0, quant Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf.

Reference: the production pure Qwen model qwen3.8-27b-IQ4_XS-pure.gguf.

Both sides used the same local llama.cpp image and production-style runtime settings: one slot, Q4_0 KV, unified KV, 24 GiB prompt cache, batch 2048, ubatch 128 and embedded MTP with draft length 3. The container-visible GPU order was CUDA0 = RTX 5080 and CUDA1 = RTX 3060. A separate Python process occupied about 1.9 GiB on the RTX 3060 throughout the run and was deliberately not disturbed.

Stable candidate matrix

Context Stable split (5080:3060) Short prefill Short decode Long tested prompt Long prefill Long decode Recall
80k 87:13 1,381 t/s 68.8 t/s 56.2k 1,127 t/s 51.8 t/s 3/3
160k 80:20 1,268 t/s 64.2 t/s 112.3k 854 t/s 40.1 t/s 3/3
192k 72:28 1,124 t/s 60.2 t/s 134.7k 677 t/s 35.3 t/s 3/3
262,144 70:30 1,076 t/s 59.6 t/s 183.8k 539 t/s 29.9 t/s 3/3

At 80k, 88:12 loaded but failed on the first real prefill; 87:13 was the maximum practical split. At 262k, 68:32 exhausted the RTX 3060 during KV allocation and 72:28 exhausted the RTX 5080 during compute-buffer allocation. 70:30 was the only tested midpoint that loaded and completed the 183.8k-token recall probe.

Fair 160k comparison

Model Split Short prefill Short decode 112k prefill 112k decode Recall
Pure IQ4_XS 85:15 1,479 t/s 104.4 t/s 954 t/s 56.4 t/s 3/3
Dirk Q4_K_XL 80:20 1,268 t/s 64.2 t/s 854 t/s 40.1 t/s 3/3

Dirk was 14% slower on short prefill, 11% slower on the long prefill, 38% slower on short decode and 29% slower on long decode.

Quality and tool use

The fixed acceptance set covered logic, evidence-based diagnosis, concurrent Python, capacity planning, prompt-injection resistance, configuration versus runtime state, safe read-only diagnostics and evidence boundaries. Both models emitted the requested native function call with the correct argument.

Model Completion tokens Total task wall time Mean decode Completed final answers
Pure IQ4_XS 17,542 224.9 s 77.2 t/s 7/9 before limit
Dirk Q4_K_XL 12,178 260.2 s 47.5 t/s 9/9

Dirk used about 31% fewer completion tokens and was more concise. It produced the cleaner proof for the impossible live-migration task. Pure reached the right conclusion but its state proof contained a source-host accounting error and hit the output limit. Both concurrent-Python answers had a subtle remaining edge case: simultaneously completed failing tasks outside the returned task were not all gathered, so neither answer was perfect.

Despite producing fewer tokens, Dirk needed about 16% more wall time for the whole quality set because decode was much slower.

Vision

The supplied F16 projector loaded at 160k with the 80:20 split. With reasoning disabled, the private synthetic image was described correctly. Image prefill was 106.7 t/s, decode was 39.7 t/s, and end-to-end latency was 11.0 seconds.

Decision

Do not replace the production Pure Qwen medium profile with Dirk. Pure is the clear speed winner and retained the same long-context recall and tool-call ability. Dirk is useful only as an optional high-context/concise profile: it can provide a verified 262k configured context on both GPUs and tends to spend fewer output tokens, but it is slower in real elapsed time.

The test container was removed after the run. Production mike-ai-llama-medium and mike-ai-llama-review were restarted and verified healthy.