Files
AI-Profile-Router/docs/DIRK_QWEN38_AB_20260901.md
T

81 lines
3.7 KiB
Markdown

# Dirk Qwen3.8-27B A/B benchmark (2026-09-01)
Candidate: `peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF`, pinned revision
`12362f2b3d7dc11044e99c9e7e99fb9f530528c0`, quant
`Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf`.
Reference: the production pure Qwen model
`qwen3.8-27b-IQ4_XS-pure.gguf`.
Both sides used the same local llama.cpp image and production-style runtime
settings: one slot, Q4_0 KV, unified KV, 24 GiB prompt cache, batch 2048,
ubatch 128 and embedded MTP with draft length 3. The container-visible GPU
order was CUDA0 = RTX 5080 and CUDA1 = RTX 3060. A separate Python process
occupied about 1.9 GiB on the RTX 3060 throughout the run and was deliberately
not disturbed.
## Stable candidate matrix
| Context | Stable split (5080:3060) | Short prefill | Short decode | Long tested prompt | Long prefill | Long decode | Recall |
|---:|---:|---:|---:|---:|---:|---:|---:|
| 80k | 87:13 | 1,381 t/s | 68.8 t/s | 56.2k | 1,127 t/s | 51.8 t/s | 3/3 |
| 160k | 80:20 | 1,268 t/s | 64.2 t/s | 112.3k | 854 t/s | 40.1 t/s | 3/3 |
| 192k | 72:28 | 1,124 t/s | 60.2 t/s | 134.7k | 677 t/s | 35.3 t/s | 3/3 |
| 262,144 | 70:30 | 1,076 t/s | 59.6 t/s | 183.8k | 539 t/s | 29.9 t/s | 3/3 |
At 80k, 88:12 loaded but failed on the first real prefill; 87:13 was the
maximum practical split. At 262k, 68:32 exhausted the RTX 3060 during KV
allocation and 72:28 exhausted the RTX 5080 during compute-buffer allocation.
70:30 was the only tested midpoint that loaded and completed the 183.8k-token
recall probe.
## Fair 160k comparison
| Model | Split | Short prefill | Short decode | 112k prefill | 112k decode | Recall |
|---|---:|---:|---:|---:|---:|---:|
| Pure IQ4_XS | 85:15 | 1,479 t/s | 104.4 t/s | 954 t/s | 56.4 t/s | 3/3 |
| Dirk Q4_K_XL | 80:20 | 1,268 t/s | 64.2 t/s | 854 t/s | 40.1 t/s | 3/3 |
Dirk was 14% slower on short prefill, 11% slower on the long prefill, 38%
slower on short decode and 29% slower on long decode.
## Quality and tool use
The fixed acceptance set covered logic, evidence-based diagnosis, concurrent
Python, capacity planning, prompt-injection resistance, configuration versus
runtime state, safe read-only diagnostics and evidence boundaries. Both models
emitted the requested native function call with the correct argument.
| Model | Completion tokens | Total task wall time | Mean decode | Completed final answers |
|---|---:|---:|---:|---:|
| Pure IQ4_XS | 17,542 | 224.9 s | 77.2 t/s | 7/9 before limit |
| Dirk Q4_K_XL | 12,178 | 260.2 s | 47.5 t/s | 9/9 |
Dirk used about 31% fewer completion tokens and was more concise. It produced
the cleaner proof for the impossible live-migration task. Pure reached the
right conclusion but its state proof contained a source-host accounting error
and hit the output limit. Both concurrent-Python answers had a subtle remaining
edge case: simultaneously completed failing tasks outside the returned task
were not all gathered, so neither answer was perfect.
Despite producing fewer tokens, Dirk needed about 16% more wall time for the
whole quality set because decode was much slower.
## Vision
The supplied F16 projector loaded at 160k with the 80:20 split. With reasoning
disabled, the private synthetic image was described correctly. Image prefill
was 106.7 t/s, decode was 39.7 t/s, and end-to-end latency was 11.0 seconds.
## Decision
Do not replace the production Pure Qwen medium profile with Dirk. Pure is the
clear speed winner and retained the same long-context recall and tool-call
ability. Dirk is useful only as an optional high-context/concise profile: it
can provide a verified 262k configured context on both GPUs and tends to spend
fewer output tokens, but it is slower in real elapsed time.
The test container was removed after the run. Production `mike-ai-llama-medium`
and `mike-ai-llama-review` were restarted and verified healthy.