81 lines
3.7 KiB
Markdown
81 lines
3.7 KiB
Markdown
# Dirk Qwen3.8-27B A/B benchmark (2026-09-01)
|
|
|
|
Candidate: `peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF`, pinned revision
|
|
`12362f2b3d7dc11044e99c9e7e99fb9f530528c0`, quant
|
|
`Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf`.
|
|
|
|
Reference: the production pure Qwen model
|
|
`qwen3.8-27b-IQ4_XS-pure.gguf`.
|
|
|
|
Both sides used the same local llama.cpp image and production-style runtime
|
|
settings: one slot, Q4_0 KV, unified KV, 24 GiB prompt cache, batch 2048,
|
|
ubatch 128 and embedded MTP with draft length 3. The container-visible GPU
|
|
order was CUDA0 = RTX 5080 and CUDA1 = RTX 3060. A separate Python process
|
|
occupied about 1.9 GiB on the RTX 3060 throughout the run and was deliberately
|
|
not disturbed.
|
|
|
|
## Stable candidate matrix
|
|
|
|
| Context | Stable split (5080:3060) | Short prefill | Short decode | Long tested prompt | Long prefill | Long decode | Recall |
|
|
|---:|---:|---:|---:|---:|---:|---:|---:|
|
|
| 80k | 87:13 | 1,381 t/s | 68.8 t/s | 56.2k | 1,127 t/s | 51.8 t/s | 3/3 |
|
|
| 160k | 80:20 | 1,268 t/s | 64.2 t/s | 112.3k | 854 t/s | 40.1 t/s | 3/3 |
|
|
| 192k | 72:28 | 1,124 t/s | 60.2 t/s | 134.7k | 677 t/s | 35.3 t/s | 3/3 |
|
|
| 262,144 | 70:30 | 1,076 t/s | 59.6 t/s | 183.8k | 539 t/s | 29.9 t/s | 3/3 |
|
|
|
|
At 80k, 88:12 loaded but failed on the first real prefill; 87:13 was the
|
|
maximum practical split. At 262k, 68:32 exhausted the RTX 3060 during KV
|
|
allocation and 72:28 exhausted the RTX 5080 during compute-buffer allocation.
|
|
70:30 was the only tested midpoint that loaded and completed the 183.8k-token
|
|
recall probe.
|
|
|
|
## Fair 160k comparison
|
|
|
|
| Model | Split | Short prefill | Short decode | 112k prefill | 112k decode | Recall |
|
|
|---|---:|---:|---:|---:|---:|---:|
|
|
| Pure IQ4_XS | 85:15 | 1,479 t/s | 104.4 t/s | 954 t/s | 56.4 t/s | 3/3 |
|
|
| Dirk Q4_K_XL | 80:20 | 1,268 t/s | 64.2 t/s | 854 t/s | 40.1 t/s | 3/3 |
|
|
|
|
Dirk was 14% slower on short prefill, 11% slower on the long prefill, 38%
|
|
slower on short decode and 29% slower on long decode.
|
|
|
|
## Quality and tool use
|
|
|
|
The fixed acceptance set covered logic, evidence-based diagnosis, concurrent
|
|
Python, capacity planning, prompt-injection resistance, configuration versus
|
|
runtime state, safe read-only diagnostics and evidence boundaries. Both models
|
|
emitted the requested native function call with the correct argument.
|
|
|
|
| Model | Completion tokens | Total task wall time | Mean decode | Completed final answers |
|
|
|---|---:|---:|---:|---:|
|
|
| Pure IQ4_XS | 17,542 | 224.9 s | 77.2 t/s | 7/9 before limit |
|
|
| Dirk Q4_K_XL | 12,178 | 260.2 s | 47.5 t/s | 9/9 |
|
|
|
|
Dirk used about 31% fewer completion tokens and was more concise. It produced
|
|
the cleaner proof for the impossible live-migration task. Pure reached the
|
|
right conclusion but its state proof contained a source-host accounting error
|
|
and hit the output limit. Both concurrent-Python answers had a subtle remaining
|
|
edge case: simultaneously completed failing tasks outside the returned task
|
|
were not all gathered, so neither answer was perfect.
|
|
|
|
Despite producing fewer tokens, Dirk needed about 16% more wall time for the
|
|
whole quality set because decode was much slower.
|
|
|
|
## Vision
|
|
|
|
The supplied F16 projector loaded at 160k with the 80:20 split. With reasoning
|
|
disabled, the private synthetic image was described correctly. Image prefill
|
|
was 106.7 t/s, decode was 39.7 t/s, and end-to-end latency was 11.0 seconds.
|
|
|
|
## Decision
|
|
|
|
Do not replace the production Pure Qwen medium profile with Dirk. Pure is the
|
|
clear speed winner and retained the same long-context recall and tool-call
|
|
ability. Dirk is useful only as an optional high-context/concise profile: it
|
|
can provide a verified 262k configured context on both GPUs and tends to spend
|
|
fewer output tokens, but it is slower in real elapsed time.
|
|
|
|
The test container was removed after the run. Production `mike-ai-llama-medium`
|
|
and `mike-ai-llama-review` were restarted and verified healthy.
|
|
|