Benchmark Dirk Qwen3.8 against production
This commit is contained in:
@@ -0,0 +1,80 @@
|
||||
# Dirk Qwen3.8-27B A/B benchmark (2026-09-01)
|
||||
|
||||
Candidate: `peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF`, pinned revision
|
||||
`12362f2b3d7dc11044e99c9e7e99fb9f530528c0`, quant
|
||||
`Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf`.
|
||||
|
||||
Reference: the production pure Qwen model
|
||||
`qwen3.8-27b-IQ4_XS-pure.gguf`.
|
||||
|
||||
Both sides used the same local llama.cpp image and production-style runtime
|
||||
settings: one slot, Q4_0 KV, unified KV, 24 GiB prompt cache, batch 2048,
|
||||
ubatch 128 and embedded MTP with draft length 3. The container-visible GPU
|
||||
order was CUDA0 = RTX 5080 and CUDA1 = RTX 3060. A separate Python process
|
||||
occupied about 1.9 GiB on the RTX 3060 throughout the run and was deliberately
|
||||
not disturbed.
|
||||
|
||||
## Stable candidate matrix
|
||||
|
||||
| Context | Stable split (5080:3060) | Short prefill | Short decode | Long tested prompt | Long prefill | Long decode | Recall |
|
||||
|---:|---:|---:|---:|---:|---:|---:|---:|
|
||||
| 80k | 87:13 | 1,381 t/s | 68.8 t/s | 56.2k | 1,127 t/s | 51.8 t/s | 3/3 |
|
||||
| 160k | 80:20 | 1,268 t/s | 64.2 t/s | 112.3k | 854 t/s | 40.1 t/s | 3/3 |
|
||||
| 192k | 72:28 | 1,124 t/s | 60.2 t/s | 134.7k | 677 t/s | 35.3 t/s | 3/3 |
|
||||
| 262,144 | 70:30 | 1,076 t/s | 59.6 t/s | 183.8k | 539 t/s | 29.9 t/s | 3/3 |
|
||||
|
||||
At 80k, 88:12 loaded but failed on the first real prefill; 87:13 was the
|
||||
maximum practical split. At 262k, 68:32 exhausted the RTX 3060 during KV
|
||||
allocation and 72:28 exhausted the RTX 5080 during compute-buffer allocation.
|
||||
70:30 was the only tested midpoint that loaded and completed the 183.8k-token
|
||||
recall probe.
|
||||
|
||||
## Fair 160k comparison
|
||||
|
||||
| Model | Split | Short prefill | Short decode | 112k prefill | 112k decode | Recall |
|
||||
|---|---:|---:|---:|---:|---:|---:|
|
||||
| Pure IQ4_XS | 85:15 | 1,479 t/s | 104.4 t/s | 954 t/s | 56.4 t/s | 3/3 |
|
||||
| Dirk Q4_K_XL | 80:20 | 1,268 t/s | 64.2 t/s | 854 t/s | 40.1 t/s | 3/3 |
|
||||
|
||||
Dirk was 14% slower on short prefill, 11% slower on the long prefill, 38%
|
||||
slower on short decode and 29% slower on long decode.
|
||||
|
||||
## Quality and tool use
|
||||
|
||||
The fixed acceptance set covered logic, evidence-based diagnosis, concurrent
|
||||
Python, capacity planning, prompt-injection resistance, configuration versus
|
||||
runtime state, safe read-only diagnostics and evidence boundaries. Both models
|
||||
emitted the requested native function call with the correct argument.
|
||||
|
||||
| Model | Completion tokens | Total task wall time | Mean decode | Completed final answers |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Pure IQ4_XS | 17,542 | 224.9 s | 77.2 t/s | 7/9 before limit |
|
||||
| Dirk Q4_K_XL | 12,178 | 260.2 s | 47.5 t/s | 9/9 |
|
||||
|
||||
Dirk used about 31% fewer completion tokens and was more concise. It produced
|
||||
the cleaner proof for the impossible live-migration task. Pure reached the
|
||||
right conclusion but its state proof contained a source-host accounting error
|
||||
and hit the output limit. Both concurrent-Python answers had a subtle remaining
|
||||
edge case: simultaneously completed failing tasks outside the returned task
|
||||
were not all gathered, so neither answer was perfect.
|
||||
|
||||
Despite producing fewer tokens, Dirk needed about 16% more wall time for the
|
||||
whole quality set because decode was much slower.
|
||||
|
||||
## Vision
|
||||
|
||||
The supplied F16 projector loaded at 160k with the 80:20 split. With reasoning
|
||||
disabled, the private synthetic image was described correctly. Image prefill
|
||||
was 106.7 t/s, decode was 39.7 t/s, and end-to-end latency was 11.0 seconds.
|
||||
|
||||
## Decision
|
||||
|
||||
Do not replace the production Pure Qwen medium profile with Dirk. Pure is the
|
||||
clear speed winner and retained the same long-context recall and tool-call
|
||||
ability. Dirk is useful only as an optional high-context/concise profile: it
|
||||
can provide a verified 262k configured context on both GPUs and tends to spend
|
||||
fewer output tokens, but it is slower in real elapsed time.
|
||||
|
||||
The test container was removed after the run. Production `mike-ai-llama-medium`
|
||||
and `mike-ai-llama-review` were restarted and verified healthy.
|
||||
|
||||
Reference in New Issue
Block a user