Prepare isolated Dirk Qwen3.8 benchmark

This commit is contained in:
Mikei386
2026-09-01 06:18:39 +02:00
parent b2ea53c383
commit ee28272999
5 changed files with 201 additions and 0 deletions
+49
View File
@@ -0,0 +1,49 @@
# Dirk Qwen3.8-27B experiment
Isolated A/B test environment for
`peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF`. It deliberately does not add a
production router profile and never stops or restarts production services.
## Candidate
- Main model: `Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf` (about 17.6 GB)
- Vision projector: `mmproj-F16.gguf`
- Runtime: existing `mike-ai/llama.cpp:local`
- Test endpoint: `127.0.0.1:5004`
- Results: `/data/benchmarks/dirk-qwen38/`
The candidate is intended to reduce unnecessary reasoning and total token use;
it is not expected to improve raw decode speed. The production Qwen model is
therefore the mandatory A/B reference.
## Safety boundary
`run-case.sh` refuses to start while any production `mike-ai-llama-*` model
container is running. It does not stop production itself. The model server is
bound to loopback only and cannot be reached from the LAN.
## Prepared matrix
Run the following only after the GPUs have explicitly been declared free:
```sh
./run-case.sh 80000 90,10 text
./run-case.sh 160000 90,10 text
./run-case.sh 160000 85,15 text
./run-case.sh 160000 80,20 text
./run-case.sh 192000 85,15 text
./run-case.sh 262144 80,20 text
```
The largest stable context is determined first. Vision is checked only after a
text winner exists:
```sh
./run-case.sh 160000 85,15 vision
```
For every case, record uncached prefill, cached prefill, decode throughput,
GPU memory, context recall, tool calling, code quality and total tokens needed
to finish the task. Do not promote Dirk unless it matches the base model on
technical correctness and improves real Hermes task completion.