53 lines
1.7 KiB
Markdown
53 lines
1.7 KiB
Markdown
# Dirk Qwen3.8-27B experiment
|
|
|
|
Isolated A/B test environment for
|
|
`peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF`. It deliberately does not add a
|
|
production router profile and never stops or restarts production services.
|
|
|
|
## Candidate
|
|
|
|
- Main model: `Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf` (about 17.6 GB)
|
|
- Vision projector: `mmproj-F16.gguf`
|
|
- Pinned Hugging Face revision: `12362f2b3d7dc11044e99c9e7e99fb9f530528c0`
|
|
- Runtime: existing `mike-ai/llama.cpp:local`
|
|
- Test endpoint: `127.0.0.1:5004`
|
|
- Results: `/data/benchmarks/dirk-qwen38/`
|
|
|
|
The candidate is intended to reduce unnecessary reasoning and total token use;
|
|
it is not expected to improve raw decode speed. The production Qwen model is
|
|
therefore the mandatory A/B reference.
|
|
|
|
## Safety boundary
|
|
|
|
`run-case.sh` refuses to start while any production `mike-ai-llama-*` model
|
|
container is running. It does not stop production itself. The model server is
|
|
bound to loopback only and cannot be reached from the LAN.
|
|
|
|
## Prepared matrix
|
|
|
|
Run the following only after the GPUs have explicitly been declared free:
|
|
|
|
```sh
|
|
./run-case.sh 80000 90,10 text
|
|
./run-case.sh 160000 90,10 text
|
|
./run-case.sh 160000 85,15 text
|
|
./run-case.sh 160000 80,20 text
|
|
./run-case.sh 192000 85,15 text
|
|
./run-case.sh 262144 80,20 text
|
|
```
|
|
|
|
The largest stable context is determined first. Vision is checked only after a
|
|
text winner exists:
|
|
|
|
```sh
|
|
./run-case.sh 160000 85,15 vision
|
|
```
|
|
|
|
For every case, record uncached prefill, cached prefill, decode throughput,
|
|
GPU memory, context recall, tool calling, code quality and total tokens needed
|
|
to finish the task. Do not promote Dirk unless it matches the base model on
|
|
technical correctness and improves real Hermes task completion.
|
|
|
|
Expected SHA-256 checksums are stored in `MODEL_ARTIFACTS.sha256`.
|
|
|