Files
AI-Profile-Router/experiments/dirk-qwen38/README.md
T

1.8 KiB

Dirk Qwen3.8-27B experiment

Isolated A/B test environment for peculiar-ragdoll/Dirk-Qwen3.8-27B-GGUF. It deliberately does not add a production router profile and never stops or restarts production services.

Candidate

  • Main model: Dirk-Qwen3.8-27B-UD-Q4_K_XL.gguf (about 17.6 GB)
  • Vision projector: mmproj-F16.gguf
  • Pinned Hugging Face revision: 12362f2b3d7dc11044e99c9e7e99fb9f530528c0
  • Runtime: existing mike-ai/llama.cpp:local
  • Test endpoint: 127.0.0.1:5004
  • Results: /data/benchmarks/dirk-qwen38/

The candidate is intended to reduce unnecessary reasoning and total token use; it is not expected to improve raw decode speed. The production Qwen model is therefore the mandatory A/B reference.

Safety boundary

run-case.sh refuses to start while any production mike-ai-llama-* model container is running. It does not stop production itself. The model server is bound to loopback only and cannot be reached from the LAN.

Measured matrix

Run the following only after the GPUs have explicitly been declared free:

./run-case.sh 80000 87,13 text
./run-case.sh 160000 80,20 text
./run-case.sh 192000 72,28 text
./run-case.sh 262144 70,30 text

The largest stable context is determined first. Vision is checked only after a text winner exists:

./run-case.sh 160000 80,20 vision

For every case, record uncached prefill, cached prefill, decode throughput, GPU memory, context recall, tool calling, code quality and total tokens needed to finish the task. Do not promote Dirk unless it matches the base model on technical correctness and improves real Hermes task completion.

Expected SHA-256 checksums are stored in MODEL_ARTIFACTS.sha256.

The completed A/B result is documented in docs/DIRK_QWEN38_AB_20260901.md. The candidate did not replace the production Pure Qwen profile.