Files
AI-Profile-Router/experiments/qwen38-20260907-ab/README.md
T

83 lines
2.5 KiB
Markdown

# Qwen3.8 September 2026 A/B preparation
This directory prepares an isolated comparison of two Qwen3.8-27B IQ4_XS
artifacts without adding router profiles or changing the running Athena stack.
## Candidates
| ID | Artifact | Purpose | Pinned revision | Size |
| --- | --- | --- | --- | ---: |
| `qwopus` | `Jackrong/Qwopus3.8-27B-Flash-GGUF` / `Qwopus3.8-27B-Flash-MTP-IQ4_XS.gguf` | Efficiency fine-tune | `e146d61e88782677805b3b68ad3adf8674dde80d` | 15,420,445,792 B |
| `bartowski` | `bartowski/Qwen3.8-27B-GGUF` / `Qwen3.8-27B-IQ4_XS.gguf` | Standard Qwen, alternative IQ4_XS quant | `f0eec4a4bb4975114a030d048952d83c0a53c034` | 15,567,824,480 B |
The production reference remains
`jpetrina/Qwen3.8-27B-IQ4_XS-pure-GGUF`. The candidates intentionally use the
same IQ4_XS quantization class so that the first comparison does not mix a
fine-tune difference with a quantization-class difference.
## Safety boundary
- Nothing in this directory is called by installation, Compose or the router.
- No production profile is added.
- Downloads happen only after explicitly running `download-candidate.sh`.
- `run-case.sh` refuses to start while a production `mike-ai-llama-*` model is
running. It never stops production itself.
- The test server binds to `127.0.0.1:5005` and is not exposed to the LAN.
- Cleanup is a dry run unless an explicit deletion flag is supplied. It only
addresses the exact test container, result directory and two pinned files.
## Later test sequence
Run these commands on Athena only after the active coding task has finished and
the GPUs have deliberately been released:
```sh
cd /opt/mike-ai/stack/experiments/qwen38-20260907-ab
./inventory.sh
./download-candidate.sh qwopus
./download-candidate.sh bartowski
./run-case.sh qwopus 160000 85,15
./wait-ready.sh
./run-benchmark.sh qwopus-160k
./stop-case.sh
./run-case.sh bartowski 160000 85,15
./wait-ready.sh
./run-benchmark.sh bartowski-160k
./stop-case.sh
```
Only after both candidates pass the 160K quality and tool-call tests should
192K and 262144 be attempted. Promotion into the router is a separate decision
and is deliberately not implemented here.
## Cleanup
Preview everything owned by this experiment:
```sh
./cleanup.sh
```
Remove only the test container and benchmark results:
```sh
./cleanup.sh --results
```
Remove only the two downloaded candidate files and their now-empty directory:
```sh
./cleanup.sh --models
```
Remove both:
```sh
./cleanup.sh --all
```
The script never touches the production Pure, Mix, Beta 1 or uncensored model
directories.