# Bonsai 2 A/B on Athena Preparation is safe for the live router: `prepare.sh` downloads the pinned 7.21 GB PQ2_0 model, checks SHA-256, verifies PrismML's official Linux CUDA 12.8 archive from release `prism-b10685-7dffb15`, builds an isolated runtime image, and **creates a stopped container**. It never starts inference. The model and image are separate from production. A running production Qwen container blocks `case.sh ... start`. | Case | Context | GPUs inside container | Split | u-batch | |---|---:|---|---:|---:| | Fast | 76,800 | 5080 | 100:0 | 64 | | Medium | 160,000 | 5080 + 3060 | 85:15 | 512 | | Ultra | 262,144 | 5080 + 3060 | 80:20 | 128 | The A/B matrix compares production Qwen against Bonsai at **the same context profile**. Each run uses the same nine acceptance tasks, one tool-use probe, matched repeated prefill, and a long answer. Save total latency, prompt/decode throughput from llama.cpp timings, objective task accuracy, tool-call correctness, actual VRAM per GPU, CPU offload and any OOM. Assess the free-form answers blind, since a single numerical score does not measure intelligence well. This is a text-only comparison; Qwen Fast/Medium vision remains a separate feature. Qwen's production MTP speculative decoder stays on; report that speed advantage explicitly rather than attributing every difference to model weights. After the user's explicit **GO**, use the official profile controller/router to switch or stop production Qwen, record the original profile, run the corresponding Bonsai case with `case.sh CASE start --go`, execute `measure.py ... --go`, and sample both physical GPUs with `gpu_monitor.py --output ... --go`. Record idle GPU memory before each model load so the model's incremental VRAM is distinguishable from other processes. Stop Bonsai, restore the original production profile via the controller, and verify the router and original model are healthy. Repeat case by case. Never run both model servers at once. The smallest context is tested first; only try Medium and Ultra if Fast fits. Results live at `/data/benchmarks/bonsai2-ab` and should be checked for OOM/offload before drawing any speed conclusion. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA `cuobjdump`: it includes `sm_86` and `sm_120a`, covering both cards. Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results, deployment directory and build-cache records identifiable as this experiment. Docker's shared runtime base layers are deliberately not globally pruned, because that could delete unrelated build caches. Sources: [PrismML model](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [PrismML llama.cpp CUDA release](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15). Stock llama.cpp is intentionally not used for this ternary model.