Bonsai 2 A/B on Athena
Preparation is safe for the live router: prepare.sh downloads the pinned 7.21 GB PQ2_0 model and optional 629 MB Q8_0 vision projector, checks both SHA-256 values, verifies PrismML's official Linux CUDA 12.8 archive from release prism-b10685-7dffb15, builds an isolated runtime image, and creates a stopped container. It never starts inference. The model and image are separate from production. A running production Qwen container blocks case.sh ... start.
| Case | Context | GPUs inside container | Split | u-batch |
|---|---|---|---|---|
| Fast | 76,800 | 5080 | 100:0 | 64 |
| Medium | 160,000 | 5080 + 3060 | 85:15 | 512 |
| Ultra | 262,144 | 5080 + 3060 | 80:20 | 128 |
| Vision | 76,800 | 5080 + projector on 3060 | 100:0 language model | 64 |
The A/B matrix compares production Qwen against Bonsai at the same context profile. Each run uses the same nine acceptance tasks, one tool-use probe, matched repeated prefill, and a long answer. Save total latency, prompt/decode throughput from llama.cpp timings, objective task accuracy, tool-call correctness, actual VRAM per GPU, CPU offload and any OOM. Assess the free-form answers blind, since a single numerical score does not measure intelligence well. This is a text-only comparison; Qwen Fast/Medium vision remains a separate feature. Qwen's production MTP speculative decoder stays on; report that speed advantage explicitly rather than attributing every difference to model weights.
After the user's explicit GO, run-go.sh --go runs the six text cases, repeat-go.sh --go repeats consequential failures and a 48K-token recall probe, and vision-go.sh --go compares the synthetic image fixture and checks PrismML's recommended sampling. Each script records the original profile, never overlaps Qwen and Bonsai, and restores the original profile from an EXIT/INT/TERM trap. Results live at /data/benchmarks/bonsai2-ab; the reviewed, non-secret JSON reports and conclusion are committed under results/. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA cuobjdump: it includes sm_86 and sm_120a, covering both cards.
Preparation and measure.py deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is /opt/mike-ai/experiments/bonsai2-ab. No production files are edited. cleanup.sh previews what will be removed; cleanup.sh --all stops the transient build (if still active), then removes the named test container, test image, model, results, deployment directory and build-cache records identifiable as this experiment. Docker's shared runtime base layers are deliberately not globally pruned, because that could delete unrelated build caches.
Measured conclusions are in RESULTS.md. Sources: PrismML model, PrismML demo and integration guide, PrismML llama.cpp CUDA release. Stock llama.cpp is intentionally not used for this ternary model.