Files
AI-Profile-Router/experiments/bonsai2-ab

Bonsai 2 A/B on Athena

Preparation is safe for the live router: prepare.sh downloads the pinned 7.21 GB PQ2_0 model and optional 629 MB Q8_0 vision projector, checks both SHA-256 values, verifies PrismML's official Linux CUDA 12.8 archive from release prism-b10685-7dffb15, builds an isolated runtime image, and creates a stopped container. It never starts inference. The model and image are separate from production. A running production Qwen container blocks case.sh ... start.

Case Context GPUs inside container Split u-batch
Fast 76,800 5080 100:0 64
Medium 160,000 5080 + 3060 85:15 512
Ultra 262,144 5080 + 3060 80:20 128
Vision 76,800 5080 + projector on 3060 100:0 language model 64

The A/B matrix compares production Qwen against Bonsai at the same context profile. Each run uses the same nine acceptance tasks, one tool-use probe, matched repeated prefill, and a long answer. Save total latency, prompt/decode throughput from llama.cpp timings, objective task accuracy, tool-call correctness, actual VRAM per GPU, CPU offload and any OOM. Assess the free-form answers blind, since a single numerical score does not measure intelligence well. This is a text-only comparison; Qwen Fast/Medium vision remains a separate feature. Qwen's production MTP speculative decoder stays on; report that speed advantage explicitly rather than attributing every difference to model weights.

After the user's explicit GO, run-go.sh --go runs the six text cases, repeat-go.sh --go repeats consequential failures and a 48K-token recall probe, and vision-go.sh --go compares the synthetic image fixture and checks PrismML's recommended sampling. Each script records the original profile, never overlaps Qwen and Bonsai, and restores the original profile from an EXIT/INT/TERM trap. Results live at /data/benchmarks/bonsai2-ab; the reviewed, non-secret JSON reports and conclusion are committed under results/. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA cuobjdump: it includes sm_86 and sm_120a, covering both cards.

Preparation and measure.py deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is /opt/mike-ai/experiments/bonsai2-ab. No production files are edited. cleanup.sh previews what will be removed; cleanup.sh --all stops the transient build (if still active), then removes the named test container, test image, model, results, deployment directory and build-cache records identifiable as this experiment. Docker's shared runtime base layers are deliberately not globally pruned, because that could delete unrelated build caches.

Measured conclusions are in RESULTS.md. Sources: PrismML model, PrismML demo and integration guide, PrismML llama.cpp CUDA release. Stock llama.cpp is intentionally not used for this ternary model.