Files
AI-Profile-Router/experiments/bonsai2-ab/README.md
T

19 lines
3.1 KiB
Markdown

# Bonsai 2 A/B on Athena
Preparation is safe for the live router: `prepare.sh` downloads the pinned 7.21 GB PQ2_0 model and optional 629 MB Q8_0 vision projector, checks both SHA-256 values, verifies PrismML's official Linux CUDA 12.8 archive from release `prism-b10685-7dffb15`, builds an isolated runtime image, and **creates a stopped container**. It never starts inference. The model and image are separate from production. A running production Qwen container blocks `case.sh ... start`.
| Case | Context | GPUs inside container | Split | u-batch |
|---|---:|---|---:|---:|
| Fast | 76,800 | 5080 | 100:0 | 64 |
| Medium | 160,000 | 5080 + 3060 | 85:15 | 512 |
| Ultra | 262,144 | 5080 + 3060 | 80:20 | 128 |
| Vision | 76,800 | 5080 + projector on 3060 | 100:0 language model | 64 |
The A/B matrix compares production Qwen against Bonsai at **the same context profile**. Each run uses the same nine acceptance tasks, one tool-use probe, matched repeated prefill, and a long answer. Save total latency, prompt/decode throughput from llama.cpp timings, objective task accuracy, tool-call correctness, actual VRAM per GPU, CPU offload and any OOM. Assess the free-form answers blind, since a single numerical score does not measure intelligence well. This is a text-only comparison; Qwen Fast/Medium vision remains a separate feature. Qwen's production MTP speculative decoder stays on; report that speed advantage explicitly rather than attributing every difference to model weights.
After the user's explicit **GO**, `run-go.sh --go` runs the six text cases, `repeat-go.sh --go` repeats consequential failures and a 48K-token recall probe, and `vision-go.sh --go` compares the synthetic image fixture and checks PrismML's recommended sampling. Each script records the original profile, never overlaps Qwen and Bonsai, and restores the original profile from an EXIT/INT/TERM trap. Results live at `/data/benchmarks/bonsai2-ab`; the reviewed, non-secret JSON reports and conclusion are committed under `results/`. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA `cuobjdump`: it includes `sm_86` and `sm_120a`, covering both cards.
Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results, deployment directory and build-cache records identifiable as this experiment. Docker's shared runtime base layers are deliberately not globally pruned, because that could delete unrelated build caches.
Measured conclusions are in [`RESULTS.md`](results/RESULTS.md). Sources: [PrismML model](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [PrismML demo and integration guide](https://github.com/PrismML-Eng/Bonsai-demo), [PrismML llama.cpp CUDA release](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15). Stock llama.cpp is intentionally not used for this ternary model.