Files

48 lines
4.6 KiB
Markdown

# Bonsai 2 27B A/B result on Athena
Measured 2026-09-19 against the production Qwen3.8-27B profiles. The same nine text tasks, tool schema, 27,013-token prefill, 2,048-token decode and physical GPU sampling were used. Lower text-task time is better; prefill and decode are tokens/s. Qwen used its production MTP speculative decoder. Bonsai used PrismML's pinned PQ2_0 runtime.
| Profile | Model | Nine tasks | Prefill | Decode | Weather tool |
|---|---|---:|---:|---:|---|
| Fast 76.8K | Qwen | 191.1 s | 1,305 | 84.7 | correct |
| Fast 76.8K | Bonsai | 236.7 s | 1,480 | 83.9 | correct |
| Medium 160K | Qwen | 219.6 s | 1,415 | 65.5 | correct |
| Medium 160K | Bonsai | 247.5 s | 2,250 | 67.6 | correct |
| Ultra 262K | Qwen | 265.2 s | 1,632 | 62.6 | correct |
| Ultra 262K | Bonsai | 268.8 s | 1,506 | 63.6 | correct |
The total task time includes the model's chosen answer length, so it represents actual waiting time rather than pure kernel speed. Bonsai's Medium prefill was 59% faster, but its nine answers still took 13% longer. Fast took 24% longer. Ultra was effectively tied. Repeating the same long prompt hit the cache for both models.
## GPU memory measured while generating
Values are total board allocation and include the already-running Athena services. They remain directly comparable because each Qwen/Bonsai pair was sampled in the same test window.
| Profile | Model | RTX 5080 | RTX 3060 |
|---|---|---:|---:|
| Fast | Qwen | 15,832 MiB | 5,926 MiB |
| Fast | Bonsai text | 8,728 MiB | 4,679 MiB |
| Fast | Bonsai vision | 8,728 MiB | 5,676 MiB |
| Medium | Qwen | 15,700 MiB | 11,468 MiB |
| Medium | Bonsai | 9,786 MiB | 7,686 MiB |
| Ultra | Qwen | 15,780 MiB | 11,180 MiB |
| Ultra | Bonsai | 10,598 MiB | 8,528 MiB |
Bonsai is the clear memory winner. It saved about 7.1 GiB on the 5080 in Fast and about 5.9/3.8 GiB across the 5080/3060 in Medium. No case OOMed. Ultra left about 5.3 GiB free on the 5080 and 3.4 GiB on the 3060 at model start.
## Correctness and stability
All six main runs solved the basic logic, evidence, capacity, prompt-injection, refusal and missing-tool tasks. Both models produced the required `get_weather({"city":"Rastatt"})` call. Both recalled `KIESEL-7319` from a 48,644-token prompt; Bonsai took 36.3 s and Qwen 40.5 s. Both identified the synthetic vision fixture as a red circle on the left and a blue square on the right.
Bonsai Fast was not reliable enough to replace Qwen:
- In the first coding run it stopped after the first completed async task even when that task had failed. Across three Fast runs at temperature 0.2, only one supplied correct first-*successful*-task logic.
- In the Home Assistant state task it twice claimed that an `off` automation entity merely meant "idle". Home Assistant's official `automation.turn_off` documentation says an off automation is disabled and no longer listens for triggers. Across three Fast runs at temperature 0.2, only one was correct.
- Repeating both tasks with PrismML's recommended thinking sampling (`temperature=1.0`, `top_p=0.95`, `top_k=20`) did not fix the variance: one of three answers was correct for each consequential task. It also answered more slowly.
- Bonsai Medium and Ultra answered both main-run critical cases correctly, but the model weights are identical. The difference is generation variance, not evidence that the larger context profile makes the model smarter.
Qwen returned correct answers for all of these critical runs. The practical decision is therefore **keep Qwen as Athena's production model**. Bonsai is useful as a stopped, optional memory-saving profile for experiments or workloads where freeing VRAM matters more than maximum agent reliability. Do not silently replace Fast, Medium or Ultra with it.
The official model card reports 98.2% of FP16 aggregate benchmark performance and provides a separate Q8_0 vision projector. Our result does not contradict that aggregate score; it shows that a small average loss can still appear as a consequential intermittent error in an agent workflow. Sources: [PrismML model card](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [Bonsai demo guide](https://github.com/PrismML-Eng/Bonsai-demo), [Home Assistant turn-off semantics](https://www.home-assistant.io/actions/automation.turn_off/).
After every run, the trap restored the original `medium` profile. Final verification: controller reported `active_profile=medium`, router and model containers were healthy, and Qwen Medium returned `OK` to a live completion request. The Bonsai container is stopped. `cleanup.sh --all` removes the experiment container, image, model, projector, raw server results and deployment staging without touching production.