Document Bonsai 2 A/B results
This commit is contained in:
@@ -0,0 +1,47 @@
|
||||
# Bonsai 2 27B A/B result on Athena
|
||||
|
||||
Measured 2026-09-19 against the production Qwen3.8-27B profiles. The same nine text tasks, tool schema, 27,013-token prefill, 2,048-token decode and physical GPU sampling were used. Lower text-task time is better; prefill and decode are tokens/s. Qwen used its production MTP speculative decoder. Bonsai used PrismML's pinned PQ2_0 runtime.
|
||||
|
||||
| Profile | Model | Nine tasks | Prefill | Decode | Weather tool |
|
||||
|---|---|---:|---:|---:|---|
|
||||
| Fast 76.8K | Qwen | 191.1 s | 1,305 | 84.7 | correct |
|
||||
| Fast 76.8K | Bonsai | 236.7 s | 1,480 | 83.9 | correct |
|
||||
| Medium 160K | Qwen | 219.6 s | 1,415 | 65.5 | correct |
|
||||
| Medium 160K | Bonsai | 247.5 s | 2,250 | 67.6 | correct |
|
||||
| Ultra 262K | Qwen | 265.2 s | 1,632 | 62.6 | correct |
|
||||
| Ultra 262K | Bonsai | 268.8 s | 1,506 | 63.6 | correct |
|
||||
|
||||
The total task time includes the model's chosen answer length, so it represents actual waiting time rather than pure kernel speed. Bonsai's Medium prefill was 59% faster, but its nine answers still took 13% longer. Fast took 24% longer. Ultra was effectively tied. Repeating the same long prompt hit the cache for both models.
|
||||
|
||||
## GPU memory measured while generating
|
||||
|
||||
Values are total board allocation and include the already-running Athena services. They remain directly comparable because each Qwen/Bonsai pair was sampled in the same test window.
|
||||
|
||||
| Profile | Model | RTX 5080 | RTX 3060 |
|
||||
|---|---|---:|---:|
|
||||
| Fast | Qwen | 15,832 MiB | 5,926 MiB |
|
||||
| Fast | Bonsai text | 8,728 MiB | 4,679 MiB |
|
||||
| Fast | Bonsai vision | 8,728 MiB | 5,676 MiB |
|
||||
| Medium | Qwen | 15,700 MiB | 11,468 MiB |
|
||||
| Medium | Bonsai | 9,786 MiB | 7,686 MiB |
|
||||
| Ultra | Qwen | 15,780 MiB | 11,180 MiB |
|
||||
| Ultra | Bonsai | 10,598 MiB | 8,528 MiB |
|
||||
|
||||
Bonsai is the clear memory winner. It saved about 7.1 GiB on the 5080 in Fast and about 5.9/3.8 GiB across the 5080/3060 in Medium. No case OOMed. Ultra left about 5.3 GiB free on the 5080 and 3.4 GiB on the 3060 at model start.
|
||||
|
||||
## Correctness and stability
|
||||
|
||||
All six main runs solved the basic logic, evidence, capacity, prompt-injection, refusal and missing-tool tasks. Both models produced the required `get_weather({"city":"Rastatt"})` call. Both recalled `KIESEL-7319` from a 48,644-token prompt; Bonsai took 36.3 s and Qwen 40.5 s. Both identified the synthetic vision fixture as a red circle on the left and a blue square on the right.
|
||||
|
||||
Bonsai Fast was not reliable enough to replace Qwen:
|
||||
|
||||
- In the first coding run it stopped after the first completed async task even when that task had failed. Across three Fast runs at temperature 0.2, only one supplied correct first-*successful*-task logic.
|
||||
- In the Home Assistant state task it twice claimed that an `off` automation entity merely meant "idle". Home Assistant's official `automation.turn_off` documentation says an off automation is disabled and no longer listens for triggers. Across three Fast runs at temperature 0.2, only one was correct.
|
||||
- Repeating both tasks with PrismML's recommended thinking sampling (`temperature=1.0`, `top_p=0.95`, `top_k=20`) did not fix the variance: one of three answers was correct for each consequential task. It also answered more slowly.
|
||||
- Bonsai Medium and Ultra answered both main-run critical cases correctly, but the model weights are identical. The difference is generation variance, not evidence that the larger context profile makes the model smarter.
|
||||
|
||||
Qwen returned correct answers for all of these critical runs. The practical decision is therefore **keep Qwen as Athena's production model**. Bonsai is useful as a stopped, optional memory-saving profile for experiments or workloads where freeing VRAM matters more than maximum agent reliability. Do not silently replace Fast, Medium or Ultra with it.
|
||||
|
||||
The official model card reports 98.2% of FP16 aggregate benchmark performance and provides a separate Q8_0 vision projector. Our result does not contradict that aggregate score; it shows that a small average loss can still appear as a consequential intermittent error in an agent workflow. Sources: [PrismML model card](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [Bonsai demo guide](https://github.com/PrismML-Eng/Bonsai-demo), [Home Assistant turn-off semantics](https://www.home-assistant.io/actions/automation.turn_off/).
|
||||
|
||||
After every run, the trap restored the original `medium` profile. Final verification: controller reported `active_profile=medium`, router and model containers were healthy, and Qwen Medium returned `OK` to a live completion request. The Bonsai container is stopped. `cleanup.sh --all` removes the experiment container, image, model, projector, raw server results and deployment staging without touching production.
|
||||
Reference in New Issue
Block a user