Document Bonsai 2 A/B results

This commit is contained in:
Mikei386
2026-09-19 02:00:40 +02:00
parent de9b17f054
commit 9978e5b7e6
21 changed files with 3141 additions and 6 deletions
+4 -3
View File
@@ -1,17 +1,18 @@
# Bonsai 2 A/B on Athena # Bonsai 2 A/B on Athena
Preparation is safe for the live router: `prepare.sh` downloads the pinned 7.21 GB PQ2_0 model, checks SHA-256, verifies PrismML's official Linux CUDA 12.8 archive from release `prism-b10685-7dffb15`, builds an isolated runtime image, and **creates a stopped container**. It never starts inference. The model and image are separate from production. A running production Qwen container blocks `case.sh ... start`. Preparation is safe for the live router: `prepare.sh` downloads the pinned 7.21 GB PQ2_0 model and optional 629 MB Q8_0 vision projector, checks both SHA-256 values, verifies PrismML's official Linux CUDA 12.8 archive from release `prism-b10685-7dffb15`, builds an isolated runtime image, and **creates a stopped container**. It never starts inference. The model and image are separate from production. A running production Qwen container blocks `case.sh ... start`.
| Case | Context | GPUs inside container | Split | u-batch | | Case | Context | GPUs inside container | Split | u-batch |
|---|---:|---|---:|---:| |---|---:|---|---:|---:|
| Fast | 76,800 | 5080 | 100:0 | 64 | | Fast | 76,800 | 5080 | 100:0 | 64 |
| Medium | 160,000 | 5080 + 3060 | 85:15 | 512 | | Medium | 160,000 | 5080 + 3060 | 85:15 | 512 |
| Ultra | 262,144 | 5080 + 3060 | 80:20 | 128 | | Ultra | 262,144 | 5080 + 3060 | 80:20 | 128 |
| Vision | 76,800 | 5080 + projector on 3060 | 100:0 language model | 64 |
The A/B matrix compares production Qwen against Bonsai at **the same context profile**. Each run uses the same nine acceptance tasks, one tool-use probe, matched repeated prefill, and a long answer. Save total latency, prompt/decode throughput from llama.cpp timings, objective task accuracy, tool-call correctness, actual VRAM per GPU, CPU offload and any OOM. Assess the free-form answers blind, since a single numerical score does not measure intelligence well. This is a text-only comparison; Qwen Fast/Medium vision remains a separate feature. Qwen's production MTP speculative decoder stays on; report that speed advantage explicitly rather than attributing every difference to model weights. The A/B matrix compares production Qwen against Bonsai at **the same context profile**. Each run uses the same nine acceptance tasks, one tool-use probe, matched repeated prefill, and a long answer. Save total latency, prompt/decode throughput from llama.cpp timings, objective task accuracy, tool-call correctness, actual VRAM per GPU, CPU offload and any OOM. Assess the free-form answers blind, since a single numerical score does not measure intelligence well. This is a text-only comparison; Qwen Fast/Medium vision remains a separate feature. Qwen's production MTP speculative decoder stays on; report that speed advantage explicitly rather than attributing every difference to model weights.
After the user's explicit **GO**, use the official profile controller/router to switch or stop production Qwen, record the original profile, run the corresponding Bonsai case with `case.sh CASE start --go`, execute `measure.py ... --go`, and sample both physical GPUs with `gpu_monitor.py --output ... --go`. Record idle GPU memory before each model load so the model's incremental VRAM is distinguishable from other processes. Stop Bonsai, restore the original production profile via the controller, and verify the router and original model are healthy. Repeat case by case. Never run both model servers at once. The smallest context is tested first; only try Medium and Ultra if Fast fits. Results live at `/data/benchmarks/bonsai2-ab` and should be checked for OOM/offload before drawing any speed conclusion. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA `cuobjdump`: it includes `sm_86` and `sm_120a`, covering both cards. After the user's explicit **GO**, `run-go.sh --go` runs the six text cases, `repeat-go.sh --go` repeats consequential failures and a 48K-token recall probe, and `vision-go.sh --go` compares the synthetic image fixture and checks PrismML's recommended sampling. Each script records the original profile, never overlaps Qwen and Bonsai, and restores the original profile from an EXIT/INT/TERM trap. Results live at `/data/benchmarks/bonsai2-ab`; the reviewed, non-secret JSON reports and conclusion are committed under `results/`. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA `cuobjdump`: it includes `sm_86` and `sm_120a`, covering both cards.
Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results, deployment directory and build-cache records identifiable as this experiment. Docker's shared runtime base layers are deliberately not globally pruned, because that could delete unrelated build caches. Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results, deployment directory and build-cache records identifiable as this experiment. Docker's shared runtime base layers are deliberately not globally pruned, because that could delete unrelated build caches.
Sources: [PrismML model](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [PrismML llama.cpp CUDA release](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15). Stock llama.cpp is intentionally not used for this ternary model. Measured conclusions are in [`RESULTS.md`](results/RESULTS.md). Sources: [PrismML model](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [PrismML demo and integration guide](https://github.com/PrismML-Eng/Bonsai-demo), [PrismML llama.cpp CUDA release](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15). Stock llama.cpp is intentionally not used for this ternary model.
Binary file not shown.

After

Width:  |  Height:  |  Size: 1.6 KiB

+11 -3
View File
@@ -8,8 +8,8 @@ IMAGE=mike-ai/bonsai2-ab:prism-b10685
MODEL=/data/models/bonsai2-ab/Ternary-Bonsai-2-27B-PQ2_0.gguf MODEL=/data/models/bonsai2-ab/Ternary-Bonsai-2-27B-PQ2_0.gguf
GPU_5080=GPU-8ad38c6c-5a01-9d8e-1dfa-ed662ad78fbe GPU_5080=GPU-8ad38c6c-5a01-9d8e-1dfa-ed662ad78fbe
GPU_3060=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b GPU_3060=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
[[ $CASE =~ ^(fast|medium|ultra)$ && $ACTION =~ ^(create|start|stop)$ ]] || { [[ $CASE =~ ^(fast|medium|ultra|vision)$ && $ACTION =~ ^(create|start|stop)$ ]] || {
echo 'Usage: case.sh {fast|medium|ultra} {create|start|stop}' >&2; exit 2; echo 'Usage: case.sh {fast|medium|ultra|vision} {create|start|stop}' >&2; exit 2;
} }
test "$(hostname)" = athena || exit 2 test "$(hostname)" = athena || exit 2
if [[ $ACTION == stop ]]; then if [[ $ACTION == stop ]]; then
@@ -36,7 +36,15 @@ case "$CASE" in
fast) CONTEXT=76800; UBATCH=64; DEVICES="$GPU_5080"; SPLIT=(--device CUDA0 --split-mode none) ;; fast) CONTEXT=76800; UBATCH=64; DEVICES="$GPU_5080"; SPLIT=(--device CUDA0 --split-mode none) ;;
medium) CONTEXT=160000; UBATCH=512; DEVICES="$GPU_5080,$GPU_3060"; SPLIT=(--device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 85,15) ;; medium) CONTEXT=160000; UBATCH=512; DEVICES="$GPU_5080,$GPU_3060"; SPLIT=(--device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 85,15) ;;
ultra) CONTEXT=262144; UBATCH=128; DEVICES="$GPU_5080,$GPU_3060"; SPLIT=(--device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20) ;; ultra) CONTEXT=262144; UBATCH=128; DEVICES="$GPU_5080,$GPU_3060"; SPLIT=(--device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20) ;;
vision) CONTEXT=76800; UBATCH=64; DEVICES="$GPU_5080,$GPU_3060"; SPLIT=(--device CUDA0 --split-mode none) ;;
esac esac
VISION=()
if [[ $CASE == vision ]]; then
test -s /data/models/bonsai2-ab/Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf || {
echo 'Verified vision projector is missing' >&2; exit 1;
}
VISION=(--mmproj /models/Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf --mmproj-offload --mmproj-device CUDA1)
fi
docker create --name "$NAME" --gpus "\"device=$DEVICES\"" \ docker create --name "$NAME" --gpus "\"device=$DEVICES\"" \
--network bridge -p 127.0.0.1:5006:8080 \ --network bridge -p 127.0.0.1:5006:8080 \
--read-only --tmpfs /tmp:rw,noexec,nosuid,nodev,size=256m \ --read-only --tmpfs /tmp:rw,noexec,nosuid,nodev,size=256m \
@@ -50,5 +58,5 @@ docker create --name "$NAME" --gpus "\"device=$DEVICES\"" \
--cache-prompt --threads 6 --threads-batch 6 --batch-size 2048 \ --cache-prompt --threads 6 --threads-batch 6 --batch-size 2048 \
--ubatch-size "$UBATCH" --parallel 1 --jinja --reasoning auto \ --ubatch-size "$UBATCH" --parallel 1 --jinja --reasoning auto \
--host 0.0.0.0 --port 8080 --metrics --fit off --n-gpu-layers all \ --host 0.0.0.0 --port 8080 --metrics --fit off --n-gpu-layers all \
--temperature 1.0 --top-p 0.95 --top-k 20 "${SPLIT[@]}" >/dev/null --temperature 1.0 --top-p 0.95 --top-k 20 "${VISION[@]}" "${SPLIT[@]}" >/dev/null
echo "Created stopped Bonsai $CASE case ($CONTEXT context)." echo "Created stopped Bonsai $CASE case ($CONTEXT context)."
+9
View File
@@ -6,6 +6,8 @@ MODEL_DIR=/data/models/bonsai2-ab
FILE=Ternary-Bonsai-2-27B-PQ2_0.gguf FILE=Ternary-Bonsai-2-27B-PQ2_0.gguf
REVISION=6ed5e12bf84b7a63069882c91dd9e9218647d17b REVISION=6ed5e12bf84b7a63069882c91dd9e9218647d17b
SHA256=3907dc1658db1f78a9826bf8d5bcb8dc65db0d466388937af57f2294fae62ec1 SHA256=3907dc1658db1f78a9826bf8d5bcb8dc65db0d466388937af57f2294fae62ec1
PROJECTOR=Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf
PROJECTOR_SHA256=6807ede61d570bb86ba34b756a0fa109edc33668604de867c6ea6d8f1d631903
IMAGE=mike-ai/bonsai2-ab:prism-b10685 IMAGE=mike-ai/bonsai2-ab:prism-b10685
ARCHIVE=prism-cuda.tar.gz ARCHIVE=prism-cuda.tar.gz
ARCHIVE_SHA256=4ec1572702fa3fd359653528fa5625fdd9c7ae02dedab7146da7773dae48cf2c ARCHIVE_SHA256=4ec1572702fa3fd359653528fa5625fdd9c7ae02dedab7146da7773dae48cf2c
@@ -20,6 +22,13 @@ if ! (cd "$MODEL_DIR" && printf '%s %s\n' "$SHA256" "$FILE" | sha256sum -c --st
(cd "$MODEL_DIR" && printf '%s %s\n' "$SHA256" "$FILE.part" | sha256sum -c) (cd "$MODEL_DIR" && printf '%s %s\n' "$SHA256" "$FILE.part" | sha256sum -c)
mv "$MODEL_DIR/$FILE.part" "$MODEL_DIR/$FILE" mv "$MODEL_DIR/$FILE.part" "$MODEL_DIR/$FILE"
fi fi
if ! (cd "$MODEL_DIR" && printf '%s %s\n' "$PROJECTOR_SHA256" "$PROJECTOR" | sha256sum -c --status); then
curl --fail --location --retry 5 --continue-at - \
--output "$MODEL_DIR/$PROJECTOR.part" \
"https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/$REVISION/$PROJECTOR"
(cd "$MODEL_DIR" && printf '%s %s\n' "$PROJECTOR_SHA256" "$PROJECTOR.part" | sha256sum -c)
mv "$MODEL_DIR/$PROJECTOR.part" "$MODEL_DIR/$PROJECTOR"
fi
if ! printf '%s %s\n' "$ARCHIVE_SHA256" "$ARCHIVE" | sha256sum -c --status; then if ! printf '%s %s\n' "$ARCHIVE_SHA256" "$ARCHIVE" | sha256sum -c --status; then
curl --fail --location --retry 5 --continue-at - --output "$ARCHIVE.part" "$ARCHIVE_URL" curl --fail --location --retry 5 --continue-at - --output "$ARCHIVE.part" "$ARCHIVE_URL"
printf '%s %s\n' "$ARCHIVE_SHA256" "$ARCHIVE.part" | sha256sum -c printf '%s %s\n' "$ARCHIVE_SHA256" "$ARCHIVE.part" | sha256sum -c
+43
View File
@@ -0,0 +1,43 @@
#!/usr/bin/env bash
set -Eeuo pipefail
[[ ${1:-} == --go && $(hostname) == athena ]] || { echo 'Usage on Athena: repeat-go.sh --go' >&2; exit 2; }
cd /opt/mike-ai/experiments/bonsai2-ab
OUT=/data/benchmarks/bonsai2-ab
ORIGINAL=$(docker ps --format '{{.Names}}' | sed -n 's/^mike-ai-llama-\(fast\|medium\|large\|ultra\|uncensored\)$/\1/p')
[[ $(wc -w <<<"$ORIGINAL") == 1 ]] || exit 1
controller() {
docker exec mike-ai-profile-controller python3 -c '
import os, sys, urllib.request
token = os.environ.get("CONTROLLER_TOKEN", "").strip()
if not token:
token = open(os.environ.get("CONTROLLER_TOKEN_FILE", "/run/secrets/controller-token"), encoding="utf-8").read().strip()
req = urllib.request.Request("http://127.0.0.1:8090" + sys.argv[1], data=b"{}", headers={"Authorization": "Bearer " + token, "Content-Type": "application/json"}, method="POST")
print(urllib.request.urlopen(req, timeout=180).read().decode())
' "$1"
}
finish() {
local rc=$?
trap - EXIT INT TERM
./case.sh fast stop || true
controller "/profiles/$ORIGINAL/activate" || true
echo "REPEAT_EXIT=$rc RESTORED=$ORIGINAL"
exit "$rc"
}
trap finish EXIT INT TERM
wait_health() {
local url=$1
for ((i=0;i<120;i++)); do
curl -fsS --max-time 2 "$url/health" >/dev/null 2>&1 && return 0
sleep 2
done
return 1
}
controller /profiles/fast/activate
ip=$(docker inspect mike-ai-llama-fast --format '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}')
wait_health "http://$ip:8080"
python3 repeat_critical.py --base "http://$ip:8080" --model qwen-fast --output "$OUT/repeat-qwen-fast.json" --go
controller /inference/stop
./case.sh fast create
./case.sh fast start --go
wait_health http://127.0.0.1:5006
python3 repeat_critical.py --base http://127.0.0.1:5006 --model bonsai2-test --output "$OUT/repeat-bonsai-fast.json" --go
+52
View File
@@ -0,0 +1,52 @@
#!/usr/bin/env python3
"""Repeat the two consequential failures and a 49k-token recall probe."""
import argparse
import json
import pathlib
import time
from measure import request
def main():
p = argparse.ArgumentParser()
p.add_argument("--base", required=True)
p.add_argument("--model", required=True)
p.add_argument("--output", type=pathlib.Path, required=True)
p.add_argument("--temperature", type=float, default=0.2)
p.add_argument("--seeds", default="7,99")
p.add_argument("--skip-long", action="store_true")
p.add_argument("--go", action="store_true")
args = p.parse_args()
if not args.go:
p.error("Inference requires explicit --go")
tasks = {item["id"]: item for item in json.loads(pathlib.Path(__file__).with_name("tasks.json").read_text())}
seeds = [int(value) for value in args.seeds.split(",")]
report = {"model": args.model, "temperature": args.temperature,
"started": time.time(), "repeats": []}
for task_id in ("i3_code_debugging", "i6_state_vs_configuration"):
for seed in seeds:
payload = {"model": args.model, "stream": False, "temperature": args.temperature,
"seed": seed, "reasoning_effort": "medium", "max_tokens": 4096,
"messages": [{"role": "user", "content": tasks[task_id]["prompt"]}]}
answer = request(args.base, None, payload)
report["repeats"].append({"id": task_id, "seed": seed, **answer})
print(task_id, seed, answer["wall_seconds"], flush=True)
if not args.skip_long:
filler = ("In one sentence, explain why a 10 mm through-hole in a 40 mm cube "
"does not change its external dimensions. " * 1800)
prompt = "Merke dir diesen einmaligen Prüfcode: KIESEL-7319.\n" + filler + \
"\nNenne ausschließlich den Prüfcode aus dem Anfang der Nachricht."
payload = {"model": args.model, "stream": False, "temperature": 0,
"seed": 42, "reasoning_effort": "none", "max_tokens": 96,
"messages": [{"role": "user", "content": prompt}]}
report["long_recall"] = request(args.base, None, payload)
report["finished"] = time.time()
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(json.dumps(report, ensure_ascii=False, indent=2) + "\n")
if not args.skip_long:
print("long_recall", report["long_recall"]["wall_seconds"], flush=True)
if __name__ == "__main__":
main()
+47
View File
@@ -0,0 +1,47 @@
# Bonsai 2 27B A/B result on Athena
Measured 2026-09-19 against the production Qwen3.8-27B profiles. The same nine text tasks, tool schema, 27,013-token prefill, 2,048-token decode and physical GPU sampling were used. Lower text-task time is better; prefill and decode are tokens/s. Qwen used its production MTP speculative decoder. Bonsai used PrismML's pinned PQ2_0 runtime.
| Profile | Model | Nine tasks | Prefill | Decode | Weather tool |
|---|---|---:|---:|---:|---|
| Fast 76.8K | Qwen | 191.1 s | 1,305 | 84.7 | correct |
| Fast 76.8K | Bonsai | 236.7 s | 1,480 | 83.9 | correct |
| Medium 160K | Qwen | 219.6 s | 1,415 | 65.5 | correct |
| Medium 160K | Bonsai | 247.5 s | 2,250 | 67.6 | correct |
| Ultra 262K | Qwen | 265.2 s | 1,632 | 62.6 | correct |
| Ultra 262K | Bonsai | 268.8 s | 1,506 | 63.6 | correct |
The total task time includes the model's chosen answer length, so it represents actual waiting time rather than pure kernel speed. Bonsai's Medium prefill was 59% faster, but its nine answers still took 13% longer. Fast took 24% longer. Ultra was effectively tied. Repeating the same long prompt hit the cache for both models.
## GPU memory measured while generating
Values are total board allocation and include the already-running Athena services. They remain directly comparable because each Qwen/Bonsai pair was sampled in the same test window.
| Profile | Model | RTX 5080 | RTX 3060 |
|---|---|---:|---:|
| Fast | Qwen | 15,832 MiB | 5,926 MiB |
| Fast | Bonsai text | 8,728 MiB | 4,679 MiB |
| Fast | Bonsai vision | 8,728 MiB | 5,676 MiB |
| Medium | Qwen | 15,700 MiB | 11,468 MiB |
| Medium | Bonsai | 9,786 MiB | 7,686 MiB |
| Ultra | Qwen | 15,780 MiB | 11,180 MiB |
| Ultra | Bonsai | 10,598 MiB | 8,528 MiB |
Bonsai is the clear memory winner. It saved about 7.1 GiB on the 5080 in Fast and about 5.9/3.8 GiB across the 5080/3060 in Medium. No case OOMed. Ultra left about 5.3 GiB free on the 5080 and 3.4 GiB on the 3060 at model start.
## Correctness and stability
All six main runs solved the basic logic, evidence, capacity, prompt-injection, refusal and missing-tool tasks. Both models produced the required `get_weather({"city":"Rastatt"})` call. Both recalled `KIESEL-7319` from a 48,644-token prompt; Bonsai took 36.3 s and Qwen 40.5 s. Both identified the synthetic vision fixture as a red circle on the left and a blue square on the right.
Bonsai Fast was not reliable enough to replace Qwen:
- In the first coding run it stopped after the first completed async task even when that task had failed. Across three Fast runs at temperature 0.2, only one supplied correct first-*successful*-task logic.
- In the Home Assistant state task it twice claimed that an `off` automation entity merely meant "idle". Home Assistant's official `automation.turn_off` documentation says an off automation is disabled and no longer listens for triggers. Across three Fast runs at temperature 0.2, only one was correct.
- Repeating both tasks with PrismML's recommended thinking sampling (`temperature=1.0`, `top_p=0.95`, `top_k=20`) did not fix the variance: one of three answers was correct for each consequential task. It also answered more slowly.
- Bonsai Medium and Ultra answered both main-run critical cases correctly, but the model weights are identical. The difference is generation variance, not evidence that the larger context profile makes the model smarter.
Qwen returned correct answers for all of these critical runs. The practical decision is therefore **keep Qwen as Athena's production model**. Bonsai is useful as a stopped, optional memory-saving profile for experiments or workloads where freeing VRAM matters more than maximum agent reliability. Do not silently replace Fast, Medium or Ultra with it.
The official model card reports 98.2% of FP16 aggregate benchmark performance and provides a separate Q8_0 vision projector. Our result does not contradict that aggregate score; it shows that a small average loss can still appear as a consequential intermittent error in an agent workflow. Sources: [PrismML model card](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [Bonsai demo guide](https://github.com/PrismML-Eng/Bonsai-demo), [Home Assistant turn-off semantics](https://www.home-assistant.io/actions/automation.turn_off/).
After every run, the trap restored the original `medium` profile. Final verification: controller reported `active_profile=medium`, router and model containers were healthy, and Qwen Medium returned `OK` to a live completion request. The Bonsai container is stopped. `cleanup.sh --all` removes the experiment container, image, model, projector, raw server results and deployment staging without touching production.
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,32 @@
{
"model": "bonsai2-test",
"prompt": "Was ist links und was ist rechts im Bild? Nenne Form und Farbe beider Objekte.",
"expected": "Links ein roter Kreis, rechts ein blaues Quadrat.",
"started": 1789775633.9445844,
"response": {
"wall_seconds": 3.703,
"usage": {
"completion_tokens": 40,
"prompt_tokens": 162,
"total_tokens": 202,
"prompt_tokens_details": {
"cached_tokens": 0
}
},
"timings": {
"cache_n": 0,
"prompt_n": 162,
"prompt_ms": 2738.932,
"prompt_per_token_ms": 16.906987654320986,
"prompt_per_second": 59.1471420247016,
"predicted_n": 40,
"predicted_ms": 952.303,
"predicted_per_token_ms": 24.41802564102564,
"predicted_per_second": 40.95335203186381
},
"finish_reason": "stop",
"content": "**Links:**\n* **Form:** Kreis\n* **Farbe:** Rot\n\n**Rechts:**\n* **Form:** Quadrat\n* **Farbe:** Blau",
"reasoning_content": "",
"tool_calls": []
}
}
@@ -0,0 +1,34 @@
{
"model": "qwen-fast",
"prompt": "Was ist links und was ist rechts im Bild? Nenne Form und Farbe beider Objekte.",
"expected": "Links ein roter Kreis, rechts ein blaues Quadrat.",
"started": 1789775622.0384746,
"response": {
"wall_seconds": 2.664,
"usage": {
"completion_tokens": 84,
"prompt_tokens": 162,
"total_tokens": 246,
"prompt_tokens_details": {
"cached_tokens": 0
}
},
"timings": {
"cache_n": 0,
"prompt_n": 162,
"prompt_ms": 1725.751,
"prompt_per_token_ms": 10.652783950617284,
"prompt_per_second": 93.87217507044758,
"predicted_n": 84,
"predicted_ms": 926.95,
"predicted_per_token_ms": 11.168072289156626,
"predicted_per_second": 89.54096768973514,
"draft_n": 64,
"draft_n_accepted": 51
},
"finish_reason": "stop",
"content": "Im Bild sind zwei geometrische Objekte zu sehen:\n\n- **Links**: Ein **roter Kreis**.\n - Form: Kreis\n - Farbe: Rot\n\n- **Rechts**: Ein **blaues Quadrat**.\n - Form: Quadrat\n - Farbe: Blau\n\nZusammenfassend:\n> Links ist ein roter Kreis, rechts ist ein blaues Quadrat.",
"reasoning_content": "",
"tool_calls": []
}
}
+85
View File
@@ -0,0 +1,85 @@
#!/usr/bin/env bash
# Run after the user has explicitly approved GPU inference. Restore the initial
# production profile even if a benchmark or model load fails.
set -Eeuo pipefail
[[ ${1:-} == --go && $(hostname) == athena ]] || { echo 'Usage on Athena: run-go.sh --go' >&2; exit 2; }
cd /opt/mike-ai/experiments/bonsai2-ab
OUT=/data/benchmarks/bonsai2-ab
mkdir -p "$OUT"
ORIGINAL=$(docker ps --format '{{.Names}}' | sed -n 's/^mike-ai-llama-\(fast\|medium\|large\|ultra\|uncensored\)$/\1/p')
[[ $(wc -w <<<"$ORIGINAL") == 1 ]] || { echo 'Expected exactly one active production profile' >&2; exit 1; }
echo "$ORIGINAL" > "$OUT/original-profile.txt"
controller() {
docker exec mike-ai-profile-controller python3 -c '
import os, sys, urllib.request
token = os.environ.get("CONTROLLER_TOKEN", "").strip()
if not token:
token = open(os.environ.get("CONTROLLER_TOKEN_FILE", "/run/secrets/controller-token"), encoding="utf-8").read().strip()
req = urllib.request.Request("http://127.0.0.1:8090" + sys.argv[1], data=b"{}", headers={"Authorization": "Bearer " + token, "Content-Type": "application/json"}, method="POST")
print(urllib.request.urlopen(req, timeout=180).read().decode())
' "$1"
}
monitor_pid=''
finish() {
local rc=$?
trap - EXIT INT TERM
if [[ -n "$monitor_pid" ]]; then kill "$monitor_pid" 2>/dev/null || true; wait "$monitor_pid" 2>/dev/null || true; fi
./case.sh fast stop || true
controller "/profiles/$ORIGINAL/activate" || true
docker ps --format '{{.Names}} {{.Status}}' | grep -E 'mike-ai-(llama-|router|profile-controller|bonsai2-ab)' || true
echo "AB_RUN_EXIT=$rc ORIGINAL=$ORIGINAL" | tee -a "$OUT/run-go.log"
exit "$rc"
}
trap finish EXIT INT TERM
wait_health() {
local url=$1
for ((i=0;i<360;i++)); do
if curl -fsS --max-time 2 "$url/health" >/dev/null 2>&1; then return 0; fi
sleep 2
done
echo "Health timeout: $url" >&2
return 1
}
record() {
local label=$1 base=$2 model=$3
echo "START $label $(date -Is)" | tee -a "$OUT/run-go.log"
nvidia-smi --query-gpu=uuid,name,memory.used,memory.free --format=csv,noheader,nounits > "$OUT/$label-before-gpu.csv"
python3 gpu_monitor.py --output "$OUT/$label-gpu.jsonl" --go >/dev/null 2>&1 &
monitor_pid=$!
python3 measure.py --label "$label" --base "$base" --model "$model" --output "$OUT/$label.json" --go 2>&1 | tee "$OUT/$label.log"
kill "$monitor_pid" 2>/dev/null || true
wait "$monitor_pid" 2>/dev/null || true
monitor_pid=''
echo "END $label $(date -Is)" | tee -a "$OUT/run-go.log"
}
qwen() {
local case=$1 ip
controller "/profiles/$case/activate"
ip=$(docker inspect "mike-ai-llama-$case" --format '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}')
wait_health "http://$ip:8080"
record "qwen-$case" "http://$ip:8080" "qwen-$case"
}
bonsai() {
local case=$1
controller /inference/stop
nvidia-smi --query-gpu=uuid,name,memory.used,memory.free --format=csv,noheader,nounits > "$OUT/bonsai-$case-idle-gpu.csv"
./case.sh "$case" create
./case.sh "$case" start --go
wait_health http://127.0.0.1:5006
record "bonsai-$case" http://127.0.0.1:5006 bonsai2-test
./case.sh "$case" stop
}
# Qwen Medium was measured first while already active, before this script.
[[ -s "$OUT/qwen-medium.json" ]] || { echo 'Qwen Medium baseline incomplete' >&2; exit 1; }
qwen fast
bonsai fast
bonsai medium
qwen ultra
bonsai ultra
+48
View File
@@ -0,0 +1,48 @@
#!/usr/bin/env bash
set -Eeuo pipefail
[[ ${1:-} == --go && $(hostname) == athena ]] || { echo 'Usage on Athena: vision-go.sh --go' >&2; exit 2; }
cd /opt/mike-ai/experiments/bonsai2-ab
OUT=/data/benchmarks/bonsai2-ab
ORIGINAL=$(docker ps --format '{{.Names}}' | sed -n 's/^mike-ai-llama-\(fast\|medium\|large\|ultra\|uncensored\)$/\1/p')
[[ $(wc -w <<<"$ORIGINAL") == 1 ]] || exit 1
controller() {
docker exec mike-ai-profile-controller python3 -c '
import os, sys, urllib.request
token = os.environ.get("CONTROLLER_TOKEN", "").strip()
if not token:
token = open(os.environ.get("CONTROLLER_TOKEN_FILE", "/run/secrets/controller-token"), encoding="utf-8").read().strip()
req = urllib.request.Request("http://127.0.0.1:8090" + sys.argv[1], data=b"{}", headers={"Authorization": "Bearer " + token, "Content-Type": "application/json"}, method="POST")
print(urllib.request.urlopen(req, timeout=180).read().decode())
' "$1"
}
finish() {
local rc=$?
trap - EXIT INT TERM
./case.sh vision stop || true
controller "/profiles/$ORIGINAL/activate" || true
echo "VISION_EXIT=$rc RESTORED=$ORIGINAL"
exit "$rc"
}
trap finish EXIT INT TERM
wait_health() {
local url=$1
for ((i=0;i<120;i++)); do
curl -fsS --max-time 2 "$url/health" >/dev/null 2>&1 && return 0
sleep 2
done
return 1
}
controller /profiles/fast/activate
ip=$(docker inspect mike-ai-llama-fast --format '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}')
wait_health "http://$ip:8080"
python3 vision_probe.py --base "http://$ip:8080" --model qwen-fast --output "$OUT/vision-qwen-fast.json" --go
controller /inference/stop
./case.sh vision create
./case.sh vision start --go
wait_health http://127.0.0.1:5006
nvidia-smi --query-gpu=uuid,memory.used,memory.free --format=csv,noheader,nounits > "$OUT/vision-bonsai-before-gpu.csv"
python3 vision_probe.py --base http://127.0.0.1:5006 --model bonsai2-test --output "$OUT/vision-bonsai-fast.json" --go
nvidia-smi --query-gpu=uuid,memory.used,memory.free --format=csv,noheader,nounits > "$OUT/vision-bonsai-after-gpu.csv"
python3 repeat_critical.py --base http://127.0.0.1:5006 --model bonsai2-test \
--temperature 1.0 --seeds 7,42,99 --skip-long \
--output "$OUT/recommended-bonsai-fast.json" --go
+38
View File
@@ -0,0 +1,38 @@
#!/usr/bin/env python3
"""Small, fully synthetic vision smoke test for the isolated A/B servers."""
import argparse
import base64
import json
import pathlib
import time
from measure import request
def main():
p = argparse.ArgumentParser()
p.add_argument("--base", required=True)
p.add_argument("--model", required=True)
p.add_argument("--output", type=pathlib.Path, required=True)
p.add_argument("--go", action="store_true")
args = p.parse_args()
if not args.go:
p.error("Inference requires explicit --go")
image = pathlib.Path(__file__).with_name("assets") / "vision-fixture.png"
url = "data:image/png;base64," + base64.b64encode(image.read_bytes()).decode("ascii")
prompt = "Was ist links und was ist rechts im Bild? Nenne Form und Farbe beider Objekte."
payload = {"model": args.model, "stream": False, "temperature": 0,
"reasoning_effort": "none", "max_tokens": 512,
"messages": [{"role": "user", "content": [
{"type": "text", "text": prompt},
{"type": "image_url", "image_url": {"url": url}}]}]}
result = {"model": args.model, "prompt": prompt, "expected":
"Links ein roter Kreis, rechts ein blaues Quadrat.",
"started": time.time(), "response": request(args.base, None, payload)}
args.output.parent.mkdir(parents=True, exist_ok=True)
args.output.write_text(json.dumps(result, ensure_ascii=False, indent=2) + "\n")
print(result["response"]["content"])
if __name__ == "__main__":
main()