Document Bonsai 2 A/B results
This commit is contained in:
@@ -1,17 +1,18 @@
|
||||
# Bonsai 2 A/B on Athena
|
||||
|
||||
Preparation is safe for the live router: `prepare.sh` downloads the pinned 7.21 GB PQ2_0 model, checks SHA-256, verifies PrismML's official Linux CUDA 12.8 archive from release `prism-b10685-7dffb15`, builds an isolated runtime image, and **creates a stopped container**. It never starts inference. The model and image are separate from production. A running production Qwen container blocks `case.sh ... start`.
|
||||
Preparation is safe for the live router: `prepare.sh` downloads the pinned 7.21 GB PQ2_0 model and optional 629 MB Q8_0 vision projector, checks both SHA-256 values, verifies PrismML's official Linux CUDA 12.8 archive from release `prism-b10685-7dffb15`, builds an isolated runtime image, and **creates a stopped container**. It never starts inference. The model and image are separate from production. A running production Qwen container blocks `case.sh ... start`.
|
||||
|
||||
| Case | Context | GPUs inside container | Split | u-batch |
|
||||
|---|---:|---|---:|---:|
|
||||
| Fast | 76,800 | 5080 | 100:0 | 64 |
|
||||
| Medium | 160,000 | 5080 + 3060 | 85:15 | 512 |
|
||||
| Ultra | 262,144 | 5080 + 3060 | 80:20 | 128 |
|
||||
| Vision | 76,800 | 5080 + projector on 3060 | 100:0 language model | 64 |
|
||||
|
||||
The A/B matrix compares production Qwen against Bonsai at **the same context profile**. Each run uses the same nine acceptance tasks, one tool-use probe, matched repeated prefill, and a long answer. Save total latency, prompt/decode throughput from llama.cpp timings, objective task accuracy, tool-call correctness, actual VRAM per GPU, CPU offload and any OOM. Assess the free-form answers blind, since a single numerical score does not measure intelligence well. This is a text-only comparison; Qwen Fast/Medium vision remains a separate feature. Qwen's production MTP speculative decoder stays on; report that speed advantage explicitly rather than attributing every difference to model weights.
|
||||
|
||||
After the user's explicit **GO**, use the official profile controller/router to switch or stop production Qwen, record the original profile, run the corresponding Bonsai case with `case.sh CASE start --go`, execute `measure.py ... --go`, and sample both physical GPUs with `gpu_monitor.py --output ... --go`. Record idle GPU memory before each model load so the model's incremental VRAM is distinguishable from other processes. Stop Bonsai, restore the original production profile via the controller, and verify the router and original model are healthy. Repeat case by case. Never run both model servers at once. The smallest context is tested first; only try Medium and Ultra if Fast fits. Results live at `/data/benchmarks/bonsai2-ab` and should be checked for OOM/offload before drawing any speed conclusion. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA `cuobjdump`: it includes `sm_86` and `sm_120a`, covering both cards.
|
||||
After the user's explicit **GO**, `run-go.sh --go` runs the six text cases, `repeat-go.sh --go` repeats consequential failures and a 48K-token recall probe, and `vision-go.sh --go` compares the synthetic image fixture and checks PrismML's recommended sampling. Each script records the original profile, never overlaps Qwen and Bonsai, and restores the original profile from an EXIT/INT/TERM trap. Results live at `/data/benchmarks/bonsai2-ab`; the reviewed, non-secret JSON reports and conclusion are committed under `results/`. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA `cuobjdump`: it includes `sm_86` and `sm_120a`, covering both cards.
|
||||
|
||||
Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results, deployment directory and build-cache records identifiable as this experiment. Docker's shared runtime base layers are deliberately not globally pruned, because that could delete unrelated build caches.
|
||||
|
||||
Sources: [PrismML model](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [PrismML llama.cpp CUDA release](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15). Stock llama.cpp is intentionally not used for this ternary model.
|
||||
Measured conclusions are in [`RESULTS.md`](results/RESULTS.md). Sources: [PrismML model](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [PrismML demo and integration guide](https://github.com/PrismML-Eng/Bonsai-demo), [PrismML llama.cpp CUDA release](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15). Stock llama.cpp is intentionally not used for this ternary model.
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 1.6 KiB |
@@ -8,8 +8,8 @@ IMAGE=mike-ai/bonsai2-ab:prism-b10685
|
||||
MODEL=/data/models/bonsai2-ab/Ternary-Bonsai-2-27B-PQ2_0.gguf
|
||||
GPU_5080=GPU-8ad38c6c-5a01-9d8e-1dfa-ed662ad78fbe
|
||||
GPU_3060=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
|
||||
[[ $CASE =~ ^(fast|medium|ultra)$ && $ACTION =~ ^(create|start|stop)$ ]] || {
|
||||
echo 'Usage: case.sh {fast|medium|ultra} {create|start|stop}' >&2; exit 2;
|
||||
[[ $CASE =~ ^(fast|medium|ultra|vision)$ && $ACTION =~ ^(create|start|stop)$ ]] || {
|
||||
echo 'Usage: case.sh {fast|medium|ultra|vision} {create|start|stop}' >&2; exit 2;
|
||||
}
|
||||
test "$(hostname)" = athena || exit 2
|
||||
if [[ $ACTION == stop ]]; then
|
||||
@@ -36,7 +36,15 @@ case "$CASE" in
|
||||
fast) CONTEXT=76800; UBATCH=64; DEVICES="$GPU_5080"; SPLIT=(--device CUDA0 --split-mode none) ;;
|
||||
medium) CONTEXT=160000; UBATCH=512; DEVICES="$GPU_5080,$GPU_3060"; SPLIT=(--device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 85,15) ;;
|
||||
ultra) CONTEXT=262144; UBATCH=128; DEVICES="$GPU_5080,$GPU_3060"; SPLIT=(--device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20) ;;
|
||||
vision) CONTEXT=76800; UBATCH=64; DEVICES="$GPU_5080,$GPU_3060"; SPLIT=(--device CUDA0 --split-mode none) ;;
|
||||
esac
|
||||
VISION=()
|
||||
if [[ $CASE == vision ]]; then
|
||||
test -s /data/models/bonsai2-ab/Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf || {
|
||||
echo 'Verified vision projector is missing' >&2; exit 1;
|
||||
}
|
||||
VISION=(--mmproj /models/Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf --mmproj-offload --mmproj-device CUDA1)
|
||||
fi
|
||||
docker create --name "$NAME" --gpus "\"device=$DEVICES\"" \
|
||||
--network bridge -p 127.0.0.1:5006:8080 \
|
||||
--read-only --tmpfs /tmp:rw,noexec,nosuid,nodev,size=256m \
|
||||
@@ -50,5 +58,5 @@ docker create --name "$NAME" --gpus "\"device=$DEVICES\"" \
|
||||
--cache-prompt --threads 6 --threads-batch 6 --batch-size 2048 \
|
||||
--ubatch-size "$UBATCH" --parallel 1 --jinja --reasoning auto \
|
||||
--host 0.0.0.0 --port 8080 --metrics --fit off --n-gpu-layers all \
|
||||
--temperature 1.0 --top-p 0.95 --top-k 20 "${SPLIT[@]}" >/dev/null
|
||||
--temperature 1.0 --top-p 0.95 --top-k 20 "${VISION[@]}" "${SPLIT[@]}" >/dev/null
|
||||
echo "Created stopped Bonsai $CASE case ($CONTEXT context)."
|
||||
|
||||
@@ -6,6 +6,8 @@ MODEL_DIR=/data/models/bonsai2-ab
|
||||
FILE=Ternary-Bonsai-2-27B-PQ2_0.gguf
|
||||
REVISION=6ed5e12bf84b7a63069882c91dd9e9218647d17b
|
||||
SHA256=3907dc1658db1f78a9826bf8d5bcb8dc65db0d466388937af57f2294fae62ec1
|
||||
PROJECTOR=Ternary-Bonsai-2-27B-mmproj-Q8_0.gguf
|
||||
PROJECTOR_SHA256=6807ede61d570bb86ba34b756a0fa109edc33668604de867c6ea6d8f1d631903
|
||||
IMAGE=mike-ai/bonsai2-ab:prism-b10685
|
||||
ARCHIVE=prism-cuda.tar.gz
|
||||
ARCHIVE_SHA256=4ec1572702fa3fd359653528fa5625fdd9c7ae02dedab7146da7773dae48cf2c
|
||||
@@ -20,6 +22,13 @@ if ! (cd "$MODEL_DIR" && printf '%s %s\n' "$SHA256" "$FILE" | sha256sum -c --st
|
||||
(cd "$MODEL_DIR" && printf '%s %s\n' "$SHA256" "$FILE.part" | sha256sum -c)
|
||||
mv "$MODEL_DIR/$FILE.part" "$MODEL_DIR/$FILE"
|
||||
fi
|
||||
if ! (cd "$MODEL_DIR" && printf '%s %s\n' "$PROJECTOR_SHA256" "$PROJECTOR" | sha256sum -c --status); then
|
||||
curl --fail --location --retry 5 --continue-at - \
|
||||
--output "$MODEL_DIR/$PROJECTOR.part" \
|
||||
"https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf/resolve/$REVISION/$PROJECTOR"
|
||||
(cd "$MODEL_DIR" && printf '%s %s\n' "$PROJECTOR_SHA256" "$PROJECTOR.part" | sha256sum -c)
|
||||
mv "$MODEL_DIR/$PROJECTOR.part" "$MODEL_DIR/$PROJECTOR"
|
||||
fi
|
||||
if ! printf '%s %s\n' "$ARCHIVE_SHA256" "$ARCHIVE" | sha256sum -c --status; then
|
||||
curl --fail --location --retry 5 --continue-at - --output "$ARCHIVE.part" "$ARCHIVE_URL"
|
||||
printf '%s %s\n' "$ARCHIVE_SHA256" "$ARCHIVE.part" | sha256sum -c
|
||||
|
||||
@@ -0,0 +1,43 @@
|
||||
#!/usr/bin/env bash
|
||||
set -Eeuo pipefail
|
||||
[[ ${1:-} == --go && $(hostname) == athena ]] || { echo 'Usage on Athena: repeat-go.sh --go' >&2; exit 2; }
|
||||
cd /opt/mike-ai/experiments/bonsai2-ab
|
||||
OUT=/data/benchmarks/bonsai2-ab
|
||||
ORIGINAL=$(docker ps --format '{{.Names}}' | sed -n 's/^mike-ai-llama-\(fast\|medium\|large\|ultra\|uncensored\)$/\1/p')
|
||||
[[ $(wc -w <<<"$ORIGINAL") == 1 ]] || exit 1
|
||||
controller() {
|
||||
docker exec mike-ai-profile-controller python3 -c '
|
||||
import os, sys, urllib.request
|
||||
token = os.environ.get("CONTROLLER_TOKEN", "").strip()
|
||||
if not token:
|
||||
token = open(os.environ.get("CONTROLLER_TOKEN_FILE", "/run/secrets/controller-token"), encoding="utf-8").read().strip()
|
||||
req = urllib.request.Request("http://127.0.0.1:8090" + sys.argv[1], data=b"{}", headers={"Authorization": "Bearer " + token, "Content-Type": "application/json"}, method="POST")
|
||||
print(urllib.request.urlopen(req, timeout=180).read().decode())
|
||||
' "$1"
|
||||
}
|
||||
finish() {
|
||||
local rc=$?
|
||||
trap - EXIT INT TERM
|
||||
./case.sh fast stop || true
|
||||
controller "/profiles/$ORIGINAL/activate" || true
|
||||
echo "REPEAT_EXIT=$rc RESTORED=$ORIGINAL"
|
||||
exit "$rc"
|
||||
}
|
||||
trap finish EXIT INT TERM
|
||||
wait_health() {
|
||||
local url=$1
|
||||
for ((i=0;i<120;i++)); do
|
||||
curl -fsS --max-time 2 "$url/health" >/dev/null 2>&1 && return 0
|
||||
sleep 2
|
||||
done
|
||||
return 1
|
||||
}
|
||||
controller /profiles/fast/activate
|
||||
ip=$(docker inspect mike-ai-llama-fast --format '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}')
|
||||
wait_health "http://$ip:8080"
|
||||
python3 repeat_critical.py --base "http://$ip:8080" --model qwen-fast --output "$OUT/repeat-qwen-fast.json" --go
|
||||
controller /inference/stop
|
||||
./case.sh fast create
|
||||
./case.sh fast start --go
|
||||
wait_health http://127.0.0.1:5006
|
||||
python3 repeat_critical.py --base http://127.0.0.1:5006 --model bonsai2-test --output "$OUT/repeat-bonsai-fast.json" --go
|
||||
@@ -0,0 +1,52 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Repeat the two consequential failures and a 49k-token recall probe."""
|
||||
import argparse
|
||||
import json
|
||||
import pathlib
|
||||
import time
|
||||
|
||||
from measure import request
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser()
|
||||
p.add_argument("--base", required=True)
|
||||
p.add_argument("--model", required=True)
|
||||
p.add_argument("--output", type=pathlib.Path, required=True)
|
||||
p.add_argument("--temperature", type=float, default=0.2)
|
||||
p.add_argument("--seeds", default="7,99")
|
||||
p.add_argument("--skip-long", action="store_true")
|
||||
p.add_argument("--go", action="store_true")
|
||||
args = p.parse_args()
|
||||
if not args.go:
|
||||
p.error("Inference requires explicit --go")
|
||||
tasks = {item["id"]: item for item in json.loads(pathlib.Path(__file__).with_name("tasks.json").read_text())}
|
||||
seeds = [int(value) for value in args.seeds.split(",")]
|
||||
report = {"model": args.model, "temperature": args.temperature,
|
||||
"started": time.time(), "repeats": []}
|
||||
for task_id in ("i3_code_debugging", "i6_state_vs_configuration"):
|
||||
for seed in seeds:
|
||||
payload = {"model": args.model, "stream": False, "temperature": args.temperature,
|
||||
"seed": seed, "reasoning_effort": "medium", "max_tokens": 4096,
|
||||
"messages": [{"role": "user", "content": tasks[task_id]["prompt"]}]}
|
||||
answer = request(args.base, None, payload)
|
||||
report["repeats"].append({"id": task_id, "seed": seed, **answer})
|
||||
print(task_id, seed, answer["wall_seconds"], flush=True)
|
||||
if not args.skip_long:
|
||||
filler = ("In one sentence, explain why a 10 mm through-hole in a 40 mm cube "
|
||||
"does not change its external dimensions. " * 1800)
|
||||
prompt = "Merke dir diesen einmaligen Prüfcode: KIESEL-7319.\n" + filler + \
|
||||
"\nNenne ausschließlich den Prüfcode aus dem Anfang der Nachricht."
|
||||
payload = {"model": args.model, "stream": False, "temperature": 0,
|
||||
"seed": 42, "reasoning_effort": "none", "max_tokens": 96,
|
||||
"messages": [{"role": "user", "content": prompt}]}
|
||||
report["long_recall"] = request(args.base, None, payload)
|
||||
report["finished"] = time.time()
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
args.output.write_text(json.dumps(report, ensure_ascii=False, indent=2) + "\n")
|
||||
if not args.skip_long:
|
||||
print("long_recall", report["long_recall"]["wall_seconds"], flush=True)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,47 @@
|
||||
# Bonsai 2 27B A/B result on Athena
|
||||
|
||||
Measured 2026-09-19 against the production Qwen3.8-27B profiles. The same nine text tasks, tool schema, 27,013-token prefill, 2,048-token decode and physical GPU sampling were used. Lower text-task time is better; prefill and decode are tokens/s. Qwen used its production MTP speculative decoder. Bonsai used PrismML's pinned PQ2_0 runtime.
|
||||
|
||||
| Profile | Model | Nine tasks | Prefill | Decode | Weather tool |
|
||||
|---|---|---:|---:|---:|---|
|
||||
| Fast 76.8K | Qwen | 191.1 s | 1,305 | 84.7 | correct |
|
||||
| Fast 76.8K | Bonsai | 236.7 s | 1,480 | 83.9 | correct |
|
||||
| Medium 160K | Qwen | 219.6 s | 1,415 | 65.5 | correct |
|
||||
| Medium 160K | Bonsai | 247.5 s | 2,250 | 67.6 | correct |
|
||||
| Ultra 262K | Qwen | 265.2 s | 1,632 | 62.6 | correct |
|
||||
| Ultra 262K | Bonsai | 268.8 s | 1,506 | 63.6 | correct |
|
||||
|
||||
The total task time includes the model's chosen answer length, so it represents actual waiting time rather than pure kernel speed. Bonsai's Medium prefill was 59% faster, but its nine answers still took 13% longer. Fast took 24% longer. Ultra was effectively tied. Repeating the same long prompt hit the cache for both models.
|
||||
|
||||
## GPU memory measured while generating
|
||||
|
||||
Values are total board allocation and include the already-running Athena services. They remain directly comparable because each Qwen/Bonsai pair was sampled in the same test window.
|
||||
|
||||
| Profile | Model | RTX 5080 | RTX 3060 |
|
||||
|---|---|---:|---:|
|
||||
| Fast | Qwen | 15,832 MiB | 5,926 MiB |
|
||||
| Fast | Bonsai text | 8,728 MiB | 4,679 MiB |
|
||||
| Fast | Bonsai vision | 8,728 MiB | 5,676 MiB |
|
||||
| Medium | Qwen | 15,700 MiB | 11,468 MiB |
|
||||
| Medium | Bonsai | 9,786 MiB | 7,686 MiB |
|
||||
| Ultra | Qwen | 15,780 MiB | 11,180 MiB |
|
||||
| Ultra | Bonsai | 10,598 MiB | 8,528 MiB |
|
||||
|
||||
Bonsai is the clear memory winner. It saved about 7.1 GiB on the 5080 in Fast and about 5.9/3.8 GiB across the 5080/3060 in Medium. No case OOMed. Ultra left about 5.3 GiB free on the 5080 and 3.4 GiB on the 3060 at model start.
|
||||
|
||||
## Correctness and stability
|
||||
|
||||
All six main runs solved the basic logic, evidence, capacity, prompt-injection, refusal and missing-tool tasks. Both models produced the required `get_weather({"city":"Rastatt"})` call. Both recalled `KIESEL-7319` from a 48,644-token prompt; Bonsai took 36.3 s and Qwen 40.5 s. Both identified the synthetic vision fixture as a red circle on the left and a blue square on the right.
|
||||
|
||||
Bonsai Fast was not reliable enough to replace Qwen:
|
||||
|
||||
- In the first coding run it stopped after the first completed async task even when that task had failed. Across three Fast runs at temperature 0.2, only one supplied correct first-*successful*-task logic.
|
||||
- In the Home Assistant state task it twice claimed that an `off` automation entity merely meant "idle". Home Assistant's official `automation.turn_off` documentation says an off automation is disabled and no longer listens for triggers. Across three Fast runs at temperature 0.2, only one was correct.
|
||||
- Repeating both tasks with PrismML's recommended thinking sampling (`temperature=1.0`, `top_p=0.95`, `top_k=20`) did not fix the variance: one of three answers was correct for each consequential task. It also answered more slowly.
|
||||
- Bonsai Medium and Ultra answered both main-run critical cases correctly, but the model weights are identical. The difference is generation variance, not evidence that the larger context profile makes the model smarter.
|
||||
|
||||
Qwen returned correct answers for all of these critical runs. The practical decision is therefore **keep Qwen as Athena's production model**. Bonsai is useful as a stopped, optional memory-saving profile for experiments or workloads where freeing VRAM matters more than maximum agent reliability. Do not silently replace Fast, Medium or Ultra with it.
|
||||
|
||||
The official model card reports 98.2% of FP16 aggregate benchmark performance and provides a separate Q8_0 vision projector. Our result does not contradict that aggregate score; it shows that a small average loss can still appear as a consequential intermittent error in an agent workflow. Sources: [PrismML model card](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [Bonsai demo guide](https://github.com/PrismML-Eng/Bonsai-demo), [Home Assistant turn-off semantics](https://www.home-assistant.io/actions/automation.turn_off/).
|
||||
|
||||
After every run, the trap restored the original `medium` profile. Final verification: controller reported `active_profile=medium`, router and model containers were healthy, and Qwen Medium returned `OK` to a live completion request. The Bonsai container is stopped. `cleanup.sh --all` removes the experiment container, image, model, projector, raw server results and deployment staging without touching production.
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,32 @@
|
||||
{
|
||||
"model": "bonsai2-test",
|
||||
"prompt": "Was ist links und was ist rechts im Bild? Nenne Form und Farbe beider Objekte.",
|
||||
"expected": "Links ein roter Kreis, rechts ein blaues Quadrat.",
|
||||
"started": 1789775633.9445844,
|
||||
"response": {
|
||||
"wall_seconds": 3.703,
|
||||
"usage": {
|
||||
"completion_tokens": 40,
|
||||
"prompt_tokens": 162,
|
||||
"total_tokens": 202,
|
||||
"prompt_tokens_details": {
|
||||
"cached_tokens": 0
|
||||
}
|
||||
},
|
||||
"timings": {
|
||||
"cache_n": 0,
|
||||
"prompt_n": 162,
|
||||
"prompt_ms": 2738.932,
|
||||
"prompt_per_token_ms": 16.906987654320986,
|
||||
"prompt_per_second": 59.1471420247016,
|
||||
"predicted_n": 40,
|
||||
"predicted_ms": 952.303,
|
||||
"predicted_per_token_ms": 24.41802564102564,
|
||||
"predicted_per_second": 40.95335203186381
|
||||
},
|
||||
"finish_reason": "stop",
|
||||
"content": "**Links:**\n* **Form:** Kreis\n* **Farbe:** Rot\n\n**Rechts:**\n* **Form:** Quadrat\n* **Farbe:** Blau",
|
||||
"reasoning_content": "",
|
||||
"tool_calls": []
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,34 @@
|
||||
{
|
||||
"model": "qwen-fast",
|
||||
"prompt": "Was ist links und was ist rechts im Bild? Nenne Form und Farbe beider Objekte.",
|
||||
"expected": "Links ein roter Kreis, rechts ein blaues Quadrat.",
|
||||
"started": 1789775622.0384746,
|
||||
"response": {
|
||||
"wall_seconds": 2.664,
|
||||
"usage": {
|
||||
"completion_tokens": 84,
|
||||
"prompt_tokens": 162,
|
||||
"total_tokens": 246,
|
||||
"prompt_tokens_details": {
|
||||
"cached_tokens": 0
|
||||
}
|
||||
},
|
||||
"timings": {
|
||||
"cache_n": 0,
|
||||
"prompt_n": 162,
|
||||
"prompt_ms": 1725.751,
|
||||
"prompt_per_token_ms": 10.652783950617284,
|
||||
"prompt_per_second": 93.87217507044758,
|
||||
"predicted_n": 84,
|
||||
"predicted_ms": 926.95,
|
||||
"predicted_per_token_ms": 11.168072289156626,
|
||||
"predicted_per_second": 89.54096768973514,
|
||||
"draft_n": 64,
|
||||
"draft_n_accepted": 51
|
||||
},
|
||||
"finish_reason": "stop",
|
||||
"content": "Im Bild sind zwei geometrische Objekte zu sehen:\n\n- **Links**: Ein **roter Kreis**.\n - Form: Kreis\n - Farbe: Rot\n\n- **Rechts**: Ein **blaues Quadrat**.\n - Form: Quadrat\n - Farbe: Blau\n\nZusammenfassend:\n> Links ist ein roter Kreis, rechts ist ein blaues Quadrat.",
|
||||
"reasoning_content": "",
|
||||
"tool_calls": []
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,85 @@
|
||||
#!/usr/bin/env bash
|
||||
# Run after the user has explicitly approved GPU inference. Restore the initial
|
||||
# production profile even if a benchmark or model load fails.
|
||||
set -Eeuo pipefail
|
||||
[[ ${1:-} == --go && $(hostname) == athena ]] || { echo 'Usage on Athena: run-go.sh --go' >&2; exit 2; }
|
||||
cd /opt/mike-ai/experiments/bonsai2-ab
|
||||
OUT=/data/benchmarks/bonsai2-ab
|
||||
mkdir -p "$OUT"
|
||||
ORIGINAL=$(docker ps --format '{{.Names}}' | sed -n 's/^mike-ai-llama-\(fast\|medium\|large\|ultra\|uncensored\)$/\1/p')
|
||||
[[ $(wc -w <<<"$ORIGINAL") == 1 ]] || { echo 'Expected exactly one active production profile' >&2; exit 1; }
|
||||
echo "$ORIGINAL" > "$OUT/original-profile.txt"
|
||||
|
||||
controller() {
|
||||
docker exec mike-ai-profile-controller python3 -c '
|
||||
import os, sys, urllib.request
|
||||
token = os.environ.get("CONTROLLER_TOKEN", "").strip()
|
||||
if not token:
|
||||
token = open(os.environ.get("CONTROLLER_TOKEN_FILE", "/run/secrets/controller-token"), encoding="utf-8").read().strip()
|
||||
req = urllib.request.Request("http://127.0.0.1:8090" + sys.argv[1], data=b"{}", headers={"Authorization": "Bearer " + token, "Content-Type": "application/json"}, method="POST")
|
||||
print(urllib.request.urlopen(req, timeout=180).read().decode())
|
||||
' "$1"
|
||||
}
|
||||
|
||||
monitor_pid=''
|
||||
finish() {
|
||||
local rc=$?
|
||||
trap - EXIT INT TERM
|
||||
if [[ -n "$monitor_pid" ]]; then kill "$monitor_pid" 2>/dev/null || true; wait "$monitor_pid" 2>/dev/null || true; fi
|
||||
./case.sh fast stop || true
|
||||
controller "/profiles/$ORIGINAL/activate" || true
|
||||
docker ps --format '{{.Names}} {{.Status}}' | grep -E 'mike-ai-(llama-|router|profile-controller|bonsai2-ab)' || true
|
||||
echo "AB_RUN_EXIT=$rc ORIGINAL=$ORIGINAL" | tee -a "$OUT/run-go.log"
|
||||
exit "$rc"
|
||||
}
|
||||
trap finish EXIT INT TERM
|
||||
|
||||
wait_health() {
|
||||
local url=$1
|
||||
for ((i=0;i<360;i++)); do
|
||||
if curl -fsS --max-time 2 "$url/health" >/dev/null 2>&1; then return 0; fi
|
||||
sleep 2
|
||||
done
|
||||
echo "Health timeout: $url" >&2
|
||||
return 1
|
||||
}
|
||||
|
||||
record() {
|
||||
local label=$1 base=$2 model=$3
|
||||
echo "START $label $(date -Is)" | tee -a "$OUT/run-go.log"
|
||||
nvidia-smi --query-gpu=uuid,name,memory.used,memory.free --format=csv,noheader,nounits > "$OUT/$label-before-gpu.csv"
|
||||
python3 gpu_monitor.py --output "$OUT/$label-gpu.jsonl" --go >/dev/null 2>&1 &
|
||||
monitor_pid=$!
|
||||
python3 measure.py --label "$label" --base "$base" --model "$model" --output "$OUT/$label.json" --go 2>&1 | tee "$OUT/$label.log"
|
||||
kill "$monitor_pid" 2>/dev/null || true
|
||||
wait "$monitor_pid" 2>/dev/null || true
|
||||
monitor_pid=''
|
||||
echo "END $label $(date -Is)" | tee -a "$OUT/run-go.log"
|
||||
}
|
||||
|
||||
qwen() {
|
||||
local case=$1 ip
|
||||
controller "/profiles/$case/activate"
|
||||
ip=$(docker inspect "mike-ai-llama-$case" --format '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}')
|
||||
wait_health "http://$ip:8080"
|
||||
record "qwen-$case" "http://$ip:8080" "qwen-$case"
|
||||
}
|
||||
|
||||
bonsai() {
|
||||
local case=$1
|
||||
controller /inference/stop
|
||||
nvidia-smi --query-gpu=uuid,name,memory.used,memory.free --format=csv,noheader,nounits > "$OUT/bonsai-$case-idle-gpu.csv"
|
||||
./case.sh "$case" create
|
||||
./case.sh "$case" start --go
|
||||
wait_health http://127.0.0.1:5006
|
||||
record "bonsai-$case" http://127.0.0.1:5006 bonsai2-test
|
||||
./case.sh "$case" stop
|
||||
}
|
||||
|
||||
# Qwen Medium was measured first while already active, before this script.
|
||||
[[ -s "$OUT/qwen-medium.json" ]] || { echo 'Qwen Medium baseline incomplete' >&2; exit 1; }
|
||||
qwen fast
|
||||
bonsai fast
|
||||
bonsai medium
|
||||
qwen ultra
|
||||
bonsai ultra
|
||||
@@ -0,0 +1,48 @@
|
||||
#!/usr/bin/env bash
|
||||
set -Eeuo pipefail
|
||||
[[ ${1:-} == --go && $(hostname) == athena ]] || { echo 'Usage on Athena: vision-go.sh --go' >&2; exit 2; }
|
||||
cd /opt/mike-ai/experiments/bonsai2-ab
|
||||
OUT=/data/benchmarks/bonsai2-ab
|
||||
ORIGINAL=$(docker ps --format '{{.Names}}' | sed -n 's/^mike-ai-llama-\(fast\|medium\|large\|ultra\|uncensored\)$/\1/p')
|
||||
[[ $(wc -w <<<"$ORIGINAL") == 1 ]] || exit 1
|
||||
controller() {
|
||||
docker exec mike-ai-profile-controller python3 -c '
|
||||
import os, sys, urllib.request
|
||||
token = os.environ.get("CONTROLLER_TOKEN", "").strip()
|
||||
if not token:
|
||||
token = open(os.environ.get("CONTROLLER_TOKEN_FILE", "/run/secrets/controller-token"), encoding="utf-8").read().strip()
|
||||
req = urllib.request.Request("http://127.0.0.1:8090" + sys.argv[1], data=b"{}", headers={"Authorization": "Bearer " + token, "Content-Type": "application/json"}, method="POST")
|
||||
print(urllib.request.urlopen(req, timeout=180).read().decode())
|
||||
' "$1"
|
||||
}
|
||||
finish() {
|
||||
local rc=$?
|
||||
trap - EXIT INT TERM
|
||||
./case.sh vision stop || true
|
||||
controller "/profiles/$ORIGINAL/activate" || true
|
||||
echo "VISION_EXIT=$rc RESTORED=$ORIGINAL"
|
||||
exit "$rc"
|
||||
}
|
||||
trap finish EXIT INT TERM
|
||||
wait_health() {
|
||||
local url=$1
|
||||
for ((i=0;i<120;i++)); do
|
||||
curl -fsS --max-time 2 "$url/health" >/dev/null 2>&1 && return 0
|
||||
sleep 2
|
||||
done
|
||||
return 1
|
||||
}
|
||||
controller /profiles/fast/activate
|
||||
ip=$(docker inspect mike-ai-llama-fast --format '{{range .NetworkSettings.Networks}}{{.IPAddress}}{{end}}')
|
||||
wait_health "http://$ip:8080"
|
||||
python3 vision_probe.py --base "http://$ip:8080" --model qwen-fast --output "$OUT/vision-qwen-fast.json" --go
|
||||
controller /inference/stop
|
||||
./case.sh vision create
|
||||
./case.sh vision start --go
|
||||
wait_health http://127.0.0.1:5006
|
||||
nvidia-smi --query-gpu=uuid,memory.used,memory.free --format=csv,noheader,nounits > "$OUT/vision-bonsai-before-gpu.csv"
|
||||
python3 vision_probe.py --base http://127.0.0.1:5006 --model bonsai2-test --output "$OUT/vision-bonsai-fast.json" --go
|
||||
nvidia-smi --query-gpu=uuid,memory.used,memory.free --format=csv,noheader,nounits > "$OUT/vision-bonsai-after-gpu.csv"
|
||||
python3 repeat_critical.py --base http://127.0.0.1:5006 --model bonsai2-test \
|
||||
--temperature 1.0 --seeds 7,42,99 --skip-long \
|
||||
--output "$OUT/recommended-bonsai-fast.json" --go
|
||||
@@ -0,0 +1,38 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Small, fully synthetic vision smoke test for the isolated A/B servers."""
|
||||
import argparse
|
||||
import base64
|
||||
import json
|
||||
import pathlib
|
||||
import time
|
||||
|
||||
from measure import request
|
||||
|
||||
|
||||
def main():
|
||||
p = argparse.ArgumentParser()
|
||||
p.add_argument("--base", required=True)
|
||||
p.add_argument("--model", required=True)
|
||||
p.add_argument("--output", type=pathlib.Path, required=True)
|
||||
p.add_argument("--go", action="store_true")
|
||||
args = p.parse_args()
|
||||
if not args.go:
|
||||
p.error("Inference requires explicit --go")
|
||||
image = pathlib.Path(__file__).with_name("assets") / "vision-fixture.png"
|
||||
url = "data:image/png;base64," + base64.b64encode(image.read_bytes()).decode("ascii")
|
||||
prompt = "Was ist links und was ist rechts im Bild? Nenne Form und Farbe beider Objekte."
|
||||
payload = {"model": args.model, "stream": False, "temperature": 0,
|
||||
"reasoning_effort": "none", "max_tokens": 512,
|
||||
"messages": [{"role": "user", "content": [
|
||||
{"type": "text", "text": prompt},
|
||||
{"type": "image_url", "image_url": {"url": url}}]}]}
|
||||
result = {"model": args.model, "prompt": prompt, "expected":
|
||||
"Links ein roter Kreis, rechts ein blaues Quadrat.",
|
||||
"started": time.time(), "response": request(args.base, None, payload)}
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
args.output.write_text(json.dumps(result, ensure_ascii=False, indent=2) + "\n")
|
||||
print(result["response"]["content"])
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
Reference in New Issue
Block a user