From de9b17f054acd3a4ffd062934d1bd87adc0e1334 Mon Sep 17 00:00:00 2001 From: Mikei386 <44135113+Mikei386@users.noreply.github.com> Date: Sat, 19 Sep 2026 01:07:53 +0200 Subject: [PATCH] Scope Bonsai benchmark cache cleanup --- experiments/bonsai2-ab/README.md | 2 +- experiments/bonsai2-ab/cleanup.sh | 8 ++++++++ 2 files changed, 9 insertions(+), 1 deletion(-) diff --git a/experiments/bonsai2-ab/README.md b/experiments/bonsai2-ab/README.md index 63d5762..b775c63 100644 --- a/experiments/bonsai2-ab/README.md +++ b/experiments/bonsai2-ab/README.md @@ -12,6 +12,6 @@ The A/B matrix compares production Qwen against Bonsai at **the same context pro After the user's explicit **GO**, use the official profile controller/router to switch or stop production Qwen, record the original profile, run the corresponding Bonsai case with `case.sh CASE start --go`, execute `measure.py ... --go`, and sample both physical GPUs with `gpu_monitor.py --output ... --go`. Record idle GPU memory before each model load so the model's incremental VRAM is distinguishable from other processes. Stop Bonsai, restore the original production profile via the controller, and verify the router and original model are healthy. Repeat case by case. Never run both model servers at once. The smallest context is tested first; only try Medium and Ultra if Fast fits. Results live at `/data/benchmarks/bonsai2-ab` and should be checked for OOM/offload before drawing any speed conclusion. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA `cuobjdump`: it includes `sm_86` and `sm_120a`, covering both cards. -Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results and deployment directory. Docker's shared build cache and shared NVIDIA base layers are deliberately not globally pruned, because that could delete unrelated build caches. +Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results, deployment directory and build-cache records identifiable as this experiment. Docker's shared runtime base layers are deliberately not globally pruned, because that could delete unrelated build caches. Sources: [PrismML model](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [PrismML llama.cpp CUDA release](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15). Stock llama.cpp is intentionally not used for this ternary model. diff --git a/experiments/bonsai2-ab/cleanup.sh b/experiments/bonsai2-ab/cleanup.sh index b48066a..fece86e 100755 --- a/experiments/bonsai2-ab/cleanup.sh +++ b/experiments/bonsai2-ab/cleanup.sh @@ -11,6 +11,14 @@ test "$(hostname)" = athena || exit 2 systemctl stop mike-ai-bonsai2-prepare.service 2>/dev/null || true docker rm -f mike-ai-bonsai2-ab 2>/dev/null || true docker image rm mike-ai/bonsai2-ab:prism-b10685 2>/dev/null || true +for description in \ + 'PrismML-Eng/llama.cpp' \ + '/src/llama.cpp/build' \ + 'ADD prism-cuda.tar.gz' \ + 'useradd --system --uid 10013' \ + 'nvidia/cuda:12.8.1-devel-ubuntu24.04'; do + docker buildx prune --force --filter "description~=$description" >/dev/null 2>&1 || true +done rm -rf -- /data/models/bonsai2-ab /data/benchmarks/bonsai2-ab echo 'Bonsai experiment container, image, weights, results and staging removed.' rm -rf -- /opt/mike-ai/experiments/bonsai2-ab