Scope Bonsai benchmark cache cleanup

This commit is contained in:
Mikei386
2026-09-19 01:07:53 +02:00
parent 8c0084755b
commit de9b17f054
2 changed files with 9 additions and 1 deletions
+1 -1
View File
@@ -12,6 +12,6 @@ The A/B matrix compares production Qwen against Bonsai at **the same context pro
After the user's explicit **GO**, use the official profile controller/router to switch or stop production Qwen, record the original profile, run the corresponding Bonsai case with `case.sh CASE start --go`, execute `measure.py ... --go`, and sample both physical GPUs with `gpu_monitor.py --output ... --go`. Record idle GPU memory before each model load so the model's incremental VRAM is distinguishable from other processes. Stop Bonsai, restore the original production profile via the controller, and verify the router and original model are healthy. Repeat case by case. Never run both model servers at once. The smallest context is tested first; only try Medium and Ultra if Fast fits. Results live at `/data/benchmarks/bonsai2-ab` and should be checked for OOM/offload before drawing any speed conclusion. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA `cuobjdump`: it includes `sm_86` and `sm_120a`, covering both cards. After the user's explicit **GO**, use the official profile controller/router to switch or stop production Qwen, record the original profile, run the corresponding Bonsai case with `case.sh CASE start --go`, execute `measure.py ... --go`, and sample both physical GPUs with `gpu_monitor.py --output ... --go`. Record idle GPU memory before each model load so the model's incremental VRAM is distinguishable from other processes. Stop Bonsai, restore the original production profile via the controller, and verify the router and original model are healthy. Repeat case by case. Never run both model servers at once. The smallest context is tested first; only try Medium and Ultra if Fast fits. Results live at `/data/benchmarks/bonsai2-ab` and should be checked for OOM/offload before drawing any speed conclusion. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA `cuobjdump`: it includes `sm_86` and `sm_120a`, covering both cards.
Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results and deployment directory. Docker's shared build cache and shared NVIDIA base layers are deliberately not globally pruned, because that could delete unrelated build caches. Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results, deployment directory and build-cache records identifiable as this experiment. Docker's shared runtime base layers are deliberately not globally pruned, because that could delete unrelated build caches.
Sources: [PrismML model](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [PrismML llama.cpp CUDA release](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15). Stock llama.cpp is intentionally not used for this ternary model. Sources: [PrismML model](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [PrismML llama.cpp CUDA release](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15). Stock llama.cpp is intentionally not used for this ternary model.
+8
View File
@@ -11,6 +11,14 @@ test "$(hostname)" = athena || exit 2
systemctl stop mike-ai-bonsai2-prepare.service 2>/dev/null || true systemctl stop mike-ai-bonsai2-prepare.service 2>/dev/null || true
docker rm -f mike-ai-bonsai2-ab 2>/dev/null || true docker rm -f mike-ai-bonsai2-ab 2>/dev/null || true
docker image rm mike-ai/bonsai2-ab:prism-b10685 2>/dev/null || true docker image rm mike-ai/bonsai2-ab:prism-b10685 2>/dev/null || true
for description in \
'PrismML-Eng/llama.cpp' \
'/src/llama.cpp/build' \
'ADD prism-cuda.tar.gz' \
'useradd --system --uid 10013' \
'nvidia/cuda:12.8.1-devel-ubuntu24.04'; do
docker buildx prune --force --filter "description~=$description" >/dev/null 2>&1 || true
done
rm -rf -- /data/models/bonsai2-ab /data/benchmarks/bonsai2-ab rm -rf -- /data/models/bonsai2-ab /data/benchmarks/bonsai2-ab
echo 'Bonsai experiment container, image, weights, results and staging removed.' echo 'Bonsai experiment container, image, weights, results and staging removed.'
rm -rf -- /opt/mike-ai/experiments/bonsai2-ab rm -rf -- /opt/mike-ai/experiments/bonsai2-ab