Scope Bonsai benchmark cache cleanup
This commit is contained in:
@@ -12,6 +12,6 @@ The A/B matrix compares production Qwen against Bonsai at **the same context pro
|
||||
|
||||
After the user's explicit **GO**, use the official profile controller/router to switch or stop production Qwen, record the original profile, run the corresponding Bonsai case with `case.sh CASE start --go`, execute `measure.py ... --go`, and sample both physical GPUs with `gpu_monitor.py --output ... --go`. Record idle GPU memory before each model load so the model's incremental VRAM is distinguishable from other processes. Stop Bonsai, restore the original production profile via the controller, and verify the router and original model are healthy. Repeat case by case. Never run both model servers at once. The smallest context is tested first; only try Medium and Ultra if Fast fits. Results live at `/data/benchmarks/bonsai2-ab` and should be checked for OOM/offload before drawing any speed conclusion. GPU UUIDs, not host GPU indexes, identify the 5080 and 3060. The official archive was statically inspected with NVIDIA `cuobjdump`: it includes `sm_86` and `sm_120a`, covering both cards.
|
||||
|
||||
Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results and deployment directory. Docker's shared build cache and shared NVIDIA base layers are deliberately not globally pruned, because that could delete unrelated build caches.
|
||||
Preparation and `measure.py` deliberately refuse to infer without an explicit GO. The pre-GO deployment directory is `/opt/mike-ai/experiments/bonsai2-ab`. No production files are edited. `cleanup.sh` previews what will be removed; `cleanup.sh --all` stops the transient build (if still active), then removes the named test container, test image, model, results, deployment directory and build-cache records identifiable as this experiment. Docker's shared runtime base layers are deliberately not globally pruned, because that could delete unrelated build caches.
|
||||
|
||||
Sources: [PrismML model](https://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf), [PrismML llama.cpp CUDA release](https://github.com/PrismML-Eng/llama.cpp/releases/tag/prism-b10685-7dffb15). Stock llama.cpp is intentionally not used for this ternary model.
|
||||
|
||||
@@ -11,6 +11,14 @@ test "$(hostname)" = athena || exit 2
|
||||
systemctl stop mike-ai-bonsai2-prepare.service 2>/dev/null || true
|
||||
docker rm -f mike-ai-bonsai2-ab 2>/dev/null || true
|
||||
docker image rm mike-ai/bonsai2-ab:prism-b10685 2>/dev/null || true
|
||||
for description in \
|
||||
'PrismML-Eng/llama.cpp' \
|
||||
'/src/llama.cpp/build' \
|
||||
'ADD prism-cuda.tar.gz' \
|
||||
'useradd --system --uid 10013' \
|
||||
'nvidia/cuda:12.8.1-devel-ubuntu24.04'; do
|
||||
docker buildx prune --force --filter "description~=$description" >/dev/null 2>&1 || true
|
||||
done
|
||||
rm -rf -- /data/models/bonsai2-ab /data/benchmarks/bonsai2-ab
|
||||
echo 'Bonsai experiment container, image, weights, results and staging removed.'
|
||||
rm -rf -- /opt/mike-ai/experiments/bonsai2-ab
|
||||
|
||||
Reference in New Issue
Block a user