Files
AI-Profile-Router/experiments/gsq-rco-iq3s-ab-v2/README.md
T

782 B

GSQ-RCO IQ3_S profile A/B test

run-case.sh starts one isolated llama.cpp test container for either the current Q4 reference or GSQ-RCO IQ3_S. It refuses to start while a production LLM container is running. Context, GPU split, vision projector, MTP depth and batch sizes are explicit command-line arguments so every candidate can use the same settings as its corresponding production profile.

The 2026-09-08 run used the existing dependency-free benchmark programs:

  • experiments/dirk-qwen38/bench-case.py
  • experiments/dirk-qwen38/quality-ab.py
  • dev/QWEN38-FINAL-ACCEPTANCE-v1.json

Results are retained on Athena in /data/model-benchmarks/gsq-rco-iq3s-ab-v2-20260908/. The conclusions and aggregate measurements are documented in docs/GSQ_RCO_BETA1_20260904.md.