Files
AI-Profile-Router/platform/hermes/skills/athena-operator/references/model-evaluation.md
T

1.8 KiB

Model evaluation

Before downloading

  1. Read all of docs/TESTED_MODELS.md; it is the no-repeat register.
  2. Record the exact repository, revision, filename, base model, fine-tune, quantization, license, format and estimated disk/VRAM/RAM requirements.
  3. Confirm the model adds a genuinely new candidate rather than a renamed artifact already tested.
  4. Do not stop the active production model merely to download or prepare the candidate. Put temporary scripts under /tmp; durable code belongs in Git.

Quality rule

Standard text profiles must not fall below Q4. A smaller Q3 quantization may be kept only when an A/B test documents that its quality loss is negligible for Mike's intended workload. Smaller files are not presumed faster: GPU split, memory bandwidth, kernels, cache formats and cross-GPU traffic must be measured.

Fair A/B test

Hold these equal wherever the models permit it:

  • prompt set and conversation history;
  • context size and filled-context test point;
  • KV-cache quantization, slots, batch/uBatch, MTP and sampling;
  • GPU visibility and tensor split;
  • warm-up state and output-token limit.

Measure prompt processing, short decode, long-context decode, peak VRAM and wall time. Test meaning preservation, uncertainty, negation, ordered safety constraints, German language consistency, tool-call schema and long-context recall. Do not replace a model on synthetic benchmark scores alone.

Completion

Write the exact artifact, settings, raw result path, interpretation and decision to docs/TESTED_MODELS.md in the same commit. If promoted, update the manifest, Compose/env examples, profile matrix and user documentation. If rejected, remove candidate-only weights, images and containers after preserving the result. Restore and functionally test the previous profile, then publish.