1.8 KiB
Model evaluation
Before downloading
- Read all of
docs/TESTED_MODELS.md; it is the no-repeat register. - Record the exact repository, revision, filename, base model, fine-tune, quantization, license, format and estimated disk/VRAM/RAM requirements.
- Confirm the model adds a genuinely new candidate rather than a renamed artifact already tested.
- Do not stop the active production model merely to download or prepare the
candidate. Put temporary scripts under
/tmp; durable code belongs in Git.
Quality rule
Standard text profiles must not fall below Q4. A smaller Q3 quantization may be kept only when an A/B test documents that its quality loss is negligible for Mike's intended workload. Smaller files are not presumed faster: GPU split, memory bandwidth, kernels, cache formats and cross-GPU traffic must be measured.
Fair A/B test
Hold these equal wherever the models permit it:
- prompt set and conversation history;
- context size and filled-context test point;
- KV-cache quantization, slots, batch/uBatch, MTP and sampling;
- GPU visibility and tensor split;
- warm-up state and output-token limit.
Measure prompt processing, short decode, long-context decode, peak VRAM and wall time. Test meaning preservation, uncertainty, negation, ordered safety constraints, German language consistency, tool-call schema and long-context recall. Do not replace a model on synthetic benchmark scores alone.
Completion
Write the exact artifact, settings, raw result path, interpretation and decision
to docs/TESTED_MODELS.md in the same commit. If promoted, update the manifest,
Compose/env examples, profile matrix and user documentation. If rejected,
remove candidate-only weights, images and containers after preserving the
result. Restore and functionally test the previous profile, then publish.