Files
AI-Profile-Router/experiments/ornith15-ab/README.md
T

50 lines
2.7 KiB
Markdown

# Ornith 1.5 35B-A3B A/B on Athena
This experiment compares the official `ornith-ai/Ornith-1.5-35B-A3B-GGUF`
Q4_K_M build with Athena's production Qwen3.8-27B profiles. It is isolated
from production: the weights, container, port, raw results and deployment
directory use experiment-specific names.
The official checkpoint is pinned to revision
`12393612fd4f730ff5aadc23e9b8f9648aa49ceb`. The 21.7 GB Q4_K_M file and the
optional BF16 vision projector are verified with their published SHA-256
hashes. The experiment reuses Athena's pinned `mike-ai/llama.cpp:local`
runtime so model quality, rather than a different inference engine, is being
compared.
The server caps medium-effort reasoning at 2,048 tokens. Without that cap,
Ornith consumed the complete 4,096-token response allowance on several tasks
and returned no visible final answer. The cap matches the published llama.cpp
community configuration for this exact GGUF and preserves room for the answer.
Ornith does not fit completely on the RTX 5080 at Q4 while retaining a useful
context window. All text cases therefore use both GPUs. Each context profile
uses the highest 5080-heavy split that passes model load and generation without
OOM. The resident Qwen-TTS process occupies about 4.7 GiB on the
RTX 3060. `run-go.sh` records the initial state of Qwen-TTS and Whisper, stops
them only for the measured window, and restores them from an EXIT/HUP/INT/TERM
trap. The same trap restores the initial production LLM profile.
| Case | Context | GPUs | Layer split | u-batch |
|---|---:|---|---:|---:|
| Fast | 76,800 | 5080 + 3060 | 70:30 | 64 |
| Medium | 160,000 | 5080 + 3060 | 68:32 | 512 |
| Ultra | 262,144 | 5080 + 3060 | 65:35 | 128 |
| Vision | 76,800 | 5080 + 3060 | 70:30 | 64 |
`prepare.sh` only downloads and verifies artifacts and creates a stopped
container. `validate-splits.sh --go` first proves that each 5080-heavy split
can load and generate without OOM. `run-go.sh --go`, `repeat-go.sh --go` and `vision-go.sh --go` are
the only entry points that run inference. They reuse the same nine acceptance
tasks, tool-call probe, long-context recall probe and synthetic vision fixture
as the Bonsai/Qwen comparison. Raw measurements live under
`/data/benchmarks/ornith15-ab`; reviewed reports belong in `results/`.
`cleanup.sh` previews the exact experiment artifacts. `cleanup.sh --all`
removes only the named Ornith container, downloaded weights, raw results and
deployment staging. It does not prune shared images or touch production.
Sources: [official model](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF),
[official evaluation](https://ornith.ai/ornith_1_5.html),
[published llama.cpp configuration](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF/discussions/14).