Ornith 1.5 35B-A3B A/B on Athena
This experiment compares the official ornith-ai/Ornith-1.5-35B-A3B-GGUF
Q4_K_M build with Athena's production Qwen3.8-27B profiles. It is isolated
from production: the weights, container, port, raw results and deployment
directory use experiment-specific names.
The official checkpoint is pinned to revision
12393612fd4f730ff5aadc23e9b8f9648aa49ceb. The 21.7 GB Q4_K_M file and the
optional BF16 vision projector are verified with their published SHA-256
hashes. The experiment reuses Athena's pinned mike-ai/llama.cpp:local
runtime so model quality, rather than a different inference engine, is being
compared.
The server caps medium-effort reasoning at 2,048 tokens. Without that cap, Ornith consumed the complete 4,096-token response allowance on several tasks and returned no visible final answer. The cap matches the published llama.cpp community configuration for this exact GGUF and preserves room for the answer.
Ornith does not fit completely on the RTX 5080 at Q4 while retaining a useful
context window. All text cases therefore use both GPUs. Each context profile
uses the highest 5080-heavy split that passes model load and generation without
OOM. The resident Qwen-TTS process occupies about 4.7 GiB on the
RTX 3060. run-go.sh records the initial state of Qwen-TTS and Whisper, stops
them only for the measured window, and restores them from an EXIT/HUP/INT/TERM
trap. The same trap restores the initial production LLM profile.
| Case | Context | GPUs | Layer split | u-batch |
|---|---|---|---|---|
| Fast | 76,800 | 5080 + 3060 | 70:30 | 64 |
| Medium | 160,000 | 5080 + 3060 | 68:32 | 512 |
| Ultra | 262,144 | 5080 + 3060 | 65:35 | 128 |
| Vision | 76,800 | 5080 + 3060 | 70:30 | 64 |
prepare.sh only downloads and verifies artifacts and creates a stopped
container. validate-splits.sh --go first proves that each 5080-heavy split
can load and generate without OOM. run-go.sh --go, repeat-go.sh --go and vision-go.sh --go are
the only entry points that run inference. They reuse the same nine acceptance
tasks, tool-call probe, long-context recall probe and synthetic vision fixture
as the Bonsai/Qwen comparison. Raw measurements live under
/data/benchmarks/ornith15-ab; reviewed reports belong in results/.
cleanup.sh previews the exact experiment artifacts. cleanup.sh --all
removes only the named Ornith container, downloaded weights, raw results and
deployment staging. It does not prune shared images or touch production.
Sources: official model, official evaluation, published llama.cpp configuration.