50 lines
2.7 KiB
Markdown
50 lines
2.7 KiB
Markdown
# Ornith 1.5 35B-A3B A/B on Athena
|
|
|
|
This experiment compares the official `ornith-ai/Ornith-1.5-35B-A3B-GGUF`
|
|
Q4_K_M build with Athena's production Qwen3.8-27B profiles. It is isolated
|
|
from production: the weights, container, port, raw results and deployment
|
|
directory use experiment-specific names.
|
|
|
|
The official checkpoint is pinned to revision
|
|
`12393612fd4f730ff5aadc23e9b8f9648aa49ceb`. The 21.7 GB Q4_K_M file and the
|
|
optional BF16 vision projector are verified with their published SHA-256
|
|
hashes. The experiment reuses Athena's pinned `mike-ai/llama.cpp:local`
|
|
runtime so model quality, rather than a different inference engine, is being
|
|
compared.
|
|
|
|
The server caps medium-effort reasoning at 2,048 tokens. Without that cap,
|
|
Ornith consumed the complete 4,096-token response allowance on several tasks
|
|
and returned no visible final answer. The cap matches the published llama.cpp
|
|
community configuration for this exact GGUF and preserves room for the answer.
|
|
|
|
Ornith does not fit completely on the RTX 5080 at Q4 while retaining a useful
|
|
context window. All text cases therefore use both GPUs. Each context profile
|
|
uses the highest 5080-heavy split that passes model load and generation without
|
|
OOM. The resident Qwen-TTS process occupies about 4.7 GiB on the
|
|
RTX 3060. `run-go.sh` records the initial state of Qwen-TTS and Whisper, stops
|
|
them only for the measured window, and restores them from an EXIT/HUP/INT/TERM
|
|
trap. The same trap restores the initial production LLM profile.
|
|
|
|
| Case | Context | GPUs | Layer split | u-batch |
|
|
|---|---:|---|---:|---:|
|
|
| Fast | 76,800 | 5080 + 3060 | 70:30 | 64 |
|
|
| Medium | 160,000 | 5080 + 3060 | 68:32 | 512 |
|
|
| Ultra | 262,144 | 5080 + 3060 | 65:35 | 128 |
|
|
| Vision | 76,800 | 5080 + 3060 | 70:30 | 64 |
|
|
|
|
`prepare.sh` only downloads and verifies artifacts and creates a stopped
|
|
container. `validate-splits.sh --go` first proves that each 5080-heavy split
|
|
can load and generate without OOM. `run-go.sh --go`, `repeat-go.sh --go` and `vision-go.sh --go` are
|
|
the only entry points that run inference. They reuse the same nine acceptance
|
|
tasks, tool-call probe, long-context recall probe and synthetic vision fixture
|
|
as the Bonsai/Qwen comparison. Raw measurements live under
|
|
`/data/benchmarks/ornith15-ab`; reviewed reports belong in `results/`.
|
|
|
|
`cleanup.sh` previews the exact experiment artifacts. `cleanup.sh --all`
|
|
removes only the named Ornith container, downloaded weights, raw results and
|
|
deployment staging. It does not prune shared images or touch production.
|
|
|
|
Sources: [official model](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF),
|
|
[official evaluation](https://ornith.ai/ornith_1_5.html),
|
|
[published llama.cpp configuration](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF/discussions/14).
|