Document Ornith 1.5 A/B benchmark
This commit is contained in:
1 parent
9978e5b7e6
commit
9af719674b
27 files changed
+3503
-1
No files matched your search
@@ -0,0 +1,49 @@
|
||||
# Ornith 1.5 35B-A3B A/B on Athena
|
||||
|
||||
This experiment compares the official `ornith-ai/Ornith-1.5-35B-A3B-GGUF`
|
||||
Q4_K_M build with Athena's production Qwen3.8-27B profiles. It is isolated
|
||||
from production: the weights, container, port, raw results and deployment
|
||||
directory use experiment-specific names.
|
||||
|
||||
The official checkpoint is pinned to revision
|
||||
`12393612fd4f730ff5aadc23e9b8f9648aa49ceb`. The 21.7 GB Q4_K_M file and the
|
||||
optional BF16 vision projector are verified with their published SHA-256
|
||||
hashes. The experiment reuses Athena's pinned `mike-ai/llama.cpp:local`
|
||||
runtime so model quality, rather than a different inference engine, is being
|
||||
compared.
|
||||
|
||||
The server caps medium-effort reasoning at 2,048 tokens. Without that cap,
|
||||
Ornith consumed the complete 4,096-token response allowance on several tasks
|
||||
and returned no visible final answer. The cap matches the published llama.cpp
|
||||
community configuration for this exact GGUF and preserves room for the answer.
|
||||
|
||||
Ornith does not fit completely on the RTX 5080 at Q4 while retaining a useful
|
||||
context window. All text cases therefore use both GPUs. Each context profile
|
||||
uses the highest 5080-heavy split that passes model load and generation without
|
||||
OOM. The resident Qwen-TTS process occupies about 4.7 GiB on the
|
||||
RTX 3060. `run-go.sh` records the initial state of Qwen-TTS and Whisper, stops
|
||||
them only for the measured window, and restores them from an EXIT/HUP/INT/TERM
|
||||
trap. The same trap restores the initial production LLM profile.
|
||||
|
||||
| Case | Context | GPUs | Layer split | u-batch |
|
||||
|---|---:|---|---:|---:|
|
||||
| Fast | 76,800 | 5080 + 3060 | 70:30 | 64 |
|
||||
| Medium | 160,000 | 5080 + 3060 | 68:32 | 512 |
|
||||
| Ultra | 262,144 | 5080 + 3060 | 65:35 | 128 |
|
||||
| Vision | 76,800 | 5080 + 3060 | 70:30 | 64 |
|
||||
|
||||
`prepare.sh` only downloads and verifies artifacts and creates a stopped
|
||||
container. `validate-splits.sh --go` first proves that each 5080-heavy split
|
||||
can load and generate without OOM. `run-go.sh --go`, `repeat-go.sh --go` and `vision-go.sh --go` are
|
||||
the only entry points that run inference. They reuse the same nine acceptance
|
||||
tasks, tool-call probe, long-context recall probe and synthetic vision fixture
|
||||
as the Bonsai/Qwen comparison. Raw measurements live under
|
||||
`/data/benchmarks/ornith15-ab`; reviewed reports belong in `results/`.
|
||||
|
||||
`cleanup.sh` previews the exact experiment artifacts. `cleanup.sh --all`
|
||||
removes only the named Ornith container, downloaded weights, raw results and
|
||||
deployment staging. It does not prune shared images or touch production.
|
||||
|
||||
Sources: [official model](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF),
|
||||
[official evaluation](https://ornith.ai/ornith_1_5.html),
|
||||
[published llama.cpp configuration](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF/discussions/14).
|
||||
Reference in new issue
Block a user