Document Ornith 1.5 A/B benchmark
This commit is contained in:
1 parent
9978e5b7e6
commit
9af719674b
27 files changed
+3503
-1
No files matched your search
@@ -0,0 +1,100 @@
|
||||
# Ornith 1.5 35B-A3B A/B result on Athena
|
||||
|
||||
Measured 2026-09-19 against Athena's production Qwen3.8-27B profiles. The
|
||||
same nine text tasks, tool schema, 27,013-token prefill, 2,048-token decode,
|
||||
long-context recall and synthetic vision fixture were used. Ornith ran from
|
||||
the official Q4_K_M GGUF at pinned revision
|
||||
`12393612fd4f730ff5aadc23e9b8f9648aa49ceb`.
|
||||
|
||||
Ornith requires a 2,048-token reasoning cap with medium effort. Without it,
|
||||
the model exhausted the complete 4,096-token response allowance on several
|
||||
tasks and returned no visible final answer. This report uses the corrected
|
||||
configuration published for this exact GGUF.
|
||||
|
||||
## Performance
|
||||
|
||||
| Profile | Model | Nine tasks | Task decode | 27K prefill | Long decode | Weather tool |
|
||||
|---|---|---:|---:|---:|---:|---:|
|
||||
| Fast 76.8K | Qwen | 191.1 s | 95.1 tok/s | 1,306 tok/s | 84.7 tok/s | 0.68 s, correct |
|
||||
| Fast 76.8K | Ornith | **152.9 s** | **124.2 tok/s** | 1,225 tok/s | **118.0 tok/s** | **0.58 s, correct** |
|
||||
| Medium 160K | Qwen | 220.4 s | 75.7 tok/s | 1,836 tok/s | 65.5 tok/s | 0.68 s, correct |
|
||||
| Medium 160K | Ornith | **155.2 s** | **123.7 tok/s** | **3,320 tok/s** | **111.9 tok/s** | **0.55 s, correct** |
|
||||
| Ultra 262K | Qwen | 265.3 s | 68.8 tok/s | **1,632 tok/s** | 63.1 tok/s | 0.81 s, correct |
|
||||
| Ultra 262K | Ornith | **163.0 s** | **122.8 tok/s** | 1,411 tok/s | **111.7 tok/s** | **0.65 s, correct** |
|
||||
|
||||
Across the nine actual answers, Ornith reduced waiting time by 20% in Fast,
|
||||
30% in Medium and 39% in Ultra. Long decode improved by 39%, 71% and 77%.
|
||||
Prefill was profile-dependent: 6% slower in Fast, 81% faster in Medium and
|
||||
14% slower in Ultra. Repeating the same 27K prompt hit llama.cpp's cache for
|
||||
both models.
|
||||
|
||||
## GPU distribution and stability
|
||||
|
||||
The split was tuned empirically, one whole model layer at a time. A 73:27 and
|
||||
72:28 Fast split loaded most weights but OOMed on the first CUDA calculation.
|
||||
The following are the fastest splits that both loaded and generated:
|
||||
|
||||
| Profile | Tensor split 5080:3060 | RTX 5080 after generation | RTX 3060 after generation | 5080 reserve |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Fast | 70:30 | 15,536 MiB | 6,685 MiB | 393 MiB |
|
||||
| Medium | 68:32 | 15,720 MiB | 8,471 MiB | 209 MiB |
|
||||
| Ultra | 65:35 | 15,842 MiB | 9,173 MiB | 87 MiB |
|
||||
|
||||
All three complete runs passed without OOM. `--n-gpu-layers all` keeps all
|
||||
model layers on CUDA; no CPU model-layer offload was configured. The lower
|
||||
Ultra ratio is required because its larger KV cache also consumes GPU memory.
|
||||
Qwen-TTS and Whisper were stopped only during each isolated Ornith window and
|
||||
restored by a signal-safe trap.
|
||||
|
||||
## Correctness
|
||||
|
||||
The fixed rubric awards 0 (wrong/missing), 1 (partly correct) or 2 (correct
|
||||
and supported) for nine answers plus the tool call.
|
||||
|
||||
| Profile | Qwen | Ornith | Material difference |
|
||||
|---|---:|---:|---|
|
||||
| Fast | 19/20 | 18/20 | Ornith's async solution was wrong; Qwen's migration proof was truncated |
|
||||
| Medium | **20/20** | 18/20 | Ornith's async solution was wrong |
|
||||
| Ultra | 17/20 | 16/20 | both had long-answer truncation; Ornith's async solution was wrong |
|
||||
|
||||
Ornith correctly solved logic, evidence diagnosis, capacity planning,
|
||||
instruction injection, runtime-state interpretation, read-only diagnosis,
|
||||
destructive-action handling and evidence boundaries. It also emitted the
|
||||
required `get_weather({"city":"Rastatt"})` call without inventing a result.
|
||||
|
||||
The blocking defect is reproducible async-agent logic. In all three main
|
||||
profiles and both repeated Fast seeds, Ornith failed the requirement "return
|
||||
the first *successful* concurrent task and cancel/await the rest":
|
||||
|
||||
- one answer used `asyncio.wait` without `FIRST_COMPLETED` and even passed the
|
||||
unsupported `return_exceptions` argument;
|
||||
- repeated answers used `asyncio.gather`, which waits for every task and then
|
||||
chooses by input order instead of completion order;
|
||||
- another answer returned the entire `gather` result rather than the first
|
||||
successful result.
|
||||
|
||||
Qwen Fast and Medium produced the correct `create_task` + `FIRST_COMPLETED` +
|
||||
cancel + `gather(..., return_exceptions=True)` pattern in the main run and both
|
||||
repeated Fast seeds. This is directly relevant to OpenClaw's long tool runs,
|
||||
so Ornith is not a safe production replacement despite its speed.
|
||||
|
||||
Both models recalled the exact marker `KIESEL-7319` from a 48,644-token prompt.
|
||||
Qwen took 40.9 s and Ornith 42.2 s. Both read the vision fixture correctly as
|
||||
a red circle on the left and a blue square on the right; Qwen took 3.18 s and
|
||||
Ornith 9.80 s. Ornith's projector works, but this small vision case was about
|
||||
three times slower.
|
||||
|
||||
## Decision
|
||||
|
||||
**Keep Qwen3.8-27B as Athena's production model.** Ornith is a compelling
|
||||
speed experiment and a useful stopped candidate for ordinary text workloads,
|
||||
but it is measurably less reliable on the kind of concurrent control logic an
|
||||
agent must generate. Do not replace Fast, Medium or Ultra silently.
|
||||
|
||||
The experiment remains isolated and reversible. The Ornith container is
|
||||
stopped. Final verification showed `active_profile=medium`; router, Qwen Medium,
|
||||
Whisper and Qwen3-TTS were healthy, and Qwen Medium returned exactly `OK` to a
|
||||
live request. `cleanup.sh --all` removes only the experiment container,
|
||||
weights, raw results and deployment staging. Sources: [official GGUF](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF),
|
||||
[official evaluation](https://ornith.ai/ornith_1_5.html),
|
||||
[published llama.cpp configuration](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF/discussions/14).
|
||||
Reference in new issue
Block a user