Files

101 lines
5.4 KiB
Markdown

# Ornith 1.5 35B-A3B A/B result on Athena
Measured 2026-09-19 against Athena's production Qwen3.8-27B profiles. The
same nine text tasks, tool schema, 27,013-token prefill, 2,048-token decode,
long-context recall and synthetic vision fixture were used. Ornith ran from
the official Q4_K_M GGUF at pinned revision
`12393612fd4f730ff5aadc23e9b8f9648aa49ceb`.
Ornith requires a 2,048-token reasoning cap with medium effort. Without it,
the model exhausted the complete 4,096-token response allowance on several
tasks and returned no visible final answer. This report uses the corrected
configuration published for this exact GGUF.
## Performance
| Profile | Model | Nine tasks | Task decode | 27K prefill | Long decode | Weather tool |
|---|---|---:|---:|---:|---:|---:|
| Fast 76.8K | Qwen | 191.1 s | 95.1 tok/s | 1,306 tok/s | 84.7 tok/s | 0.68 s, correct |
| Fast 76.8K | Ornith | **152.9 s** | **124.2 tok/s** | 1,225 tok/s | **118.0 tok/s** | **0.58 s, correct** |
| Medium 160K | Qwen | 220.4 s | 75.7 tok/s | 1,836 tok/s | 65.5 tok/s | 0.68 s, correct |
| Medium 160K | Ornith | **155.2 s** | **123.7 tok/s** | **3,320 tok/s** | **111.9 tok/s** | **0.55 s, correct** |
| Ultra 262K | Qwen | 265.3 s | 68.8 tok/s | **1,632 tok/s** | 63.1 tok/s | 0.81 s, correct |
| Ultra 262K | Ornith | **163.0 s** | **122.8 tok/s** | 1,411 tok/s | **111.7 tok/s** | **0.65 s, correct** |
Across the nine actual answers, Ornith reduced waiting time by 20% in Fast,
30% in Medium and 39% in Ultra. Long decode improved by 39%, 71% and 77%.
Prefill was profile-dependent: 6% slower in Fast, 81% faster in Medium and
14% slower in Ultra. Repeating the same 27K prompt hit llama.cpp's cache for
both models.
## GPU distribution and stability
The split was tuned empirically, one whole model layer at a time. A 73:27 and
72:28 Fast split loaded most weights but OOMed on the first CUDA calculation.
The following are the fastest splits that both loaded and generated:
| Profile | Tensor split 5080:3060 | RTX 5080 after generation | RTX 3060 after generation | 5080 reserve |
|---|---:|---:|---:|---:|
| Fast | 70:30 | 15,536 MiB | 6,685 MiB | 393 MiB |
| Medium | 68:32 | 15,720 MiB | 8,471 MiB | 209 MiB |
| Ultra | 65:35 | 15,842 MiB | 9,173 MiB | 87 MiB |
All three complete runs passed without OOM. `--n-gpu-layers all` keeps all
model layers on CUDA; no CPU model-layer offload was configured. The lower
Ultra ratio is required because its larger KV cache also consumes GPU memory.
Qwen-TTS and Whisper were stopped only during each isolated Ornith window and
restored by a signal-safe trap.
## Correctness
The fixed rubric awards 0 (wrong/missing), 1 (partly correct) or 2 (correct
and supported) for nine answers plus the tool call.
| Profile | Qwen | Ornith | Material difference |
|---|---:|---:|---|
| Fast | 19/20 | 18/20 | Ornith's async solution was wrong; Qwen's migration proof was truncated |
| Medium | **20/20** | 18/20 | Ornith's async solution was wrong |
| Ultra | 17/20 | 16/20 | both had long-answer truncation; Ornith's async solution was wrong |
Ornith correctly solved logic, evidence diagnosis, capacity planning,
instruction injection, runtime-state interpretation, read-only diagnosis,
destructive-action handling and evidence boundaries. It also emitted the
required `get_weather({"city":"Rastatt"})` call without inventing a result.
The blocking defect is reproducible async-agent logic. In all three main
profiles and both repeated Fast seeds, Ornith failed the requirement "return
the first *successful* concurrent task and cancel/await the rest":
- one answer used `asyncio.wait` without `FIRST_COMPLETED` and even passed the
unsupported `return_exceptions` argument;
- repeated answers used `asyncio.gather`, which waits for every task and then
chooses by input order instead of completion order;
- another answer returned the entire `gather` result rather than the first
successful result.
Qwen Fast and Medium produced the correct `create_task` + `FIRST_COMPLETED` +
cancel + `gather(..., return_exceptions=True)` pattern in the main run and both
repeated Fast seeds. This is directly relevant to OpenClaw's long tool runs,
so Ornith is not a safe production replacement despite its speed.
Both models recalled the exact marker `KIESEL-7319` from a 48,644-token prompt.
Qwen took 40.9 s and Ornith 42.2 s. Both read the vision fixture correctly as
a red circle on the left and a blue square on the right; Qwen took 3.18 s and
Ornith 9.80 s. Ornith's projector works, but this small vision case was about
three times slower.
## Decision
**Keep Qwen3.8-27B as Athena's production model.** Ornith is a compelling
speed experiment and a useful stopped candidate for ordinary text workloads,
but it is measurably less reliable on the kind of concurrent control logic an
agent must generate. Do not replace Fast, Medium or Ultra silently.
The experiment remains isolated and reversible. The Ornith container is
stopped. Final verification showed `active_profile=medium`; router, Qwen Medium,
Whisper and Qwen3-TTS were healthy, and Qwen Medium returned exactly `OK` to a
live request. `cleanup.sh --all` removes only the experiment container,
weights, raw results and deployment staging. Sources: [official GGUF](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF),
[official evaluation](https://ornith.ai/ornith_1_5.html),
[published llama.cpp configuration](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF/discussions/14).