101 lines
5.4 KiB
Markdown
101 lines
5.4 KiB
Markdown
# Ornith 1.5 35B-A3B A/B result on Athena
|
|
|
|
Measured 2026-09-19 against Athena's production Qwen3.8-27B profiles. The
|
|
same nine text tasks, tool schema, 27,013-token prefill, 2,048-token decode,
|
|
long-context recall and synthetic vision fixture were used. Ornith ran from
|
|
the official Q4_K_M GGUF at pinned revision
|
|
`12393612fd4f730ff5aadc23e9b8f9648aa49ceb`.
|
|
|
|
Ornith requires a 2,048-token reasoning cap with medium effort. Without it,
|
|
the model exhausted the complete 4,096-token response allowance on several
|
|
tasks and returned no visible final answer. This report uses the corrected
|
|
configuration published for this exact GGUF.
|
|
|
|
## Performance
|
|
|
|
| Profile | Model | Nine tasks | Task decode | 27K prefill | Long decode | Weather tool |
|
|
|---|---|---:|---:|---:|---:|---:|
|
|
| Fast 76.8K | Qwen | 191.1 s | 95.1 tok/s | 1,306 tok/s | 84.7 tok/s | 0.68 s, correct |
|
|
| Fast 76.8K | Ornith | **152.9 s** | **124.2 tok/s** | 1,225 tok/s | **118.0 tok/s** | **0.58 s, correct** |
|
|
| Medium 160K | Qwen | 220.4 s | 75.7 tok/s | 1,836 tok/s | 65.5 tok/s | 0.68 s, correct |
|
|
| Medium 160K | Ornith | **155.2 s** | **123.7 tok/s** | **3,320 tok/s** | **111.9 tok/s** | **0.55 s, correct** |
|
|
| Ultra 262K | Qwen | 265.3 s | 68.8 tok/s | **1,632 tok/s** | 63.1 tok/s | 0.81 s, correct |
|
|
| Ultra 262K | Ornith | **163.0 s** | **122.8 tok/s** | 1,411 tok/s | **111.7 tok/s** | **0.65 s, correct** |
|
|
|
|
Across the nine actual answers, Ornith reduced waiting time by 20% in Fast,
|
|
30% in Medium and 39% in Ultra. Long decode improved by 39%, 71% and 77%.
|
|
Prefill was profile-dependent: 6% slower in Fast, 81% faster in Medium and
|
|
14% slower in Ultra. Repeating the same 27K prompt hit llama.cpp's cache for
|
|
both models.
|
|
|
|
## GPU distribution and stability
|
|
|
|
The split was tuned empirically, one whole model layer at a time. A 73:27 and
|
|
72:28 Fast split loaded most weights but OOMed on the first CUDA calculation.
|
|
The following are the fastest splits that both loaded and generated:
|
|
|
|
| Profile | Tensor split 5080:3060 | RTX 5080 after generation | RTX 3060 after generation | 5080 reserve |
|
|
|---|---:|---:|---:|---:|
|
|
| Fast | 70:30 | 15,536 MiB | 6,685 MiB | 393 MiB |
|
|
| Medium | 68:32 | 15,720 MiB | 8,471 MiB | 209 MiB |
|
|
| Ultra | 65:35 | 15,842 MiB | 9,173 MiB | 87 MiB |
|
|
|
|
All three complete runs passed without OOM. `--n-gpu-layers all` keeps all
|
|
model layers on CUDA; no CPU model-layer offload was configured. The lower
|
|
Ultra ratio is required because its larger KV cache also consumes GPU memory.
|
|
Qwen-TTS and Whisper were stopped only during each isolated Ornith window and
|
|
restored by a signal-safe trap.
|
|
|
|
## Correctness
|
|
|
|
The fixed rubric awards 0 (wrong/missing), 1 (partly correct) or 2 (correct
|
|
and supported) for nine answers plus the tool call.
|
|
|
|
| Profile | Qwen | Ornith | Material difference |
|
|
|---|---:|---:|---|
|
|
| Fast | 19/20 | 18/20 | Ornith's async solution was wrong; Qwen's migration proof was truncated |
|
|
| Medium | **20/20** | 18/20 | Ornith's async solution was wrong |
|
|
| Ultra | 17/20 | 16/20 | both had long-answer truncation; Ornith's async solution was wrong |
|
|
|
|
Ornith correctly solved logic, evidence diagnosis, capacity planning,
|
|
instruction injection, runtime-state interpretation, read-only diagnosis,
|
|
destructive-action handling and evidence boundaries. It also emitted the
|
|
required `get_weather({"city":"Rastatt"})` call without inventing a result.
|
|
|
|
The blocking defect is reproducible async-agent logic. In all three main
|
|
profiles and both repeated Fast seeds, Ornith failed the requirement "return
|
|
the first *successful* concurrent task and cancel/await the rest":
|
|
|
|
- one answer used `asyncio.wait` without `FIRST_COMPLETED` and even passed the
|
|
unsupported `return_exceptions` argument;
|
|
- repeated answers used `asyncio.gather`, which waits for every task and then
|
|
chooses by input order instead of completion order;
|
|
- another answer returned the entire `gather` result rather than the first
|
|
successful result.
|
|
|
|
Qwen Fast and Medium produced the correct `create_task` + `FIRST_COMPLETED` +
|
|
cancel + `gather(..., return_exceptions=True)` pattern in the main run and both
|
|
repeated Fast seeds. This is directly relevant to OpenClaw's long tool runs,
|
|
so Ornith is not a safe production replacement despite its speed.
|
|
|
|
Both models recalled the exact marker `KIESEL-7319` from a 48,644-token prompt.
|
|
Qwen took 40.9 s and Ornith 42.2 s. Both read the vision fixture correctly as
|
|
a red circle on the left and a blue square on the right; Qwen took 3.18 s and
|
|
Ornith 9.80 s. Ornith's projector works, but this small vision case was about
|
|
three times slower.
|
|
|
|
## Decision
|
|
|
|
**Keep Qwen3.8-27B as Athena's production model.** Ornith is a compelling
|
|
speed experiment and a useful stopped candidate for ordinary text workloads,
|
|
but it is measurably less reliable on the kind of concurrent control logic an
|
|
agent must generate. Do not replace Fast, Medium or Ultra silently.
|
|
|
|
The experiment remains isolated and reversible. The Ornith container is
|
|
stopped. Final verification showed `active_profile=medium`; router, Qwen Medium,
|
|
Whisper and Qwen3-TTS were healthy, and Qwen Medium returned exactly `OK` to a
|
|
live request. `cleanup.sh --all` removes only the experiment container,
|
|
weights, raw results and deployment staging. Sources: [official GGUF](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF),
|
|
[official evaluation](https://ornith.ai/ornith_1_5.html),
|
|
[published llama.cpp configuration](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF/discussions/14).
|