Files

5.4 KiB

Ornith 1.5 35B-A3B A/B result on Athena

Measured 2026-09-19 against Athena's production Qwen3.8-27B profiles. The same nine text tasks, tool schema, 27,013-token prefill, 2,048-token decode, long-context recall and synthetic vision fixture were used. Ornith ran from the official Q4_K_M GGUF at pinned revision 12393612fd4f730ff5aadc23e9b8f9648aa49ceb.

Ornith requires a 2,048-token reasoning cap with medium effort. Without it, the model exhausted the complete 4,096-token response allowance on several tasks and returned no visible final answer. This report uses the corrected configuration published for this exact GGUF.

Performance

Profile Model Nine tasks Task decode 27K prefill Long decode Weather tool
Fast 76.8K Qwen 191.1 s 95.1 tok/s 1,306 tok/s 84.7 tok/s 0.68 s, correct
Fast 76.8K Ornith 152.9 s 124.2 tok/s 1,225 tok/s 118.0 tok/s 0.58 s, correct
Medium 160K Qwen 220.4 s 75.7 tok/s 1,836 tok/s 65.5 tok/s 0.68 s, correct
Medium 160K Ornith 155.2 s 123.7 tok/s 3,320 tok/s 111.9 tok/s 0.55 s, correct
Ultra 262K Qwen 265.3 s 68.8 tok/s 1,632 tok/s 63.1 tok/s 0.81 s, correct
Ultra 262K Ornith 163.0 s 122.8 tok/s 1,411 tok/s 111.7 tok/s 0.65 s, correct

Across the nine actual answers, Ornith reduced waiting time by 20% in Fast, 30% in Medium and 39% in Ultra. Long decode improved by 39%, 71% and 77%. Prefill was profile-dependent: 6% slower in Fast, 81% faster in Medium and 14% slower in Ultra. Repeating the same 27K prompt hit llama.cpp's cache for both models.

GPU distribution and stability

The split was tuned empirically, one whole model layer at a time. A 73:27 and 72:28 Fast split loaded most weights but OOMed on the first CUDA calculation. The following are the fastest splits that both loaded and generated:

Profile Tensor split 5080:3060 RTX 5080 after generation RTX 3060 after generation 5080 reserve
Fast 70:30 15,536 MiB 6,685 MiB 393 MiB
Medium 68:32 15,720 MiB 8,471 MiB 209 MiB
Ultra 65:35 15,842 MiB 9,173 MiB 87 MiB

All three complete runs passed without OOM. --n-gpu-layers all keeps all model layers on CUDA; no CPU model-layer offload was configured. The lower Ultra ratio is required because its larger KV cache also consumes GPU memory. Qwen-TTS and Whisper were stopped only during each isolated Ornith window and restored by a signal-safe trap.

Correctness

The fixed rubric awards 0 (wrong/missing), 1 (partly correct) or 2 (correct and supported) for nine answers plus the tool call.

Profile Qwen Ornith Material difference
Fast 19/20 18/20 Ornith's async solution was wrong; Qwen's migration proof was truncated
Medium 20/20 18/20 Ornith's async solution was wrong
Ultra 17/20 16/20 both had long-answer truncation; Ornith's async solution was wrong

Ornith correctly solved logic, evidence diagnosis, capacity planning, instruction injection, runtime-state interpretation, read-only diagnosis, destructive-action handling and evidence boundaries. It also emitted the required get_weather({"city":"Rastatt"}) call without inventing a result.

The blocking defect is reproducible async-agent logic. In all three main profiles and both repeated Fast seeds, Ornith failed the requirement "return the first successful concurrent task and cancel/await the rest":

  • one answer used asyncio.wait without FIRST_COMPLETED and even passed the unsupported return_exceptions argument;
  • repeated answers used asyncio.gather, which waits for every task and then chooses by input order instead of completion order;
  • another answer returned the entire gather result rather than the first successful result.

Qwen Fast and Medium produced the correct create_task + FIRST_COMPLETED + cancel + gather(..., return_exceptions=True) pattern in the main run and both repeated Fast seeds. This is directly relevant to OpenClaw's long tool runs, so Ornith is not a safe production replacement despite its speed.

Both models recalled the exact marker KIESEL-7319 from a 48,644-token prompt. Qwen took 40.9 s and Ornith 42.2 s. Both read the vision fixture correctly as a red circle on the left and a blue square on the right; Qwen took 3.18 s and Ornith 9.80 s. Ornith's projector works, but this small vision case was about three times slower.

Decision

Keep Qwen3.8-27B as Athena's production model. Ornith is a compelling speed experiment and a useful stopped candidate for ordinary text workloads, but it is measurably less reliable on the kind of concurrent control logic an agent must generate. Do not replace Fast, Medium or Ultra silently.

The experiment remains isolated and reversible. The Ornith container is stopped. Final verification showed active_profile=medium; router, Qwen Medium, Whisper and Qwen3-TTS were healthy, and Qwen Medium returned exactly OK to a live request. cleanup.sh --all removes only the experiment container, weights, raw results and deployment staging. Sources: official GGUF, official evaluation, published llama.cpp configuration.