Document Ornith 1.5 A/B benchmark
This commit is contained in:
@@ -0,0 +1,100 @@
|
||||
# Ornith 1.5 35B-A3B A/B result on Athena
|
||||
|
||||
Measured 2026-09-19 against Athena's production Qwen3.8-27B profiles. The
|
||||
same nine text tasks, tool schema, 27,013-token prefill, 2,048-token decode,
|
||||
long-context recall and synthetic vision fixture were used. Ornith ran from
|
||||
the official Q4_K_M GGUF at pinned revision
|
||||
`12393612fd4f730ff5aadc23e9b8f9648aa49ceb`.
|
||||
|
||||
Ornith requires a 2,048-token reasoning cap with medium effort. Without it,
|
||||
the model exhausted the complete 4,096-token response allowance on several
|
||||
tasks and returned no visible final answer. This report uses the corrected
|
||||
configuration published for this exact GGUF.
|
||||
|
||||
## Performance
|
||||
|
||||
| Profile | Model | Nine tasks | Task decode | 27K prefill | Long decode | Weather tool |
|
||||
|---|---|---:|---:|---:|---:|---:|
|
||||
| Fast 76.8K | Qwen | 191.1 s | 95.1 tok/s | 1,306 tok/s | 84.7 tok/s | 0.68 s, correct |
|
||||
| Fast 76.8K | Ornith | **152.9 s** | **124.2 tok/s** | 1,225 tok/s | **118.0 tok/s** | **0.58 s, correct** |
|
||||
| Medium 160K | Qwen | 220.4 s | 75.7 tok/s | 1,836 tok/s | 65.5 tok/s | 0.68 s, correct |
|
||||
| Medium 160K | Ornith | **155.2 s** | **123.7 tok/s** | **3,320 tok/s** | **111.9 tok/s** | **0.55 s, correct** |
|
||||
| Ultra 262K | Qwen | 265.3 s | 68.8 tok/s | **1,632 tok/s** | 63.1 tok/s | 0.81 s, correct |
|
||||
| Ultra 262K | Ornith | **163.0 s** | **122.8 tok/s** | 1,411 tok/s | **111.7 tok/s** | **0.65 s, correct** |
|
||||
|
||||
Across the nine actual answers, Ornith reduced waiting time by 20% in Fast,
|
||||
30% in Medium and 39% in Ultra. Long decode improved by 39%, 71% and 77%.
|
||||
Prefill was profile-dependent: 6% slower in Fast, 81% faster in Medium and
|
||||
14% slower in Ultra. Repeating the same 27K prompt hit llama.cpp's cache for
|
||||
both models.
|
||||
|
||||
## GPU distribution and stability
|
||||
|
||||
The split was tuned empirically, one whole model layer at a time. A 73:27 and
|
||||
72:28 Fast split loaded most weights but OOMed on the first CUDA calculation.
|
||||
The following are the fastest splits that both loaded and generated:
|
||||
|
||||
| Profile | Tensor split 5080:3060 | RTX 5080 after generation | RTX 3060 after generation | 5080 reserve |
|
||||
|---|---:|---:|---:|---:|
|
||||
| Fast | 70:30 | 15,536 MiB | 6,685 MiB | 393 MiB |
|
||||
| Medium | 68:32 | 15,720 MiB | 8,471 MiB | 209 MiB |
|
||||
| Ultra | 65:35 | 15,842 MiB | 9,173 MiB | 87 MiB |
|
||||
|
||||
All three complete runs passed without OOM. `--n-gpu-layers all` keeps all
|
||||
model layers on CUDA; no CPU model-layer offload was configured. The lower
|
||||
Ultra ratio is required because its larger KV cache also consumes GPU memory.
|
||||
Qwen-TTS and Whisper were stopped only during each isolated Ornith window and
|
||||
restored by a signal-safe trap.
|
||||
|
||||
## Correctness
|
||||
|
||||
The fixed rubric awards 0 (wrong/missing), 1 (partly correct) or 2 (correct
|
||||
and supported) for nine answers plus the tool call.
|
||||
|
||||
| Profile | Qwen | Ornith | Material difference |
|
||||
|---|---:|---:|---|
|
||||
| Fast | 19/20 | 18/20 | Ornith's async solution was wrong; Qwen's migration proof was truncated |
|
||||
| Medium | **20/20** | 18/20 | Ornith's async solution was wrong |
|
||||
| Ultra | 17/20 | 16/20 | both had long-answer truncation; Ornith's async solution was wrong |
|
||||
|
||||
Ornith correctly solved logic, evidence diagnosis, capacity planning,
|
||||
instruction injection, runtime-state interpretation, read-only diagnosis,
|
||||
destructive-action handling and evidence boundaries. It also emitted the
|
||||
required `get_weather({"city":"Rastatt"})` call without inventing a result.
|
||||
|
||||
The blocking defect is reproducible async-agent logic. In all three main
|
||||
profiles and both repeated Fast seeds, Ornith failed the requirement "return
|
||||
the first *successful* concurrent task and cancel/await the rest":
|
||||
|
||||
- one answer used `asyncio.wait` without `FIRST_COMPLETED` and even passed the
|
||||
unsupported `return_exceptions` argument;
|
||||
- repeated answers used `asyncio.gather`, which waits for every task and then
|
||||
chooses by input order instead of completion order;
|
||||
- another answer returned the entire `gather` result rather than the first
|
||||
successful result.
|
||||
|
||||
Qwen Fast and Medium produced the correct `create_task` + `FIRST_COMPLETED` +
|
||||
cancel + `gather(..., return_exceptions=True)` pattern in the main run and both
|
||||
repeated Fast seeds. This is directly relevant to OpenClaw's long tool runs,
|
||||
so Ornith is not a safe production replacement despite its speed.
|
||||
|
||||
Both models recalled the exact marker `KIESEL-7319` from a 48,644-token prompt.
|
||||
Qwen took 40.9 s and Ornith 42.2 s. Both read the vision fixture correctly as
|
||||
a red circle on the left and a blue square on the right; Qwen took 3.18 s and
|
||||
Ornith 9.80 s. Ornith's projector works, but this small vision case was about
|
||||
three times slower.
|
||||
|
||||
## Decision
|
||||
|
||||
**Keep Qwen3.8-27B as Athena's production model.** Ornith is a compelling
|
||||
speed experiment and a useful stopped candidate for ordinary text workloads,
|
||||
but it is measurably less reliable on the kind of concurrent control logic an
|
||||
agent must generate. Do not replace Fast, Medium or Ultra silently.
|
||||
|
||||
The experiment remains isolated and reversible. The Ornith container is
|
||||
stopped. Final verification showed `active_profile=medium`; router, Qwen Medium,
|
||||
Whisper and Qwen3-TTS were healthy, and Qwen Medium returned exactly `OK` to a
|
||||
live request. `cleanup.sh --all` removes only the experiment container,
|
||||
weights, raw results and deployment staging. Sources: [official GGUF](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF),
|
||||
[official evaluation](https://ornith.ai/ornith_1_5.html),
|
||||
[published llama.cpp configuration](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-GGUF/discussions/14).
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
@@ -0,0 +1,34 @@
|
||||
{
|
||||
"model": "ornith15-test",
|
||||
"prompt": "Was ist links und was ist rechts im Bild? Nenne Form und Farbe beider Objekte.",
|
||||
"expected": "Links ein roter Kreis, rechts ein blaues Quadrat.",
|
||||
"started": 1789825609.351162,
|
||||
"response": {
|
||||
"wall_seconds": 9.802,
|
||||
"usage": {
|
||||
"completion_tokens": 53,
|
||||
"prompt_tokens": 162,
|
||||
"total_tokens": 215,
|
||||
"prompt_tokens_details": {
|
||||
"cached_tokens": 0
|
||||
}
|
||||
},
|
||||
"timings": {
|
||||
"cache_n": 0,
|
||||
"prompt_n": 162,
|
||||
"prompt_ms": 9112.496,
|
||||
"prompt_per_token_ms": 56.24997530864197,
|
||||
"prompt_per_second": 17.777785581469665,
|
||||
"predicted_n": 53,
|
||||
"predicted_ms": 674.115,
|
||||
"predicted_per_token_ms": 12.963750000000001,
|
||||
"predicted_per_second": 77.13817375373637,
|
||||
"draft_n": 60,
|
||||
"draft_n_accepted": 32
|
||||
},
|
||||
"finish_reason": "stop",
|
||||
"content": "Im Bild sind zwei Objekte zu sehen:\n\n**Links:**\n- **Form:** Kreis\n- **Farbe:** Rot\n\n**Rechts:**\n- **Form:** Quadrat (bzw. Rechteck)\n- **Farbe:** Blau",
|
||||
"reasoning_content": "",
|
||||
"tool_calls": []
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,34 @@
|
||||
{
|
||||
"model": "qwen-fast",
|
||||
"prompt": "Was ist links und was ist rechts im Bild? Nenne Form und Farbe beider Objekte.",
|
||||
"expected": "Links ein roter Kreis, rechts ein blaues Quadrat.",
|
||||
"started": 1789825590.7140028,
|
||||
"response": {
|
||||
"wall_seconds": 3.177,
|
||||
"usage": {
|
||||
"completion_tokens": 84,
|
||||
"prompt_tokens": 162,
|
||||
"total_tokens": 246,
|
||||
"prompt_tokens_details": {
|
||||
"cached_tokens": 0
|
||||
}
|
||||
},
|
||||
"timings": {
|
||||
"cache_n": 0,
|
||||
"prompt_n": 162,
|
||||
"prompt_ms": 2185.727,
|
||||
"prompt_per_token_ms": 13.492141975308641,
|
||||
"prompt_per_second": 74.11721591946295,
|
||||
"predicted_n": 84,
|
||||
"predicted_ms": 977.602,
|
||||
"predicted_per_token_ms": 11.77833734939759,
|
||||
"predicted_per_second": 84.90162663333341,
|
||||
"draft_n": 64,
|
||||
"draft_n_accepted": 51
|
||||
},
|
||||
"finish_reason": "stop",
|
||||
"content": "Im Bild sind zwei geometrische Objekte zu sehen:\n\n- **Links**: Ein **roter Kreis**.\n - Form: Kreis\n - Farbe: Rot\n\n- **Rechts**: Ein **blaues Quadrat**.\n - Form: Quadrat\n - Farbe: Blau\n\nZusammenfassend:\n> Links ist ein roter Kreis, rechts ist ein blaues Quadrat.",
|
||||
"reasoning_content": "",
|
||||
"tool_calls": []
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user