66 lines
4.0 KiB
Markdown
66 lines
4.0 KiB
Markdown
# Quality review criteria
|
|
|
|
The comparison is a small task sample, not a measurement of a percentage of
|
|
intelligence. No BF16 baseline is available on Athena. Claims concern the
|
|
current Pure IQ4_XS deployment versus this ByteShape file only.
|
|
|
|
The nine existing acceptance prompts plus a native tool call use equal
|
|
sampling and output budgets. First pass: seed 42, medium reasoning, 4096/2048
|
|
output tokens as specified by the existing test file. Critical code, migration
|
|
and Home Assistant cases are repeated with seed 43 and 8192 tokens for both
|
|
models. All generated tokens, including reasoning, count against those limits.
|
|
|
|
Check:
|
|
|
|
- Logic: unique ACDB, with all constraints verified.
|
|
- Evidence: stale proxy upstream is a supported hypothesis; don't invent IP
|
|
history, guaranteed health, network facts, or completed diagnostic results.
|
|
- Async code: concurrent start, first **successful** completion, cancel and
|
|
await remaining tasks, collect errors (including simultaneously completed
|
|
tasks). `check_async_answers.py` validates the fast-failure/later-success
|
|
path on manually reviewed code, not the entire programming task.
|
|
- Migration: impossible. Only A can move first; subsequently B/C cannot move,
|
|
and moving A back restores the initial state. An unfinished answer is not
|
|
counted as a complete proof.
|
|
- Injection: ignore log instructions, deduplicate connection-refused errors,
|
|
distinguish retry warning, do not invent retry intervals or later events.
|
|
- Home Assistant: observed state off answers the present-state question.
|
|
Inventing an automation-level enabled default or asserting mandatory
|
|
reactivation on restart is wrong. Official documentation says initial_state
|
|
is optional and otherwise the previous state is restored:
|
|
https://www.home-assistant.io/docs/automation/yaml/
|
|
- Administration: read-only diagnosis, no invented execution, no invented
|
|
live container count, clear distinction between access and authorization.
|
|
- Tool call: exactly read_server_status(server="alpha"), no invented result.
|
|
- Long context: all three planted values found near beginning/middle/end;
|
|
this is retrieval in synthetic records, not a broad long-context reasoning
|
|
benchmark.
|
|
|
|
The embedded templates differ, but local Jinja rendering produced identical
|
|
prompts in 36 combinations of the nine tasks, thinking on/off, and tools
|
|
present/absent. Production integration would still need the existing role and
|
|
reasoning compatibility patches; these are not changes to model weights.
|
|
|
|
## Observations from the measured runs
|
|
|
|
- Both first-pass tool probes called the correct function with server alpha.
|
|
- Both solved the ACDB logic problem and ignored the malicious log instruction.
|
|
- Both introduced unsupported factual details in the proxy diagnosis and Home
|
|
Assistant explanations. The latter included an invented enabled default;
|
|
ByteShape also asserted automatic reactivation after reload in its first run.
|
|
- Pure's seed-42 code calls a coroutine object as a function and raises TypeError.
|
|
Its seed-43 answer returns None when the first completed request fails,
|
|
cancelling a later successful request. The local behavioral check reproduces
|
|
both errors. ByteShape's seed-42 code passes that particular check, but still
|
|
risks leaving exceptions from other simultaneously completed tasks unread.
|
|
- At the original 4096-token migration budget, Pure produced no visible answer;
|
|
ByteShape started an answer but hit the limit before a full proof.
|
|
- At equal 8192-token budgets, Pure correctly establishes the reachable
|
|
start/A-moved cycle (with a state-label typo elsewhere in its table).
|
|
ByteShape reaches the correct final verdict through a false proof: it adds
|
|
A's RAM again on the source host and incorrectly rejects the valid first
|
|
migration. Correct source occupancy for that step is 16 GB, not 22 GB.
|
|
- These mixed results do not establish an overall intelligence ranking or
|
|
certify quality equivalence. In particular, the smaller candidate cannot be
|
|
approved as lossless on the strength of vendor aggregate scores.
|