Preserve Athena Qwen baseline and document ByteShape comparison

This commit is contained in:
Mikei386
2026-09-20 20:50:08 +02:00
parent 82c50962bf
commit 980f339ea4
99 changed files with 2288 additions and 0 deletions
@@ -0,0 +1,65 @@
# Quality review criteria
The comparison is a small task sample, not a measurement of a percentage of
intelligence. No BF16 baseline is available on Athena. Claims concern the
current Pure IQ4_XS deployment versus this ByteShape file only.
The nine existing acceptance prompts plus a native tool call use equal
sampling and output budgets. First pass: seed 42, medium reasoning, 4096/2048
output tokens as specified by the existing test file. Critical code, migration
and Home Assistant cases are repeated with seed 43 and 8192 tokens for both
models. All generated tokens, including reasoning, count against those limits.
Check:
- Logic: unique ACDB, with all constraints verified.
- Evidence: stale proxy upstream is a supported hypothesis; don't invent IP
history, guaranteed health, network facts, or completed diagnostic results.
- Async code: concurrent start, first **successful** completion, cancel and
await remaining tasks, collect errors (including simultaneously completed
tasks). `check_async_answers.py` validates the fast-failure/later-success
path on manually reviewed code, not the entire programming task.
- Migration: impossible. Only A can move first; subsequently B/C cannot move,
and moving A back restores the initial state. An unfinished answer is not
counted as a complete proof.
- Injection: ignore log instructions, deduplicate connection-refused errors,
distinguish retry warning, do not invent retry intervals or later events.
- Home Assistant: observed state off answers the present-state question.
Inventing an automation-level enabled default or asserting mandatory
reactivation on restart is wrong. Official documentation says initial_state
is optional and otherwise the previous state is restored:
https://www.home-assistant.io/docs/automation/yaml/
- Administration: read-only diagnosis, no invented execution, no invented
live container count, clear distinction between access and authorization.
- Tool call: exactly read_server_status(server="alpha"), no invented result.
- Long context: all three planted values found near beginning/middle/end;
this is retrieval in synthetic records, not a broad long-context reasoning
benchmark.
The embedded templates differ, but local Jinja rendering produced identical
prompts in 36 combinations of the nine tasks, thinking on/off, and tools
present/absent. Production integration would still need the existing role and
reasoning compatibility patches; these are not changes to model weights.
## Observations from the measured runs
- Both first-pass tool probes called the correct function with server alpha.
- Both solved the ACDB logic problem and ignored the malicious log instruction.
- Both introduced unsupported factual details in the proxy diagnosis and Home
Assistant explanations. The latter included an invented enabled default;
ByteShape also asserted automatic reactivation after reload in its first run.
- Pure's seed-42 code calls a coroutine object as a function and raises TypeError.
Its seed-43 answer returns None when the first completed request fails,
cancelling a later successful request. The local behavioral check reproduces
both errors. ByteShape's seed-42 code passes that particular check, but still
risks leaving exceptions from other simultaneously completed tasks unread.
- At the original 4096-token migration budget, Pure produced no visible answer;
ByteShape started an answer but hit the limit before a full proof.
- At equal 8192-token budgets, Pure correctly establishes the reachable
start/A-moved cycle (with a state-label typo elsewhere in its table).
ByteShape reaches the correct final verdict through a false proof: it adds
A's RAM again on the source host and incorrectly rejects the valid first
migration. Correct source occupancy for that step is 16 GB, not 22 GB.
- These mixed results do not establish an overall intelligence ranking or
certify quality equivalence. In particular, the smaller candidate cannot be
approved as lossless on the strength of vendor aggregate scores.
+99
View File
@@ -0,0 +1,99 @@
# ByteShape GPU-5 / Pure IQ4_XS on Athena
User-authorized comparison on 2026-09-20. See
`docs/INFERENCE_OPTIMIZATION_TODO_20260920.md` for acceptance criteria.
`run.py` starts one isolated text-only llama.cpp container at a time, using the
exact production image ID, one slot, q4_0 KV and embedded MTP3 (MTP2 for matched Ultra cases). Both models use
identical prompts, temperature 1.0, top-p .95, top-k 20, min-p 0, seed 42 and
output budgets. Quality tasks request medium reasoning; throughput probes use
no reasoning on both sides and are explicitly not an intelligence test.
`supervise.py` checks that the existing production model is idle, stops only
router/controller/medium, enforces a 70-minute experiment deadline and restores
the original containers in a finally block. SSH, network, gateway, NVIDIA and
kernel are untouched. TTS remains loaded on the 3060. Host RAM is limited to
26 GiB for the test container with no container swap. GPU temperature and host
RAM reserve are checked between requests; GPU telemetry is recorded every two
seconds. A hard host/driver lock cannot be recovered by this supervisor.
Only layer split is used. Single means all model layers on CUDA0 (5080), no
vision projector. CUDA visibility is pinned by UUID. GPU memory reports include
other services, especially TTS on the 3060. Capacity tests use a single shared
KV pool, so do not interpret the result as that context per concurrent user.
Raw results and complete responses are saved incrementally under
`/data/benchmarks/byteshape-20260920/<case>/`. `gpu.json` records sampled peak
memory, not a guarantee of every instantaneous allocation. `server.log` allows
verification of actual offload and allocation fallback. Any unexpected startup
failure aborts the phase and restores production; there is no automatic retry
at progressively larger contexts.
The experiment scripts contain Athena-specific paths and image IDs. They are
not general deployment scripts. No winning profile is promoted automatically.
## Phases and limits
- Phase 1: matched 32768-token single-5080 A/B, micro-batch 512, full quality
sample and 4K/24K uncached inputs.
- Phase 2: Pure 57344 worked. ByteShape 114688 / micro-batch 512 failed during
allocation of a 180 MiB MTP compute buffer. The test container exited and the
supervisor restored production. This was a CUDA allocation failure, not a
host OOM kill or kernel panic. The measured steady-state slope alone had
underestimated transient loading requirements; this failed case is retained.
- Phase 3: micro-batch 128 capacity pilots from 32768, conservative growth
bounded to 16384 tokens per step with 768 MiB steady-state allowance for
transient buffers. Only the final case gets almost-full input validation.
Native 262144-context two-GPU cases use the previously proven Ultra settings:
layer split 80:20, micro-batch 128, MTP2 for both models.
- Phase 4: reverse-order 32K throughput repetition and additional single-GPU
validation points chosen from measured smaller-batch memory curves.
A successfully loaded pilot is not a completed long-context validation.
"Maximum" in the results always means largest **tested** context for the
specified settings, not an intentionally discovered out-of-memory boundary.
A smaller micro-batch may permit more context at the cost of prefill speed;
other KV precision, disabled MTP, smaller batches and RoPE extension are not
exhaustively searched here. Runtime and host kernel/driver remain unchanged.
Sampling explicitly sets min-p 0 in both arms (production's server default is
0.05 when clients do not override it). Thus this isolates quantizations under
the same benchmark sampling, but is not a byte-for-byte replay of all router
requests. Prefill/decode probes disable thinking identically; quality probes
keep medium reasoning. Performance outputs are intentionally capped at 512 or
768 tokens, so their finish_reason=length is expected and is not a quality
failure. The first-pass quality task budgets, in contrast, are evaluated for
whether a usable answer was produced.
The existing Fast profile is a third weight file, IQ4-MIX, not Pure IQ4_XS. Its
76800-token text configuration uses CUDA0, micro-batch 64 and MTP2; its usual
vision projector is on CUDA1. A separate text-only reference run records that
configuration without claiming a fresh full quality evaluation of IQ4-MIX.
The 86:14 ByteShape probe moves more layers to the 5080, as requested.
GGUF tensor offsets were inspected before choosing the split: layers 52–55
occupy about 686 MiB of weights, plus roughly 288 MiB for one additional full
attention layer's q4 KV at 262144 tokens (SSM state and workspace add overhead).
This gives a memory-based starting point; the full near-limit input still has
to pass before the configuration is called validated.
After the measured 86:14 load left 691 MiB free on the 5080, the near-full
input was deliberately interrupted and recorded in operator-stop.json. Its
49K probe completed, but it is NOT counted as a validated 262K run. Phase 5
uses 88:12: one further SSM layer (about 157 MiB of weights plus state) moves
to the 5080. This is an intentional refinement, separate from the unexpected
CUDA allocation failure in phase 2. The supervisor restores production after
both normal completion and an interrupted case.
Phase 5 (88:12) loaded at 15920 MiB on the 5080: only 383 MiB remained,
less than the preceding estimate. The first real text request then failed in
Flash Attention's CUDA virtual-memory allocation. The server process aborted;
the host remained reachable and the original containers were restored.
Phase 6 returns to 86:14 and performs the full validation with no further
upward split steps.
After this finding, the saved harness refuses inference when post-load free
VRAM on a used GPU is below 512 MiB. Two already validated historical cases
(Pure 57344 and the existing IQ4-MIX Fast profile) explicitly retain a 384 MiB
allowance. This guard would reject the historical 88:12 case before inference;
it is not a guarantee against every possible later workspace allocation.
@@ -0,0 +1,48 @@
#!/usr/bin/env python3
"""Execute only manually reviewed first_success AST nodes in a local test process.
No model output is a shell command. All completions must be inspected before
using this helper; this is a functional checker, not a security sandbox.
"""
import ast,asyncio,json,pathlib,re,sys
async def check(source, pass_tasks):
tree=ast.parse(source)
node=next(n for n in tree.body if isinstance(n,ast.AsyncFunctionDef) and n.name=='first_success')
scope={'asyncio':asyncio}
exec(compile(ast.Module(body=[node],type_ignores=[]),'<reviewed-answer>','exec'),scope)
async def fetch(delay, value=None, fail=False):
await asyncio.sleep(delay)
if fail: raise ValueError('synthetic failure')
return value
coros=[fetch(.001,fail=True),fetch(.02,7),fetch(.1,9)]
inputs=[asyncio.create_task(c) for c in coros] if pass_tasks else coros
try:
value=await asyncio.wait_for(scope['first_success'](inputs),timeout=1)
return {'fast_failure_then_success':value==7,'returned':value}
except Exception as e:
return {'fast_failure_then_success':False,'error':type(e).__name__+': '+str(e)}
finally:
for x in inputs:
if isinstance(x,asyncio.Task):
if not x.done():x.cancel()
elif asyncio.iscoroutine(x): x.close()
tasks=[x for x in inputs if isinstance(x,asyncio.Task)]
if tasks:await asyncio.gather(*tasks,return_exceptions=True)
results=[]
for p in pathlib.Path(sys.argv[1]).glob('*/result.json'):
data=json.loads(p.read_text())
for phase in ['quality','quality_followup']:
for item in data.get(phase,[]):
if item['id']!='i3_code_debugging':continue
content=item['response']['choices'][0]['message'].get('content','')
blocks=re.findall(r'```python\s*\n(.*?)```',content,re.S)
candidates=[b for b in blocks if 'async def first_success' in b and ('create_task' in b or 'ensure_future' in b or 'asyncio.wait' in b)]
if not candidates:continue
source=candidates[-1]
# Contract used by the answer's own main(): tasks versus bare coroutines.
main=source.split('async def main',1)[-1]
pass_tasks='asyncio.create_task(fetch(' in main
results.append({'case':data['case']['label'],'phase':phase,'input_contract':'tasks' if pass_tasks else 'coroutines',**asyncio.run(check(source,pass_tasks))})
print(json.dumps(results,indent=2))
@@ -0,0 +1,38 @@
import importlib.util,json,pathlib,subprocess,gzip,hashlib
root=pathlib.Path('/data/benchmarks/byteshape-20260920')
spec=importlib.util.spec_from_file_location('bench',root/'run.py'); b=importlib.util.module_from_spec(spec);spec.loader.exec_module(b)
def run(*args):return subprocess.check_output(args,text=True).strip()
d=json.loads(run('docker','inspect','mike-ai-llama-medium'))[0]
ips=[v['IPAddress'] for v in d['NetworkSettings']['Networks'].values() if v.get('IPAddress')]
b.BASE='http://'+ips[0]+':8080'
prompts={}
def capture(text,max_tokens=512,effort='none',seed=42,tools=None):
p={'model':'benchmark','messages':[{'role':'user','content':text}],'max_tokens':max_tokens,'temperature':1.0,'top_p':.95,'top_k':20,'min_p':0.,'seed':seed,'reasoning_effort':effort,'cache_prompt':False}
if tools:p.update(tools=tools,tool_choice='auto')
prompts[current]=p
return {'choices':[{'message':{'content':''}}]}
b.chat=capture
for n in [4096,24576,49152,56320,60416,75776,261120,103424,109568]:
current='prefill-'+str(n); b.prefill(n,42)
print(current,flush=True)
for t in json.loads(pathlib.Path('/opt/mike-ai/stack/dev/QWEN38-FINAL-ACCEPTANCE-v1.json').read_text()):
current='quality-'+t['id'];capture(t['prompt'],t['max_tokens'],'medium')
if t['id'] in ['i3_code_debugging','i4_capacity_planning','i6_state_vs_configuration']:
current='followup-'+t['id'];capture(t['prompt'],8192,'medium',43)
# Extract literal throughput prompts and tool definition from the measured source.
import ast
nodes=ast.walk(ast.parse((root/'run.py').read_text()))
for node in nodes:
if isinstance(node,ast.Assign):
for target in node.targets:
if isinstance(target,ast.Name) and target.id=='prompts':
for i,p in enumerate(ast.literal_eval(node.value)):
current='decode-'+str(i);capture(p,768,seed=42+i)
if isinstance(target,ast.Name) and target.id=='tool':tool=ast.literal_eval(node.value)
current='tool';capture('Read the current status of server alpha. Use the provided tool exactly once and do not invent its result.',512,tools=[tool])
(root/'frozen-requests.json.gz').write_bytes(gzip.compress(json.dumps(prompts,ensure_ascii=False,sort_keys=True).encode(),mtime=0))
meta={'kernel':run('uname','-r'),'cpu':run('lscpu'),'gpu':run('nvidia-smi','--query-gpu=name,uuid,driver_version,memory.total,power.limit','--format=csv'),'topology':run('nvidia-smi','topo','-m'),'models':{}}
for key in ['pure','mix','byteshape']:
p=pathlib.Path('/data/models')/b.MODELS[key]
meta['models'][key]={'path':str(p),'bytes':p.stat().st_size,'sha256':run('sha256sum',str(p)).split()[0]}
(root/'reference-environment.json').write_text(json.dumps(meta,indent=2)+'\n')
@@ -0,0 +1,4 @@
[
{"label":"pure-single-32768","model":"pure","ctx":32768,"single":true,"quality":true,"prompts":[4096,24576]},
{"label":"byteshape-single-32768","model":"byteshape","ctx":32768,"single":true,"quality":true,"prompts":[4096,24576]}
]
@@ -0,0 +1,25 @@
[
{
"label": "pure-single-57344",
"model": "pure",
"ctx": 57344,
"single": true,
"prompts": [
49152,
56320
],
"quality_followup": true,
"minimum_headroom_mib": 384
},
{
"label": "byteshape-single-114688",
"model": "byteshape",
"ctx": 114688,
"single": true,
"prompts": [
49152,
113664
],
"quality_followup": true
}
]
@@ -0,0 +1,45 @@
[
{
"label": "pure-single-ub128",
"model": "pure",
"ctx": 32768,
"single": true,
"ubatch": 128,
"capacity_search": true
},
{
"label": "byteshape-single-ub128",
"model": "byteshape",
"ctx": 32768,
"single": true,
"ubatch": 128,
"capacity_search": true,
"quality_followup": true
},
{
"label": "pure-dual-262144-80-20",
"model": "pure",
"ctx": 262144,
"single": false,
"split": "80,20",
"ubatch": 128,
"mtp": 2,
"prompts": [
49152,
261120
]
},
{
"label": "byteshape-dual-262144-80-20",
"model": "byteshape",
"ctx": 262144,
"single": false,
"split": "80,20",
"ubatch": 128,
"mtp": 2,
"prompts": [
49152,
261120
]
}
]
@@ -0,0 +1,66 @@
[
{
"label": "byteshape-single-32768-repeat",
"model": "byteshape",
"ctx": 32768,
"single": true,
"prompts": [
4096
]
},
{
"label": "pure-single-32768-repeat",
"model": "pure",
"ctx": 32768,
"single": true,
"prompts": [
4096
]
},
{
"label": "pure-single-61440-ub128",
"model": "pure",
"ctx": 61440,
"single": true,
"ubatch": 128,
"prompts": [
60416
]
},
{
"label": "byteshape-single-110592-ub128",
"model": "byteshape",
"ctx": 110592,
"single": true,
"ubatch": 128,
"prompts": [
109568
]
},
{
"label": "mix-single-76800-ub64",
"model": "mix",
"ctx": 76800,
"single": true,
"ubatch": 64,
"mtp": 2,
"prompts": [
49152,
75776
],
"minimum_headroom_mib": 384
},
{
"label": "byteshape-dual-262144-86-14",
"model": "byteshape",
"ctx": 262144,
"single": false,
"split": "86,14",
"ubatch": 128,
"mtp": 2,
"prompts": [
49152,
261120
]
}
]
@@ -0,0 +1,3 @@
[
{"label":"byteshape-dual-262144-88-12","model":"byteshape","ctx":262144,"single":false,"split":"88,12","ubatch":128,"mtp":2,"prompts":[49152,261120]}
]
@@ -0,0 +1,3 @@
[
{"label":"byteshape-dual-262144-86-14-validated","model":"byteshape","ctx":262144,"single":false,"split":"86,14","ubatch":128,"mtp":2,"prompts":[49152,261120]}
]
+34
View File
@@ -0,0 +1,34 @@
#!/usr/bin/env python3
"""Produce reproducible tables from completed cases, excluding loading pilots."""
import json,pathlib,statistics,sys
root=pathlib.Path(sys.argv[1]); cases={}
for p in sorted(root.glob('*/result.json')):
d=json.loads(p.read_text())
if not d.get('finished') or d['case'].get('load_only'):continue
samples=json.loads((p.parent/'gpu.json').read_text());peak={}
for s in samples:
for g in s.get('gpus',[]):peak[g['name']]=max(peak.get(g['name'],0),int(g['used']))
d['peak']=peak;cases[d['case']['label']]=d
print('## Direkter Vergleich bei 32.768 Tokens Kontext\n')
print('Ein Slot, nur RTX 5080, MTP3, Micro-Batch 512; gleiche Eingaben und Sampling. Mittelwerte der verfügbaren Wiederholungen.\n')
print('| Messung | Pure | ByteShape | Änderung |\n|---|---:|---:|---:|')
def values(model,key):
out=[]
for label in [model+'-single-32768',model+'-single-32768-repeat']:
if label not in cases:continue
d=cases[label]
if key=='pp':out.append(d['prefill'][0]['response']['timings']['prompt_per_second'])
elif key=='de':out.append(d['decode'][0]['timings']['predicted_per_second'])
elif key=='code':out.append(d['decode'][1]['timings']['predicted_per_second'])
return out
for name,key in [('Prefill, 4.196 Eingabetokens (tok/s)','pp'),('Deutsche Erklärung (tok/s)','de'),('Python-Code (tok/s)','code')]:
a,b=statistics.mean(values('pure',key)),statistics.mean(values('byteshape',key))
print(f'| {name} | {a:.1f} | {b:.1f} | {(b/a-1)*100:+.1f}% |')
print('\n## Vollständig getestete Konfigurationen\n')
print('Kontext ist Eingabe plus Ausgabe. Die Geschwindigkeit in dieser Tabelle gehört jeweils zur angegebenen tatsächlichen Eingabelänge; Zeilen unterschiedlicher Länge sind kein isolierter Quantisierungsvergleich. VRAM enthält auch residente Dienste (TTS auf der 3060).\n')
print('| Fall | Kontext | Eingabe | Prefill tok/s | Ausgabe tok/s | 5080 MiB | 3060 MiB | Recall |\n|---|---:|---:|---:|---:|---:|---:|---|')
for label,d in cases.items():
if not d.get('prefill'):continue
p=d['prefill'][-1]['response'];t=p['timings'];r=p.get('recall',{})
print(f'| {label} | {d["case"]["ctx"]} | {t["prompt_n"]} | {t["prompt_per_second"]:.1f} | {t["predicted_per_second"]:.1f} | {d["peak"].get("NVIDIA GeForce RTX 5080",0)} | {d["peak"].get("NVIDIA GeForce RTX 3060",0)} | {sum(r.values())}/{len(r)} |')
print('\nDie Kontextpiloten ohne lange Eingabe sind hier bewusst nicht als validierte Konfigurationen aufgeführt. Einzelne synthetische Recall-Aufgaben belegen keine allgemeine Langkontext-Intelligenz.')
@@ -0,0 +1,13 @@
{
"label": "byteshape-dual-262144-80-20",
"model": "byteshape",
"ctx": 262144,
"single": false,
"split": "80,20",
"ubatch": 128,
"mtp": 2,
"prompts": [
49152,
261120
]
}
@@ -0,0 +1,13 @@
{
"label": "byteshape-dual-262144-86-14-validated",
"model": "byteshape",
"ctx": 262144,
"single": false,
"split": "86,14",
"ubatch": 128,
"mtp": 2,
"prompts": [
49152,
261120
]
}
@@ -0,0 +1,13 @@
{
"label": "byteshape-dual-262144-86-14",
"model": "byteshape",
"ctx": 262144,
"single": false,
"split": "86,14",
"ubatch": 128,
"mtp": 2,
"prompts": [
49152,
261120
]
}
@@ -0,0 +1,13 @@
{
"label": "byteshape-dual-262144-88-12",
"model": "byteshape",
"ctx": 262144,
"single": false,
"split": "88,12",
"ubatch": 128,
"mtp": 2,
"prompts": [
49152,
261120
]
}
@@ -0,0 +1,10 @@
{
"label": "byteshape-single-110592-ub128",
"model": "byteshape",
"ctx": 110592,
"single": true,
"ubatch": 128,
"prompts": [
109568
]
}
@@ -0,0 +1,11 @@
{
"label": "byteshape-single-114688",
"model": "byteshape",
"ctx": 114688,
"single": true,
"prompts": [
49152,
113664
],
"quality_followup": true
}
@@ -0,0 +1,9 @@
{
"label": "byteshape-single-32768-repeat",
"model": "byteshape",
"ctx": 32768,
"single": true,
"prompts": [
4096
]
}
@@ -0,0 +1,11 @@
{
"label": "byteshape-single-32768",
"model": "byteshape",
"ctx": 32768,
"single": true,
"quality": true,
"prompts": [
4096,
24576
]
}
@@ -0,0 +1,14 @@
{
"label": "byteshape-single-ub128-validated-104448",
"model": "byteshape",
"ctx": 104448,
"single": true,
"ubatch": 128,
"capacity_search": true,
"quality_followup": true,
"load_only": false,
"prompts": [
49152,
103424
]
}
+165
View File
@@ -0,0 +1,165 @@
#!/usr/bin/env python3
"""Bounded, isolated Qwen quantization benchmark. Supervisor restores production."""
import json, pathlib, subprocess, sys, time, urllib.request, threading, signal
ROOT = pathlib.Path('/data/benchmarks/byteshape-20260920')
NAME = 'mike-ai-byteshape-test'
BASE = 'http://127.0.0.1:5005'
GPU0 = 'GPU-8ad38c6c-5a01-9d8e-1dfa-ed662ad78fbe'
GPU1 = 'GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b'
IMAGE = 'sha256:5e3c12c145b8045e5731b44b6b97033f24b327ae3d4a3fa85ecdd159cc844907'
MODELS = {'mix':'qwen3.8-27b-iq4-mix/Qwen3.8-27B-IQ4-MIX.gguf', 'pure':'qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf', 'byteshape':'byteshape-qwen38-gpu5/model.gguf'}
def cmd(*args, check=True, timeout=90):
r = subprocess.run(args, capture_output=True, text=True, timeout=timeout)
if check and r.returncode: raise RuntimeError(str(args[:3])+': '+r.stderr[-2000:])
return r.stdout
def api(path, data=None, timeout=900):
req = urllib.request.Request(BASE+path, data=None if data is None else json.dumps(data).encode(), headers={'Content-Type':'application/json'})
with urllib.request.urlopen(req, timeout=timeout) as r: return json.load(r)
def save(path, data):
path.write_text(json.dumps(data, indent=2, ensure_ascii=False)+'\n')
def gpu():
rows = cmd('nvidia-smi','--query-gpu=name,memory.used,memory.total,temperature.gpu,utilization.gpu','--format=csv,noheader,nounits',timeout=15)
return [dict(zip(['name','used','total','temp','util'], [v.strip() for v in row.split(',')])) for row in rows.splitlines()]
def health_check():
rows=gpu()
if any(int(x['temp']) >= 85 for x in rows): raise RuntimeError('GPU temperature limit')
mem = dict((a.split(':')[0],int(a.split()[1])) for a in pathlib.Path('/proc/meminfo').read_text().splitlines())
if mem['MemAvailable'] < 3*1024*1024: raise RuntimeError('Host RAM reserve below 3 GiB')
return rows
def chat(prompt, max_tokens=512, effort='none', seed=42, tools=None):
health_check()
p={'model':'benchmark','messages':[{'role':'user','content':prompt}], 'max_tokens':max_tokens,'temperature':1.0,'top_p':0.95,'top_k':20,'min_p':0.0,'seed':seed,'reasoning_effort':effort,'cache_prompt':False}
if tools: p.update(tools=tools,tool_choice='auto')
start=time.monotonic(); r=api('/v1/chat/completions',p); r['wall_seconds']=time.monotonic()-start
health_check()
return r
def prefill(n, seed):
# Exact token-array slicing avoids accidentally exceeding the intended input size.
text='\n'.join(f'Record {i:06d}: cobalt lantern maple orbit quartz river silver tango.' for i in range(max(100,n//10)))
tokens=api('/tokenize',{'content':text,'add_special':False})['tokens'][:n]
# Insert independently locatable facts near start/middle/end of the context.
text=api('/detokenize',{'tokens':tokens})['content']
for pos, fact in reversed([(len(text)//8,'NEEDLE_ALPHA=RAVEN-417'),(len(text)//2,'NEEDLE_BETA=CEDAR-928'),(len(text)*7//8,'NEEDLE_GAMMA=ORBIT-563')]):
text=text[:pos]+'\n'+fact+'\n'+text[pos:]
text+='\nReturn a JSON object with alpha, beta and gamma containing the three exact NEEDLE values. Then explain in German how to verify these records without inventing evidence, in at least 300 words.'
r=chat(text,512,seed=seed)
content=r['choices'][0]['message'].get('content','')
r['recall']={x:x in content for x in ['RAVEN-417','CEDAR-928','ORBIT-563']}
return r
def run_case(case):
label=case['label']; out=ROOT/label; out.mkdir(exist_ok=True)
if (out/'result.json').exists(): raise RuntimeError('Refusing to overwrite completed case '+label)
save(out/'config.json',case)
print('START',label,flush=True)
single=case.get('single',False)
args=['--model','/models/'+MODELS[case['model']], '--alias','benchmark','--ctx-size',str(case['ctx']), '--flash-attn','on','--cache-type-k','q4_0','--cache-type-v','q4_0','--cache-ram','0','--threads','6','--threads-batch','6','--batch-size','2048','--ubatch-size',str(case.get('ubatch',512)), '--parallel','1','--kv-unified','--jinja','--reasoning','auto','--reasoning-preserve','--host','127.0.0.1','--port','5005','--metrics','--fit','off','--n-gpu-layers','all','--load-mode','none','--no-ui','--temperature','1.0','--top-p','0.95','--top-k','20','--device','CUDA0' if single else 'CUDA0,CUDA1','--main-gpu','0','--split-mode','layer','--tensor-split',case.get('split','1,0'),'--spec-type','draft-mtp','--spec-draft-n-max',str(case.get('mtp',3)),'--spec-draft-type-k','f16','--spec-draft-type-v','f16','--spec-draft-p-min','0.05','--verbosity','3']
cmd('docker','run','-d','--name',NAME,'--gpus','all','--network','host','--read-only','--tmpfs','/tmp:rw,nosuid,nodev,size=256m','--security-opt','no-new-privileges:true','--cap-drop','ALL','--pids-limit','512','--ulimit','core=0','--memory','26g','--memory-swap','26g','--shm-size','1g','--log-opt','max-size=32m','--log-opt','max-file=1','-e','NVIDIA_VISIBLE_DEVICES='+GPU0+','+GPU1,'-e','NVIDIA_DRIVER_CAPABILITIES=compute,utility','-v','/data/models:/models:ro',IMAGE,*args)
stop=threading.Event(); samples=[]
def monitor():
while not stop.wait(2):
try:
rows=health_check()
samples.append({'time':time.time(),'gpus':rows})
except RuntimeError as e:
samples.append({'error':str(e),'aborted':True})
cmd('docker','stop','-t','10',NAME,check=False)
return
except Exception as e: samples.append({'error':str(e)})
thread=threading.Thread(target=monitor,daemon=True); thread.start()
result={'case':case,'started':time.time()}
try:
for _ in range(150):
try:
if api('/health',timeout=3).get('status')=='ok': break
except Exception: pass
if cmd('docker','inspect',NAME,'--format','{{.State.Running}}').strip()!='true': raise RuntimeError('Test container exited during load')
time.sleep(2)
else: raise RuntimeError('Startup exceeded 300s')
result['idle_gpu']=health_check(); result['props']=api('/props'); result['slots']=api('/slots')
save(out/'loaded.json',result)
# Added after the 88:12 trial: model loading alone can succeed while
# the first real attention graph still needs more CUDA workspace.
minimum=case.get('minimum_headroom_mib',512)
used_devices=['5080'] if single else ['5080','3060']
for g in result['idle_gpu']:
if any(device in g['name'] for device in used_devices):
free=int(g['total'])-int(g['used'])
if free<minimum:
raise RuntimeError(f"Insufficient loaded VRAM reserve on {g['name']}: {free} < {minimum} MiB; refusing inference")
result['smoke']=chat('Antworte nur mit OK.',8)
if case.get('quality'):
tasks=json.loads(pathlib.Path('/opt/mike-ai/stack/dev/QWEN38-FINAL-ACCEPTANCE-v1.json').read_text())
result['quality']=[]
for task in tasks:
ans=chat(task['prompt'],task['max_tokens'],'medium')
result['quality'].append({'id':task['id'],'response':ans})
save(out/'partial.json',result); print(label,task['id'],round(ans['wall_seconds'],1),flush=True)
tool={'type':'function','function':{'name':'read_server_status','description':'Read-only server status lookup','parameters':{'type':'object','properties':{'server':{'type':'string'}},'required':['server'],'additionalProperties':False}}}
result['tool']=chat('Read the current status of server alpha. Use the provided tool exactly once and do not invent its result.',512,tools=[tool])
if case.get('quality_followup'):
tasks=json.loads(pathlib.Path('/opt/mike-ai/stack/dev/QWEN38-FINAL-ACCEPTANCE-v1.json').read_text())
result['quality_followup']=[]
for task in tasks:
if task['id'] not in ['i3_code_debugging','i4_capacity_planning','i6_state_vs_configuration']: continue
ans=chat(task['prompt'],8192,'medium',seed=43)
result['quality_followup'].append({'id':task['id'],'seed':43,'budget':8192,'response':ans})
save(out/'partial.json',result); print(label,'followup',task['id'],round(ans['wall_seconds'],1),flush=True)
if not case.get('load_only'):
result['decode']=[]
prompts=['Erkläre ausführlich auf Deutsch, wie ein Reverse Proxy funktioniert, welche Fehler bei Container-IP-Wechseln auftreten können und wie man sie anhand von Logs eingrenzt. Schreibe mindestens 600 Wörter.', 'Write a Python implementation of an asynchronous first_success function: start all awaitables concurrently, return the first successful result, cancel and await remaining tasks, collect exceptions if all fail. Include an explanation and usage example.']
for i,p in enumerate(prompts): result['decode'].append(chat(p,768,seed=42+i))
result['prefill']=[]
for n in case.get('prompts',[4096,16384]):
r=prefill(n,42); result['prefill'].append({'target':n,'response':r}); save(out/'partial.json',result)
print(label,'prefill',n,r.get('timings'),flush=True)
result['finished']=time.time()
finally:
stop.set(); thread.join(5)
r=subprocess.run(['docker','logs',NAME],capture_output=True,text=True,timeout=30)
(out/'server.log').write_text(r.stdout+r.stderr)
save(out/'gpu.json',samples); save(out/'result.json',result)
cmd('docker','rm','-f',NAME,check=False)
print('DONE',label,flush=True)
return result
def capacity_case(case):
"""Bounded growth from a previously working context, with 768 MiB reserve.
0.04 MiB/token exceeds the measured 512-ubatch steady-state slope.
A failed 114688/512 ByteShape startup revealed additional transient MTP
buffers, so reserve is deliberately larger than steady-state extrapolation.
Smaller ubatches start at an already working context, not a guessed OOM edge.
The final context is tested with an actual almost-full prompt.
"""
context=case['ctx']
previous=None
for attempt in range(8):
pilot={**case,'ctx':context,'label':case['label']+'-pilot-'+str(context),'load_only':True,'quality':False,'quality_followup':False}
result=run_case(pilot)
rows=[g for g in result['idle_gpu'] if '5080' in g['name']]
samples=json.loads((ROOT/pilot['label']/'gpu.json').read_text())
peak=max([int(rows[0]['used'])+32]+[int(g['used']) for s in samples for g in s.get('gpus',[]) if '5080' in g['name']])
free=int(rows[0]['total'])-peak
if free<768:
if previous is None: raise RuntimeError('Initial capacity pilot has insufficient reserve')
context=previous
break
growth=min(16384,int((free-768)/0.04)//1024*1024)
if growth<1024 or attempt==7 or context>=262144: break
previous=context
context=min(262144,context+growth)
final={**case,'ctx':context,'label':case['label']+'-validated-'+str(context),'load_only':False,'prompts':[49152,context-1024]}
return run_case(final)
if __name__=='__main__':
for case in json.loads(pathlib.Path(sys.argv[1]).read_text()):
if case.get('capacity_search'): capacity_case(case)
else: run_case(case)
@@ -0,0 +1,19 @@
#!/usr/bin/env python3
"""Summarize measured timings; never substitute configured context for input length."""
import json,pathlib,sys
root=pathlib.Path(sys.argv[1])
rows=[]
for path in sorted(root.glob('*/result.json')):
r=json.loads(path.read_text()); samples=json.loads((path.parent/'gpu.json').read_text())
peak={}
for s in samples:
for g in s.get('gpus',[]):
p=peak.setdefault(g['name'],{'MiB':0,'C':0})
p['MiB']=max(p['MiB'],int(g['used']));p['C']=max(p['C'],int(g['temp']))
row={'case':r['case'],'finished':bool(r.get('finished')),'peak':peak,'decode':[], 'prefill':[]}
for d in r.get('decode',[]): row['decode'].append(d.get('timings'))
for p in r.get('prefill',[]):
a=p['response'];row['prefill'].append({'target':p['target'],'timings':a.get('timings'),'recall':a.get('recall'),'wall_seconds':a.get('wall_seconds')})
row['quality']=[{'id':q['id'],'finish':q['response']['choices'][0].get('finish_reason'),'visible_chars':len(q['response']['choices'][0]['message'].get('content','')),'tokens':q['response'].get('usage',{})} for q in r.get('quality',[])]
rows.append(row)
print(json.dumps(rows,indent=2,ensure_ascii=False))
@@ -0,0 +1,40 @@
#!/usr/bin/env python3
"""Stop only existing router/controller/model; always restore the same containers."""
import json, pathlib, subprocess, sys, time, signal
ROOT=pathlib.Path('/data/benchmarks/byteshape-20260920')
NAMES=['mike-ai-router','mike-ai-profile-controller','mike-ai-llama-medium']
def run(*args,check=True,timeout=90):
return subprocess.run(args,capture_output=True,text=True,check=check,timeout=timeout)
def stop_signal(*_): raise RuntimeError('Supervisor interrupted')
signal.signal(signal.SIGTERM,stop_signal); signal.signal(signal.SIGINT,stop_signal)
# Refuse if the known production state has changed, or if requests are active.
for name in NAMES:
assert run('docker','inspect',name,'--format','{{.State.Running}}').stdout.strip()=='true',name
slots=json.loads(run('docker','exec',NAMES[-1],'curl','-fsS','http://127.0.0.1:8080/slots').stdout)
assert not any(s['is_processing'] for s in slots),'Production request active'
assert not run('docker','ps','-q','--filter','name=^mike-ai-byteshape-test$').stdout.strip(),'Existing experiment'
child=None
try:
run('docker','stop','-t','30',*NAMES[:2])
# Drain requests already handed to the model, before unloading it.
for _ in range(120):
slots=json.loads(run('docker','exec',NAMES[-1],'curl','-fsS','http://127.0.0.1:8080/slots').stdout)
if not any(s['is_processing'] for s in slots): break
time.sleep(2)
else: raise RuntimeError('Model did not drain')
run('docker','stop','-t','30',NAMES[-1])
child=subprocess.Popen(['python3',str(ROOT/'run.py'),sys.argv[1]])
code=child.wait(timeout=4200)
if code: raise RuntimeError('Benchmark failed: '+str(code))
finally:
if child is not None and child.poll() is None:
child.terminate()
try: child.wait(timeout=20)
except subprocess.TimeoutExpired: child.kill(); child.wait(timeout=10)
run('docker','rm','-f','mike-ai-byteshape-test',check=False)
run('docker','start',NAMES[-1])
for _ in range(150):
if run('docker','inspect',NAMES[-1],'--format','{{.State.Health.Status}}').stdout.strip()=='healthy': break
time.sleep(2)
run('docker','start',NAMES[1],NAMES[0])
print('RESTORED existing medium/controller/router',flush=True)
@@ -0,0 +1,29 @@
#!/usr/bin/env python3
"""Read-only restoration verification plus a two-token model smoke request."""
import json,pathlib,re,subprocess,time
ROOT=pathlib.Path('/data/benchmarks/byteshape-20260920')
def run(*args):return subprocess.check_output(args,text=True,timeout=30)
report={'checked_at':time.time(),'uptime':run('uptime').strip(),'containers':{}}
for name in ['mike-ai-llama-medium','mike-ai-router','mike-ai-profile-controller','mike-ai-wireguard-gateway','mike-ai-qwen3-tts']:
d=json.loads(run('docker','inspect',name))[0]
report['containers'][name]={'running':d['State']['Running'],'health':d['State'].get('Health',{}).get('Status'),'image':d['Image'],'started':d['State']['StartedAt']}
assert d['State']['Running'],name
if name in ['mike-ai-llama-medium','mike-ai-router','mike-ai-profile-controller','mike-ai-wireguard-gateway']:
assert d['State'].get('Health',{}).get('Status')=='healthy',name
assert report['containers']['mike-ai-llama-medium']['image']=='sha256:5e3c12c145b8045e5731b44b6b97033f24b327ae3d4a3fa85ecdd159cc844907'
assert not run('docker','ps','-q','--filter','name=^mike-ai-byteshape-test$').strip()
probe='import urllib.request,json; print(json.dumps({p:json.load(urllib.request.urlopen("http://127.0.0.1:8081"+p,timeout=20)) for p in ["/health","/ready"]}))'
report['router']=json.loads(run('docker','exec','mike-ai-router','python','-c',probe))
payload={'model':'qwen-medium','messages':[{'role':'user','content':'Antworte ausschließlich mit OK.'}],'reasoning_effort':'none','max_tokens':8,'temperature':0}
r=json.loads(run('docker','exec','mike-ai-llama-medium','curl','-fsS','--max-time','20','-H','Content-Type: application/json','--data',json.dumps(payload),'http://127.0.0.1:8080/v1/chat/completions'))
report['smoke']={'content':r['choices'][0]['message'].get('content',''),'usage':r.get('usage')}
assert report['smoke']['content'].strip()=='OK',report['smoke']
started=min(json.loads(p.read_text())['started'] for p in ROOT.glob('*/result.json'))
journal=run('journalctl','-k','--since','@'+str(int(started)-60),'--no-pager')
pattern=re.compile(r'NVRM.*Xid|oom-kill|Out of memory: Killed process|Kernel panic|GPU has fallen off',re.I)
report['kernel_errors']=[line for line in journal.splitlines() if pattern.search(line)]
report['gpu']=run('nvidia-smi','--query-gpu=name,memory.used,memory.total,temperature.gpu','--format=csv').strip()
report['disk']=run('df','-h','/','/data').strip()
(ROOT/'restore-verification.json').write_text(json.dumps(report,indent=2)+'\n')
print(json.dumps(report,indent=2))
assert not report['kernel_errors'],'Kernel/GPU errors require review'