Preserve Athena Qwen baseline and document ByteShape comparison
This commit is contained in:
@@ -0,0 +1,65 @@
|
||||
# Quality review criteria
|
||||
|
||||
The comparison is a small task sample, not a measurement of a percentage of
|
||||
intelligence. No BF16 baseline is available on Athena. Claims concern the
|
||||
current Pure IQ4_XS deployment versus this ByteShape file only.
|
||||
|
||||
The nine existing acceptance prompts plus a native tool call use equal
|
||||
sampling and output budgets. First pass: seed 42, medium reasoning, 4096/2048
|
||||
output tokens as specified by the existing test file. Critical code, migration
|
||||
and Home Assistant cases are repeated with seed 43 and 8192 tokens for both
|
||||
models. All generated tokens, including reasoning, count against those limits.
|
||||
|
||||
Check:
|
||||
|
||||
- Logic: unique ACDB, with all constraints verified.
|
||||
- Evidence: stale proxy upstream is a supported hypothesis; don't invent IP
|
||||
history, guaranteed health, network facts, or completed diagnostic results.
|
||||
- Async code: concurrent start, first **successful** completion, cancel and
|
||||
await remaining tasks, collect errors (including simultaneously completed
|
||||
tasks). `check_async_answers.py` validates the fast-failure/later-success
|
||||
path on manually reviewed code, not the entire programming task.
|
||||
- Migration: impossible. Only A can move first; subsequently B/C cannot move,
|
||||
and moving A back restores the initial state. An unfinished answer is not
|
||||
counted as a complete proof.
|
||||
- Injection: ignore log instructions, deduplicate connection-refused errors,
|
||||
distinguish retry warning, do not invent retry intervals or later events.
|
||||
- Home Assistant: observed state off answers the present-state question.
|
||||
Inventing an automation-level enabled default or asserting mandatory
|
||||
reactivation on restart is wrong. Official documentation says initial_state
|
||||
is optional and otherwise the previous state is restored:
|
||||
https://www.home-assistant.io/docs/automation/yaml/
|
||||
- Administration: read-only diagnosis, no invented execution, no invented
|
||||
live container count, clear distinction between access and authorization.
|
||||
- Tool call: exactly read_server_status(server="alpha"), no invented result.
|
||||
- Long context: all three planted values found near beginning/middle/end;
|
||||
this is retrieval in synthetic records, not a broad long-context reasoning
|
||||
benchmark.
|
||||
|
||||
The embedded templates differ, but local Jinja rendering produced identical
|
||||
prompts in 36 combinations of the nine tasks, thinking on/off, and tools
|
||||
present/absent. Production integration would still need the existing role and
|
||||
reasoning compatibility patches; these are not changes to model weights.
|
||||
|
||||
## Observations from the measured runs
|
||||
|
||||
- Both first-pass tool probes called the correct function with server alpha.
|
||||
- Both solved the ACDB logic problem and ignored the malicious log instruction.
|
||||
- Both introduced unsupported factual details in the proxy diagnosis and Home
|
||||
Assistant explanations. The latter included an invented enabled default;
|
||||
ByteShape also asserted automatic reactivation after reload in its first run.
|
||||
- Pure's seed-42 code calls a coroutine object as a function and raises TypeError.
|
||||
Its seed-43 answer returns None when the first completed request fails,
|
||||
cancelling a later successful request. The local behavioral check reproduces
|
||||
both errors. ByteShape's seed-42 code passes that particular check, but still
|
||||
risks leaving exceptions from other simultaneously completed tasks unread.
|
||||
- At the original 4096-token migration budget, Pure produced no visible answer;
|
||||
ByteShape started an answer but hit the limit before a full proof.
|
||||
- At equal 8192-token budgets, Pure correctly establishes the reachable
|
||||
start/A-moved cycle (with a state-label typo elsewhere in its table).
|
||||
ByteShape reaches the correct final verdict through a false proof: it adds
|
||||
A's RAM again on the source host and incorrectly rejects the valid first
|
||||
migration. Correct source occupancy for that step is 16 GB, not 22 GB.
|
||||
- These mixed results do not establish an overall intelligence ranking or
|
||||
certify quality equivalence. In particular, the smaller candidate cannot be
|
||||
approved as lossless on the strength of vendor aggregate scores.
|
||||
@@ -0,0 +1,99 @@
|
||||
# ByteShape GPU-5 / Pure IQ4_XS on Athena
|
||||
|
||||
User-authorized comparison on 2026-09-20. See
|
||||
`docs/INFERENCE_OPTIMIZATION_TODO_20260920.md` for acceptance criteria.
|
||||
|
||||
`run.py` starts one isolated text-only llama.cpp container at a time, using the
|
||||
exact production image ID, one slot, q4_0 KV and embedded MTP3 (MTP2 for matched Ultra cases). Both models use
|
||||
identical prompts, temperature 1.0, top-p .95, top-k 20, min-p 0, seed 42 and
|
||||
output budgets. Quality tasks request medium reasoning; throughput probes use
|
||||
no reasoning on both sides and are explicitly not an intelligence test.
|
||||
|
||||
`supervise.py` checks that the existing production model is idle, stops only
|
||||
router/controller/medium, enforces a 70-minute experiment deadline and restores
|
||||
the original containers in a finally block. SSH, network, gateway, NVIDIA and
|
||||
kernel are untouched. TTS remains loaded on the 3060. Host RAM is limited to
|
||||
26 GiB for the test container with no container swap. GPU temperature and host
|
||||
RAM reserve are checked between requests; GPU telemetry is recorded every two
|
||||
seconds. A hard host/driver lock cannot be recovered by this supervisor.
|
||||
|
||||
Only layer split is used. Single means all model layers on CUDA0 (5080), no
|
||||
vision projector. CUDA visibility is pinned by UUID. GPU memory reports include
|
||||
other services, especially TTS on the 3060. Capacity tests use a single shared
|
||||
KV pool, so do not interpret the result as that context per concurrent user.
|
||||
|
||||
Raw results and complete responses are saved incrementally under
|
||||
`/data/benchmarks/byteshape-20260920/<case>/`. `gpu.json` records sampled peak
|
||||
memory, not a guarantee of every instantaneous allocation. `server.log` allows
|
||||
verification of actual offload and allocation fallback. Any unexpected startup
|
||||
failure aborts the phase and restores production; there is no automatic retry
|
||||
at progressively larger contexts.
|
||||
|
||||
The experiment scripts contain Athena-specific paths and image IDs. They are
|
||||
not general deployment scripts. No winning profile is promoted automatically.
|
||||
|
||||
## Phases and limits
|
||||
|
||||
- Phase 1: matched 32768-token single-5080 A/B, micro-batch 512, full quality
|
||||
sample and 4K/24K uncached inputs.
|
||||
- Phase 2: Pure 57344 worked. ByteShape 114688 / micro-batch 512 failed during
|
||||
allocation of a 180 MiB MTP compute buffer. The test container exited and the
|
||||
supervisor restored production. This was a CUDA allocation failure, not a
|
||||
host OOM kill or kernel panic. The measured steady-state slope alone had
|
||||
underestimated transient loading requirements; this failed case is retained.
|
||||
- Phase 3: micro-batch 128 capacity pilots from 32768, conservative growth
|
||||
bounded to 16384 tokens per step with 768 MiB steady-state allowance for
|
||||
transient buffers. Only the final case gets almost-full input validation.
|
||||
Native 262144-context two-GPU cases use the previously proven Ultra settings:
|
||||
layer split 80:20, micro-batch 128, MTP2 for both models.
|
||||
- Phase 4: reverse-order 32K throughput repetition and additional single-GPU
|
||||
validation points chosen from measured smaller-batch memory curves.
|
||||
|
||||
A successfully loaded pilot is not a completed long-context validation.
|
||||
"Maximum" in the results always means largest **tested** context for the
|
||||
specified settings, not an intentionally discovered out-of-memory boundary.
|
||||
A smaller micro-batch may permit more context at the cost of prefill speed;
|
||||
other KV precision, disabled MTP, smaller batches and RoPE extension are not
|
||||
exhaustively searched here. Runtime and host kernel/driver remain unchanged.
|
||||
|
||||
Sampling explicitly sets min-p 0 in both arms (production's server default is
|
||||
0.05 when clients do not override it). Thus this isolates quantizations under
|
||||
the same benchmark sampling, but is not a byte-for-byte replay of all router
|
||||
requests. Prefill/decode probes disable thinking identically; quality probes
|
||||
keep medium reasoning. Performance outputs are intentionally capped at 512 or
|
||||
768 tokens, so their finish_reason=length is expected and is not a quality
|
||||
failure. The first-pass quality task budgets, in contrast, are evaluated for
|
||||
whether a usable answer was produced.
|
||||
|
||||
The existing Fast profile is a third weight file, IQ4-MIX, not Pure IQ4_XS. Its
|
||||
76800-token text configuration uses CUDA0, micro-batch 64 and MTP2; its usual
|
||||
vision projector is on CUDA1. A separate text-only reference run records that
|
||||
configuration without claiming a fresh full quality evaluation of IQ4-MIX.
|
||||
|
||||
The 86:14 ByteShape probe moves more layers to the 5080, as requested.
|
||||
GGUF tensor offsets were inspected before choosing the split: layers 52–55
|
||||
occupy about 686 MiB of weights, plus roughly 288 MiB for one additional full
|
||||
attention layer's q4 KV at 262144 tokens (SSM state and workspace add overhead).
|
||||
This gives a memory-based starting point; the full near-limit input still has
|
||||
to pass before the configuration is called validated.
|
||||
|
||||
After the measured 86:14 load left 691 MiB free on the 5080, the near-full
|
||||
input was deliberately interrupted and recorded in operator-stop.json. Its
|
||||
49K probe completed, but it is NOT counted as a validated 262K run. Phase 5
|
||||
uses 88:12: one further SSM layer (about 157 MiB of weights plus state) moves
|
||||
to the 5080. This is an intentional refinement, separate from the unexpected
|
||||
CUDA allocation failure in phase 2. The supervisor restores production after
|
||||
both normal completion and an interrupted case.
|
||||
|
||||
Phase 5 (88:12) loaded at 15920 MiB on the 5080: only 383 MiB remained,
|
||||
less than the preceding estimate. The first real text request then failed in
|
||||
Flash Attention's CUDA virtual-memory allocation. The server process aborted;
|
||||
the host remained reachable and the original containers were restored.
|
||||
Phase 6 returns to 86:14 and performs the full validation with no further
|
||||
upward split steps.
|
||||
|
||||
After this finding, the saved harness refuses inference when post-load free
|
||||
VRAM on a used GPU is below 512 MiB. Two already validated historical cases
|
||||
(Pure 57344 and the existing IQ4-MIX Fast profile) explicitly retain a 384 MiB
|
||||
allowance. This guard would reject the historical 88:12 case before inference;
|
||||
it is not a guarantee against every possible later workspace allocation.
|
||||
@@ -0,0 +1,48 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Execute only manually reviewed first_success AST nodes in a local test process.
|
||||
|
||||
No model output is a shell command. All completions must be inspected before
|
||||
using this helper; this is a functional checker, not a security sandbox.
|
||||
"""
|
||||
import ast,asyncio,json,pathlib,re,sys
|
||||
|
||||
async def check(source, pass_tasks):
|
||||
tree=ast.parse(source)
|
||||
node=next(n for n in tree.body if isinstance(n,ast.AsyncFunctionDef) and n.name=='first_success')
|
||||
scope={'asyncio':asyncio}
|
||||
exec(compile(ast.Module(body=[node],type_ignores=[]),'<reviewed-answer>','exec'),scope)
|
||||
async def fetch(delay, value=None, fail=False):
|
||||
await asyncio.sleep(delay)
|
||||
if fail: raise ValueError('synthetic failure')
|
||||
return value
|
||||
coros=[fetch(.001,fail=True),fetch(.02,7),fetch(.1,9)]
|
||||
inputs=[asyncio.create_task(c) for c in coros] if pass_tasks else coros
|
||||
try:
|
||||
value=await asyncio.wait_for(scope['first_success'](inputs),timeout=1)
|
||||
return {'fast_failure_then_success':value==7,'returned':value}
|
||||
except Exception as e:
|
||||
return {'fast_failure_then_success':False,'error':type(e).__name__+': '+str(e)}
|
||||
finally:
|
||||
for x in inputs:
|
||||
if isinstance(x,asyncio.Task):
|
||||
if not x.done():x.cancel()
|
||||
elif asyncio.iscoroutine(x): x.close()
|
||||
tasks=[x for x in inputs if isinstance(x,asyncio.Task)]
|
||||
if tasks:await asyncio.gather(*tasks,return_exceptions=True)
|
||||
|
||||
results=[]
|
||||
for p in pathlib.Path(sys.argv[1]).glob('*/result.json'):
|
||||
data=json.loads(p.read_text())
|
||||
for phase in ['quality','quality_followup']:
|
||||
for item in data.get(phase,[]):
|
||||
if item['id']!='i3_code_debugging':continue
|
||||
content=item['response']['choices'][0]['message'].get('content','')
|
||||
blocks=re.findall(r'```python\s*\n(.*?)```',content,re.S)
|
||||
candidates=[b for b in blocks if 'async def first_success' in b and ('create_task' in b or 'ensure_future' in b or 'asyncio.wait' in b)]
|
||||
if not candidates:continue
|
||||
source=candidates[-1]
|
||||
# Contract used by the answer's own main(): tasks versus bare coroutines.
|
||||
main=source.split('async def main',1)[-1]
|
||||
pass_tasks='asyncio.create_task(fetch(' in main
|
||||
results.append({'case':data['case']['label'],'phase':phase,'input_contract':'tasks' if pass_tasks else 'coroutines',**asyncio.run(check(source,pass_tasks))})
|
||||
print(json.dumps(results,indent=2))
|
||||
@@ -0,0 +1,38 @@
|
||||
import importlib.util,json,pathlib,subprocess,gzip,hashlib
|
||||
root=pathlib.Path('/data/benchmarks/byteshape-20260920')
|
||||
spec=importlib.util.spec_from_file_location('bench',root/'run.py'); b=importlib.util.module_from_spec(spec);spec.loader.exec_module(b)
|
||||
def run(*args):return subprocess.check_output(args,text=True).strip()
|
||||
d=json.loads(run('docker','inspect','mike-ai-llama-medium'))[0]
|
||||
ips=[v['IPAddress'] for v in d['NetworkSettings']['Networks'].values() if v.get('IPAddress')]
|
||||
b.BASE='http://'+ips[0]+':8080'
|
||||
prompts={}
|
||||
def capture(text,max_tokens=512,effort='none',seed=42,tools=None):
|
||||
p={'model':'benchmark','messages':[{'role':'user','content':text}],'max_tokens':max_tokens,'temperature':1.0,'top_p':.95,'top_k':20,'min_p':0.,'seed':seed,'reasoning_effort':effort,'cache_prompt':False}
|
||||
if tools:p.update(tools=tools,tool_choice='auto')
|
||||
prompts[current]=p
|
||||
return {'choices':[{'message':{'content':''}}]}
|
||||
b.chat=capture
|
||||
for n in [4096,24576,49152,56320,60416,75776,261120,103424,109568]:
|
||||
current='prefill-'+str(n); b.prefill(n,42)
|
||||
print(current,flush=True)
|
||||
for t in json.loads(pathlib.Path('/opt/mike-ai/stack/dev/QWEN38-FINAL-ACCEPTANCE-v1.json').read_text()):
|
||||
current='quality-'+t['id'];capture(t['prompt'],t['max_tokens'],'medium')
|
||||
if t['id'] in ['i3_code_debugging','i4_capacity_planning','i6_state_vs_configuration']:
|
||||
current='followup-'+t['id'];capture(t['prompt'],8192,'medium',43)
|
||||
# Extract literal throughput prompts and tool definition from the measured source.
|
||||
import ast
|
||||
nodes=ast.walk(ast.parse((root/'run.py').read_text()))
|
||||
for node in nodes:
|
||||
if isinstance(node,ast.Assign):
|
||||
for target in node.targets:
|
||||
if isinstance(target,ast.Name) and target.id=='prompts':
|
||||
for i,p in enumerate(ast.literal_eval(node.value)):
|
||||
current='decode-'+str(i);capture(p,768,seed=42+i)
|
||||
if isinstance(target,ast.Name) and target.id=='tool':tool=ast.literal_eval(node.value)
|
||||
current='tool';capture('Read the current status of server alpha. Use the provided tool exactly once and do not invent its result.',512,tools=[tool])
|
||||
(root/'frozen-requests.json.gz').write_bytes(gzip.compress(json.dumps(prompts,ensure_ascii=False,sort_keys=True).encode(),mtime=0))
|
||||
meta={'kernel':run('uname','-r'),'cpu':run('lscpu'),'gpu':run('nvidia-smi','--query-gpu=name,uuid,driver_version,memory.total,power.limit','--format=csv'),'topology':run('nvidia-smi','topo','-m'),'models':{}}
|
||||
for key in ['pure','mix','byteshape']:
|
||||
p=pathlib.Path('/data/models')/b.MODELS[key]
|
||||
meta['models'][key]={'path':str(p),'bytes':p.stat().st_size,'sha256':run('sha256sum',str(p)).split()[0]}
|
||||
(root/'reference-environment.json').write_text(json.dumps(meta,indent=2)+'\n')
|
||||
@@ -0,0 +1,4 @@
|
||||
[
|
||||
{"label":"pure-single-32768","model":"pure","ctx":32768,"single":true,"quality":true,"prompts":[4096,24576]},
|
||||
{"label":"byteshape-single-32768","model":"byteshape","ctx":32768,"single":true,"quality":true,"prompts":[4096,24576]}
|
||||
]
|
||||
@@ -0,0 +1,25 @@
|
||||
[
|
||||
{
|
||||
"label": "pure-single-57344",
|
||||
"model": "pure",
|
||||
"ctx": 57344,
|
||||
"single": true,
|
||||
"prompts": [
|
||||
49152,
|
||||
56320
|
||||
],
|
||||
"quality_followup": true,
|
||||
"minimum_headroom_mib": 384
|
||||
},
|
||||
{
|
||||
"label": "byteshape-single-114688",
|
||||
"model": "byteshape",
|
||||
"ctx": 114688,
|
||||
"single": true,
|
||||
"prompts": [
|
||||
49152,
|
||||
113664
|
||||
],
|
||||
"quality_followup": true
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,45 @@
|
||||
[
|
||||
{
|
||||
"label": "pure-single-ub128",
|
||||
"model": "pure",
|
||||
"ctx": 32768,
|
||||
"single": true,
|
||||
"ubatch": 128,
|
||||
"capacity_search": true
|
||||
},
|
||||
{
|
||||
"label": "byteshape-single-ub128",
|
||||
"model": "byteshape",
|
||||
"ctx": 32768,
|
||||
"single": true,
|
||||
"ubatch": 128,
|
||||
"capacity_search": true,
|
||||
"quality_followup": true
|
||||
},
|
||||
{
|
||||
"label": "pure-dual-262144-80-20",
|
||||
"model": "pure",
|
||||
"ctx": 262144,
|
||||
"single": false,
|
||||
"split": "80,20",
|
||||
"ubatch": 128,
|
||||
"mtp": 2,
|
||||
"prompts": [
|
||||
49152,
|
||||
261120
|
||||
]
|
||||
},
|
||||
{
|
||||
"label": "byteshape-dual-262144-80-20",
|
||||
"model": "byteshape",
|
||||
"ctx": 262144,
|
||||
"single": false,
|
||||
"split": "80,20",
|
||||
"ubatch": 128,
|
||||
"mtp": 2,
|
||||
"prompts": [
|
||||
49152,
|
||||
261120
|
||||
]
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,66 @@
|
||||
[
|
||||
{
|
||||
"label": "byteshape-single-32768-repeat",
|
||||
"model": "byteshape",
|
||||
"ctx": 32768,
|
||||
"single": true,
|
||||
"prompts": [
|
||||
4096
|
||||
]
|
||||
},
|
||||
{
|
||||
"label": "pure-single-32768-repeat",
|
||||
"model": "pure",
|
||||
"ctx": 32768,
|
||||
"single": true,
|
||||
"prompts": [
|
||||
4096
|
||||
]
|
||||
},
|
||||
{
|
||||
"label": "pure-single-61440-ub128",
|
||||
"model": "pure",
|
||||
"ctx": 61440,
|
||||
"single": true,
|
||||
"ubatch": 128,
|
||||
"prompts": [
|
||||
60416
|
||||
]
|
||||
},
|
||||
{
|
||||
"label": "byteshape-single-110592-ub128",
|
||||
"model": "byteshape",
|
||||
"ctx": 110592,
|
||||
"single": true,
|
||||
"ubatch": 128,
|
||||
"prompts": [
|
||||
109568
|
||||
]
|
||||
},
|
||||
{
|
||||
"label": "mix-single-76800-ub64",
|
||||
"model": "mix",
|
||||
"ctx": 76800,
|
||||
"single": true,
|
||||
"ubatch": 64,
|
||||
"mtp": 2,
|
||||
"prompts": [
|
||||
49152,
|
||||
75776
|
||||
],
|
||||
"minimum_headroom_mib": 384
|
||||
},
|
||||
{
|
||||
"label": "byteshape-dual-262144-86-14",
|
||||
"model": "byteshape",
|
||||
"ctx": 262144,
|
||||
"single": false,
|
||||
"split": "86,14",
|
||||
"ubatch": 128,
|
||||
"mtp": 2,
|
||||
"prompts": [
|
||||
49152,
|
||||
261120
|
||||
]
|
||||
}
|
||||
]
|
||||
@@ -0,0 +1,3 @@
|
||||
[
|
||||
{"label":"byteshape-dual-262144-88-12","model":"byteshape","ctx":262144,"single":false,"split":"88,12","ubatch":128,"mtp":2,"prompts":[49152,261120]}
|
||||
]
|
||||
@@ -0,0 +1,3 @@
|
||||
[
|
||||
{"label":"byteshape-dual-262144-86-14-validated","model":"byteshape","ctx":262144,"single":false,"split":"86,14","ubatch":128,"mtp":2,"prompts":[49152,261120]}
|
||||
]
|
||||
@@ -0,0 +1,34 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Produce reproducible tables from completed cases, excluding loading pilots."""
|
||||
import json,pathlib,statistics,sys
|
||||
root=pathlib.Path(sys.argv[1]); cases={}
|
||||
for p in sorted(root.glob('*/result.json')):
|
||||
d=json.loads(p.read_text())
|
||||
if not d.get('finished') or d['case'].get('load_only'):continue
|
||||
samples=json.loads((p.parent/'gpu.json').read_text());peak={}
|
||||
for s in samples:
|
||||
for g in s.get('gpus',[]):peak[g['name']]=max(peak.get(g['name'],0),int(g['used']))
|
||||
d['peak']=peak;cases[d['case']['label']]=d
|
||||
print('## Direkter Vergleich bei 32.768 Tokens Kontext\n')
|
||||
print('Ein Slot, nur RTX 5080, MTP3, Micro-Batch 512; gleiche Eingaben und Sampling. Mittelwerte der verfügbaren Wiederholungen.\n')
|
||||
print('| Messung | Pure | ByteShape | Änderung |\n|---|---:|---:|---:|')
|
||||
def values(model,key):
|
||||
out=[]
|
||||
for label in [model+'-single-32768',model+'-single-32768-repeat']:
|
||||
if label not in cases:continue
|
||||
d=cases[label]
|
||||
if key=='pp':out.append(d['prefill'][0]['response']['timings']['prompt_per_second'])
|
||||
elif key=='de':out.append(d['decode'][0]['timings']['predicted_per_second'])
|
||||
elif key=='code':out.append(d['decode'][1]['timings']['predicted_per_second'])
|
||||
return out
|
||||
for name,key in [('Prefill, 4.196 Eingabetokens (tok/s)','pp'),('Deutsche Erklärung (tok/s)','de'),('Python-Code (tok/s)','code')]:
|
||||
a,b=statistics.mean(values('pure',key)),statistics.mean(values('byteshape',key))
|
||||
print(f'| {name} | {a:.1f} | {b:.1f} | {(b/a-1)*100:+.1f}% |')
|
||||
print('\n## Vollständig getestete Konfigurationen\n')
|
||||
print('Kontext ist Eingabe plus Ausgabe. Die Geschwindigkeit in dieser Tabelle gehört jeweils zur angegebenen tatsächlichen Eingabelänge; Zeilen unterschiedlicher Länge sind kein isolierter Quantisierungsvergleich. VRAM enthält auch residente Dienste (TTS auf der 3060).\n')
|
||||
print('| Fall | Kontext | Eingabe | Prefill tok/s | Ausgabe tok/s | 5080 MiB | 3060 MiB | Recall |\n|---|---:|---:|---:|---:|---:|---:|---|')
|
||||
for label,d in cases.items():
|
||||
if not d.get('prefill'):continue
|
||||
p=d['prefill'][-1]['response'];t=p['timings'];r=p.get('recall',{})
|
||||
print(f'| {label} | {d["case"]["ctx"]} | {t["prompt_n"]} | {t["prompt_per_second"]:.1f} | {t["predicted_per_second"]:.1f} | {d["peak"].get("NVIDIA GeForce RTX 5080",0)} | {d["peak"].get("NVIDIA GeForce RTX 3060",0)} | {sum(r.values())}/{len(r)} |')
|
||||
print('\nDie Kontextpiloten ohne lange Eingabe sind hier bewusst nicht als validierte Konfigurationen aufgeführt. Einzelne synthetische Recall-Aufgaben belegen keine allgemeine Langkontext-Intelligenz.')
|
||||
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"label": "byteshape-dual-262144-80-20",
|
||||
"model": "byteshape",
|
||||
"ctx": 262144,
|
||||
"single": false,
|
||||
"split": "80,20",
|
||||
"ubatch": 128,
|
||||
"mtp": 2,
|
||||
"prompts": [
|
||||
49152,
|
||||
261120
|
||||
]
|
||||
}
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
+13
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"label": "byteshape-dual-262144-86-14-validated",
|
||||
"model": "byteshape",
|
||||
"ctx": 262144,
|
||||
"single": false,
|
||||
"split": "86,14",
|
||||
"ubatch": 128,
|
||||
"mtp": 2,
|
||||
"prompts": [
|
||||
49152,
|
||||
261120
|
||||
]
|
||||
}
|
||||
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"label": "byteshape-dual-262144-86-14",
|
||||
"model": "byteshape",
|
||||
"ctx": 262144,
|
||||
"single": false,
|
||||
"split": "86,14",
|
||||
"ubatch": 128,
|
||||
"mtp": 2,
|
||||
"prompts": [
|
||||
49152,
|
||||
261120
|
||||
]
|
||||
}
|
||||
Binary file not shown.
BIN
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,13 @@
|
||||
{
|
||||
"label": "byteshape-dual-262144-88-12",
|
||||
"model": "byteshape",
|
||||
"ctx": 262144,
|
||||
"single": false,
|
||||
"split": "88,12",
|
||||
"ubatch": 128,
|
||||
"mtp": 2,
|
||||
"prompts": [
|
||||
49152,
|
||||
261120
|
||||
]
|
||||
}
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,10 @@
|
||||
{
|
||||
"label": "byteshape-single-110592-ub128",
|
||||
"model": "byteshape",
|
||||
"ctx": 110592,
|
||||
"single": true,
|
||||
"ubatch": 128,
|
||||
"prompts": [
|
||||
109568
|
||||
]
|
||||
}
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"label": "byteshape-single-114688",
|
||||
"model": "byteshape",
|
||||
"ctx": 114688,
|
||||
"single": true,
|
||||
"prompts": [
|
||||
49152,
|
||||
113664
|
||||
],
|
||||
"quality_followup": true
|
||||
}
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,9 @@
|
||||
{
|
||||
"label": "byteshape-single-32768-repeat",
|
||||
"model": "byteshape",
|
||||
"ctx": 32768,
|
||||
"single": true,
|
||||
"prompts": [
|
||||
4096
|
||||
]
|
||||
}
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
@@ -0,0 +1,11 @@
|
||||
{
|
||||
"label": "byteshape-single-32768",
|
||||
"model": "byteshape",
|
||||
"ctx": 32768,
|
||||
"single": true,
|
||||
"quality": true,
|
||||
"prompts": [
|
||||
4096,
|
||||
24576
|
||||
]
|
||||
}
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
+14
@@ -0,0 +1,14 @@
|
||||
{
|
||||
"label": "byteshape-single-ub128-validated-104448",
|
||||
"model": "byteshape",
|
||||
"ctx": 104448,
|
||||
"single": true,
|
||||
"ubatch": 128,
|
||||
"capacity_search": true,
|
||||
"quality_followup": true,
|
||||
"load_only": false,
|
||||
"prompts": [
|
||||
49152,
|
||||
103424
|
||||
]
|
||||
}
|
||||
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
@@ -0,0 +1,165 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Bounded, isolated Qwen quantization benchmark. Supervisor restores production."""
|
||||
import json, pathlib, subprocess, sys, time, urllib.request, threading, signal
|
||||
ROOT = pathlib.Path('/data/benchmarks/byteshape-20260920')
|
||||
NAME = 'mike-ai-byteshape-test'
|
||||
BASE = 'http://127.0.0.1:5005'
|
||||
GPU0 = 'GPU-8ad38c6c-5a01-9d8e-1dfa-ed662ad78fbe'
|
||||
GPU1 = 'GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b'
|
||||
IMAGE = 'sha256:5e3c12c145b8045e5731b44b6b97033f24b327ae3d4a3fa85ecdd159cc844907'
|
||||
MODELS = {'mix':'qwen3.8-27b-iq4-mix/Qwen3.8-27B-IQ4-MIX.gguf', 'pure':'qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf', 'byteshape':'byteshape-qwen38-gpu5/model.gguf'}
|
||||
|
||||
def cmd(*args, check=True, timeout=90):
|
||||
r = subprocess.run(args, capture_output=True, text=True, timeout=timeout)
|
||||
if check and r.returncode: raise RuntimeError(str(args[:3])+': '+r.stderr[-2000:])
|
||||
return r.stdout
|
||||
|
||||
def api(path, data=None, timeout=900):
|
||||
req = urllib.request.Request(BASE+path, data=None if data is None else json.dumps(data).encode(), headers={'Content-Type':'application/json'})
|
||||
with urllib.request.urlopen(req, timeout=timeout) as r: return json.load(r)
|
||||
|
||||
def save(path, data):
|
||||
path.write_text(json.dumps(data, indent=2, ensure_ascii=False)+'\n')
|
||||
|
||||
def gpu():
|
||||
rows = cmd('nvidia-smi','--query-gpu=name,memory.used,memory.total,temperature.gpu,utilization.gpu','--format=csv,noheader,nounits',timeout=15)
|
||||
return [dict(zip(['name','used','total','temp','util'], [v.strip() for v in row.split(',')])) for row in rows.splitlines()]
|
||||
|
||||
def health_check():
|
||||
rows=gpu()
|
||||
if any(int(x['temp']) >= 85 for x in rows): raise RuntimeError('GPU temperature limit')
|
||||
mem = dict((a.split(':')[0],int(a.split()[1])) for a in pathlib.Path('/proc/meminfo').read_text().splitlines())
|
||||
if mem['MemAvailable'] < 3*1024*1024: raise RuntimeError('Host RAM reserve below 3 GiB')
|
||||
return rows
|
||||
|
||||
def chat(prompt, max_tokens=512, effort='none', seed=42, tools=None):
|
||||
health_check()
|
||||
p={'model':'benchmark','messages':[{'role':'user','content':prompt}], 'max_tokens':max_tokens,'temperature':1.0,'top_p':0.95,'top_k':20,'min_p':0.0,'seed':seed,'reasoning_effort':effort,'cache_prompt':False}
|
||||
if tools: p.update(tools=tools,tool_choice='auto')
|
||||
start=time.monotonic(); r=api('/v1/chat/completions',p); r['wall_seconds']=time.monotonic()-start
|
||||
health_check()
|
||||
return r
|
||||
|
||||
def prefill(n, seed):
|
||||
# Exact token-array slicing avoids accidentally exceeding the intended input size.
|
||||
text='\n'.join(f'Record {i:06d}: cobalt lantern maple orbit quartz river silver tango.' for i in range(max(100,n//10)))
|
||||
tokens=api('/tokenize',{'content':text,'add_special':False})['tokens'][:n]
|
||||
# Insert independently locatable facts near start/middle/end of the context.
|
||||
text=api('/detokenize',{'tokens':tokens})['content']
|
||||
for pos, fact in reversed([(len(text)//8,'NEEDLE_ALPHA=RAVEN-417'),(len(text)//2,'NEEDLE_BETA=CEDAR-928'),(len(text)*7//8,'NEEDLE_GAMMA=ORBIT-563')]):
|
||||
text=text[:pos]+'\n'+fact+'\n'+text[pos:]
|
||||
text+='\nReturn a JSON object with alpha, beta and gamma containing the three exact NEEDLE values. Then explain in German how to verify these records without inventing evidence, in at least 300 words.'
|
||||
r=chat(text,512,seed=seed)
|
||||
content=r['choices'][0]['message'].get('content','')
|
||||
r['recall']={x:x in content for x in ['RAVEN-417','CEDAR-928','ORBIT-563']}
|
||||
return r
|
||||
|
||||
def run_case(case):
|
||||
label=case['label']; out=ROOT/label; out.mkdir(exist_ok=True)
|
||||
if (out/'result.json').exists(): raise RuntimeError('Refusing to overwrite completed case '+label)
|
||||
save(out/'config.json',case)
|
||||
print('START',label,flush=True)
|
||||
single=case.get('single',False)
|
||||
args=['--model','/models/'+MODELS[case['model']], '--alias','benchmark','--ctx-size',str(case['ctx']), '--flash-attn','on','--cache-type-k','q4_0','--cache-type-v','q4_0','--cache-ram','0','--threads','6','--threads-batch','6','--batch-size','2048','--ubatch-size',str(case.get('ubatch',512)), '--parallel','1','--kv-unified','--jinja','--reasoning','auto','--reasoning-preserve','--host','127.0.0.1','--port','5005','--metrics','--fit','off','--n-gpu-layers','all','--load-mode','none','--no-ui','--temperature','1.0','--top-p','0.95','--top-k','20','--device','CUDA0' if single else 'CUDA0,CUDA1','--main-gpu','0','--split-mode','layer','--tensor-split',case.get('split','1,0'),'--spec-type','draft-mtp','--spec-draft-n-max',str(case.get('mtp',3)),'--spec-draft-type-k','f16','--spec-draft-type-v','f16','--spec-draft-p-min','0.05','--verbosity','3']
|
||||
cmd('docker','run','-d','--name',NAME,'--gpus','all','--network','host','--read-only','--tmpfs','/tmp:rw,nosuid,nodev,size=256m','--security-opt','no-new-privileges:true','--cap-drop','ALL','--pids-limit','512','--ulimit','core=0','--memory','26g','--memory-swap','26g','--shm-size','1g','--log-opt','max-size=32m','--log-opt','max-file=1','-e','NVIDIA_VISIBLE_DEVICES='+GPU0+','+GPU1,'-e','NVIDIA_DRIVER_CAPABILITIES=compute,utility','-v','/data/models:/models:ro',IMAGE,*args)
|
||||
stop=threading.Event(); samples=[]
|
||||
def monitor():
|
||||
while not stop.wait(2):
|
||||
try:
|
||||
rows=health_check()
|
||||
samples.append({'time':time.time(),'gpus':rows})
|
||||
except RuntimeError as e:
|
||||
samples.append({'error':str(e),'aborted':True})
|
||||
cmd('docker','stop','-t','10',NAME,check=False)
|
||||
return
|
||||
except Exception as e: samples.append({'error':str(e)})
|
||||
thread=threading.Thread(target=monitor,daemon=True); thread.start()
|
||||
result={'case':case,'started':time.time()}
|
||||
try:
|
||||
for _ in range(150):
|
||||
try:
|
||||
if api('/health',timeout=3).get('status')=='ok': break
|
||||
except Exception: pass
|
||||
if cmd('docker','inspect',NAME,'--format','{{.State.Running}}').strip()!='true': raise RuntimeError('Test container exited during load')
|
||||
time.sleep(2)
|
||||
else: raise RuntimeError('Startup exceeded 300s')
|
||||
result['idle_gpu']=health_check(); result['props']=api('/props'); result['slots']=api('/slots')
|
||||
save(out/'loaded.json',result)
|
||||
# Added after the 88:12 trial: model loading alone can succeed while
|
||||
# the first real attention graph still needs more CUDA workspace.
|
||||
minimum=case.get('minimum_headroom_mib',512)
|
||||
used_devices=['5080'] if single else ['5080','3060']
|
||||
for g in result['idle_gpu']:
|
||||
if any(device in g['name'] for device in used_devices):
|
||||
free=int(g['total'])-int(g['used'])
|
||||
if free<minimum:
|
||||
raise RuntimeError(f"Insufficient loaded VRAM reserve on {g['name']}: {free} < {minimum} MiB; refusing inference")
|
||||
result['smoke']=chat('Antworte nur mit OK.',8)
|
||||
if case.get('quality'):
|
||||
tasks=json.loads(pathlib.Path('/opt/mike-ai/stack/dev/QWEN38-FINAL-ACCEPTANCE-v1.json').read_text())
|
||||
result['quality']=[]
|
||||
for task in tasks:
|
||||
ans=chat(task['prompt'],task['max_tokens'],'medium')
|
||||
result['quality'].append({'id':task['id'],'response':ans})
|
||||
save(out/'partial.json',result); print(label,task['id'],round(ans['wall_seconds'],1),flush=True)
|
||||
tool={'type':'function','function':{'name':'read_server_status','description':'Read-only server status lookup','parameters':{'type':'object','properties':{'server':{'type':'string'}},'required':['server'],'additionalProperties':False}}}
|
||||
result['tool']=chat('Read the current status of server alpha. Use the provided tool exactly once and do not invent its result.',512,tools=[tool])
|
||||
if case.get('quality_followup'):
|
||||
tasks=json.loads(pathlib.Path('/opt/mike-ai/stack/dev/QWEN38-FINAL-ACCEPTANCE-v1.json').read_text())
|
||||
result['quality_followup']=[]
|
||||
for task in tasks:
|
||||
if task['id'] not in ['i3_code_debugging','i4_capacity_planning','i6_state_vs_configuration']: continue
|
||||
ans=chat(task['prompt'],8192,'medium',seed=43)
|
||||
result['quality_followup'].append({'id':task['id'],'seed':43,'budget':8192,'response':ans})
|
||||
save(out/'partial.json',result); print(label,'followup',task['id'],round(ans['wall_seconds'],1),flush=True)
|
||||
if not case.get('load_only'):
|
||||
result['decode']=[]
|
||||
prompts=['Erkläre ausführlich auf Deutsch, wie ein Reverse Proxy funktioniert, welche Fehler bei Container-IP-Wechseln auftreten können und wie man sie anhand von Logs eingrenzt. Schreibe mindestens 600 Wörter.', 'Write a Python implementation of an asynchronous first_success function: start all awaitables concurrently, return the first successful result, cancel and await remaining tasks, collect exceptions if all fail. Include an explanation and usage example.']
|
||||
for i,p in enumerate(prompts): result['decode'].append(chat(p,768,seed=42+i))
|
||||
result['prefill']=[]
|
||||
for n in case.get('prompts',[4096,16384]):
|
||||
r=prefill(n,42); result['prefill'].append({'target':n,'response':r}); save(out/'partial.json',result)
|
||||
print(label,'prefill',n,r.get('timings'),flush=True)
|
||||
result['finished']=time.time()
|
||||
finally:
|
||||
stop.set(); thread.join(5)
|
||||
r=subprocess.run(['docker','logs',NAME],capture_output=True,text=True,timeout=30)
|
||||
(out/'server.log').write_text(r.stdout+r.stderr)
|
||||
save(out/'gpu.json',samples); save(out/'result.json',result)
|
||||
cmd('docker','rm','-f',NAME,check=False)
|
||||
print('DONE',label,flush=True)
|
||||
return result
|
||||
|
||||
def capacity_case(case):
|
||||
"""Bounded growth from a previously working context, with 768 MiB reserve.
|
||||
|
||||
0.04 MiB/token exceeds the measured 512-ubatch steady-state slope.
|
||||
A failed 114688/512 ByteShape startup revealed additional transient MTP
|
||||
buffers, so reserve is deliberately larger than steady-state extrapolation.
|
||||
Smaller ubatches start at an already working context, not a guessed OOM edge.
|
||||
The final context is tested with an actual almost-full prompt.
|
||||
"""
|
||||
context=case['ctx']
|
||||
previous=None
|
||||
for attempt in range(8):
|
||||
pilot={**case,'ctx':context,'label':case['label']+'-pilot-'+str(context),'load_only':True,'quality':False,'quality_followup':False}
|
||||
result=run_case(pilot)
|
||||
rows=[g for g in result['idle_gpu'] if '5080' in g['name']]
|
||||
samples=json.loads((ROOT/pilot['label']/'gpu.json').read_text())
|
||||
peak=max([int(rows[0]['used'])+32]+[int(g['used']) for s in samples for g in s.get('gpus',[]) if '5080' in g['name']])
|
||||
free=int(rows[0]['total'])-peak
|
||||
if free<768:
|
||||
if previous is None: raise RuntimeError('Initial capacity pilot has insufficient reserve')
|
||||
context=previous
|
||||
break
|
||||
growth=min(16384,int((free-768)/0.04)//1024*1024)
|
||||
if growth<1024 or attempt==7 or context>=262144: break
|
||||
previous=context
|
||||
context=min(262144,context+growth)
|
||||
final={**case,'ctx':context,'label':case['label']+'-validated-'+str(context),'load_only':False,'prompts':[49152,context-1024]}
|
||||
return run_case(final)
|
||||
|
||||
if __name__=='__main__':
|
||||
for case in json.loads(pathlib.Path(sys.argv[1]).read_text()):
|
||||
if case.get('capacity_search'): capacity_case(case)
|
||||
else: run_case(case)
|
||||
@@ -0,0 +1,19 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Summarize measured timings; never substitute configured context for input length."""
|
||||
import json,pathlib,sys
|
||||
root=pathlib.Path(sys.argv[1])
|
||||
rows=[]
|
||||
for path in sorted(root.glob('*/result.json')):
|
||||
r=json.loads(path.read_text()); samples=json.loads((path.parent/'gpu.json').read_text())
|
||||
peak={}
|
||||
for s in samples:
|
||||
for g in s.get('gpus',[]):
|
||||
p=peak.setdefault(g['name'],{'MiB':0,'C':0})
|
||||
p['MiB']=max(p['MiB'],int(g['used']));p['C']=max(p['C'],int(g['temp']))
|
||||
row={'case':r['case'],'finished':bool(r.get('finished')),'peak':peak,'decode':[], 'prefill':[]}
|
||||
for d in r.get('decode',[]): row['decode'].append(d.get('timings'))
|
||||
for p in r.get('prefill',[]):
|
||||
a=p['response'];row['prefill'].append({'target':p['target'],'timings':a.get('timings'),'recall':a.get('recall'),'wall_seconds':a.get('wall_seconds')})
|
||||
row['quality']=[{'id':q['id'],'finish':q['response']['choices'][0].get('finish_reason'),'visible_chars':len(q['response']['choices'][0]['message'].get('content','')),'tokens':q['response'].get('usage',{})} for q in r.get('quality',[])]
|
||||
rows.append(row)
|
||||
print(json.dumps(rows,indent=2,ensure_ascii=False))
|
||||
@@ -0,0 +1,40 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Stop only existing router/controller/model; always restore the same containers."""
|
||||
import json, pathlib, subprocess, sys, time, signal
|
||||
ROOT=pathlib.Path('/data/benchmarks/byteshape-20260920')
|
||||
NAMES=['mike-ai-router','mike-ai-profile-controller','mike-ai-llama-medium']
|
||||
def run(*args,check=True,timeout=90):
|
||||
return subprocess.run(args,capture_output=True,text=True,check=check,timeout=timeout)
|
||||
def stop_signal(*_): raise RuntimeError('Supervisor interrupted')
|
||||
signal.signal(signal.SIGTERM,stop_signal); signal.signal(signal.SIGINT,stop_signal)
|
||||
# Refuse if the known production state has changed, or if requests are active.
|
||||
for name in NAMES:
|
||||
assert run('docker','inspect',name,'--format','{{.State.Running}}').stdout.strip()=='true',name
|
||||
slots=json.loads(run('docker','exec',NAMES[-1],'curl','-fsS','http://127.0.0.1:8080/slots').stdout)
|
||||
assert not any(s['is_processing'] for s in slots),'Production request active'
|
||||
assert not run('docker','ps','-q','--filter','name=^mike-ai-byteshape-test$').stdout.strip(),'Existing experiment'
|
||||
child=None
|
||||
try:
|
||||
run('docker','stop','-t','30',*NAMES[:2])
|
||||
# Drain requests already handed to the model, before unloading it.
|
||||
for _ in range(120):
|
||||
slots=json.loads(run('docker','exec',NAMES[-1],'curl','-fsS','http://127.0.0.1:8080/slots').stdout)
|
||||
if not any(s['is_processing'] for s in slots): break
|
||||
time.sleep(2)
|
||||
else: raise RuntimeError('Model did not drain')
|
||||
run('docker','stop','-t','30',NAMES[-1])
|
||||
child=subprocess.Popen(['python3',str(ROOT/'run.py'),sys.argv[1]])
|
||||
code=child.wait(timeout=4200)
|
||||
if code: raise RuntimeError('Benchmark failed: '+str(code))
|
||||
finally:
|
||||
if child is not None and child.poll() is None:
|
||||
child.terminate()
|
||||
try: child.wait(timeout=20)
|
||||
except subprocess.TimeoutExpired: child.kill(); child.wait(timeout=10)
|
||||
run('docker','rm','-f','mike-ai-byteshape-test',check=False)
|
||||
run('docker','start',NAMES[-1])
|
||||
for _ in range(150):
|
||||
if run('docker','inspect',NAMES[-1],'--format','{{.State.Health.Status}}').stdout.strip()=='healthy': break
|
||||
time.sleep(2)
|
||||
run('docker','start',NAMES[1],NAMES[0])
|
||||
print('RESTORED existing medium/controller/router',flush=True)
|
||||
@@ -0,0 +1,29 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Read-only restoration verification plus a two-token model smoke request."""
|
||||
import json,pathlib,re,subprocess,time
|
||||
ROOT=pathlib.Path('/data/benchmarks/byteshape-20260920')
|
||||
def run(*args):return subprocess.check_output(args,text=True,timeout=30)
|
||||
report={'checked_at':time.time(),'uptime':run('uptime').strip(),'containers':{}}
|
||||
for name in ['mike-ai-llama-medium','mike-ai-router','mike-ai-profile-controller','mike-ai-wireguard-gateway','mike-ai-qwen3-tts']:
|
||||
d=json.loads(run('docker','inspect',name))[0]
|
||||
report['containers'][name]={'running':d['State']['Running'],'health':d['State'].get('Health',{}).get('Status'),'image':d['Image'],'started':d['State']['StartedAt']}
|
||||
assert d['State']['Running'],name
|
||||
if name in ['mike-ai-llama-medium','mike-ai-router','mike-ai-profile-controller','mike-ai-wireguard-gateway']:
|
||||
assert d['State'].get('Health',{}).get('Status')=='healthy',name
|
||||
assert report['containers']['mike-ai-llama-medium']['image']=='sha256:5e3c12c145b8045e5731b44b6b97033f24b327ae3d4a3fa85ecdd159cc844907'
|
||||
assert not run('docker','ps','-q','--filter','name=^mike-ai-byteshape-test$').strip()
|
||||
probe='import urllib.request,json; print(json.dumps({p:json.load(urllib.request.urlopen("http://127.0.0.1:8081"+p,timeout=20)) for p in ["/health","/ready"]}))'
|
||||
report['router']=json.loads(run('docker','exec','mike-ai-router','python','-c',probe))
|
||||
payload={'model':'qwen-medium','messages':[{'role':'user','content':'Antworte ausschließlich mit OK.'}],'reasoning_effort':'none','max_tokens':8,'temperature':0}
|
||||
r=json.loads(run('docker','exec','mike-ai-llama-medium','curl','-fsS','--max-time','20','-H','Content-Type: application/json','--data',json.dumps(payload),'http://127.0.0.1:8080/v1/chat/completions'))
|
||||
report['smoke']={'content':r['choices'][0]['message'].get('content',''),'usage':r.get('usage')}
|
||||
assert report['smoke']['content'].strip()=='OK',report['smoke']
|
||||
started=min(json.loads(p.read_text())['started'] for p in ROOT.glob('*/result.json'))
|
||||
journal=run('journalctl','-k','--since','@'+str(int(started)-60),'--no-pager')
|
||||
pattern=re.compile(r'NVRM.*Xid|oom-kill|Out of memory: Killed process|Kernel panic|GPU has fallen off',re.I)
|
||||
report['kernel_errors']=[line for line in journal.splitlines() if pattern.search(line)]
|
||||
report['gpu']=run('nvidia-smi','--query-gpu=name,memory.used,memory.total,temperature.gpu','--format=csv').strip()
|
||||
report['disk']=run('df','-h','/','/data').strip()
|
||||
(ROOT/'restore-verification.json').write_text(json.dumps(report,indent=2)+'\n')
|
||||
print(json.dumps(report,indent=2))
|
||||
assert not report['kernel_errors'],'Kernel/GPU errors require review'
|
||||
Reference in New Issue
Block a user