Test ordered one-slot DFlash2 placements and preserve results
This commit is contained in:
@@ -1,5 +1,7 @@
|
|||||||
# DFlash2: kurzer Medium-Test auf Athena
|
# DFlash2: kurzer Medium-Test auf Athena
|
||||||
|
|
||||||
|
Nachtrag: [Ein-Slot-Folgeversuch mit Draft auf 3060 funktioniert](DFLASH2_ONE_SLOT_20260920.md). Der folgende Abschnitt dokumentiert den ursprünglichen Zwei-Slot-Test.
|
||||||
|
|
||||||
20.09.2026. **Ergebnis: Laden fehlgeschlagen, keine Durchsatzmessung möglich.**
|
20.09.2026. **Ergebnis: Laden fehlgeschlagen, keine Durchsatzmessung möglich.**
|
||||||
|
|
||||||
Die vorhandene llama.cpp-Version unterstützt DFlash2 einschließlich Selector und
|
Die vorhandene llama.cpp-Version unterstützt DFlash2 einschließlich Selector und
|
||||||
|
|||||||
@@ -0,0 +1,99 @@
|
|||||||
|
# DFlash2 auf Athena: Ein-Slot-Vergleich
|
||||||
|
|
||||||
|
20.09.2026, im Anschluss an den gescheiterten Zwei-Slot-Kurztest.
|
||||||
|
|
||||||
|
**Ein Slot mit Draft auf der RTX 3060 funktioniert im Kurztest. Ein klarer
|
||||||
|
Geschwindigkeitsvorteil gegenüber der gespeicherten MTP-Referenz ist nicht
|
||||||
|
sichtbar. Kein produktiver Wechsel.**
|
||||||
|
|
||||||
|
## Gewünschte Reihenfolge und Ergebnis
|
||||||
|
|
||||||
|
| Priorität | Konfiguration | Ergebnis |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | Ein Slot, Qwen 80:20, Draft auf 5080 | Initialisierung scheitert an Gerätezuordnung der geteilten Ausgabematrix |
|
||||||
|
| 2 | Ein Slot, Qwen 80:20, Draft auf 3060 | Laden, Deutsch, Code und kurze Prefill-/Recall-Probe bestanden |
|
||||||
|
| 3 | Ein Slot, mehr Qwen auf 3060 (70:30), Draft auf 5080 | Nicht ausgeführt: nur als Rückfall bei Fehlschlag von Priorität 2 vorgesehen |
|
||||||
|
|
||||||
|
Beim ersten Ein-Slot-Lauf scheitert die Laufzeit mit
|
||||||
|
`pre-allocated tensor (output.weight) in a buffer (CUDA1) that cannot run the operation (NONE)`.
|
||||||
|
Der Draft ist auf CUDA0 begrenzt, während die verwendete Ausgabematrix auf CUDA1
|
||||||
|
liegt. Anders als beim ursprünglichen Zwei-Slot-Versuch ist dies kein gemeldeter
|
||||||
|
CUDA-Malloc-Fehler. Der Testprozess brach ab, der Host blieb erreichbar.
|
||||||
|
|
||||||
|
## Durchsatz
|
||||||
|
|
||||||
|
| Messung | DFlash2, ein Slot, 160K Kontext | Gespeicherte Pure/MTP-Zwei-GPU-Referenz |
|
||||||
|
|---|---:|---:|
|
||||||
|
| Deutsche Erklärung, Ausgabe | 37,8 tok/s | 57,3 tok/s |
|
||||||
|
| Python-Code, Ausgabe | 75,5 tok/s | 73,4 tok/s |
|
||||||
|
| Prefill, 4.196 Eingabetokens | 1.437,6 tok/s | kein gleicher 4K-Messpunkt dieses Referenzprofils |
|
||||||
|
| Ausgabe nach dieser 4K-Eingabe | 41,1 tok/s | kein gleicher 4K-Messpunkt dieses Referenzprofils |
|
||||||
|
| Recall der drei eingebauten Fakten | 3/3 | — |
|
||||||
|
|
||||||
|
Die Referenz verwendet dasselbe Pure-Modell, 80:20-Layer-Split und Microbatch 128,
|
||||||
|
aber **262.144 statt 160.000 Kontext, MTP2 statt DFlash2**. Das sind Vergleiche
|
||||||
|
vollständiger Konfigurationen, keine isolierten DFlash-Beschleunigungsfaktoren.
|
||||||
|
Es wurde kein neuer Qwen-Referenzlauf durchgeführt.
|
||||||
|
|
||||||
|
Deutsch erzeugte 768 Tokens in 20,29 Sekunden, Code endete nach 671 Tokens in
|
||||||
|
8,87 Sekunden; die Referenz erzeugte in beiden Aufgaben 768 Tokens. Gleiche
|
||||||
|
Eingaben und Seeds bedeuten bei unterschiedlicher spekulativer Verarbeitung
|
||||||
|
keine identischen Ausgaben. Ein Durchlauf je Aufgabe; geringe Unterschiede
|
||||||
|
wie die rund 3 % beim Code sind kein belastbarer allgemeiner Geschwindigkeitsgewinn.
|
||||||
|
|
||||||
|
DFlash-Draft-Akzeptanz: Deutsch 417/2444 = 17,1 %, Code 517/1071 = 48,3 %, kurze
|
||||||
|
Recall-Aufgabe 300/1460 = 20,5 %. Das sind angenommene Draft-Tokens, keine
|
||||||
|
Qualitätswerte; der Zusatzaufwand lohnt bei geringer Annahme weniger.
|
||||||
|
|
||||||
|
## Speicher und Kontext
|
||||||
|
|
||||||
|
Getestet: Pure IQ4_XS, bestehendes llama.cpp-Image b29c606, Q4_K_M-DFlash2,
|
||||||
|
160.000 Kontext, ein Slot, q4_0-K/V für Hauptmodell und Draft, Microbatch 128,
|
||||||
|
Batch 2048, Draft-Länge 7. Text ohne Vision-Projektor, TTS auf 3060 resident.
|
||||||
|
|
||||||
|
| GPU | Belegung nach Laden | Höchste gesampelte Belegung | Freier Speicher am gemessenen Peak |
|
||||||
|
|---|---:|---:|---:|
|
||||||
|
| RTX 5080 | 14.632 MiB | 14.660 MiB | 1.643 MiB |
|
||||||
|
| RTX 3060, inklusive TTS | 10.320 MiB | 10.378 MiB | 1.910 MiB |
|
||||||
|
|
||||||
|
Die Konfiguration mit 160K Kontext wurde geladen. **Die tatsächlich größte
|
||||||
|
Testeingabe hatte 4.196 Tokens.** Damit ist weder ein voller 160K-Langkontexttest
|
||||||
|
noch die Stabilität bei Vision oder paralleler Last erbracht. Die gemessenen
|
||||||
|
Reserven rechtfertigen keine automatische Hochrechnung einer maximalen Kapazität.
|
||||||
|
|
||||||
|
## Antwortqualität
|
||||||
|
|
||||||
|
Alle drei synthetischen Fakten wurden gefunden. Das Codebeispiel besteht den
|
||||||
|
geprüften Fehlerpfad „erste Aufgabe scheitert, spätere gelingt“. Wenn alle
|
||||||
|
Aufgaben scheitern, erzeugt es jedoch einen **NameError**, weil `builtins` ohne
|
||||||
|
Import verwendet wird. Die manuell geprüfte Antwort und der reproduzierbare
|
||||||
|
Teiltest sind archiviert. Die Deutschantwort enthält sprachliche und sachliche
|
||||||
|
Unschärfen; die Ausgabe endet am festgelegten Tokenlimit.
|
||||||
|
|
||||||
|
Diese kleinen Durchsatzaufgaben erlauben keine Freigabe als qualitativ gleichwertig
|
||||||
|
und belegen auch keinen durch DFlash verursachten Intelligenzverlust. Eine
|
||||||
|
allgemeine Qualitätsbewertung oder deterministische Äquivalenzprüfung wurde
|
||||||
|
in diesem kurzen Versuch nicht durchgeführt.
|
||||||
|
|
||||||
|
## Abschluss
|
||||||
|
|
||||||
|
Das bisherige Medium-Profil mit seinen ursprünglichen zwei Slots wurde nach
|
||||||
|
jedem Versuch wiederhergestellt; die Ein-Slot-Änderung war nur im isolierten
|
||||||
|
Testcontainer aktiv. Medium, Router, Controller, Gateway und TTS gesund,
|
||||||
|
Router readiness und minimale Modellantwort geprüft. Keine Xid-, OOM-Kill-,
|
||||||
|
Kernel-Panic- oder GPU-fallen-off-Meldungen im geprüften Kerneljournal.
|
||||||
|
Keine Treiber-, Kernel- oder Netzwerkänderungen, kein Neustart des Hosts.
|
||||||
|
|
||||||
|
Ein aktiver Produktionsrequest verzögerte beide Teststarts; der zweite musste
|
||||||
|
auf die Verarbeitung einer langen Eingabe warten. Keine laufende Anfrage wurde
|
||||||
|
absichtlich abgebrochen. Der gespeicherte Supervisor wartet künftig zusätzlich
|
||||||
|
auf gesunde Router-/Controller-Checks, bevor er die Wiederherstellung meldet.
|
||||||
|
|
||||||
|
**Empfehlung:** DFlash2 ist in dieser Ein-Slot-Anordnung technisch nutzbar,
|
||||||
|
aber nach diesem Kurztest kein überzeugender Ersatz für das vorhandene MTP.
|
||||||
|
Weitere Arbeit wäre gezieltes Draft-/Platzierungs-Tuning mit anschließender
|
||||||
|
Qualitäts- und Langkontextprüfung, keine produktive Aktivierung des jetzigen Falls.
|
||||||
|
|
||||||
|
Konfigurationen, Rohantworten, Serverlogs, GPU-Samples und Prüfbelege:
|
||||||
|
`experiments/dflash2-medium-20260920/` im Repository, entsprechend
|
||||||
|
`/data/benchmarks/dflash2-medium-20260920/` auf Athena.
|
||||||
@@ -35,3 +35,5 @@ Stand: 20. September 2026. Modellfamilie bleibt Qwen3.8-27B.
|
|||||||
## DFlash2-Kurztest
|
## DFlash2-Kurztest
|
||||||
|
|
||||||
20.09.2026: [Medium-Ladetest](DFLASH2_MEDIUM_QUICK_20260920.md) mit 160.000 Kontext/zwei Slots scheitert an zusätzlichem Draft-VRAM auf der 5080. Keine Durchsatz- oder Qualitätswerte. Bisheriges Medium wiederhergestellt; Punkt 3 bleibt offen und benötigt ein anderes Speicherlayout.
|
20.09.2026: [Medium-Ladetest](DFLASH2_MEDIUM_QUICK_20260920.md) mit 160.000 Kontext/zwei Slots scheitert an zusätzlichem Draft-VRAM auf der 5080. Keine Durchsatz- oder Qualitätswerte. Bisheriges Medium wiederhergestellt; Punkt 3 bleibt offen und benötigt ein anderes Speicherlayout.
|
||||||
|
|
||||||
|
Ein-Slot-Folgeversuche: [Bericht](DFLASH2_ONE_SLOT_20260920.md). Draft auf 5080 scheitert an Gerätezuordnung; Draft auf 3060 besteht kurze Texttests. Deutsch 37,8 tok/s, Code 75,5 tok/s: kein klarer Vorteil gegenüber gespeicherter MTP-Referenz. Dritter Rückfalltest nicht nötig. Vollständige Bewertung/Tuning in Punkt 3 bleibt offen.
|
||||||
|
|||||||
@@ -55,3 +55,32 @@ Target loaded; draft allocation of1079.61MiB onCUDA0 failed with CUDA malloc
|
|||||||
OOM. No benchmark inference executed. Supervisor restored original production;
|
OOM. No benchmark inference executed. Supervisor restored original production;
|
||||||
health/smoke/kernel checks passed. No performance or quality conclusion can be
|
health/smoke/kernel checks passed. No performance or quality conclusion can be
|
||||||
drawn. See results/ and restore-verification.json.
|
drawn. See results/ and restore-verification.json.
|
||||||
|
|
||||||
|
## User-ordered one-slot follow-up
|
||||||
|
|
||||||
|
The user authorized this order after the initial two-slot allocation failure:
|
||||||
|
|
||||||
|
1. One slot, target80:20, draft on5080.
|
||||||
|
2. One slot, target80:20, draft on3060, regardless of whether the first works.
|
||||||
|
3. Only if step2 fails: one slot, target70:30, draft back on5080. Moving both
|
||||||
|
additional target layers and the draft onto3060 would compete for its memory;
|
||||||
|
this fallback instead makes space on5080 for the draft.
|
||||||
|
|
||||||
|
Context remains160000 and other quick-probe settings are unchanged. These tests
|
||||||
|
are temporary; production containers retain their original two-slot arguments.
|
||||||
|
Case JSON now controls slot count and draft device (historical quick.json keeps
|
||||||
|
its original two-slot meaning). Failed attempts include an explicit error field.
|
||||||
|
|
||||||
|
An active production request postpones startup for up to180 seconds; no live
|
||||||
|
request is deliberately interrupted. The existing drain check still protects
|
||||||
|
requests racing with the transition. Each attempt has its own result directory;
|
||||||
|
previous results and the frozen Qwen reference are never overwritten.
|
||||||
|
|
||||||
|
## One-slot results
|
||||||
|
|
||||||
|
Priority1 fails during draft graph initialization: output.weight resides onCUDA1
|
||||||
|
but draft backends are restricted toCUDA0. Priority2 works; priority3 therefore
|
||||||
|
not run. German37.79 tok/s, code75.51 tok/s, 4196-token prefill1437.56 tok/s,
|
||||||
|
recall3/3. Allocation160000 is not full-context validation. Original two-slot
|
||||||
|
production restored and checked after both cases. See
|
||||||
|
../../docs/DFLASH2_ONE_SLOT_20260920.md for comparison caveats and quality findings.
|
||||||
|
|||||||
@@ -0,0 +1,20 @@
|
|||||||
|
import ast,asyncio,json,pathlib,re,typing,gzip
|
||||||
|
p=pathlib.Path(__file__).resolve().parent/'results/medium-slot1-priority2-draft-cuda1/result.json.gz'
|
||||||
|
d=json.loads(gzip.decompress(p.read_bytes()));content=d['decode'][1]['choices'][0]['message']['content']
|
||||||
|
source=re.search(r'```python\s*\n(.*?)```',content,re.S).group(1)
|
||||||
|
# Reviewed above: only asyncio task orchestration and exception handling, no I/O.
|
||||||
|
node=next(n for n in ast.parse(source).body if isinstance(n,ast.AsyncFunctionDef) and n.name=='first_success')
|
||||||
|
scope={'asyncio':asyncio,'Any':typing.Any,'List':typing.List,'Awaitable':typing.Awaitable}
|
||||||
|
exec(compile(ast.Module(body=[node],type_ignores=[]),'<reviewed-dflash-answer>','exec'),scope)
|
||||||
|
async def check(all_fail):
|
||||||
|
async def task(delay,fail):
|
||||||
|
await asyncio.sleep(delay)
|
||||||
|
if fail:raise ValueError('synthetic failure')
|
||||||
|
return 7
|
||||||
|
try:
|
||||||
|
r=await asyncio.wait_for(scope['first_success']([task(.001,True),task(.005,all_fail)]),timeout=1)
|
||||||
|
return {'returned':r,'pass':not all_fail and r==7}
|
||||||
|
except Exception as e:
|
||||||
|
return {'error':type(e).__name__+': '+str(e),'pass':all_fail and type(e).__name__=='ExceptionGroup'}
|
||||||
|
out={'scope':'Two checks of manually reviewed decode answer, not broad quality evaluation','fast_failure_then_success':asyncio.run(check(False)),'all_fail':asyncio.run(check(True))}
|
||||||
|
print(json.dumps(out,indent=2))
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
{
|
||||||
|
"scope": "Two checks of manually reviewed decode answer, not broad quality evaluation",
|
||||||
|
"fast_failure_then_success": {
|
||||||
|
"returned": 7,
|
||||||
|
"pass": true
|
||||||
|
},
|
||||||
|
"all_fail": {
|
||||||
|
"error": "NameError: name 'builtins' is not defined",
|
||||||
|
"pass": false
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,61 @@
|
|||||||
|
{
|
||||||
|
"checked_at": 1789931025.6105647,
|
||||||
|
"uptime": "21:03:45 up 3 days, 9:58, 1 user, load average: 1.34, 0.91, 0.73",
|
||||||
|
"containers": {
|
||||||
|
"mike-ai-llama-medium": {
|
||||||
|
"running": true,
|
||||||
|
"health": "healthy",
|
||||||
|
"image": "sha256:5e3c12c145b8045e5731b44b6b97033f24b327ae3d4a3fa85ecdd159cc844907",
|
||||||
|
"started": "2026-09-20T19:03:00.899996682Z"
|
||||||
|
},
|
||||||
|
"mike-ai-router": {
|
||||||
|
"running": true,
|
||||||
|
"health": "healthy",
|
||||||
|
"image": "sha256:0448758bec6968b29263bcac0f8b4682c6d3029d626f7c12334698c11ef096bb",
|
||||||
|
"started": "2026-09-20T19:03:11.359298419Z"
|
||||||
|
},
|
||||||
|
"mike-ai-profile-controller": {
|
||||||
|
"running": true,
|
||||||
|
"health": "healthy",
|
||||||
|
"image": "sha256:a5f156d94c4921e671fafa524cf9c0fe91e1cbec0113c1f19c136204d63f243f",
|
||||||
|
"started": "2026-09-20T19:03:11.213886021Z"
|
||||||
|
},
|
||||||
|
"mike-ai-wireguard-gateway": {
|
||||||
|
"running": true,
|
||||||
|
"health": "healthy",
|
||||||
|
"image": "sha256:0d24e93c85a1c420b52b17666ede5fbd4672d92ab8dc28ea4ac5ba144ebc41d0",
|
||||||
|
"started": "2026-09-17T09:05:07.164533904Z"
|
||||||
|
},
|
||||||
|
"mike-ai-qwen3-tts": {
|
||||||
|
"running": true,
|
||||||
|
"health": "healthy",
|
||||||
|
"image": "sha256:b363a01d08b1bbecbfc3ca6f585368fae2cfdc591f9ecca6643738369f9a9d98",
|
||||||
|
"started": "2026-09-19T13:47:00.227813668Z"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"router": {
|
||||||
|
"/health": {
|
||||||
|
"status": "ok",
|
||||||
|
"router": "alive"
|
||||||
|
},
|
||||||
|
"/ready": {
|
||||||
|
"status": "ok",
|
||||||
|
"router": "alive",
|
||||||
|
"upstream": "ready"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"smoke": {
|
||||||
|
"content": "OK",
|
||||||
|
"usage": {
|
||||||
|
"completion_tokens": 2,
|
||||||
|
"prompt_tokens": 19,
|
||||||
|
"total_tokens": 21,
|
||||||
|
"prompt_tokens_details": {
|
||||||
|
"cached_tokens": 0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"kernel_errors": [],
|
||||||
|
"gpu": "name, memory.used [MiB], memory.total [MiB], temperature.gpu\nNVIDIA GeForce RTX 3060, 10930 MiB, 12288 MiB, 56\nNVIDIA GeForce RTX 5080, 15714 MiB, 16303 MiB, 61",
|
||||||
|
"disk": "Filesystem Size Used Avail Use% Mounted on\n/dev/nvme0n1p2 868G 187G 637G 23% /\n/dev/nvme1n1p1 916G 821G 49G 95% /data"
|
||||||
|
}
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
[
|
||||||
|
{
|
||||||
|
"label": "medium-slot1-priority1-draft-cuda0",
|
||||||
|
"model": "pure",
|
||||||
|
"ctx": 160000,
|
||||||
|
"single": false,
|
||||||
|
"split": "80,20",
|
||||||
|
"ubatch": 128,
|
||||||
|
"parallel": 1,
|
||||||
|
"draft_max": 7,
|
||||||
|
"draft_kv": "q4_0",
|
||||||
|
"draft_device": "CUDA0",
|
||||||
|
"quality": false,
|
||||||
|
"prompts": [
|
||||||
|
4096
|
||||||
|
],
|
||||||
|
"minimum_headroom_mib": 512
|
||||||
|
}
|
||||||
|
]
|
||||||
@@ -0,0 +1,61 @@
|
|||||||
|
{
|
||||||
|
"checked_at": 1789931253.8391275,
|
||||||
|
"uptime": "21:07:33 up 3 days, 10:02, 1 user, load average: 0.93, 0.95, 0.80",
|
||||||
|
"containers": {
|
||||||
|
"mike-ai-llama-medium": {
|
||||||
|
"running": true,
|
||||||
|
"health": "healthy",
|
||||||
|
"image": "sha256:5e3c12c145b8045e5731b44b6b97033f24b327ae3d4a3fa85ecdd159cc844907",
|
||||||
|
"started": "2026-09-20T19:07:06.261497776Z"
|
||||||
|
},
|
||||||
|
"mike-ai-router": {
|
||||||
|
"running": true,
|
||||||
|
"health": "healthy",
|
||||||
|
"image": "sha256:0448758bec6968b29263bcac0f8b4682c6d3029d626f7c12334698c11ef096bb",
|
||||||
|
"started": "2026-09-20T19:07:16.676662725Z"
|
||||||
|
},
|
||||||
|
"mike-ai-profile-controller": {
|
||||||
|
"running": true,
|
||||||
|
"health": "healthy",
|
||||||
|
"image": "sha256:a5f156d94c4921e671fafa524cf9c0fe91e1cbec0113c1f19c136204d63f243f",
|
||||||
|
"started": "2026-09-20T19:07:16.548271903Z"
|
||||||
|
},
|
||||||
|
"mike-ai-wireguard-gateway": {
|
||||||
|
"running": true,
|
||||||
|
"health": "healthy",
|
||||||
|
"image": "sha256:0d24e93c85a1c420b52b17666ede5fbd4672d92ab8dc28ea4ac5ba144ebc41d0",
|
||||||
|
"started": "2026-09-17T09:05:07.164533904Z"
|
||||||
|
},
|
||||||
|
"mike-ai-qwen3-tts": {
|
||||||
|
"running": true,
|
||||||
|
"health": "healthy",
|
||||||
|
"image": "sha256:b363a01d08b1bbecbfc3ca6f585368fae2cfdc591f9ecca6643738369f9a9d98",
|
||||||
|
"started": "2026-09-19T13:47:00.227813668Z"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"router": {
|
||||||
|
"/health": {
|
||||||
|
"status": "ok",
|
||||||
|
"router": "alive"
|
||||||
|
},
|
||||||
|
"/ready": {
|
||||||
|
"status": "ok",
|
||||||
|
"router": "alive",
|
||||||
|
"upstream": "ready"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"smoke": {
|
||||||
|
"content": "OK",
|
||||||
|
"usage": {
|
||||||
|
"completion_tokens": 2,
|
||||||
|
"prompt_tokens": 19,
|
||||||
|
"total_tokens": 21,
|
||||||
|
"prompt_tokens_details": {
|
||||||
|
"cached_tokens": 0
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"kernel_errors": [],
|
||||||
|
"gpu": "name, memory.used [MiB], memory.total [MiB], temperature.gpu\nNVIDIA GeForce RTX 3060, 10920 MiB, 12288 MiB, 45\nNVIDIA GeForce RTX 5080, 15714 MiB, 16303 MiB, 49",
|
||||||
|
"disk": "Filesystem Size Used Avail Use% Mounted on\n/dev/nvme0n1p2 868G 187G 637G 23% /\n/dev/nvme1n1p1 916G 821G 49G 95% /data"
|
||||||
|
}
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
[
|
||||||
|
{
|
||||||
|
"label": "medium-slot1-priority2-draft-cuda1",
|
||||||
|
"model": "pure",
|
||||||
|
"ctx": 160000,
|
||||||
|
"single": false,
|
||||||
|
"split": "80,20",
|
||||||
|
"ubatch": 128,
|
||||||
|
"parallel": 1,
|
||||||
|
"draft_max": 7,
|
||||||
|
"draft_kv": "q4_0",
|
||||||
|
"draft_device": "CUDA1",
|
||||||
|
"quality": false,
|
||||||
|
"prompts": [
|
||||||
|
4096
|
||||||
|
],
|
||||||
|
"minimum_headroom_mib": 512
|
||||||
|
}
|
||||||
|
]
|
||||||
@@ -0,0 +1,19 @@
|
|||||||
|
[
|
||||||
|
{
|
||||||
|
"label": "medium-slot1-priority3-draft-cuda0",
|
||||||
|
"model": "pure",
|
||||||
|
"ctx": 160000,
|
||||||
|
"single": false,
|
||||||
|
"split": "70,30",
|
||||||
|
"ubatch": 128,
|
||||||
|
"parallel": 1,
|
||||||
|
"draft_max": 7,
|
||||||
|
"draft_kv": "q4_0",
|
||||||
|
"draft_device": "CUDA0",
|
||||||
|
"quality": false,
|
||||||
|
"prompts": [
|
||||||
|
4096
|
||||||
|
],
|
||||||
|
"minimum_headroom_mib": 512
|
||||||
|
}
|
||||||
|
]
|
||||||
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
BIN
Binary file not shown.
@@ -54,8 +54,8 @@ def run_case(case):
|
|||||||
save(out/'config.json',case)
|
save(out/'config.json',case)
|
||||||
print('START',label,flush=True)
|
print('START',label,flush=True)
|
||||||
single=case.get('single',False)
|
single=case.get('single',False)
|
||||||
args=['--model','/models/'+MODELS[case['model']], '--alias','benchmark','--ctx-size',str(case['ctx']), '--flash-attn','on','--cache-type-k','q4_0','--cache-type-v','q4_0','--cache-ram','0','--threads','6','--threads-batch','6','--batch-size','2048','--ubatch-size',str(case.get('ubatch',512)), '--parallel','2','--kv-unified','--jinja','--reasoning','auto','--reasoning-preserve','--host','127.0.0.1','--port','5005','--metrics','--fit','off','--n-gpu-layers','all','--load-mode','none','--no-ui','--temperature','1.0','--top-p','0.95','--top-k','20','--device','CUDA0' if single else 'CUDA0,CUDA1','--main-gpu','0','--split-mode','layer','--tensor-split',case.get('split','1,0'),'--spec-type','draft-dflash','--spec-draft-n-max','7','--spec-draft-type-k','q4_0','--spec-draft-type-v','q4_0','--spec-draft-p-min','0.05','--verbosity','3']
|
args=['--model','/models/'+MODELS[case['model']], '--alias','benchmark','--ctx-size',str(case['ctx']), '--flash-attn','on','--cache-type-k','q4_0','--cache-type-v','q4_0','--cache-ram','0','--threads','6','--threads-batch','6','--batch-size','2048','--ubatch-size',str(case.get('ubatch',512)), '--parallel',str(case.get('parallel',2)),'--kv-unified','--jinja','--reasoning','auto','--reasoning-preserve','--host','127.0.0.1','--port','5005','--metrics','--fit','off','--n-gpu-layers','all','--load-mode','none','--no-ui','--temperature','1.0','--top-p','0.95','--top-k','20','--device','CUDA0' if single else 'CUDA0,CUDA1','--main-gpu','0','--split-mode','layer','--tensor-split',case.get('split','1,0'),'--spec-type','draft-dflash','--spec-draft-n-max','7','--spec-draft-type-k','q4_0','--spec-draft-type-v','q4_0','--spec-draft-p-min','0.05','--verbosity','3']
|
||||||
args += ['--spec-draft-model','/models/qwen38-dflash2/draft.gguf','--spec-draft-device','CUDA0','--spec-draft-ngl','all']
|
args += ['--spec-draft-model','/models/qwen38-dflash2/draft.gguf','--spec-draft-device',case.get('draft_device','CUDA0'),'--spec-draft-ngl','all']
|
||||||
save(out/'server-args.json',args)
|
save(out/'server-args.json',args)
|
||||||
cmd('docker','run','-d','--name',NAME,'--gpus','all','--network','host','--read-only','--tmpfs','/tmp:rw,nosuid,nodev,size=256m','--security-opt','no-new-privileges:true','--cap-drop','ALL','--pids-limit','512','--ulimit','core=0','--memory','26g','--memory-swap','26g','--shm-size','1g','--log-opt','max-size=32m','--log-opt','max-file=1','-e','NVIDIA_VISIBLE_DEVICES='+GPU0+','+GPU1,'-e','NVIDIA_DRIVER_CAPABILITIES=compute,utility','-v','/data/models:/models:ro',IMAGE,*args)
|
cmd('docker','run','-d','--name',NAME,'--gpus','all','--network','host','--read-only','--tmpfs','/tmp:rw,nosuid,nodev,size=256m','--security-opt','no-new-privileges:true','--cap-drop','ALL','--pids-limit','512','--ulimit','core=0','--memory','26g','--memory-swap','26g','--shm-size','1g','--log-opt','max-size=32m','--log-opt','max-file=1','-e','NVIDIA_VISIBLE_DEVICES='+GPU0+','+GPU1,'-e','NVIDIA_DRIVER_CAPABILITIES=compute,utility','-v','/data/models:/models:ro',IMAGE,*args)
|
||||||
stop=threading.Event(); samples=[]
|
stop=threading.Event(); samples=[]
|
||||||
@@ -117,6 +117,9 @@ def run_case(case):
|
|||||||
r=prefill(n,42); result['prefill'].append({'target':n,'response':r}); save(out/'partial.json',result)
|
r=prefill(n,42); result['prefill'].append({'target':n,'response':r}); save(out/'partial.json',result)
|
||||||
print(label,'prefill',n,r.get('timings'),flush=True)
|
print(label,'prefill',n,r.get('timings'),flush=True)
|
||||||
result['finished']=time.time()
|
result['finished']=time.time()
|
||||||
|
except Exception as exc:
|
||||||
|
result['error']=str(exc)
|
||||||
|
raise
|
||||||
finally:
|
finally:
|
||||||
stop.set(); thread.join(5)
|
stop.set(); thread.join(5)
|
||||||
r=subprocess.run(['docker','logs',NAME],capture_output=True,text=True,timeout=30)
|
r=subprocess.run(['docker','logs',NAME],capture_output=True,text=True,timeout=30)
|
||||||
|
|||||||
@@ -10,8 +10,12 @@ signal.signal(signal.SIGTERM,stop_signal); signal.signal(signal.SIGINT,stop_sign
|
|||||||
# Refuse if the known production state has changed, or if requests are active.
|
# Refuse if the known production state has changed, or if requests are active.
|
||||||
for name in NAMES:
|
for name in NAMES:
|
||||||
assert run('docker','inspect',name,'--format','{{.State.Running}}').stdout.strip()=='true',name
|
assert run('docker','inspect',name,'--format','{{.State.Running}}').stdout.strip()=='true',name
|
||||||
|
for attempt in range(60):
|
||||||
slots=json.loads(run('docker','exec',NAMES[-1],'curl','-fsS','http://127.0.0.1:8080/slots').stdout)
|
slots=json.loads(run('docker','exec',NAMES[-1],'curl','-fsS','http://127.0.0.1:8080/slots').stdout)
|
||||||
assert not any(s['is_processing'] for s in slots),'Production request active'
|
if not any(s['is_processing'] for s in slots): break
|
||||||
|
if attempt==0: print('WAIT production request active; no interruption',flush=True)
|
||||||
|
time.sleep(3)
|
||||||
|
else: raise RuntimeError('Production remained busy for 180s; no services stopped')
|
||||||
assert not run('docker','ps','-q','--filter','name=^mike-ai-dflash2-test$').stdout.strip(),'Existing experiment'
|
assert not run('docker','ps','-q','--filter','name=^mike-ai-dflash2-test$').stdout.strip(),'Existing experiment'
|
||||||
child=None
|
child=None
|
||||||
try:
|
try:
|
||||||
@@ -36,5 +40,11 @@ finally:
|
|||||||
for _ in range(150):
|
for _ in range(150):
|
||||||
if run('docker','inspect',NAMES[-1],'--format','{{.State.Health.Status}}').stdout.strip()=='healthy': break
|
if run('docker','inspect',NAMES[-1],'--format','{{.State.Health.Status}}').stdout.strip()=='healthy': break
|
||||||
time.sleep(2)
|
time.sleep(2)
|
||||||
|
else: raise RuntimeError('Restored Medium did not become healthy')
|
||||||
run('docker','start',NAMES[1],NAMES[0])
|
run('docker','start',NAMES[1],NAMES[0])
|
||||||
print('RESTORED existing medium/controller/router',flush=True)
|
for _ in range(60):
|
||||||
|
statuses=[run('docker','inspect',name,'--format','{{.State.Health.Status}}').stdout.strip() for name in NAMES[:2]]
|
||||||
|
if all(status=='healthy' for status in statuses): break
|
||||||
|
time.sleep(2)
|
||||||
|
else: raise RuntimeError('Restored router/controller did not become healthy')
|
||||||
|
print('RESTORED existing medium/controller/router; all healthy',flush=True)
|
||||||
|
|||||||
Reference in New Issue
Block a user