87 lines
4.6 KiB
Markdown
87 lines
4.6 KiB
Markdown
# DFlash2 Medium-only feasibility probe, 2026-09-20
|
|
|
|
User requested one quick indication of whether deeper tests are worthwhile.
|
|
No new Qwen baseline run. Reference: `benchmarks/athena-qwen38-reference-20260920/`.
|
|
|
|
Target: existing Pure IQ4_XS, exact existing llama.cpp image b29c606. Runtime
|
|
help advertises draft-dflash; libllama contains build_dflash2_conv and
|
|
build_dflash2_selector. No rebuild, driver or kernel changes.
|
|
|
|
Draft: incoai/Qwen3.8-27B-DFlash2-GGUF, revision
|
|
`51962825493a48b846b40126d35c799ac4093ad0`, Q4_K_M, 1,143,006,816 bytes.
|
|
SHA256: `1a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd`.
|
|
Source: https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF
|
|
Runtime source inspected:
|
|
https://github.com/ggml-org/llama.cpp/blob/b29c606/common/speculative.cpp
|
|
|
|
## Scope
|
|
|
|
160000 shared context, two configured slots, **one sequential request at a time**.
|
|
Pure target split 80:20, draft entirely on CUDA0/5080, q4_0 KV for target and
|
|
draft, microbatch128, draft7, batch2048, text only/no projector, cache RAM0.
|
|
TTS remains resident on 3060. These are changes from production Medium's
|
|
85:15/MTP3/ub512/vision/cache RAM32768. The test is not full Medium equivalence.
|
|
Maximum context is allocated, not validated by a full-length input in this quick
|
|
probe. No concurrency, vision or broad quality validation is claimed.
|
|
|
|
Two existing decode prompts (German/code,768 output-token cap) and one frozen
|
|
4196-input-token retrieval/prefill prompt (512 output-token cap), plus health
|
|
smoke. Sampling matches the frozen requests: temp1/top-p.95/top-k20/min-p0,
|
|
seed42 (code43), no reasoning. Save all responses, timings and GPU samples.
|
|
|
|
The saved dual-GPU Pure/MTP reference is a useful orientation, but uses 262144
|
|
context and one slot. The single-GPU reference uses 32768 context/ub512/MTP3.
|
|
Any ratios are whole-configuration comparisons, not isolated DFlash speedups.
|
|
Same seed under speculative sampling does not guarantee identical output text.
|
|
|
|
## Safety / restoration
|
|
|
|
Adapted from the prior bounded isolated harness. No changes to production
|
|
container arguments or images. Supervisor drains and stops router/controller/
|
|
Medium only and restores the same containers in finally. Ten-minute child limit;
|
|
read-only models,26GiB container RAM limit without extra swap, no privileges,
|
|
no core files. At least512MiB free per used GPU required before inference;
|
|
85C temperature and3GiB host RAM limits monitored. A failed case is not retried
|
|
with increasingly aggressive settings. No automated invocation on repo checkout.
|
|
|
|
Commands, case configuration and results are archived here. Scripts have fixed
|
|
Athena paths; do not run them as generic unit tests or rerun without a concrete
|
|
benchmark task. This probe does not prove quality equivalence even if responses
|
|
look reasonable; that requires subsequent verification.
|
|
|
|
## Observed outcome
|
|
|
|
Target loaded; draft allocation of1079.61MiB onCUDA0 failed with CUDA malloc
|
|
OOM. No benchmark inference executed. Supervisor restored original production;
|
|
health/smoke/kernel checks passed. No performance or quality conclusion can be
|
|
drawn. See results/ and restore-verification.json.
|
|
|
|
## User-ordered one-slot follow-up
|
|
|
|
The user authorized this order after the initial two-slot allocation failure:
|
|
|
|
1. One slot, target80:20, draft on5080.
|
|
2. One slot, target80:20, draft on3060, regardless of whether the first works.
|
|
3. Only if step2 fails: one slot, target70:30, draft back on5080. Moving both
|
|
additional target layers and the draft onto3060 would compete for its memory;
|
|
this fallback instead makes space on5080 for the draft.
|
|
|
|
Context remains160000 and other quick-probe settings are unchanged. These tests
|
|
are temporary; production containers retain their original two-slot arguments.
|
|
Case JSON now controls slot count and draft device (historical quick.json keeps
|
|
its original two-slot meaning). Failed attempts include an explicit error field.
|
|
|
|
An active production request postpones startup for up to180 seconds; no live
|
|
request is deliberately interrupted. The existing drain check still protects
|
|
requests racing with the transition. Each attempt has its own result directory;
|
|
previous results and the frozen Qwen reference are never overwritten.
|
|
|
|
## One-slot results
|
|
|
|
Priority1 fails during draft graph initialization: output.weight resides onCUDA1
|
|
but draft backends are restricted toCUDA0. Priority2 works; priority3 therefore
|
|
not run. German37.79 tok/s, code75.51 tok/s, 4196-token prefill1437.56 tok/s,
|
|
recall3/3. Allocation160000 is not full-context validation. Original two-slot
|
|
production restored and checked after both cases. See
|
|
../../docs/DFLASH2_ONE_SLOT_20260920.md for comparison caveats and quality findings.
|