Files
..

DFlash2 Medium-only feasibility probe, 2026-09-20

User requested one quick indication of whether deeper tests are worthwhile. No new Qwen baseline run. Reference: benchmarks/athena-qwen38-reference-20260920/.

Target: existing Pure IQ4_XS, exact existing llama.cpp image b29c606. Runtime help advertises draft-dflash; libllama contains build_dflash2_conv and build_dflash2_selector. No rebuild, driver or kernel changes.

Draft: incoai/Qwen3.8-27B-DFlash2-GGUF, revision 51962825493a48b846b40126d35c799ac4093ad0, Q4_K_M, 1,143,006,816 bytes. SHA256: 1a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd. Source: https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF Runtime source inspected: https://github.com/ggml-org/llama.cpp/blob/b29c606/common/speculative.cpp

Scope

160000 shared context, two configured slots, one sequential request at a time. Pure target split 80:20, draft entirely on CUDA0/5080, q4_0 KV for target and draft, microbatch128, draft7, batch2048, text only/no projector, cache RAM0. TTS remains resident on 3060. These are changes from production Medium's 85:15/MTP3/ub512/vision/cache RAM32768. The test is not full Medium equivalence. Maximum context is allocated, not validated by a full-length input in this quick probe. No concurrency, vision or broad quality validation is claimed.

Two existing decode prompts (German/code,768 output-token cap) and one frozen 4196-input-token retrieval/prefill prompt (512 output-token cap), plus health smoke. Sampling matches the frozen requests: temp1/top-p.95/top-k20/min-p0, seed42 (code43), no reasoning. Save all responses, timings and GPU samples.

The saved dual-GPU Pure/MTP reference is a useful orientation, but uses 262144 context and one slot. The single-GPU reference uses 32768 context/ub512/MTP3. Any ratios are whole-configuration comparisons, not isolated DFlash speedups. Same seed under speculative sampling does not guarantee identical output text.

Safety / restoration

Adapted from the prior bounded isolated harness. No changes to production container arguments or images. Supervisor drains and stops router/controller/ Medium only and restores the same containers in finally. Ten-minute child limit; read-only models,26GiB container RAM limit without extra swap, no privileges, no core files. At least512MiB free per used GPU required before inference; 85C temperature and3GiB host RAM limits monitored. A failed case is not retried with increasingly aggressive settings. No automated invocation on repo checkout.

Commands, case configuration and results are archived here. Scripts have fixed Athena paths; do not run them as generic unit tests or rerun without a concrete benchmark task. This probe does not prove quality equivalence even if responses look reasonable; that requires subsequent verification.

Observed outcome

Target loaded; draft allocation of1079.61MiB onCUDA0 failed with CUDA malloc OOM. No benchmark inference executed. Supervisor restored original production; health/smoke/kernel checks passed. No performance or quality conclusion can be drawn. See results/ and restore-verification.json.

User-ordered one-slot follow-up

The user authorized this order after the initial two-slot allocation failure:

  1. One slot, target80:20, draft on5080.
  2. One slot, target80:20, draft on3060, regardless of whether the first works.
  3. Only if step2 fails: one slot, target70:30, draft back on5080. Moving both additional target layers and the draft onto3060 would compete for its memory; this fallback instead makes space on5080 for the draft.

Context remains160000 and other quick-probe settings are unchanged. These tests are temporary; production containers retain their original two-slot arguments. Case JSON now controls slot count and draft device (historical quick.json keeps its original two-slot meaning). Failed attempts include an explicit error field.

An active production request postpones startup for up to180 seconds; no live request is deliberately interrupted. The existing drain check still protects requests racing with the transition. Each attempt has its own result directory; previous results and the frozen Qwen reference are never overwritten.

One-slot results

Priority1 fails during draft graph initialization: output.weight resides onCUDA1 but draft backends are restricted toCUDA0. Priority2 works; priority3 therefore not run. German37.79 tok/s, code75.51 tok/s, 4196-token prefill1437.56 tok/s, recall3/3. Allocation160000 is not full-context validation. Original two-slot production restored and checked after both cases. See ../../docs/DFLASH2_ONE_SLOT_20260920.md for comparison caveats and quality findings.