3.1 KiB
DFlash2 Medium-only feasibility probe, 2026-09-20
User requested one quick indication of whether deeper tests are worthwhile.
No new Qwen baseline run. Reference: benchmarks/athena-qwen38-reference-20260920/.
Target: existing Pure IQ4_XS, exact existing llama.cpp image b29c606. Runtime help advertises draft-dflash; libllama contains build_dflash2_conv and build_dflash2_selector. No rebuild, driver or kernel changes.
Draft: incoai/Qwen3.8-27B-DFlash2-GGUF, revision
51962825493a48b846b40126d35c799ac4093ad0, Q4_K_M, 1,143,006,816 bytes.
SHA256: 1a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd.
Source: https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF
Runtime source inspected:
https://github.com/ggml-org/llama.cpp/blob/b29c606/common/speculative.cpp
Scope
160000 shared context, two configured slots, one sequential request at a time. Pure target split 80:20, draft entirely on CUDA0/5080, q4_0 KV for target and draft, microbatch128, draft7, batch2048, text only/no projector, cache RAM0. TTS remains resident on 3060. These are changes from production Medium's 85:15/MTP3/ub512/vision/cache RAM32768. The test is not full Medium equivalence. Maximum context is allocated, not validated by a full-length input in this quick probe. No concurrency, vision or broad quality validation is claimed.
Two existing decode prompts (German/code,768 output-token cap) and one frozen 4196-input-token retrieval/prefill prompt (512 output-token cap), plus health smoke. Sampling matches the frozen requests: temp1/top-p.95/top-k20/min-p0, seed42 (code43), no reasoning. Save all responses, timings and GPU samples.
The saved dual-GPU Pure/MTP reference is a useful orientation, but uses 262144 context and one slot. The single-GPU reference uses 32768 context/ub512/MTP3. Any ratios are whole-configuration comparisons, not isolated DFlash speedups. Same seed under speculative sampling does not guarantee identical output text.
Safety / restoration
Adapted from the prior bounded isolated harness. No changes to production container arguments or images. Supervisor drains and stops router/controller/ Medium only and restores the same containers in finally. Ten-minute child limit; read-only models,26GiB container RAM limit without extra swap, no privileges, no core files. At least512MiB free per used GPU required before inference; 85C temperature and3GiB host RAM limits monitored. A failed case is not retried with increasingly aggressive settings. No automated invocation on repo checkout.
Commands, case configuration and results are archived here. Scripts have fixed Athena paths; do not run them as generic unit tests or rerun without a concrete benchmark task. This probe does not prove quality equivalence even if responses look reasonable; that requires subsequent verification.
Observed outcome
Target loaded; draft allocation of1079.61MiB onCUDA0 failed with CUDA malloc OOM. No benchmark inference executed. Supervisor restored original production; health/smoke/kernel checks passed. No performance or quality conclusion can be drawn. See results/ and restore-verification.json.