Files
AI-Profile-Router/experiments/dflash2-medium-20260920/README.md
T

3.1 KiB

DFlash2 Medium-only feasibility probe, 2026-09-20

User requested one quick indication of whether deeper tests are worthwhile. No new Qwen baseline run. Reference: benchmarks/athena-qwen38-reference-20260920/.

Target: existing Pure IQ4_XS, exact existing llama.cpp image b29c606. Runtime help advertises draft-dflash; libllama contains build_dflash2_conv and build_dflash2_selector. No rebuild, driver or kernel changes.

Draft: incoai/Qwen3.8-27B-DFlash2-GGUF, revision 51962825493a48b846b40126d35c799ac4093ad0, Q4_K_M, 1,143,006,816 bytes. SHA256: 1a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd. Source: https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF Runtime source inspected: https://github.com/ggml-org/llama.cpp/blob/b29c606/common/speculative.cpp

Scope

160000 shared context, two configured slots, one sequential request at a time. Pure target split 80:20, draft entirely on CUDA0/5080, q4_0 KV for target and draft, microbatch128, draft7, batch2048, text only/no projector, cache RAM0. TTS remains resident on 3060. These are changes from production Medium's 85:15/MTP3/ub512/vision/cache RAM32768. The test is not full Medium equivalence. Maximum context is allocated, not validated by a full-length input in this quick probe. No concurrency, vision or broad quality validation is claimed.

Two existing decode prompts (German/code,768 output-token cap) and one frozen 4196-input-token retrieval/prefill prompt (512 output-token cap), plus health smoke. Sampling matches the frozen requests: temp1/top-p.95/top-k20/min-p0, seed42 (code43), no reasoning. Save all responses, timings and GPU samples.

The saved dual-GPU Pure/MTP reference is a useful orientation, but uses 262144 context and one slot. The single-GPU reference uses 32768 context/ub512/MTP3. Any ratios are whole-configuration comparisons, not isolated DFlash speedups. Same seed under speculative sampling does not guarantee identical output text.

Safety / restoration

Adapted from the prior bounded isolated harness. No changes to production container arguments or images. Supervisor drains and stops router/controller/ Medium only and restores the same containers in finally. Ten-minute child limit; read-only models,26GiB container RAM limit without extra swap, no privileges, no core files. At least512MiB free per used GPU required before inference; 85C temperature and3GiB host RAM limits monitored. A failed case is not retried with increasingly aggressive settings. No automated invocation on repo checkout.

Commands, case configuration and results are archived here. Scripts have fixed Athena paths; do not run them as generic unit tests or rerun without a concrete benchmark task. This probe does not prove quality equivalence even if responses look reasonable; that requires subsequent verification.

Observed outcome

Target loaded; draft allocation of1079.61MiB onCUDA0 failed with CUDA malloc OOM. No benchmark inference executed. Supervisor restored original production; health/smoke/kernel checks passed. No performance or quality conclusion can be drawn. See results/ and restore-verification.json.