Record bounded Medium DFlash2 feasibility test and VRAM limit
This commit is contained in:
1 parent
980f339ea4
commit
3d5931146b
13 files changed
+422
No files matched your search
@@ -0,0 +1,57 @@
|
||||
# DFlash2 Medium-only feasibility probe, 2026-09-20
|
||||
|
||||
User requested one quick indication of whether deeper tests are worthwhile.
|
||||
No new Qwen baseline run. Reference: `benchmarks/athena-qwen38-reference-20260920/`.
|
||||
|
||||
Target: existing Pure IQ4_XS, exact existing llama.cpp image b29c606. Runtime
|
||||
help advertises draft-dflash; libllama contains build_dflash2_conv and
|
||||
build_dflash2_selector. No rebuild, driver or kernel changes.
|
||||
|
||||
Draft: incoai/Qwen3.8-27B-DFlash2-GGUF, revision
|
||||
`51962825493a48b846b40126d35c799ac4093ad0`, Q4_K_M, 1,143,006,816 bytes.
|
||||
SHA256: `1a25c56858e1ebe93f2718ac1d49d1151f9323325c1bbfd6209370f4db131ebd`.
|
||||
Source: https://huggingface.co/incoai/Qwen3.8-27B-DFlash2-GGUF
|
||||
Runtime source inspected:
|
||||
https://github.com/ggml-org/llama.cpp/blob/b29c606/common/speculative.cpp
|
||||
|
||||
## Scope
|
||||
|
||||
160000 shared context, two configured slots, **one sequential request at a time**.
|
||||
Pure target split 80:20, draft entirely on CUDA0/5080, q4_0 KV for target and
|
||||
draft, microbatch128, draft7, batch2048, text only/no projector, cache RAM0.
|
||||
TTS remains resident on 3060. These are changes from production Medium's
|
||||
85:15/MTP3/ub512/vision/cache RAM32768. The test is not full Medium equivalence.
|
||||
Maximum context is allocated, not validated by a full-length input in this quick
|
||||
probe. No concurrency, vision or broad quality validation is claimed.
|
||||
|
||||
Two existing decode prompts (German/code,768 output-token cap) and one frozen
|
||||
4196-input-token retrieval/prefill prompt (512 output-token cap), plus health
|
||||
smoke. Sampling matches the frozen requests: temp1/top-p.95/top-k20/min-p0,
|
||||
seed42 (code43), no reasoning. Save all responses, timings and GPU samples.
|
||||
|
||||
The saved dual-GPU Pure/MTP reference is a useful orientation, but uses 262144
|
||||
context and one slot. The single-GPU reference uses 32768 context/ub512/MTP3.
|
||||
Any ratios are whole-configuration comparisons, not isolated DFlash speedups.
|
||||
Same seed under speculative sampling does not guarantee identical output text.
|
||||
|
||||
## Safety / restoration
|
||||
|
||||
Adapted from the prior bounded isolated harness. No changes to production
|
||||
container arguments or images. Supervisor drains and stops router/controller/
|
||||
Medium only and restores the same containers in finally. Ten-minute child limit;
|
||||
read-only models,26GiB container RAM limit without extra swap, no privileges,
|
||||
no core files. At least512MiB free per used GPU required before inference;
|
||||
85C temperature and3GiB host RAM limits monitored. A failed case is not retried
|
||||
with increasingly aggressive settings. No automated invocation on repo checkout.
|
||||
|
||||
Commands, case configuration and results are archived here. Scripts have fixed
|
||||
Athena paths; do not run them as generic unit tests or rerun without a concrete
|
||||
benchmark task. This probe does not prove quality equivalence even if responses
|
||||
look reasonable; that requires subsequent verification.
|
||||
|
||||
## Observed outcome
|
||||
|
||||
Target loaded; draft allocation of1079.61MiB onCUDA0 failed with CUDA malloc
|
||||
OOM. No benchmark inference executed. Supervisor restored original production;
|
||||
health/smoke/kernel checks passed. No performance or quality conclusion can be
|
||||
drawn. See results/ and restore-verification.json.
|
||||
Reference in new issue
Block a user