Preserve Athena Qwen baseline and document ByteShape comparison
This commit is contained in:
@@ -0,0 +1,99 @@
|
||||
# ByteShape GPU-5 / Pure IQ4_XS on Athena
|
||||
|
||||
User-authorized comparison on 2026-09-20. See
|
||||
`docs/INFERENCE_OPTIMIZATION_TODO_20260920.md` for acceptance criteria.
|
||||
|
||||
`run.py` starts one isolated text-only llama.cpp container at a time, using the
|
||||
exact production image ID, one slot, q4_0 KV and embedded MTP3 (MTP2 for matched Ultra cases). Both models use
|
||||
identical prompts, temperature 1.0, top-p .95, top-k 20, min-p 0, seed 42 and
|
||||
output budgets. Quality tasks request medium reasoning; throughput probes use
|
||||
no reasoning on both sides and are explicitly not an intelligence test.
|
||||
|
||||
`supervise.py` checks that the existing production model is idle, stops only
|
||||
router/controller/medium, enforces a 70-minute experiment deadline and restores
|
||||
the original containers in a finally block. SSH, network, gateway, NVIDIA and
|
||||
kernel are untouched. TTS remains loaded on the 3060. Host RAM is limited to
|
||||
26 GiB for the test container with no container swap. GPU temperature and host
|
||||
RAM reserve are checked between requests; GPU telemetry is recorded every two
|
||||
seconds. A hard host/driver lock cannot be recovered by this supervisor.
|
||||
|
||||
Only layer split is used. Single means all model layers on CUDA0 (5080), no
|
||||
vision projector. CUDA visibility is pinned by UUID. GPU memory reports include
|
||||
other services, especially TTS on the 3060. Capacity tests use a single shared
|
||||
KV pool, so do not interpret the result as that context per concurrent user.
|
||||
|
||||
Raw results and complete responses are saved incrementally under
|
||||
`/data/benchmarks/byteshape-20260920/<case>/`. `gpu.json` records sampled peak
|
||||
memory, not a guarantee of every instantaneous allocation. `server.log` allows
|
||||
verification of actual offload and allocation fallback. Any unexpected startup
|
||||
failure aborts the phase and restores production; there is no automatic retry
|
||||
at progressively larger contexts.
|
||||
|
||||
The experiment scripts contain Athena-specific paths and image IDs. They are
|
||||
not general deployment scripts. No winning profile is promoted automatically.
|
||||
|
||||
## Phases and limits
|
||||
|
||||
- Phase 1: matched 32768-token single-5080 A/B, micro-batch 512, full quality
|
||||
sample and 4K/24K uncached inputs.
|
||||
- Phase 2: Pure 57344 worked. ByteShape 114688 / micro-batch 512 failed during
|
||||
allocation of a 180 MiB MTP compute buffer. The test container exited and the
|
||||
supervisor restored production. This was a CUDA allocation failure, not a
|
||||
host OOM kill or kernel panic. The measured steady-state slope alone had
|
||||
underestimated transient loading requirements; this failed case is retained.
|
||||
- Phase 3: micro-batch 128 capacity pilots from 32768, conservative growth
|
||||
bounded to 16384 tokens per step with 768 MiB steady-state allowance for
|
||||
transient buffers. Only the final case gets almost-full input validation.
|
||||
Native 262144-context two-GPU cases use the previously proven Ultra settings:
|
||||
layer split 80:20, micro-batch 128, MTP2 for both models.
|
||||
- Phase 4: reverse-order 32K throughput repetition and additional single-GPU
|
||||
validation points chosen from measured smaller-batch memory curves.
|
||||
|
||||
A successfully loaded pilot is not a completed long-context validation.
|
||||
"Maximum" in the results always means largest **tested** context for the
|
||||
specified settings, not an intentionally discovered out-of-memory boundary.
|
||||
A smaller micro-batch may permit more context at the cost of prefill speed;
|
||||
other KV precision, disabled MTP, smaller batches and RoPE extension are not
|
||||
exhaustively searched here. Runtime and host kernel/driver remain unchanged.
|
||||
|
||||
Sampling explicitly sets min-p 0 in both arms (production's server default is
|
||||
0.05 when clients do not override it). Thus this isolates quantizations under
|
||||
the same benchmark sampling, but is not a byte-for-byte replay of all router
|
||||
requests. Prefill/decode probes disable thinking identically; quality probes
|
||||
keep medium reasoning. Performance outputs are intentionally capped at 512 or
|
||||
768 tokens, so their finish_reason=length is expected and is not a quality
|
||||
failure. The first-pass quality task budgets, in contrast, are evaluated for
|
||||
whether a usable answer was produced.
|
||||
|
||||
The existing Fast profile is a third weight file, IQ4-MIX, not Pure IQ4_XS. Its
|
||||
76800-token text configuration uses CUDA0, micro-batch 64 and MTP2; its usual
|
||||
vision projector is on CUDA1. A separate text-only reference run records that
|
||||
configuration without claiming a fresh full quality evaluation of IQ4-MIX.
|
||||
|
||||
The 86:14 ByteShape probe moves more layers to the 5080, as requested.
|
||||
GGUF tensor offsets were inspected before choosing the split: layers 52–55
|
||||
occupy about 686 MiB of weights, plus roughly 288 MiB for one additional full
|
||||
attention layer's q4 KV at 262144 tokens (SSM state and workspace add overhead).
|
||||
This gives a memory-based starting point; the full near-limit input still has
|
||||
to pass before the configuration is called validated.
|
||||
|
||||
After the measured 86:14 load left 691 MiB free on the 5080, the near-full
|
||||
input was deliberately interrupted and recorded in operator-stop.json. Its
|
||||
49K probe completed, but it is NOT counted as a validated 262K run. Phase 5
|
||||
uses 88:12: one further SSM layer (about 157 MiB of weights plus state) moves
|
||||
to the 5080. This is an intentional refinement, separate from the unexpected
|
||||
CUDA allocation failure in phase 2. The supervisor restores production after
|
||||
both normal completion and an interrupted case.
|
||||
|
||||
Phase 5 (88:12) loaded at 15920 MiB on the 5080: only 383 MiB remained,
|
||||
less than the preceding estimate. The first real text request then failed in
|
||||
Flash Attention's CUDA virtual-memory allocation. The server process aborted;
|
||||
the host remained reachable and the original containers were restored.
|
||||
Phase 6 returns to 86:14 and performs the full validation with no further
|
||||
upward split steps.
|
||||
|
||||
After this finding, the saved harness refuses inference when post-load free
|
||||
VRAM on a used GPU is below 512 MiB. Two already validated historical cases
|
||||
(Pure 57344 and the existing IQ4-MIX Fast profile) explicitly retain a 384 MiB
|
||||
allowance. This guard would reject the historical 88:12 case before inference;
|
||||
it is not a guarantee against every possible later workspace allocation.
|
||||
Reference in New Issue
Block a user