Files

100 lines
6.0 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# ByteShape GPU-5 / Pure IQ4_XS on Athena
User-authorized comparison on 2026-09-20. See
`docs/INFERENCE_OPTIMIZATION_TODO_20260920.md` for acceptance criteria.
`run.py` starts one isolated text-only llama.cpp container at a time, using the
exact production image ID, one slot, q4_0 KV and embedded MTP3 (MTP2 for matched Ultra cases). Both models use
identical prompts, temperature 1.0, top-p .95, top-k 20, min-p 0, seed 42 and
output budgets. Quality tasks request medium reasoning; throughput probes use
no reasoning on both sides and are explicitly not an intelligence test.
`supervise.py` checks that the existing production model is idle, stops only
router/controller/medium, enforces a 70-minute experiment deadline and restores
the original containers in a finally block. SSH, network, gateway, NVIDIA and
kernel are untouched. TTS remains loaded on the 3060. Host RAM is limited to
26 GiB for the test container with no container swap. GPU temperature and host
RAM reserve are checked between requests; GPU telemetry is recorded every two
seconds. A hard host/driver lock cannot be recovered by this supervisor.
Only layer split is used. Single means all model layers on CUDA0 (5080), no
vision projector. CUDA visibility is pinned by UUID. GPU memory reports include
other services, especially TTS on the 3060. Capacity tests use a single shared
KV pool, so do not interpret the result as that context per concurrent user.
Raw results and complete responses are saved incrementally under
`/data/benchmarks/byteshape-20260920/<case>/`. `gpu.json` records sampled peak
memory, not a guarantee of every instantaneous allocation. `server.log` allows
verification of actual offload and allocation fallback. Any unexpected startup
failure aborts the phase and restores production; there is no automatic retry
at progressively larger contexts.
The experiment scripts contain Athena-specific paths and image IDs. They are
not general deployment scripts. No winning profile is promoted automatically.
## Phases and limits
- Phase 1: matched 32768-token single-5080 A/B, micro-batch 512, full quality
sample and 4K/24K uncached inputs.
- Phase 2: Pure 57344 worked. ByteShape 114688 / micro-batch 512 failed during
allocation of a 180 MiB MTP compute buffer. The test container exited and the
supervisor restored production. This was a CUDA allocation failure, not a
host OOM kill or kernel panic. The measured steady-state slope alone had
underestimated transient loading requirements; this failed case is retained.
- Phase 3: micro-batch 128 capacity pilots from 32768, conservative growth
bounded to 16384 tokens per step with 768 MiB steady-state allowance for
transient buffers. Only the final case gets almost-full input validation.
Native 262144-context two-GPU cases use the previously proven Ultra settings:
layer split 80:20, micro-batch 128, MTP2 for both models.
- Phase 4: reverse-order 32K throughput repetition and additional single-GPU
validation points chosen from measured smaller-batch memory curves.
A successfully loaded pilot is not a completed long-context validation.
"Maximum" in the results always means largest **tested** context for the
specified settings, not an intentionally discovered out-of-memory boundary.
A smaller micro-batch may permit more context at the cost of prefill speed;
other KV precision, disabled MTP, smaller batches and RoPE extension are not
exhaustively searched here. Runtime and host kernel/driver remain unchanged.
Sampling explicitly sets min-p 0 in both arms (production's server default is
0.05 when clients do not override it). Thus this isolates quantizations under
the same benchmark sampling, but is not a byte-for-byte replay of all router
requests. Prefill/decode probes disable thinking identically; quality probes
keep medium reasoning. Performance outputs are intentionally capped at 512 or
768 tokens, so their finish_reason=length is expected and is not a quality
failure. The first-pass quality task budgets, in contrast, are evaluated for
whether a usable answer was produced.
The existing Fast profile is a third weight file, IQ4-MIX, not Pure IQ4_XS. Its
76800-token text configuration uses CUDA0, micro-batch 64 and MTP2; its usual
vision projector is on CUDA1. A separate text-only reference run records that
configuration without claiming a fresh full quality evaluation of IQ4-MIX.
The 86:14 ByteShape probe moves more layers to the 5080, as requested.
GGUF tensor offsets were inspected before choosing the split: layers 52–55
occupy about 686 MiB of weights, plus roughly 288 MiB for one additional full
attention layer's q4 KV at 262144 tokens (SSM state and workspace add overhead).
This gives a memory-based starting point; the full near-limit input still has
to pass before the configuration is called validated.
After the measured 86:14 load left 691 MiB free on the 5080, the near-full
input was deliberately interrupted and recorded in operator-stop.json. Its
49K probe completed, but it is NOT counted as a validated 262K run. Phase 5
uses 88:12: one further SSM layer (about 157 MiB of weights plus state) moves
to the 5080. This is an intentional refinement, separate from the unexpected
CUDA allocation failure in phase 2. The supervisor restores production after
both normal completion and an interrupted case.
Phase 5 (88:12) loaded at 15920 MiB on the 5080: only 383 MiB remained,
less than the preceding estimate. The first real text request then failed in
Flash Attention's CUDA virtual-memory allocation. The server process aborted;
the host remained reachable and the original containers were restored.
Phase 6 returns to 86:14 and performs the full validation with no further
upward split steps.
After this finding, the saved harness refuses inference when post-load free
VRAM on a used GPU is below 512 MiB. Two already validated historical cases
(Pure 57344 and the existing IQ4-MIX Fast profile) explicitly retain a 384 MiB
allowance. This guard would reject the historical 88:12 case before inference;
it is not a guarantee against every possible later workspace allocation.