ByteShape GPU-5 / Pure IQ4_XS on Athena
User-authorized comparison on 2026-09-20. See
docs/INFERENCE_OPTIMIZATION_TODO_20260920.md for acceptance criteria.
run.py starts one isolated text-only llama.cpp container at a time, using the
exact production image ID, one slot, q4_0 KV and embedded MTP3 (MTP2 for matched Ultra cases). Both models use
identical prompts, temperature 1.0, top-p .95, top-k 20, min-p 0, seed 42 and
output budgets. Quality tasks request medium reasoning; throughput probes use
no reasoning on both sides and are explicitly not an intelligence test.
supervise.py checks that the existing production model is idle, stops only
router/controller/medium, enforces a 70-minute experiment deadline and restores
the original containers in a finally block. SSH, network, gateway, NVIDIA and
kernel are untouched. TTS remains loaded on the 3060. Host RAM is limited to
26 GiB for the test container with no container swap. GPU temperature and host
RAM reserve are checked between requests; GPU telemetry is recorded every two
seconds. A hard host/driver lock cannot be recovered by this supervisor.
Only layer split is used. Single means all model layers on CUDA0 (5080), no vision projector. CUDA visibility is pinned by UUID. GPU memory reports include other services, especially TTS on the 3060. Capacity tests use a single shared KV pool, so do not interpret the result as that context per concurrent user.
Raw results and complete responses are saved incrementally under
/data/benchmarks/byteshape-20260920/<case>/. gpu.json records sampled peak
memory, not a guarantee of every instantaneous allocation. server.log allows
verification of actual offload and allocation fallback. Any unexpected startup
failure aborts the phase and restores production; there is no automatic retry
at progressively larger contexts.
The experiment scripts contain Athena-specific paths and image IDs. They are not general deployment scripts. No winning profile is promoted automatically.
Phases and limits
- Phase 1: matched 32768-token single-5080 A/B, micro-batch 512, full quality sample and 4K/24K uncached inputs.
- Phase 2: Pure 57344 worked. ByteShape 114688 / micro-batch 512 failed during allocation of a 180 MiB MTP compute buffer. The test container exited and the supervisor restored production. This was a CUDA allocation failure, not a host OOM kill or kernel panic. The measured steady-state slope alone had underestimated transient loading requirements; this failed case is retained.
- Phase 3: micro-batch 128 capacity pilots from 32768, conservative growth bounded to 16384 tokens per step with 768 MiB steady-state allowance for transient buffers. Only the final case gets almost-full input validation. Native 262144-context two-GPU cases use the previously proven Ultra settings: layer split 80:20, micro-batch 128, MTP2 for both models.
- Phase 4: reverse-order 32K throughput repetition and additional single-GPU validation points chosen from measured smaller-batch memory curves.
A successfully loaded pilot is not a completed long-context validation. "Maximum" in the results always means largest tested context for the specified settings, not an intentionally discovered out-of-memory boundary. A smaller micro-batch may permit more context at the cost of prefill speed; other KV precision, disabled MTP, smaller batches and RoPE extension are not exhaustively searched here. Runtime and host kernel/driver remain unchanged.
Sampling explicitly sets min-p 0 in both arms (production's server default is 0.05 when clients do not override it). Thus this isolates quantizations under the same benchmark sampling, but is not a byte-for-byte replay of all router requests. Prefill/decode probes disable thinking identically; quality probes keep medium reasoning. Performance outputs are intentionally capped at 512 or 768 tokens, so their finish_reason=length is expected and is not a quality failure. The first-pass quality task budgets, in contrast, are evaluated for whether a usable answer was produced.
The existing Fast profile is a third weight file, IQ4-MIX, not Pure IQ4_XS. Its 76800-token text configuration uses CUDA0, micro-batch 64 and MTP2; its usual vision projector is on CUDA1. A separate text-only reference run records that configuration without claiming a fresh full quality evaluation of IQ4-MIX.
The 86:14 ByteShape probe moves more layers to the 5080, as requested. GGUF tensor offsets were inspected before choosing the split: layers 52–55 occupy about 686 MiB of weights, plus roughly 288 MiB for one additional full attention layer's q4 KV at 262144 tokens (SSM state and workspace add overhead). This gives a memory-based starting point; the full near-limit input still has to pass before the configuration is called validated.
After the measured 86:14 load left 691 MiB free on the 5080, the near-full input was deliberately interrupted and recorded in operator-stop.json. Its 49K probe completed, but it is NOT counted as a validated 262K run. Phase 5 uses 88:12: one further SSM layer (about 157 MiB of weights plus state) moves to the 5080. This is an intentional refinement, separate from the unexpected CUDA allocation failure in phase 2. The supervisor restores production after both normal completion and an interrupted case.
Phase 5 (88:12) loaded at 15920 MiB on the 5080: only 383 MiB remained, less than the preceding estimate. The first real text request then failed in Flash Attention's CUDA virtual-memory allocation. The server process aborted; the host remained reachable and the original containers were restored. Phase 6 returns to 86:14 and performs the full validation with no further upward split steps.
After this finding, the saved harness refuses inference when post-load free VRAM on a used GPU is below 512 MiB. Two already validated historical cases (Pure 57344 and the existing IQ4-MIX Fast profile) explicitly retain a 384 MiB allowance. This guard would reject the historical 88:12 case before inference; it is not a guarantee against every possible later workspace allocation.