Files

2.6 KiB

Medium microbatch quick A/B, 2026-09-20

User-authorized first-priority test: microbatch512 versus256, all other Medium model settings unchanged. Three sequential requests per arm: frozen German and code decode prompts (768-token caps), then frozen24576-target prefill/recall (24674 actual prompt tokens,512 output-token cap). No reasoning in these speed probes; same sampling/seeds as archived requests. Small matched512 control is necessary because the frozen baseline has different slot/vision/context settings; this is not a repeat of the broad Qwen benchmark or an overwrite of that baseline.

Supervisor snapshots actual production image/Args into production.json. Runner copies them, changes only microbatch and isolated bind address/port, and sets log verbosity3. Vision remains loaded on3060; context160000 shared by two slots, Pure model,85:15 split,MTP3, q4_0 target KV, f16 MTP KV, batch2048, six CPU threads, cache-RAM32768, TTS resident. Requests disable prompt cache for cold measurement. No concurrency/vision task/full-context validation in this small speed probe.

Bounded600-second child; production requests drain before stopping the existing router/controller/Medium. Isolated read-only model mount,26GiB memory cap with no additional swap allowance,512MiB loaded VRAM reserve,85C/3GiB host RAM guard, no core dumps. Supervisor restores the same production containers and waits for healthy status. Kernel/network/driver untouched. Each result directory is unique and cannot be overwritten. New scripts are not run automatically from checkout.

Check startup logs for successful allocation versus explicit fallback without pipeline parallelism. Absence of a fallback line alone is not proof of actual pipeline utilization; source/verbose logs may be needed to establish activation. Judge walltime, prompt/decode rates, VRAM and the three-needle recall together.

Outcome

512 control completed;256 loaded without the earlier fallback message but had only461MiB free on5080, below the512MiB guard. No256 inference was executed. Args differ only in --ubatch-size. Production512 restored and checked. See ../../docs/MEDIUM_MICROBATCH_20260920.md. Do not claim a256 speedup or quality result.

Explicitly authorized lower-reserve follow-up

reserve448.json lowers only this256 case to448MiB loaded free-memory minimum. It passed all three short probes: +7.1% prefill, short decodes essentially tied. Recorded peak free memory443MiB; the guard is a startup threshold, not a reserved allocation or runtime minimum. No full160K/vision/concurrency qualification. Production512 restored. See reserve448-comparison.json and reserve448-restore.json.