Fix Athena recovery and installation consistency

This commit is contained in:
Mikei386
2026-09-01 10:19:13 +02:00
parent 7e14da77f9
commit e30f4f024c
10 changed files with 179 additions and 27 deletions
+5 -2
View File
@@ -21,9 +21,12 @@ Full-context results with the selected settings:
The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. The original one-slot benchmark selected 2048 / 128 at 90:10.
## Medium two-slot benchmark
## Historischer Medium-Zwei-Slot-Test
Medium now uses two parallel slots with unified KV, so both chats dynamically share one total 160K-token pool. The model weights remain loaded only once. To fit the additional scheduler buffers, the production GPU split is 85:15 while batch / ubatch remains 2048 / 128.
This was an A/B candidate, not the current production configuration. Production
was returned to **one slot** because concurrent Hermes requests did not behave
reliably enough. The 85:15 GPU split and batch / ubatch 2048 / 128 remain in
production because they also work with the single-slot profile.
Identical fresh 100,297-token prompt with a deterministic 256-token completion: