Fix Athena recovery and installation consistency
This commit is contained in:
@@ -21,9 +21,12 @@ Full-context results with the selected settings:
|
||||
|
||||
The larger logical batches 3072 and 4096 did not improve Medium at ubatch 128. The original one-slot benchmark selected 2048 / 128 at 90:10.
|
||||
|
||||
## Medium two-slot benchmark
|
||||
## Historischer Medium-Zwei-Slot-Test
|
||||
|
||||
Medium now uses two parallel slots with unified KV, so both chats dynamically share one total 160K-token pool. The model weights remain loaded only once. To fit the additional scheduler buffers, the production GPU split is 85:15 while batch / ubatch remains 2048 / 128.
|
||||
This was an A/B candidate, not the current production configuration. Production
|
||||
was returned to **one slot** because concurrent Hermes requests did not behave
|
||||
reliably enough. The 85:15 GPU split and batch / ubatch 2048 / 128 remain in
|
||||
production because they also work with the single-slot profile.
|
||||
|
||||
Identical fresh 100,297-token prompt with a deterministic 256-token completion:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user