Make reasoning levels enforce real token budgets

This commit is contained in:
Mikei386
2026-09-01 13:55:11 +02:00
parent 5f793020b0
commit b3e86cc7ae
5 changed files with 52 additions and 21 deletions
+2 -5
View File
@@ -194,11 +194,8 @@ services:
- --jinja
- --reasoning
- auto
# Bound each individual thinking phase. Long agent jobs can still use
# many phases around tool calls, but one degenerate reasoning loop can
# no longer consume the complete response budget indefinitely.
- --reasoning-budget
- "8192"
# No fixed --reasoning-budget here: the router supplies a real budget
# per request from the client's reasoning_effort selection.
- --reasoning-preserve
- --host
- 0.0.0.0