Make reasoning levels enforce real token budgets
This commit is contained in:
+2
-5
@@ -194,11 +194,8 @@ services:
|
||||
- --jinja
|
||||
- --reasoning
|
||||
- auto
|
||||
# Bound each individual thinking phase. Long agent jobs can still use
|
||||
# many phases around tool calls, but one degenerate reasoning loop can
|
||||
# no longer consume the complete response budget indefinitely.
|
||||
- --reasoning-budget
|
||||
- "8192"
|
||||
# No fixed --reasoning-budget here: the router supplies a real budget
|
||||
# per request from the client's reasoning_effort selection.
|
||||
- --reasoning-preserve
|
||||
- --host
|
||||
- 0.0.0.0
|
||||
|
||||
Reference in New Issue
Block a user