Expose profile context metadata for llama.cpp model discovery
This commit is contained in:
1 parent
a7b30b62d2
commit
1f8ffe2e72
3 files changed
+61
-2
No files matched your search
@@ -0,0 +1,7 @@
|
||||
# Model discovery
|
||||
|
||||
The authenticated `/v1/models` chat catalog includes `context_window`, `parallel_slots`, and `context_scope: shared`. These values come from current enabled, runnable profiles rather than the model training limit. `/models` exposes the same catalog for native llama.cpp clients, including during exclusive GPU modes. In video mode `/v1/models` retains the original video API forwarding behavior.
|
||||
|
||||
`GET /props?model=<profile name>&autoload=false` exposes the configured context as `default_generation_settings.n_ctx` and `n_ctx`, slot count, vision-projector presence and supported known template capabilities. Queries never start, stop or load model workers. Unknown templates do not advertise tool or reasoning capabilities. Every discovery route requires the existing API bearer token.
|
||||
|
||||
OpenClaw's llama-cpp provider reads `/health`, `/models` and each model's `/props`. Explicit model rows override discovered rows; remove those only after verifying discovery and capability parity. Refresh then reads profile context changes from Deck. Shared context is not a guarantee that simultaneous slots can each use the full advertised budget.
|
||||
Reference in new issue
Block a user