Files
Athena-Deck/docs/model-discovery.md
T

1.1 KiB

Model discovery

The authenticated /v1/models chat catalog includes context_window, parallel_slots, and context_scope: shared. These values come from current enabled, runnable profiles rather than the model training limit. /models exposes the same catalog for native llama.cpp clients, including during exclusive GPU modes. In video mode /v1/models retains the original video API forwarding behavior.

GET /props?model=<profile name>&autoload=false exposes the configured context as default_generation_settings.n_ctx and n_ctx, slot count, vision-projector presence and supported known template capabilities. Queries never start, stop or load model workers. Unknown templates do not advertise tool or reasoning capabilities. Every discovery route requires the existing API bearer token.

OpenClaw's llama-cpp provider reads /health, /models and each model's /props. Explicit model rows override discovered rows; remove those only after verifying discovery and capability parity. Refresh then reads profile context changes from Deck. Shared context is not a guarantee that simultaneous slots can each use the full advertised budget.