2.9 KiB
Client-neutral chat compatibility
Deck exposes authenticated OpenAI-style Chat Completions at /v1/chat/completions.
It is a router, not a complete replica of OpenAI or of every llama-server route.
The compatibility layer is api_compat.py; it does not modify model profiles.
Discovery and behaviour
GET /v1/capabilities with the same Bearer token lists accepted request fields,
llama.cpp extensions, aliases and routes not implemented by Deck. This describes
router support, not a guarantee that every model/build supports every function.
Existing /v1/models remains a chat-only list. Images, speech and transcription
have their own model-list routes. Request handling and streaming share the same
normalization path; client-specific branches are not used.
max_completion_tokensbecomesmax_tokens; conflicting limits fail explicitly.- Legacy
functionsandfunction_callbecometoolsandtool_choice. - SDK
extra_bodyis flattened, validated and cannot replace model/messages. - Null optional stop/tools/stream options are omitted.
reasoning_effortaliases retain existing nearest-supported mappings.- Template options
enable_thinkingandpreserve_thinkingrequire booleans. - Additional llama.cpp options:
min_p,typical_p,repeat_last_n,mirostat,mirostat_tau,mirostat_eta,dynatemp_range,dynatemp_exponent,cache_prompt,ignore_eos,reasoning_format, andlogit_bias. Numeric types, finite values and bounded ranges are validated. - Metadata, safety identifiers and prompt-cache keys are accepted as annotations,
but not persisted or used for OpenAI-hosted safety/caching semantics. Their
removal is reported in
X-Athena-Compatibility-Ignored(also on SSE responses). store=false/nulland service tierauto/default/nullare tolerated and reported; stored completions and priority service tiers are rejected, not simulated.
Unknown fields and conflicting aliases still produce explicit HTTP 400 errors.
Only one completion (n=1) is supported. Embeddings, Responses, native completion,
slot administration and arbitrary template overrides are not added by this change.
Tools, vision, reasoning and structured output depend on model weights, template
and active runtime version. Tool calls are returned to clients, not executed by Deck.
Model capability declarations must not be inferred solely from a model name.
Verification
Tests cover alias translation, conflicts, input isolation, invalid sampling values, metadata treatment, upstream forwarding and authenticated capability discovery. Synthetic/fake-worker API tests do not establish inference quality or actual tool calling for every installed model. New upstream features require review and tests.
References: llama.cpp server API and OpenAI Chat Completions.