Files
Athena-Deck/API_COMPATIBILITY.md
T

48 lines
2.9 KiB
Markdown

# Client-neutral chat compatibility
Deck exposes authenticated OpenAI-style Chat Completions at `/v1/chat/completions`.
It is a router, not a complete replica of OpenAI or of every llama-server route.
The compatibility layer is `api_compat.py`; it does not modify model profiles.
## Discovery and behaviour
`GET /v1/capabilities` with the same Bearer token lists accepted request fields,
llama.cpp extensions, aliases and routes not implemented by Deck. This describes
router support, not a guarantee that every model/build supports every function.
Existing `/v1/models` remains a chat-only list. Images, speech and transcription
have their own model-list routes. Request handling and streaming share the same
normalization path; client-specific branches are not used.
- `max_completion_tokens` becomes `max_tokens`; conflicting limits fail explicitly.
- Legacy `functions` and `function_call` become `tools` and `tool_choice`.
- SDK `extra_body` is flattened, validated and cannot replace model/messages.
- Null optional stop/tools/stream options are omitted.
- `reasoning_effort` aliases retain existing nearest-supported mappings.
- Template options `enable_thinking` and `preserve_thinking` require booleans.
- Additional llama.cpp options: `min_p`, `typical_p`, `repeat_last_n`, `mirostat`,
`mirostat_tau`, `mirostat_eta`, `dynatemp_range`, `dynatemp_exponent`,
`cache_prompt`, `ignore_eos`, `reasoning_format`, and `logit_bias`.
Numeric types, finite values and bounded ranges are validated.
- Metadata, safety identifiers and prompt-cache keys are accepted as annotations,
but not persisted or used for OpenAI-hosted safety/caching semantics. Their
removal is reported in `X-Athena-Compatibility-Ignored` (also on SSE responses).
- `store=false/null` and service tier `auto/default/null` are tolerated and reported;
stored completions and priority service tiers are rejected, not simulated.
Unknown fields and conflicting aliases still produce explicit HTTP 400 errors.
Only one completion (`n=1`) is supported. Embeddings, Responses, native completion,
slot administration and arbitrary template overrides are not added by this change.
Tools, vision, reasoning and structured output depend on model weights, template
and active runtime version. Tool calls are returned to clients, not executed by Deck.
Model capability declarations must not be inferred solely from a model name.
## Verification
Tests cover alias translation, conflicts, input isolation, invalid sampling values,
metadata treatment, upstream forwarding and authenticated capability discovery.
Synthetic/fake-worker API tests do not establish inference quality or actual tool
calling for every installed model. New upstream features require review and tests.
References: [llama.cpp server API](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
and [OpenAI Chat Completions](https://platform.openai.com/docs/api-reference/chat/create).