48 lines
2.9 KiB
Markdown
48 lines
2.9 KiB
Markdown
# Client-neutral chat compatibility
|
|
|
|
Deck exposes authenticated OpenAI-style Chat Completions at `/v1/chat/completions`.
|
|
It is a router, not a complete replica of OpenAI or of every llama-server route.
|
|
The compatibility layer is `api_compat.py`; it does not modify model profiles.
|
|
|
|
## Discovery and behaviour
|
|
|
|
`GET /v1/capabilities` with the same Bearer token lists accepted request fields,
|
|
llama.cpp extensions, aliases and routes not implemented by Deck. This describes
|
|
router support, not a guarantee that every model/build supports every function.
|
|
Existing `/v1/models` remains a chat-only list. Images, speech and transcription
|
|
have their own model-list routes. Request handling and streaming share the same
|
|
normalization path; client-specific branches are not used.
|
|
|
|
- `max_completion_tokens` becomes `max_tokens`; conflicting limits fail explicitly.
|
|
- Legacy `functions` and `function_call` become `tools` and `tool_choice`.
|
|
- SDK `extra_body` is flattened, validated and cannot replace model/messages.
|
|
- Null optional stop/tools/stream options are omitted.
|
|
- `reasoning_effort` aliases retain existing nearest-supported mappings.
|
|
- Template options `enable_thinking` and `preserve_thinking` require booleans.
|
|
- Additional llama.cpp options: `min_p`, `typical_p`, `repeat_last_n`, `mirostat`,
|
|
`mirostat_tau`, `mirostat_eta`, `dynatemp_range`, `dynatemp_exponent`,
|
|
`cache_prompt`, `ignore_eos`, `reasoning_format`, and `logit_bias`.
|
|
Numeric types, finite values and bounded ranges are validated.
|
|
- Metadata, safety identifiers and prompt-cache keys are accepted as annotations,
|
|
but not persisted or used for OpenAI-hosted safety/caching semantics. Their
|
|
removal is reported in `X-Athena-Compatibility-Ignored` (also on SSE responses).
|
|
- `store=false/null` and service tier `auto/default/null` are tolerated and reported;
|
|
stored completions and priority service tiers are rejected, not simulated.
|
|
|
|
Unknown fields and conflicting aliases still produce explicit HTTP 400 errors.
|
|
Only one completion (`n=1`) is supported. Embeddings, Responses, native completion,
|
|
slot administration and arbitrary template overrides are not added by this change.
|
|
Tools, vision, reasoning and structured output depend on model weights, template
|
|
and active runtime version. Tool calls are returned to clients, not executed by Deck.
|
|
Model capability declarations must not be inferred solely from a model name.
|
|
|
|
## Verification
|
|
|
|
Tests cover alias translation, conflicts, input isolation, invalid sampling values,
|
|
metadata treatment, upstream forwarding and authenticated capability discovery.
|
|
Synthetic/fake-worker API tests do not establish inference quality or actual tool
|
|
calling for every installed model. New upstream features require review and tests.
|
|
|
|
References: [llama.cpp server API](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
|
|
and [OpenAI Chat Completions](https://platform.openai.com/docs/api-reference/chat/create).
|