Broaden client-neutral OpenAI and llama.cpp chat compatibility
This commit is contained in:
1 parent
c9036201f7
commit
3cd9fa5750
5 files changed
+158
-3
No files matched your search
@@ -0,0 +1,47 @@
|
||||
# Client-neutral chat compatibility
|
||||
|
||||
Deck exposes authenticated OpenAI-style Chat Completions at `/v1/chat/completions`.
|
||||
It is a router, not a complete replica of OpenAI or of every llama-server route.
|
||||
The compatibility layer is `api_compat.py`; it does not modify model profiles.
|
||||
|
||||
## Discovery and behaviour
|
||||
|
||||
`GET /v1/capabilities` with the same Bearer token lists accepted request fields,
|
||||
llama.cpp extensions, aliases and routes not implemented by Deck. This describes
|
||||
router support, not a guarantee that every model/build supports every function.
|
||||
Existing `/v1/models` remains a chat-only list. Images, speech and transcription
|
||||
have their own model-list routes. Request handling and streaming share the same
|
||||
normalization path; client-specific branches are not used.
|
||||
|
||||
- `max_completion_tokens` becomes `max_tokens`; conflicting limits fail explicitly.
|
||||
- Legacy `functions` and `function_call` become `tools` and `tool_choice`.
|
||||
- SDK `extra_body` is flattened, validated and cannot replace model/messages.
|
||||
- Null optional stop/tools/stream options are omitted.
|
||||
- `reasoning_effort` aliases retain existing nearest-supported mappings.
|
||||
- Template options `enable_thinking` and `preserve_thinking` require booleans.
|
||||
- Additional llama.cpp options: `min_p`, `typical_p`, `repeat_last_n`, `mirostat`,
|
||||
`mirostat_tau`, `mirostat_eta`, `dynatemp_range`, `dynatemp_exponent`,
|
||||
`cache_prompt`, `ignore_eos`, `reasoning_format`, and `logit_bias`.
|
||||
Numeric types, finite values and bounded ranges are validated.
|
||||
- Metadata, safety identifiers and prompt-cache keys are accepted as annotations,
|
||||
but not persisted or used for OpenAI-hosted safety/caching semantics. Their
|
||||
removal is reported in `X-Athena-Compatibility-Ignored` (also on SSE responses).
|
||||
- `store=false/null` and service tier `auto/default/null` are tolerated and reported;
|
||||
stored completions and priority service tiers are rejected, not simulated.
|
||||
|
||||
Unknown fields and conflicting aliases still produce explicit HTTP 400 errors.
|
||||
Only one completion (`n=1`) is supported. Embeddings, Responses, native completion,
|
||||
slot administration and arbitrary template overrides are not added by this change.
|
||||
Tools, vision, reasoning and structured output depend on model weights, template
|
||||
and active runtime version. Tool calls are returned to clients, not executed by Deck.
|
||||
Model capability declarations must not be inferred solely from a model name.
|
||||
|
||||
## Verification
|
||||
|
||||
Tests cover alias translation, conflicts, input isolation, invalid sampling values,
|
||||
metadata treatment, upstream forwarding and authenticated capability discovery.
|
||||
Synthetic/fake-worker API tests do not establish inference quality or actual tool
|
||||
calling for every installed model. New upstream features require review and tests.
|
||||
|
||||
References: [llama.cpp server API](https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md)
|
||||
and [OpenAI Chat Completions](https://platform.openai.com/docs/api-reference/chat/create).
|
||||
Reference in new issue
Block a user