Files
Athena-Deck/API_COMPATIBILITY.md
T

2.9 KiB

Client-neutral chat compatibility

Deck exposes authenticated OpenAI-style Chat Completions at /v1/chat/completions. It is a router, not a complete replica of OpenAI or of every llama-server route. The compatibility layer is api_compat.py; it does not modify model profiles.

Discovery and behaviour

GET /v1/capabilities with the same Bearer token lists accepted request fields, llama.cpp extensions, aliases and routes not implemented by Deck. This describes router support, not a guarantee that every model/build supports every function. Existing /v1/models remains a chat-only list. Images, speech and transcription have their own model-list routes. Request handling and streaming share the same normalization path; client-specific branches are not used.

  • max_completion_tokens becomes max_tokens; conflicting limits fail explicitly.
  • Legacy functions and function_call become tools and tool_choice.
  • SDK extra_body is flattened, validated and cannot replace model/messages.
  • Null optional stop/tools/stream options are omitted.
  • reasoning_effort aliases retain existing nearest-supported mappings.
  • Template options enable_thinking and preserve_thinking require booleans.
  • Additional llama.cpp options: min_p, typical_p, repeat_last_n, mirostat, mirostat_tau, mirostat_eta, dynatemp_range, dynatemp_exponent, cache_prompt, ignore_eos, reasoning_format, and logit_bias. Numeric types, finite values and bounded ranges are validated.
  • Metadata, safety identifiers and prompt-cache keys are accepted as annotations, but not persisted or used for OpenAI-hosted safety/caching semantics. Their removal is reported in X-Athena-Compatibility-Ignored (also on SSE responses).
  • store=false/null and service tier auto/default/null are tolerated and reported; stored completions and priority service tiers are rejected, not simulated.

Unknown fields and conflicting aliases still produce explicit HTTP 400 errors. Only one completion (n=1) is supported. Embeddings, Responses, native completion, slot administration and arbitrary template overrides are not added by this change. Tools, vision, reasoning and structured output depend on model weights, template and active runtime version. Tool calls are returned to clients, not executed by Deck. Model capability declarations must not be inferred solely from a model name.

Verification

Tests cover alias translation, conflicts, input isolation, invalid sampling values, metadata treatment, upstream forwarding and authenticated capability discovery. Synthetic/fake-worker API tests do not establish inference quality or actual tool calling for every installed model. New upstream features require review and tests.

References: llama.cpp server API and OpenAI Chat Completions.