Document Athena containers and expand Hermes operator skill
This commit is contained in:
@@ -0,0 +1,48 @@
|
||||
# Athena architecture and modes
|
||||
|
||||
Use `ATHENA.md` as the short operational truth and
|
||||
`docs/CONTAINER_INVENTORY.md` for the current mapping of container, model and
|
||||
role. If live state disagrees with documentation, report the discrepancy and
|
||||
correct the durable source when the user requested maintenance.
|
||||
|
||||
## Boundaries
|
||||
|
||||
- Hermes, chats, skills and portable specialist MCPs live on Unraid.
|
||||
- Athena is the inference host. Its canonical checkout is
|
||||
`/opt/mike-ai/stack`; models live under `/data/models`; local secrets and
|
||||
runtime configuration live under `/etc/mike-ai` and never enter Git.
|
||||
- The Profile Router is the single OpenAI-compatible address clients use.
|
||||
- The Profile Controller is the only component allowed to orchestrate approved
|
||||
model and specialist workers.
|
||||
- The Athena Operator is host-bound and is the normal maintenance interface.
|
||||
|
||||
## Exclusive states
|
||||
|
||||
Athena has five mutually exclusive persistent modes: `llm`, `music`,
|
||||
`separation`, `voice`, and `voicechange`. Image generation is a transactional
|
||||
request: it temporarily pauses the active text profile and Qwen3-TTS, runs the
|
||||
image worker, then restores the previous LLM state.
|
||||
|
||||
Only one heavy GPU path may be active. Do not manually start a second GPU
|
||||
worker around the controller. The lightweight dashboard, router, controller,
|
||||
gateway, UI, CPU-STT, Piper, backup and operator containers may remain active.
|
||||
|
||||
## Model profiles
|
||||
|
||||
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
|
||||
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
|
||||
- Large: the same Q4 model, 192,000 tokens, vision.
|
||||
- Ultra: the same Q4 model, 262,144 tokens, no vision projector.
|
||||
- Beta 1: Qwen3.8-27B GSQ-RCO IQ3_S-MTP, 112,000 tokens; experimental only.
|
||||
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
|
||||
|
||||
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
|
||||
infer allocation from model size alone; verify the profile's Compose arguments
|
||||
and live VRAM. RTX 3060 also hosts Qwen3-TTS during normal LLM operation.
|
||||
|
||||
## What is and is not stale
|
||||
|
||||
The GPU workers use `restart: "no"` and are created once, then started on
|
||||
demand. A stopped `mike-ai-llama-*`, image, music, separator, OmniVoice or X-VC
|
||||
container is expected. A candidate is stale only after checking Compose,
|
||||
labels, mounts, router/controller references, model paths and test history.
|
||||
Reference in New Issue
Block a user