Document Athena containers and expand Hermes operator skill

This commit is contained in:
Mikei386 committed 2026-09-09 16:39:34 +02:00
1 parent 87a2ae5704
commit b4f4bf37fd
13 files changed
+232 -29

No files matched your search

@@ -0,0 +1,48 @@
# Athena architecture and modes
Use `ATHENA.md` as the short operational truth and
`docs/CONTAINER_INVENTORY.md` for the current mapping of container, model and
role. If live state disagrees with documentation, report the discrepancy and
correct the durable source when the user requested maintenance.
## Boundaries
- Hermes, chats, skills and portable specialist MCPs live on Unraid.
- Athena is the inference host. Its canonical checkout is
`/opt/mike-ai/stack`; models live under `/data/models`; local secrets and
runtime configuration live under `/etc/mike-ai` and never enter Git.
- The Profile Router is the single OpenAI-compatible address clients use.
- The Profile Controller is the only component allowed to orchestrate approved
model and specialist workers.
- The Athena Operator is host-bound and is the normal maintenance interface.
## Exclusive states
Athena has five mutually exclusive persistent modes: `llm`, `music`,
`separation`, `voice`, and `voicechange`. Image generation is a transactional
request: it temporarily pauses the active text profile and Qwen3-TTS, runs the
image worker, then restores the previous LLM state.
Only one heavy GPU path may be active. Do not manually start a second GPU
worker around the controller. The lightweight dashboard, router, controller,
gateway, UI, CPU-STT, Piper, backup and operator containers may remain active.
## Model profiles
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
- Large: the same Q4 model, 192,000 tokens, vision.
- Ultra: the same Q4 model, 262,144 tokens, no vision projector.
- Beta 1: Qwen3.8-27B GSQ-RCO IQ3_S-MTP, 112,000 tokens; experimental only.
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
infer allocation from model size alone; verify the profile's Compose arguments
and live VRAM. RTX 3060 also hosts Qwen3-TTS during normal LLM operation.
## What is and is not stale
The GPU workers use `restart: "no"` and are created once, then started on
demand. A stopped `mike-ai-llama-*`, image, music, separator, OmniVoice or X-VC
container is expected. A candidate is stale only after checking Compose,
labels, mounts, router/controller references, model paths and test history.
@@ -0,0 +1,43 @@
# Durable changes, cleanup and publication
## Change flow
1. Inspect `guide` and the affected live subject.
2. Check `git_status`; preserve unrelated user changes.
3. Search the canonical checkout and edit the smallest source of truth.
4. Validate syntax and Compose before deployment.
5. Deploy only the affected service unless the requested change genuinely
spans the core stack.
6. Verify container health and one real function, not just process existence.
7. Update documentation and the container/model inventory when architecture,
models, ports, modes or ownership changed.
8. Commit and push only after checks pass. Verify the remote result.
Use an asynchronous operator job only once and poll it with
`athena_operator_job`; never launch a duplicate because a long task is quiet.
## Cleanup proof
Before deletion, prove that an item is unused by checking:
- current Compose projects and Docker labels;
- router and controller source references;
- mounts, volumes and model manifests;
- `docs/TESTED_MODELS.md` and any rollback requirement;
- whether a stopped container is an intentional on-demand worker.
Prefer precise targets. Never perform broad recursive deletion from `/`,
`/data`, `/opt` or a variable that was not resolved and printed first. Build
cache can be pruned after confirming no build is running. Do not prune named
volumes or active images generically.
After storage-affecting work, report exactly what was removed and free space,
run the existing manual backup operation, and verify the produced archive.
## Safety and secrets
Never reboot or shut down Athena or alter SSH, networking, WireGuard, firewall,
kernel, boot, partitions or mounts without a separate explicit current user
instruction. Do not print environment dumps, tokens, API keys, private keys or
the contents of `/etc/mike-ai`. Redact accidental secret material from reports
and never commit it.
@@ -0,0 +1,41 @@
# Model evaluation
## Before downloading
1. Read all of `docs/TESTED_MODELS.md`; it is the no-repeat register.
2. Record the exact repository, revision, filename, base model, fine-tune,
quantization, license, format and estimated disk/VRAM/RAM requirements.
3. Confirm the model adds a genuinely new candidate rather than a renamed
artifact already tested.
4. Do not stop the active production model merely to download or prepare the
candidate. Put temporary scripts under `/tmp`; durable code belongs in Git.
## Quality rule
Standard text profiles must not fall below Q4. A smaller Q3 quantization may be
kept only when an A/B test documents that its quality loss is negligible for
Mike's intended workload. Smaller files are not presumed faster: GPU split,
memory bandwidth, kernels, cache formats and cross-GPU traffic must be measured.
## Fair A/B test
Hold these equal wherever the models permit it:
- prompt set and conversation history;
- context size and filled-context test point;
- KV-cache quantization, slots, batch/uBatch, MTP and sampling;
- GPU visibility and tensor split;
- warm-up state and output-token limit.
Measure prompt processing, short decode, long-context decode, peak VRAM and
wall time. Test meaning preservation, uncertainty, negation, ordered safety
constraints, German language consistency, tool-call schema and long-context
recall. Do not replace a model on synthetic benchmark scores alone.
## Completion
Write the exact artifact, settings, raw result path, interpretation and decision
to `docs/TESTED_MODELS.md` in the same commit. If promoted, update the manifest,
Compose/env examples, profile matrix and user documentation. If rejected,
remove candidate-only weights, images and containers after preserving the
result. Restore and functionally test the previous profile, then publish.