Document Athena containers and expand Hermes operator skill
This commit is contained in:
1 parent
87a2ae5704
commit
b4f4bf37fd
13 files changed
+232
-29
No files matched your search
@@ -1,44 +1,53 @@
|
||||
---
|
||||
name: athena-operator
|
||||
description: Understand, operate and extend the Athena AI host.
|
||||
description: Operate and extend Mike's Athena AI host safely. Use for Athena models, profiles, Docker services, GPU allocation, inference modes, benchmarks, cleanup, deployment, backups, documentation, or when testing a new local AI model for Hermes.
|
||||
license: MIT
|
||||
metadata:
|
||||
hermes:
|
||||
version: 2.0.0
|
||||
version: 3.0.0
|
||||
author: Michael Roll
|
||||
platforms: [linux]
|
||||
tags: [athena, docker, mcp, models, backup]
|
||||
tags: [athena, docker, gpu, models, benchmark, cleanup, backup]
|
||||
---
|
||||
|
||||
# Athena Operator
|
||||
|
||||
Use this skill for work on Athena itself: Docker, MCPs, models, profiles,
|
||||
Hermes, OpenWebUI, TTS, STT, image generation, Git and backups.
|
||||
Use this skill for work on Athena itself: Docker, models, profiles, Hermes
|
||||
integration, TTS, STT, image/audio/music workers, Git and backups.
|
||||
|
||||
## Start
|
||||
|
||||
1. Call `athena_operator_inspect` with `subject=guide`; it returns the current
|
||||
`ATHENA.md` as the architectural truth.
|
||||
2. Inspect the affected live area only if needed.
|
||||
3. Search for the concrete source file, then read only the required lines.
|
||||
1. Call `athena_operator_inspect` with `subject=guide`.
|
||||
2. Choose and read the matching reference before acting:
|
||||
- architecture, containers or modes: `references/architecture-and-modes.md`
|
||||
- a model download, profile or A/B test: `references/model-evaluation.md`
|
||||
- deployment, cleanup, documentation or publication:
|
||||
`references/change-and-cleanup.md`
|
||||
3. Inspect only the affected live area with `overview`, `containers`, `models`,
|
||||
`jobs` or `git_status`.
|
||||
4. Search for the exact source path before reading or editing it.
|
||||
|
||||
Do not rediscover the complete platform for every task. Do not read entire
|
||||
large files when a bounded section is enough. Do not guess file paths.
|
||||
|
||||
## Change
|
||||
|
||||
When the user has clearly requested a change, use `athena_operator_change` to
|
||||
apply the smallest durable change. The operator owns the Git worktree and
|
||||
deployment access; do not clone another repository or request another SSH key.
|
||||
When the user clearly requests a change, use `athena_operator_change` to apply
|
||||
the smallest durable change. The operator owns the Git worktree and deployment
|
||||
access; do not clone another repository or request another SSH key.
|
||||
|
||||
Afterwards run focused checks, verify the affected service, commit and push.
|
||||
Afterwards run focused checks, verify the affected service functionally, update
|
||||
the affected documentation, commit and push.
|
||||
The scheduled Docker-data backup is automatic. After storage-affecting work,
|
||||
run one manual backup and verify its archive instead of building a special
|
||||
recovery kit.
|
||||
|
||||
For a new MCP, normally change only its server code, Dockerfile, MCP Compose
|
||||
service, env example, client registration and a focused test. Reuse an existing
|
||||
backend instead of installing a duplicate service.
|
||||
For every model candidate, preserve the currently working profile until the
|
||||
candidate is downloaded and ready. Compare like with like, record the exact
|
||||
artifact and result in `docs/TESTED_MODELS.md`, then either promote it or remove
|
||||
its weights and test-only runtime. Never silently lower a standard profile
|
||||
below Q4; Q3 is allowed only after a documented comparison shows negligible
|
||||
quality loss for the intended work.
|
||||
|
||||
## Tool discipline
|
||||
|
||||
@@ -48,6 +57,8 @@ backend instead of installing a duplicate service.
|
||||
- Prefer specialist MCPs for Home Assistant, Unraid, ARR, Navidrome and other
|
||||
external systems.
|
||||
- A healthy container is not proof; perform one bounded functional check.
|
||||
- `Created` or cleanly stopped model workers are normal on-demand services, not
|
||||
proof of garbage.
|
||||
- Never claim a write, deploy, commit, push or backup succeeded without its
|
||||
actual result.
|
||||
|
||||
|
||||
@@ -0,0 +1,48 @@
|
||||
# Athena architecture and modes
|
||||
|
||||
Use `ATHENA.md` as the short operational truth and
|
||||
`docs/CONTAINER_INVENTORY.md` for the current mapping of container, model and
|
||||
role. If live state disagrees with documentation, report the discrepancy and
|
||||
correct the durable source when the user requested maintenance.
|
||||
|
||||
## Boundaries
|
||||
|
||||
- Hermes, chats, skills and portable specialist MCPs live on Unraid.
|
||||
- Athena is the inference host. Its canonical checkout is
|
||||
`/opt/mike-ai/stack`; models live under `/data/models`; local secrets and
|
||||
runtime configuration live under `/etc/mike-ai` and never enter Git.
|
||||
- The Profile Router is the single OpenAI-compatible address clients use.
|
||||
- The Profile Controller is the only component allowed to orchestrate approved
|
||||
model and specialist workers.
|
||||
- The Athena Operator is host-bound and is the normal maintenance interface.
|
||||
|
||||
## Exclusive states
|
||||
|
||||
Athena has five mutually exclusive persistent modes: `llm`, `music`,
|
||||
`separation`, `voice`, and `voicechange`. Image generation is a transactional
|
||||
request: it temporarily pauses the active text profile and Qwen3-TTS, runs the
|
||||
image worker, then restores the previous LLM state.
|
||||
|
||||
Only one heavy GPU path may be active. Do not manually start a second GPU
|
||||
worker around the controller. The lightweight dashboard, router, controller,
|
||||
gateway, UI, CPU-STT, Piper, backup and operator containers may remain active.
|
||||
|
||||
## Model profiles
|
||||
|
||||
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
|
||||
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
|
||||
- Large: the same Q4 model, 192,000 tokens, vision.
|
||||
- Ultra: the same Q4 model, 262,144 tokens, no vision projector.
|
||||
- Beta 1: Qwen3.8-27B GSQ-RCO IQ3_S-MTP, 112,000 tokens; experimental only.
|
||||
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
|
||||
|
||||
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
|
||||
infer allocation from model size alone; verify the profile's Compose arguments
|
||||
and live VRAM. RTX 3060 also hosts Qwen3-TTS during normal LLM operation.
|
||||
|
||||
## What is and is not stale
|
||||
|
||||
The GPU workers use `restart: "no"` and are created once, then started on
|
||||
demand. A stopped `mike-ai-llama-*`, image, music, separator, OmniVoice or X-VC
|
||||
container is expected. A candidate is stale only after checking Compose,
|
||||
labels, mounts, router/controller references, model paths and test history.
|
||||
@@ -0,0 +1,43 @@
|
||||
# Durable changes, cleanup and publication
|
||||
|
||||
## Change flow
|
||||
|
||||
1. Inspect `guide` and the affected live subject.
|
||||
2. Check `git_status`; preserve unrelated user changes.
|
||||
3. Search the canonical checkout and edit the smallest source of truth.
|
||||
4. Validate syntax and Compose before deployment.
|
||||
5. Deploy only the affected service unless the requested change genuinely
|
||||
spans the core stack.
|
||||
6. Verify container health and one real function, not just process existence.
|
||||
7. Update documentation and the container/model inventory when architecture,
|
||||
models, ports, modes or ownership changed.
|
||||
8. Commit and push only after checks pass. Verify the remote result.
|
||||
|
||||
Use an asynchronous operator job only once and poll it with
|
||||
`athena_operator_job`; never launch a duplicate because a long task is quiet.
|
||||
|
||||
## Cleanup proof
|
||||
|
||||
Before deletion, prove that an item is unused by checking:
|
||||
|
||||
- current Compose projects and Docker labels;
|
||||
- router and controller source references;
|
||||
- mounts, volumes and model manifests;
|
||||
- `docs/TESTED_MODELS.md` and any rollback requirement;
|
||||
- whether a stopped container is an intentional on-demand worker.
|
||||
|
||||
Prefer precise targets. Never perform broad recursive deletion from `/`,
|
||||
`/data`, `/opt` or a variable that was not resolved and printed first. Build
|
||||
cache can be pruned after confirming no build is running. Do not prune named
|
||||
volumes or active images generically.
|
||||
|
||||
After storage-affecting work, report exactly what was removed and free space,
|
||||
run the existing manual backup operation, and verify the produced archive.
|
||||
|
||||
## Safety and secrets
|
||||
|
||||
Never reboot or shut down Athena or alter SSH, networking, WireGuard, firewall,
|
||||
kernel, boot, partitions or mounts without a separate explicit current user
|
||||
instruction. Do not print environment dumps, tokens, API keys, private keys or
|
||||
the contents of `/etc/mike-ai`. Redact accidental secret material from reports
|
||||
and never commit it.
|
||||
@@ -0,0 +1,41 @@
|
||||
# Model evaluation
|
||||
|
||||
## Before downloading
|
||||
|
||||
1. Read all of `docs/TESTED_MODELS.md`; it is the no-repeat register.
|
||||
2. Record the exact repository, revision, filename, base model, fine-tune,
|
||||
quantization, license, format and estimated disk/VRAM/RAM requirements.
|
||||
3. Confirm the model adds a genuinely new candidate rather than a renamed
|
||||
artifact already tested.
|
||||
4. Do not stop the active production model merely to download or prepare the
|
||||
candidate. Put temporary scripts under `/tmp`; durable code belongs in Git.
|
||||
|
||||
## Quality rule
|
||||
|
||||
Standard text profiles must not fall below Q4. A smaller Q3 quantization may be
|
||||
kept only when an A/B test documents that its quality loss is negligible for
|
||||
Mike's intended workload. Smaller files are not presumed faster: GPU split,
|
||||
memory bandwidth, kernels, cache formats and cross-GPU traffic must be measured.
|
||||
|
||||
## Fair A/B test
|
||||
|
||||
Hold these equal wherever the models permit it:
|
||||
|
||||
- prompt set and conversation history;
|
||||
- context size and filled-context test point;
|
||||
- KV-cache quantization, slots, batch/uBatch, MTP and sampling;
|
||||
- GPU visibility and tensor split;
|
||||
- warm-up state and output-token limit.
|
||||
|
||||
Measure prompt processing, short decode, long-context decode, peak VRAM and
|
||||
wall time. Test meaning preservation, uncertainty, negation, ordered safety
|
||||
constraints, German language consistency, tool-call schema and long-context
|
||||
recall. Do not replace a model on synthetic benchmark scores alone.
|
||||
|
||||
## Completion
|
||||
|
||||
Write the exact artifact, settings, raw result path, interpretation and decision
|
||||
to `docs/TESTED_MODELS.md` in the same commit. If promoted, update the manifest,
|
||||
Compose/env examples, profile matrix and user documentation. If rejected,
|
||||
remove candidate-only weights, images and containers after preserving the
|
||||
result. Restore and functionally test the previous profile, then publish.
|
||||
Reference in new issue
Block a user