Document Athena containers and expand Hermes operator skill

This commit is contained in:
Mikei386
2026-09-09 16:39:34 +02:00
parent 87a2ae5704
commit b4f4bf37fd
13 changed files with 232 additions and 29 deletions
+13 -4
View File
@@ -11,17 +11,26 @@ die() { printf 'FEHLER: %s\n' "$*" >&2; exit 1; }
[[ -s $PROFILE_MATRIX ]] || die "Profilmatrix fehlt: $PROFILE_MATRIX"
install_skill() {
local source=$1 root=$2 name target
local source=$1 root=$2 name source_dir target_dir target path relative
name=${source%/SKILL.md}
name=${name##*/}
target=$root/platform/$name/SKILL.md
source_dir=${source%/SKILL.md}
target_dir=$root/platform/$name
target=$target_dir/SKILL.md
grep -Fxq -- "name: $name" "$source" || \
die "Skill-Quelle hat kein gültiges Frontmatter: $source"
install -d -o 10000 -g 10000 -m 0750 "${target%/*}"
install -d -o 10000 -g 10000 -m 0750 "$target_dir"
if [[ -s $target ]] && ! cmp -s "$source" "$target"; then
cp -a "$target" "$target.before-managed-update-$(date +%Y%m%d-%H%M%S)"
fi
install -o 10000 -g 10000 -m 0640 "$source" "$target"
while IFS= read -r path; do
relative=${path#"$source_dir"/}
if [[ -d $path ]]; then
install -d -o 10000 -g 10000 -m 0750 "$target_dir/$relative"
else
install -o 10000 -g 10000 -m 0640 "$path" "$target_dir/$relative"
fi
done < <(find "$source_dir" -mindepth 1 -print)
cmp -s "$source" "$target" || die "Skill-Synchronisierung fehlgeschlagen: $target"
}
+27 -16
View File
@@ -1,44 +1,53 @@
---
name: athena-operator
description: Understand, operate and extend the Athena AI host.
description: Operate and extend Mike's Athena AI host safely. Use for Athena models, profiles, Docker services, GPU allocation, inference modes, benchmarks, cleanup, deployment, backups, documentation, or when testing a new local AI model for Hermes.
license: MIT
metadata:
hermes:
version: 2.0.0
version: 3.0.0
author: Michael Roll
platforms: [linux]
tags: [athena, docker, mcp, models, backup]
tags: [athena, docker, gpu, models, benchmark, cleanup, backup]
---
# Athena Operator
Use this skill for work on Athena itself: Docker, MCPs, models, profiles,
Hermes, OpenWebUI, TTS, STT, image generation, Git and backups.
Use this skill for work on Athena itself: Docker, models, profiles, Hermes
integration, TTS, STT, image/audio/music workers, Git and backups.
## Start
1. Call `athena_operator_inspect` with `subject=guide`; it returns the current
`ATHENA.md` as the architectural truth.
2. Inspect the affected live area only if needed.
3. Search for the concrete source file, then read only the required lines.
1. Call `athena_operator_inspect` with `subject=guide`.
2. Choose and read the matching reference before acting:
- architecture, containers or modes: `references/architecture-and-modes.md`
- a model download, profile or A/B test: `references/model-evaluation.md`
- deployment, cleanup, documentation or publication:
`references/change-and-cleanup.md`
3. Inspect only the affected live area with `overview`, `containers`, `models`,
`jobs` or `git_status`.
4. Search for the exact source path before reading or editing it.
Do not rediscover the complete platform for every task. Do not read entire
large files when a bounded section is enough. Do not guess file paths.
## Change
When the user has clearly requested a change, use `athena_operator_change` to
apply the smallest durable change. The operator owns the Git worktree and
deployment access; do not clone another repository or request another SSH key.
When the user clearly requests a change, use `athena_operator_change` to apply
the smallest durable change. The operator owns the Git worktree and deployment
access; do not clone another repository or request another SSH key.
Afterwards run focused checks, verify the affected service, commit and push.
Afterwards run focused checks, verify the affected service functionally, update
the affected documentation, commit and push.
The scheduled Docker-data backup is automatic. After storage-affecting work,
run one manual backup and verify its archive instead of building a special
recovery kit.
For a new MCP, normally change only its server code, Dockerfile, MCP Compose
service, env example, client registration and a focused test. Reuse an existing
backend instead of installing a duplicate service.
For every model candidate, preserve the currently working profile until the
candidate is downloaded and ready. Compare like with like, record the exact
artifact and result in `docs/TESTED_MODELS.md`, then either promote it or remove
its weights and test-only runtime. Never silently lower a standard profile
below Q4; Q3 is allowed only after a documented comparison shows negligible
quality loss for the intended work.
## Tool discipline
@@ -48,6 +57,8 @@ backend instead of installing a duplicate service.
- Prefer specialist MCPs for Home Assistant, Unraid, ARR, Navidrome and other
external systems.
- A healthy container is not proof; perform one bounded functional check.
- `Created` or cleanly stopped model workers are normal on-demand services, not
proof of garbage.
- Never claim a write, deploy, commit, push or backup succeeded without its
actual result.
@@ -0,0 +1,48 @@
# Athena architecture and modes
Use `ATHENA.md` as the short operational truth and
`docs/CONTAINER_INVENTORY.md` for the current mapping of container, model and
role. If live state disagrees with documentation, report the discrepancy and
correct the durable source when the user requested maintenance.
## Boundaries
- Hermes, chats, skills and portable specialist MCPs live on Unraid.
- Athena is the inference host. Its canonical checkout is
`/opt/mike-ai/stack`; models live under `/data/models`; local secrets and
runtime configuration live under `/etc/mike-ai` and never enter Git.
- The Profile Router is the single OpenAI-compatible address clients use.
- The Profile Controller is the only component allowed to orchestrate approved
model and specialist workers.
- The Athena Operator is host-bound and is the normal maintenance interface.
## Exclusive states
Athena has five mutually exclusive persistent modes: `llm`, `music`,
`separation`, `voice`, and `voicechange`. Image generation is a transactional
request: it temporarily pauses the active text profile and Qwen3-TTS, runs the
image worker, then restores the previous LLM state.
Only one heavy GPU path may be active. Do not manually start a second GPU
worker around the controller. The lightweight dashboard, router, controller,
gateway, UI, CPU-STT, Piper, backup and operator containers may remain active.
## Model profiles
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
- Large: the same Q4 model, 192,000 tokens, vision.
- Ultra: the same Q4 model, 262,144 tokens, no vision projector.
- Beta 1: Qwen3.8-27B GSQ-RCO IQ3_S-MTP, 112,000 tokens; experimental only.
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
infer allocation from model size alone; verify the profile's Compose arguments
and live VRAM. RTX 3060 also hosts Qwen3-TTS during normal LLM operation.
## What is and is not stale
The GPU workers use `restart: "no"` and are created once, then started on
demand. A stopped `mike-ai-llama-*`, image, music, separator, OmniVoice or X-VC
container is expected. A candidate is stale only after checking Compose,
labels, mounts, router/controller references, model paths and test history.
@@ -0,0 +1,43 @@
# Durable changes, cleanup and publication
## Change flow
1. Inspect `guide` and the affected live subject.
2. Check `git_status`; preserve unrelated user changes.
3. Search the canonical checkout and edit the smallest source of truth.
4. Validate syntax and Compose before deployment.
5. Deploy only the affected service unless the requested change genuinely
spans the core stack.
6. Verify container health and one real function, not just process existence.
7. Update documentation and the container/model inventory when architecture,
models, ports, modes or ownership changed.
8. Commit and push only after checks pass. Verify the remote result.
Use an asynchronous operator job only once and poll it with
`athena_operator_job`; never launch a duplicate because a long task is quiet.
## Cleanup proof
Before deletion, prove that an item is unused by checking:
- current Compose projects and Docker labels;
- router and controller source references;
- mounts, volumes and model manifests;
- `docs/TESTED_MODELS.md` and any rollback requirement;
- whether a stopped container is an intentional on-demand worker.
Prefer precise targets. Never perform broad recursive deletion from `/`,
`/data`, `/opt` or a variable that was not resolved and printed first. Build
cache can be pruned after confirming no build is running. Do not prune named
volumes or active images generically.
After storage-affecting work, report exactly what was removed and free space,
run the existing manual backup operation, and verify the produced archive.
## Safety and secrets
Never reboot or shut down Athena or alter SSH, networking, WireGuard, firewall,
kernel, boot, partitions or mounts without a separate explicit current user
instruction. Do not print environment dumps, tokens, API keys, private keys or
the contents of `/etc/mike-ai`. Redact accidental secret material from reports
and never commit it.
@@ -0,0 +1,41 @@
# Model evaluation
## Before downloading
1. Read all of `docs/TESTED_MODELS.md`; it is the no-repeat register.
2. Record the exact repository, revision, filename, base model, fine-tune,
quantization, license, format and estimated disk/VRAM/RAM requirements.
3. Confirm the model adds a genuinely new candidate rather than a renamed
artifact already tested.
4. Do not stop the active production model merely to download or prepare the
candidate. Put temporary scripts under `/tmp`; durable code belongs in Git.
## Quality rule
Standard text profiles must not fall below Q4. A smaller Q3 quantization may be
kept only when an A/B test documents that its quality loss is negligible for
Mike's intended workload. Smaller files are not presumed faster: GPU split,
memory bandwidth, kernels, cache formats and cross-GPU traffic must be measured.
## Fair A/B test
Hold these equal wherever the models permit it:
- prompt set and conversation history;
- context size and filled-context test point;
- KV-cache quantization, slots, batch/uBatch, MTP and sampling;
- GPU visibility and tensor split;
- warm-up state and output-token limit.
Measure prompt processing, short decode, long-context decode, peak VRAM and
wall time. Test meaning preservation, uncertainty, negation, ordered safety
constraints, German language consistency, tool-call schema and long-context
recall. Do not replace a model on synthetic benchmark scores alone.
## Completion
Write the exact artifact, settings, raw result path, interpretation and decision
to `docs/TESTED_MODELS.md` in the same commit. If promoted, update the manifest,
Compose/env examples, profile matrix and user documentation. If rejected,
remove candidate-only weights, images and containers after preserving the
result. Restore and functionally test the previous profile, then publish.