Document Athena containers and expand Hermes operator skill

This commit is contained in:
Mikei386
2026-09-09 16:39:34 +02:00
parent 87a2ae5704
commit b4f4bf37fd
13 changed files with 232 additions and 29 deletions
+2 -1
View File
@@ -13,7 +13,7 @@ Bild- und Sprachausgabe. **Hermes und die Fach-MCPs laufen auf Unraid.**
- FLUX.2 Klein 9B FP8 Beta für Textbilder und Referenzbild-Bearbeitung: - FLUX.2 Klein 9B FP8 Beta für Textbilder und Referenzbild-Bearbeitung:
Transformer auf RTX 5080, Qwen3-8B-NF4-Textencoder auf RTX 3060 Transformer auf RTX 5080, Qwen3-8B-NF4-Textencoder auf RTX 3060
- Qwen3-TTS 1.7B auf der RTX 3060 mit Piper als CPU-Fallback - Qwen3-TTS 1.7B auf der RTX 3060 mit Piper als CPU-Fallback
- Whisper.cpp `large-v3-turbo` auf der CPU für lokale deutsche Spracherkennung - Whisper.cpp `ggml-small` auf der CPU für lokale deutsche Spracherkennung
- Live-Dashboard mit 21 Tagen Detailhistorie auf Port 8099 - Live-Dashboard mit 21 Tagen Detailhistorie auf Port 8099
- Dashboard-Umschaltung zwischen LLM-Betrieb, ACE-Step-Musikstudio, - Dashboard-Umschaltung zwischen LLM-Betrieb, ACE-Step-Musikstudio,
BS-RoFormer-Stimmtrennung, OmniVoice und X-VC BS-RoFormer-Stimmtrennung, OmniVoice und X-VC
@@ -186,6 +186,7 @@ Der genaue Sicherungsumfang steht in [docs/RECOVERY.md](docs/RECOVERY.md).
- [ATHENA.md](ATHENA.md) – kurze Betriebsanleitung - [ATHENA.md](ATHENA.md) – kurze Betriebsanleitung
- [docs/STANDARD_PROFILE_MATRIX.md](docs/STANDARD_PROFILE_MATRIX.md) – Profile - [docs/STANDARD_PROFILE_MATRIX.md](docs/STANDARD_PROFILE_MATRIX.md) – Profile
- [docs/CONTAINER_INVENTORY.md](docs/CONTAINER_INVENTORY.md) – alle Container, Modelle und Aufgaben
- [docs/TESTED_MODELS.md](docs/TESTED_MODELS.md) – zentrale Testhistorie und Sperrliste gegen Doppeltests - [docs/TESTED_MODELS.md](docs/TESTED_MODELS.md) – zentrale Testhistorie und Sperrliste gegen Doppeltests
- [docs/MCP_SERVERS.md](docs/MCP_SERVERS.md) – produktive Werkzeuge - [docs/MCP_SERVERS.md](docs/MCP_SERVERS.md) – produktive Werkzeuge
- [docs/RECOVERY.md](docs/RECOVERY.md) – Backup und Neuaufbau - [docs/RECOVERY.md](docs/RECOVERY.md) – Backup und Neuaufbau
+1 -1
View File
@@ -89,7 +89,7 @@ MEDIUM_CONTEXT=160000
MEDIUM_BATCH_SIZE=2048 MEDIUM_BATCH_SIZE=2048
MEDIUM_UBATCH_SIZE=128 MEDIUM_UBATCH_SIZE=128
MEDIUM_TENSOR_SPLIT=85,15 MEDIUM_TENSOR_SPLIT=85,15
BETA1_CONTEXT=192000 BETA1_CONTEXT=112000
BETA1_BATCH_SIZE=2048 BETA1_BATCH_SIZE=2048
BETA1_UBATCH_SIZE=128 BETA1_UBATCH_SIZE=128
BETA1_GPU_DEVICES=GPU-8ad38c6c-5a01-9d8e-1dfa-ed662ad78fbe,GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b BETA1_GPU_DEVICES=GPU-8ad38c6c-5a01-9d8e-1dfa-ed662ad78fbe,GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
+2 -2
View File
@@ -9,7 +9,7 @@ flowchart LR
R --> I[FLUX.2 Klein 9B FP8 Beta<br/>RTX 5080 Transformer] R --> I[FLUX.2 Klein 9B FP8 Beta<br/>RTX 5080 Transformer]
I --> E[Qwen3-8B NF4 Textencoder<br/>RTX 3060 während Bildauftrag] I --> E[Qwen3-8B NF4 Textencoder<br/>RTX 3060 während Bildauftrag]
R --> T[Qwen3-TTS RTX 3060<br/>Piper CPU-Fallback] R --> T[Qwen3-TTS RTX 3060<br/>Piper CPU-Fallback]
R --> STT[Whisper.cpp large-v3-turbo<br/>CPU, lokale Spracherkennung] R --> STT[Whisper.cpp ggml-small<br/>CPU, lokale Spracherkennung]
H --> U[MUA / Unraid MCP] H --> U[MUA / Unraid MCP]
H --> A[ARR-MCP] H --> A[ARR-MCP]
@@ -50,7 +50,7 @@ Kontextgröße:
| Medium | 160.000 Token | | Medium | 160.000 Token |
| Large | 192.000 Token | | Large | 192.000 Token |
| Ultra | 262.144 Token | | Ultra | 262.144 Token |
| Beta 1 | 192.000 Token | | Beta 1 | 112.000 Token |
| Uncensored | 80.000 Token | | Uncensored | 80.000 Token |
## Exklusiver Bildmodus ## Exklusiver Bildmodus
+49
View File
@@ -0,0 +1,49 @@
# Container-Inventar auf Athena
Stand: 9. September 2026
Athena besteht derzeit aus 23 Docker-Containern. Nicht jeder Container enthält
ein KI-Modell: Router, Oberflächen, Netzwerk, Steuerung und Sicherung sind
gewöhnliche Dienste. Die rechenintensiven GPU-Worker werden absichtlich nur bei
Bedarf gestartet. Ein Container im Zustand `Created` oder `Exited (0)` ist daher
nicht automatisch ein ungenutzter Rest.
| Container | Modell oder wesentliche Komponente | Aufgabe |
|---|---|---|
| `mike-ai-backup` | kein Modell; Offen Docker Volume Backup | Sichert `/data`, `/etc/mike-ai`, den Stack und die persistenten Docker-Volumes im Fünf-Stunden-Takt. |
| `mike-ai-image-worker` | FLUX.2 Klein 9B FP8, Qwen3-8B NF4 Textencoder und VAE | Erzeugt und bearbeitet Bilder transaktional; nutzt während eines Auftrags RTX 5080 und RTX 3060. |
| `mike-ai-llama-beta1` | Qwen3.8-27B GSQ-RCO `IQ3_S-MTP`, Qwen-MMProj BF16 | Experimentelles Beta-Profil mit 112.000 Token Kontext; wird nur auf Anforderung geladen. |
| `mike-ai-llama-dashboard` | kein Modell | Zeigt Telemetrie, Profile, GPU-Nutzung und Betriebsarten an und bietet die Modusumschaltung. |
| `mike-ai-llama-fast` | Qwen3.8-27B `IQ4-MIX`, Qwen-MMProj BF16 | Schnelles Q4-Text-/Vision-Profil mit 76.800 Token Kontext. |
| `mike-ai-llama-large` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Q4-Text-/Vision-Profil mit 192.000 Token Kontext und Verteilung auf beide GPUs. |
| `mike-ai-llama-medium` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Standard-Q4-Text-/Vision-Profil mit 160.000 Token Kontext und Verteilung auf beide GPUs. |
| `mike-ai-llama-ultra` | Qwen3.8-27B `IQ4_XS-pure`, ohne Vision-Projektor | Maximales Langkontextprofil mit 262.144 Token Kontext und Verteilung auf beide GPUs. |
| `mike-ai-llama-uncensored` | Qwen3.8-27B Abliterated `Q4_K_M`, eigener MMProj F16 | Spezialprofil mit 80.000 Token Kontext und gelockerten Modellgrenzen. |
| `mike-ai-mcp-athena-operator` | kein Modell | Stellt Hermes begrenzte Werkzeuge zum Prüfen, Ändern, Testen, Sichern und Versionieren von Athena bereit. |
| `mike-ai-music-acestep-test` | ACE-Step 1.5 XL-SFT und `acestep-5Hz-lm-1.7B` | Generiert Musik im exklusiven Musikmodus auf der RTX 5080. |
| `mike-ai-music-ui` | kein Modell; `fspecii/ace-step-ui` | Community-Oberfläche für ACE-Step; bleibt als leichte UI verfügbar, während der GPU-Worker bedarfsgesteuert läuft. |
| `mike-ai-piper` | `de_DE-thorsten-high` | CPU-basierte deutsche TTS-Rückfallebene, die auch während GPU-Umschaltungen verfügbar bleibt. |
| `mike-ai-portainer` | kein Modell; Portainer CE | Optionale Docker-Verwaltungsoberfläche. |
| `mike-ai-profile-controller` | kein Modell | Startet und stoppt ausschließlich freigegebene Modellprofile und Spezialworker in einer sicheren Reihenfolge. |
| `mike-ai-qwen3-tts` | `Qwen/Qwen3-TTS-12Hz-1.7B-Base`, Stimme Serena | Hochwertige deutsche Sprachausgabe auf der RTX 3060 im LLM-Betrieb. |
| `mike-ai-router` | kein eigenes Modell | Einzige OpenAI-kompatible Modelladresse; koordiniert Profile, Bildaufträge, Sprache und Betriebsarten. |
| `mike-ai-stem-separator` | BS-RoFormer Viperx 1297, `htdemucs_ft`, `htdemucs_6s`, `MossFormer2_SE_48K` | Trennt Gesang, Instrumente oder Sprache/Hintergrundgeräusche im exklusiven Separationsmodus. |
| `mike-ai-tts-gateway` | kein eigenes Modell | Normalisiert Text, leitet TTS an Qwen3-TTS weiter und fällt bei Bedarf auf Piper zurück. |
| `mike-ai-voice-studio` | `k2-fsa/OmniVoice` 0.2.1 mit Whisper-ASR | Erzeugt Text-to-Speech mit einer Referenzstimme; kein Audio-to-Audio-Voice-Changer. |
| `mike-ai-whisper` | Whisper.cpp `ggml-small` | Lokale deutsche Spracherkennung auf der CPU über `/v1/audio/transcriptions`. |
| `mike-ai-wireguard-gateway` | kein Modell | Veröffentlicht Dashboard und Fachoberflächen ausschließlich über den privaten WireGuard-Pfad. |
| `mike-ai-xvc-studio` | `chenxie95/X-VC` und GLM-4-Voice-Tokenizer | Wandelt eine vorhandene Sprachaufnahme anhand einer Referenzstimme in Audio zu Audio um. |
## Aufräumregel
Vor dem Löschen muss ein Kandidat gegen Compose-Dateien, Docker-Labels,
Mounts, Router-/Controller-Verweise und `/data` geprüft werden. Entfernt werden
nur nachweislich abgelöste Images, Gewichte, Versuchsdaten und Build-Caches.
Gewollt gestoppte Profilcontainer, persistente Modell-Caches und die letzte
funktionierende Produktionsvariante bleiben erhalten.
Am 9. September wurden die verworfenen Vevo2-Images und -Daten, das alte
Vevo2-Projektverzeichnis, ein leeres Test-Lab sowie der Docker-Build-Cache
entfernt. Der Build-Cache allein gab 74,43 GB frei; `/data` besitzt danach rund
562 GB freien Speicher. Kein produktiver oder bedarfsgesteuerter Container
wurde gelöscht.
+1 -1
View File
@@ -40,7 +40,7 @@ akzeptiert werden. Details stehen in [FLUX_9B_BETA.md](FLUX_9B_BETA.md).
## Lokale Spracherkennung ## Lokale Spracherkennung
Athena betreibt Whisper.cpp v1.9.1 mit `large-v3-turbo` als CPU-Dienst. Der Athena betreibt Whisper.cpp v1.9.1 mit `ggml-small` als CPU-Dienst. Der
Profile Router veröffentlicht ihn als OpenAI-kompatiblen Endpunkt Profile Router veröffentlicht ihn als OpenAI-kompatiblen Endpunkt
`/v1/audio/transcriptions`; Standardsprache ist Deutsch. Modell und Download `/v1/audio/transcriptions`; Standardsprache ist Deutsch. Modell und Download
bleiben im persistenten Docker-Volume `whisper-data` erhalten. bleiben im persistenten Docker-Volume `whisper-data` erhalten.
+1 -1
View File
@@ -1,6 +1,6 @@
# Athena-Betriebsmodi # Athena-Betriebsmodi
Athena besitzt sechs gegenseitig exklusive Betriebsmodi: Athena besitzt fünf gegenseitig exklusive Betriebsmodi:
- `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt. - `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt.
- `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt. - `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt.
+3 -2
View File
@@ -55,8 +55,9 @@ Titelgenerierung und Kontextkompression in Hermes.
|---|---|---|---| |---|---|---|---|
| 05.09.2026 | Coqui XTTS v2 | deutsche Satzzeichen, Enden und Streaming-Chunks erzeugten Halluzinationen und unnatürliche Prosodie | **verworfen und entfernt** | | 05.09.2026 | Coqui XTTS v2 | deutsche Satzzeichen, Enden und Streaming-Chunks erzeugten Halluzinationen und unnatürliche Prosodie | **verworfen und entfernt** |
| 05.09.2026 | `Qwen/Qwen3-TTS-12Hz-1.7B-Base` | deutlich natürlichere deutsche Ausgabe ohne die XTTS-Endhalluzinationen | **produktiv** mit Piper-Fallback | | 05.09.2026 | `Qwen/Qwen3-TTS-12Hz-1.7B-Base` | deutlich natürlichere deutsche Ausgabe ohne die XTTS-Endhalluzinationen | **produktiv** mit Piper-Fallback |
| 06.–08.09.2026 | Whisper.cpp `large-v3-turbo` | lokaler TTS→STT-Rundlauf und OpenClaw-Transkription erfolgreich | **produktiv** | | 06.–08.09.2026 | Whisper.cpp `large-v3-turbo` | lokaler TTS→STT-Rundlauf und OpenClaw-Transkription erfolgreich; später zugunsten des kleineren Laufzeitmodells entfernt | **ersetzt** |
| 09.09.2026 | `RMSnow/Vevo2`, Amphion `26f6883110181f1dbfe95c70a7c7dbaf4de5f42a` | Technik und Geschwindigkeit funktionierten, reale deutsche Sprachwandlung mit kurzer und langer Referenz war jedoch unverständlich, halluzinierend oder musikalisch | **qualitativ verworfen**; Container ersetzt, Image und Daten vorerst nur als Rollback erhalten | | seit 03.09.2026 | Whisper.cpp `ggml-small` | tatsächlich im Compose-Stack und im laufenden Container verwendetes CPU-STT-Modell | **produktiv** |
| 09.09.2026 | `RMSnow/Vevo2`, Amphion `26f6883110181f1dbfe95c70a7c7dbaf4de5f42a` | Technik und Geschwindigkeit funktionierten, reale deutsche Sprachwandlung mit kurzer und langer Referenz war jedoch unverständlich, halluzinierend oder musikalisch | **qualitativ verworfen und entfernt**; Ergebnis bleibt hier dokumentiert, Images, Daten und altes Projekt wurden am 09.09. bereinigt |
| 09.09.2026 | `k2-fsa/OmniVoice` 0.2.1 | Offizielle Gradio-UI auf RTX 5080 gestartet; Modell plus Whisper-ASR belegen rund 3,7 GiB VRAM. `omnivoice-triton` 0.1.0 ist kompatibel im Image vorhanden, für den ersten Hörtest aber bewusst noch nicht aktiviert | **technischer Starttest bestanden**, Hörabnahme und Basis-vs.-Triton-Messung offen; Gewichte CC BY-NC und daher nur nichtkommerziell einsetzen | | 09.09.2026 | `k2-fsa/OmniVoice` 0.2.1 | Offizielle Gradio-UI auf RTX 5080 gestartet; Modell plus Whisper-ASR belegen rund 3,7 GiB VRAM. `omnivoice-triton` 0.1.0 ist kompatibel im Image vorhanden, für den ersten Hörtest aber bewusst noch nicht aktiviert | **technischer Starttest bestanden**, Hörabnahme und Basis-vs.-Triton-Messung offen; Gewichte CC BY-NC und daher nur nichtkommerziell einsetzen |
| 09.09.2026 | `chenxie95/X-VC`, Code `49df8c591eafc48b096e466d96f9839f9c0dd739`, UI-Basis `d761cd6421e85376b2656dfefd8471d7f35a42be` | Offizielles Beispiel Ende-zu-Ende gewandelt: 5,20 s Audio in 1,07 s (RTF 0,21), gültiges 16-kHz-Mono-PCM-WAV; Modell belegt rund 2,9 GiB auf der RTX 5080. GLM-4-Voice-Tokenizer dokumentiert Chinesisch und Englisch | **technischer Start- und Konvertierungstest bestanden**; deutsche Hörabnahme offen | | 09.09.2026 | `chenxie95/X-VC`, Code `49df8c591eafc48b096e466d96f9839f9c0dd739`, UI-Basis `d761cd6421e85376b2656dfefd8471d7f35a42be` | Offizielles Beispiel Ende-zu-Ende gewandelt: 5,20 s Audio in 1,07 s (RTF 0,21), gültiges 16-kHz-Mono-PCM-WAV; Modell belegt rund 2,9 GiB auf der RTX 5080. GLM-4-Voice-Tokenizer dokumentiert Chinesisch und Englisch | **technischer Start- und Konvertierungstest bestanden**; deutsche Hörabnahme offen |
+1 -1
View File
@@ -337,7 +337,7 @@ MEDIUM_CONTEXT=${MEDIUM_CONTEXT:-160000}
MEDIUM_BATCH_SIZE=${MEDIUM_BATCH_SIZE:-2048} MEDIUM_BATCH_SIZE=${MEDIUM_BATCH_SIZE:-2048}
MEDIUM_UBATCH_SIZE=${MEDIUM_UBATCH_SIZE:-128} MEDIUM_UBATCH_SIZE=${MEDIUM_UBATCH_SIZE:-128}
MEDIUM_PARALLEL_SLOTS=${MEDIUM_PARALLEL_SLOTS:-1} MEDIUM_PARALLEL_SLOTS=${MEDIUM_PARALLEL_SLOTS:-1}
BETA1_CONTEXT=${BETA1_CONTEXT:-192000} BETA1_CONTEXT=${BETA1_CONTEXT:-112000}
BETA1_BATCH_SIZE=${BETA1_BATCH_SIZE:-2048} BETA1_BATCH_SIZE=${BETA1_BATCH_SIZE:-2048}
BETA1_UBATCH_SIZE=${BETA1_UBATCH_SIZE:-128} BETA1_UBATCH_SIZE=${BETA1_UBATCH_SIZE:-128}
BETA1_PARALLEL_SLOTS=${BETA1_PARALLEL_SLOTS:-1} BETA1_PARALLEL_SLOTS=${BETA1_PARALLEL_SLOTS:-1}
+13 -4
View File
@@ -11,17 +11,26 @@ die() { printf 'FEHLER: %s\n' "$*" >&2; exit 1; }
[[ -s $PROFILE_MATRIX ]] || die "Profilmatrix fehlt: $PROFILE_MATRIX" [[ -s $PROFILE_MATRIX ]] || die "Profilmatrix fehlt: $PROFILE_MATRIX"
install_skill() { install_skill() {
local source=$1 root=$2 name target local source=$1 root=$2 name source_dir target_dir target path relative
name=${source%/SKILL.md} name=${source%/SKILL.md}
name=${name##*/} name=${name##*/}
target=$root/platform/$name/SKILL.md source_dir=${source%/SKILL.md}
target_dir=$root/platform/$name
target=$target_dir/SKILL.md
grep -Fxq -- "name: $name" "$source" || \ grep -Fxq -- "name: $name" "$source" || \
die "Skill-Quelle hat kein gültiges Frontmatter: $source" die "Skill-Quelle hat kein gültiges Frontmatter: $source"
install -d -o 10000 -g 10000 -m 0750 "${target%/*}" install -d -o 10000 -g 10000 -m 0750 "$target_dir"
if [[ -s $target ]] && ! cmp -s "$source" "$target"; then if [[ -s $target ]] && ! cmp -s "$source" "$target"; then
cp -a "$target" "$target.before-managed-update-$(date +%Y%m%d-%H%M%S)" cp -a "$target" "$target.before-managed-update-$(date +%Y%m%d-%H%M%S)"
fi fi
install -o 10000 -g 10000 -m 0640 "$source" "$target" while IFS= read -r path; do
relative=${path#"$source_dir"/}
if [[ -d $path ]]; then
install -d -o 10000 -g 10000 -m 0750 "$target_dir/$relative"
else
install -o 10000 -g 10000 -m 0640 "$path" "$target_dir/$relative"
fi
done < <(find "$source_dir" -mindepth 1 -print)
cmp -s "$source" "$target" || die "Skill-Synchronisierung fehlgeschlagen: $target" cmp -s "$source" "$target" || die "Skill-Synchronisierung fehlgeschlagen: $target"
} }
+27 -16
View File
@@ -1,44 +1,53 @@
--- ---
name: athena-operator name: athena-operator
description: Understand, operate and extend the Athena AI host. description: Operate and extend Mike's Athena AI host safely. Use for Athena models, profiles, Docker services, GPU allocation, inference modes, benchmarks, cleanup, deployment, backups, documentation, or when testing a new local AI model for Hermes.
license: MIT license: MIT
metadata: metadata:
hermes: hermes:
version: 2.0.0 version: 3.0.0
author: Michael Roll author: Michael Roll
platforms: [linux] platforms: [linux]
tags: [athena, docker, mcp, models, backup] tags: [athena, docker, gpu, models, benchmark, cleanup, backup]
--- ---
# Athena Operator # Athena Operator
Use this skill for work on Athena itself: Docker, MCPs, models, profiles, Use this skill for work on Athena itself: Docker, models, profiles, Hermes
Hermes, OpenWebUI, TTS, STT, image generation, Git and backups. integration, TTS, STT, image/audio/music workers, Git and backups.
## Start ## Start
1. Call `athena_operator_inspect` with `subject=guide`; it returns the current 1. Call `athena_operator_inspect` with `subject=guide`.
`ATHENA.md` as the architectural truth. 2. Choose and read the matching reference before acting:
2. Inspect the affected live area only if needed. - architecture, containers or modes: `references/architecture-and-modes.md`
3. Search for the concrete source file, then read only the required lines. - a model download, profile or A/B test: `references/model-evaluation.md`
- deployment, cleanup, documentation or publication:
`references/change-and-cleanup.md`
3. Inspect only the affected live area with `overview`, `containers`, `models`,
`jobs` or `git_status`.
4. Search for the exact source path before reading or editing it.
Do not rediscover the complete platform for every task. Do not read entire Do not rediscover the complete platform for every task. Do not read entire
large files when a bounded section is enough. Do not guess file paths. large files when a bounded section is enough. Do not guess file paths.
## Change ## Change
When the user has clearly requested a change, use `athena_operator_change` to When the user clearly requests a change, use `athena_operator_change` to apply
apply the smallest durable change. The operator owns the Git worktree and the smallest durable change. The operator owns the Git worktree and deployment
deployment access; do not clone another repository or request another SSH key. access; do not clone another repository or request another SSH key.
Afterwards run focused checks, verify the affected service, commit and push. Afterwards run focused checks, verify the affected service functionally, update
the affected documentation, commit and push.
The scheduled Docker-data backup is automatic. After storage-affecting work, The scheduled Docker-data backup is automatic. After storage-affecting work,
run one manual backup and verify its archive instead of building a special run one manual backup and verify its archive instead of building a special
recovery kit. recovery kit.
For a new MCP, normally change only its server code, Dockerfile, MCP Compose For every model candidate, preserve the currently working profile until the
service, env example, client registration and a focused test. Reuse an existing candidate is downloaded and ready. Compare like with like, record the exact
backend instead of installing a duplicate service. artifact and result in `docs/TESTED_MODELS.md`, then either promote it or remove
its weights and test-only runtime. Never silently lower a standard profile
below Q4; Q3 is allowed only after a documented comparison shows negligible
quality loss for the intended work.
## Tool discipline ## Tool discipline
@@ -48,6 +57,8 @@ backend instead of installing a duplicate service.
- Prefer specialist MCPs for Home Assistant, Unraid, ARR, Navidrome and other - Prefer specialist MCPs for Home Assistant, Unraid, ARR, Navidrome and other
external systems. external systems.
- A healthy container is not proof; perform one bounded functional check. - A healthy container is not proof; perform one bounded functional check.
- `Created` or cleanly stopped model workers are normal on-demand services, not
proof of garbage.
- Never claim a write, deploy, commit, push or backup succeeded without its - Never claim a write, deploy, commit, push or backup succeeded without its
actual result. actual result.
@@ -0,0 +1,48 @@
# Athena architecture and modes
Use `ATHENA.md` as the short operational truth and
`docs/CONTAINER_INVENTORY.md` for the current mapping of container, model and
role. If live state disagrees with documentation, report the discrepancy and
correct the durable source when the user requested maintenance.
## Boundaries
- Hermes, chats, skills and portable specialist MCPs live on Unraid.
- Athena is the inference host. Its canonical checkout is
`/opt/mike-ai/stack`; models live under `/data/models`; local secrets and
runtime configuration live under `/etc/mike-ai` and never enter Git.
- The Profile Router is the single OpenAI-compatible address clients use.
- The Profile Controller is the only component allowed to orchestrate approved
model and specialist workers.
- The Athena Operator is host-bound and is the normal maintenance interface.
## Exclusive states
Athena has five mutually exclusive persistent modes: `llm`, `music`,
`separation`, `voice`, and `voicechange`. Image generation is a transactional
request: it temporarily pauses the active text profile and Qwen3-TTS, runs the
image worker, then restores the previous LLM state.
Only one heavy GPU path may be active. Do not manually start a second GPU
worker around the controller. The lightweight dashboard, router, controller,
gateway, UI, CPU-STT, Piper, backup and operator containers may remain active.
## Model profiles
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
- Large: the same Q4 model, 192,000 tokens, vision.
- Ultra: the same Q4 model, 262,144 tokens, no vision projector.
- Beta 1: Qwen3.8-27B GSQ-RCO IQ3_S-MTP, 112,000 tokens; experimental only.
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
infer allocation from model size alone; verify the profile's Compose arguments
and live VRAM. RTX 3060 also hosts Qwen3-TTS during normal LLM operation.
## What is and is not stale
The GPU workers use `restart: "no"` and are created once, then started on
demand. A stopped `mike-ai-llama-*`, image, music, separator, OmniVoice or X-VC
container is expected. A candidate is stale only after checking Compose,
labels, mounts, router/controller references, model paths and test history.
@@ -0,0 +1,43 @@
# Durable changes, cleanup and publication
## Change flow
1. Inspect `guide` and the affected live subject.
2. Check `git_status`; preserve unrelated user changes.
3. Search the canonical checkout and edit the smallest source of truth.
4. Validate syntax and Compose before deployment.
5. Deploy only the affected service unless the requested change genuinely
spans the core stack.
6. Verify container health and one real function, not just process existence.
7. Update documentation and the container/model inventory when architecture,
models, ports, modes or ownership changed.
8. Commit and push only after checks pass. Verify the remote result.
Use an asynchronous operator job only once and poll it with
`athena_operator_job`; never launch a duplicate because a long task is quiet.
## Cleanup proof
Before deletion, prove that an item is unused by checking:
- current Compose projects and Docker labels;
- router and controller source references;
- mounts, volumes and model manifests;
- `docs/TESTED_MODELS.md` and any rollback requirement;
- whether a stopped container is an intentional on-demand worker.
Prefer precise targets. Never perform broad recursive deletion from `/`,
`/data`, `/opt` or a variable that was not resolved and printed first. Build
cache can be pruned after confirming no build is running. Do not prune named
volumes or active images generically.
After storage-affecting work, report exactly what was removed and free space,
run the existing manual backup operation, and verify the produced archive.
## Safety and secrets
Never reboot or shut down Athena or alter SSH, networking, WireGuard, firewall,
kernel, boot, partitions or mounts without a separate explicit current user
instruction. Do not print environment dumps, tokens, API keys, private keys or
the contents of `/etc/mike-ai`. Redact accidental secret material from reports
and never commit it.
@@ -0,0 +1,41 @@
# Model evaluation
## Before downloading
1. Read all of `docs/TESTED_MODELS.md`; it is the no-repeat register.
2. Record the exact repository, revision, filename, base model, fine-tune,
quantization, license, format and estimated disk/VRAM/RAM requirements.
3. Confirm the model adds a genuinely new candidate rather than a renamed
artifact already tested.
4. Do not stop the active production model merely to download or prepare the
candidate. Put temporary scripts under `/tmp`; durable code belongs in Git.
## Quality rule
Standard text profiles must not fall below Q4. A smaller Q3 quantization may be
kept only when an A/B test documents that its quality loss is negligible for
Mike's intended workload. Smaller files are not presumed faster: GPU split,
memory bandwidth, kernels, cache formats and cross-GPU traffic must be measured.
## Fair A/B test
Hold these equal wherever the models permit it:
- prompt set and conversation history;
- context size and filled-context test point;
- KV-cache quantization, slots, batch/uBatch, MTP and sampling;
- GPU visibility and tensor split;
- warm-up state and output-token limit.
Measure prompt processing, short decode, long-context decode, peak VRAM and
wall time. Test meaning preservation, uncertainty, negation, ordered safety
constraints, German language consistency, tool-call schema and long-context
recall. Do not replace a model on synthetic benchmark scores alone.
## Completion
Write the exact artifact, settings, raw result path, interpretation and decision
to `docs/TESTED_MODELS.md` in the same commit. If promoted, update the manifest,
Compose/env examples, profile matrix and user documentation. If rejected,
remove candidate-only weights, images and containers after preserving the
result. Restore and functionally test the previous profile, then publish.