Document Athena containers and expand Hermes operator skill
This commit is contained in:
@@ -13,7 +13,7 @@ Bild- und Sprachausgabe. **Hermes und die Fach-MCPs laufen auf Unraid.**
|
|||||||
- FLUX.2 Klein 9B FP8 Beta für Textbilder und Referenzbild-Bearbeitung:
|
- FLUX.2 Klein 9B FP8 Beta für Textbilder und Referenzbild-Bearbeitung:
|
||||||
Transformer auf RTX 5080, Qwen3-8B-NF4-Textencoder auf RTX 3060
|
Transformer auf RTX 5080, Qwen3-8B-NF4-Textencoder auf RTX 3060
|
||||||
- Qwen3-TTS 1.7B auf der RTX 3060 mit Piper als CPU-Fallback
|
- Qwen3-TTS 1.7B auf der RTX 3060 mit Piper als CPU-Fallback
|
||||||
- Whisper.cpp `large-v3-turbo` auf der CPU für lokale deutsche Spracherkennung
|
- Whisper.cpp `ggml-small` auf der CPU für lokale deutsche Spracherkennung
|
||||||
- Live-Dashboard mit 21 Tagen Detailhistorie auf Port 8099
|
- Live-Dashboard mit 21 Tagen Detailhistorie auf Port 8099
|
||||||
- Dashboard-Umschaltung zwischen LLM-Betrieb, ACE-Step-Musikstudio,
|
- Dashboard-Umschaltung zwischen LLM-Betrieb, ACE-Step-Musikstudio,
|
||||||
BS-RoFormer-Stimmtrennung, OmniVoice und X-VC
|
BS-RoFormer-Stimmtrennung, OmniVoice und X-VC
|
||||||
@@ -186,6 +186,7 @@ Der genaue Sicherungsumfang steht in [docs/RECOVERY.md](docs/RECOVERY.md).
|
|||||||
|
|
||||||
- [ATHENA.md](ATHENA.md) – kurze Betriebsanleitung
|
- [ATHENA.md](ATHENA.md) – kurze Betriebsanleitung
|
||||||
- [docs/STANDARD_PROFILE_MATRIX.md](docs/STANDARD_PROFILE_MATRIX.md) – Profile
|
- [docs/STANDARD_PROFILE_MATRIX.md](docs/STANDARD_PROFILE_MATRIX.md) – Profile
|
||||||
|
- [docs/CONTAINER_INVENTORY.md](docs/CONTAINER_INVENTORY.md) – alle Container, Modelle und Aufgaben
|
||||||
- [docs/TESTED_MODELS.md](docs/TESTED_MODELS.md) – zentrale Testhistorie und Sperrliste gegen Doppeltests
|
- [docs/TESTED_MODELS.md](docs/TESTED_MODELS.md) – zentrale Testhistorie und Sperrliste gegen Doppeltests
|
||||||
- [docs/MCP_SERVERS.md](docs/MCP_SERVERS.md) – produktive Werkzeuge
|
- [docs/MCP_SERVERS.md](docs/MCP_SERVERS.md) – produktive Werkzeuge
|
||||||
- [docs/RECOVERY.md](docs/RECOVERY.md) – Backup und Neuaufbau
|
- [docs/RECOVERY.md](docs/RECOVERY.md) – Backup und Neuaufbau
|
||||||
|
|||||||
@@ -89,7 +89,7 @@ MEDIUM_CONTEXT=160000
|
|||||||
MEDIUM_BATCH_SIZE=2048
|
MEDIUM_BATCH_SIZE=2048
|
||||||
MEDIUM_UBATCH_SIZE=128
|
MEDIUM_UBATCH_SIZE=128
|
||||||
MEDIUM_TENSOR_SPLIT=85,15
|
MEDIUM_TENSOR_SPLIT=85,15
|
||||||
BETA1_CONTEXT=192000
|
BETA1_CONTEXT=112000
|
||||||
BETA1_BATCH_SIZE=2048
|
BETA1_BATCH_SIZE=2048
|
||||||
BETA1_UBATCH_SIZE=128
|
BETA1_UBATCH_SIZE=128
|
||||||
BETA1_GPU_DEVICES=GPU-8ad38c6c-5a01-9d8e-1dfa-ed662ad78fbe,GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
|
BETA1_GPU_DEVICES=GPU-8ad38c6c-5a01-9d8e-1dfa-ed662ad78fbe,GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
|
||||||
|
|||||||
@@ -9,7 +9,7 @@ flowchart LR
|
|||||||
R --> I[FLUX.2 Klein 9B FP8 Beta<br/>RTX 5080 Transformer]
|
R --> I[FLUX.2 Klein 9B FP8 Beta<br/>RTX 5080 Transformer]
|
||||||
I --> E[Qwen3-8B NF4 Textencoder<br/>RTX 3060 während Bildauftrag]
|
I --> E[Qwen3-8B NF4 Textencoder<br/>RTX 3060 während Bildauftrag]
|
||||||
R --> T[Qwen3-TTS RTX 3060<br/>Piper CPU-Fallback]
|
R --> T[Qwen3-TTS RTX 3060<br/>Piper CPU-Fallback]
|
||||||
R --> STT[Whisper.cpp large-v3-turbo<br/>CPU, lokale Spracherkennung]
|
R --> STT[Whisper.cpp ggml-small<br/>CPU, lokale Spracherkennung]
|
||||||
|
|
||||||
H --> U[MUA / Unraid MCP]
|
H --> U[MUA / Unraid MCP]
|
||||||
H --> A[ARR-MCP]
|
H --> A[ARR-MCP]
|
||||||
@@ -50,7 +50,7 @@ Kontextgröße:
|
|||||||
| Medium | 160.000 Token |
|
| Medium | 160.000 Token |
|
||||||
| Large | 192.000 Token |
|
| Large | 192.000 Token |
|
||||||
| Ultra | 262.144 Token |
|
| Ultra | 262.144 Token |
|
||||||
| Beta 1 | 192.000 Token |
|
| Beta 1 | 112.000 Token |
|
||||||
| Uncensored | 80.000 Token |
|
| Uncensored | 80.000 Token |
|
||||||
|
|
||||||
## Exklusiver Bildmodus
|
## Exklusiver Bildmodus
|
||||||
|
|||||||
@@ -0,0 +1,49 @@
|
|||||||
|
# Container-Inventar auf Athena
|
||||||
|
|
||||||
|
Stand: 9. September 2026
|
||||||
|
|
||||||
|
Athena besteht derzeit aus 23 Docker-Containern. Nicht jeder Container enthält
|
||||||
|
ein KI-Modell: Router, Oberflächen, Netzwerk, Steuerung und Sicherung sind
|
||||||
|
gewöhnliche Dienste. Die rechenintensiven GPU-Worker werden absichtlich nur bei
|
||||||
|
Bedarf gestartet. Ein Container im Zustand `Created` oder `Exited (0)` ist daher
|
||||||
|
nicht automatisch ein ungenutzter Rest.
|
||||||
|
|
||||||
|
| Container | Modell oder wesentliche Komponente | Aufgabe |
|
||||||
|
|---|---|---|
|
||||||
|
| `mike-ai-backup` | kein Modell; Offen Docker Volume Backup | Sichert `/data`, `/etc/mike-ai`, den Stack und die persistenten Docker-Volumes im Fünf-Stunden-Takt. |
|
||||||
|
| `mike-ai-image-worker` | FLUX.2 Klein 9B FP8, Qwen3-8B NF4 Textencoder und VAE | Erzeugt und bearbeitet Bilder transaktional; nutzt während eines Auftrags RTX 5080 und RTX 3060. |
|
||||||
|
| `mike-ai-llama-beta1` | Qwen3.8-27B GSQ-RCO `IQ3_S-MTP`, Qwen-MMProj BF16 | Experimentelles Beta-Profil mit 112.000 Token Kontext; wird nur auf Anforderung geladen. |
|
||||||
|
| `mike-ai-llama-dashboard` | kein Modell | Zeigt Telemetrie, Profile, GPU-Nutzung und Betriebsarten an und bietet die Modusumschaltung. |
|
||||||
|
| `mike-ai-llama-fast` | Qwen3.8-27B `IQ4-MIX`, Qwen-MMProj BF16 | Schnelles Q4-Text-/Vision-Profil mit 76.800 Token Kontext. |
|
||||||
|
| `mike-ai-llama-large` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Q4-Text-/Vision-Profil mit 192.000 Token Kontext und Verteilung auf beide GPUs. |
|
||||||
|
| `mike-ai-llama-medium` | Qwen3.8-27B `IQ4_XS-pure`, Qwen-MMProj BF16 | Standard-Q4-Text-/Vision-Profil mit 160.000 Token Kontext und Verteilung auf beide GPUs. |
|
||||||
|
| `mike-ai-llama-ultra` | Qwen3.8-27B `IQ4_XS-pure`, ohne Vision-Projektor | Maximales Langkontextprofil mit 262.144 Token Kontext und Verteilung auf beide GPUs. |
|
||||||
|
| `mike-ai-llama-uncensored` | Qwen3.8-27B Abliterated `Q4_K_M`, eigener MMProj F16 | Spezialprofil mit 80.000 Token Kontext und gelockerten Modellgrenzen. |
|
||||||
|
| `mike-ai-mcp-athena-operator` | kein Modell | Stellt Hermes begrenzte Werkzeuge zum Prüfen, Ändern, Testen, Sichern und Versionieren von Athena bereit. |
|
||||||
|
| `mike-ai-music-acestep-test` | ACE-Step 1.5 XL-SFT und `acestep-5Hz-lm-1.7B` | Generiert Musik im exklusiven Musikmodus auf der RTX 5080. |
|
||||||
|
| `mike-ai-music-ui` | kein Modell; `fspecii/ace-step-ui` | Community-Oberfläche für ACE-Step; bleibt als leichte UI verfügbar, während der GPU-Worker bedarfsgesteuert läuft. |
|
||||||
|
| `mike-ai-piper` | `de_DE-thorsten-high` | CPU-basierte deutsche TTS-Rückfallebene, die auch während GPU-Umschaltungen verfügbar bleibt. |
|
||||||
|
| `mike-ai-portainer` | kein Modell; Portainer CE | Optionale Docker-Verwaltungsoberfläche. |
|
||||||
|
| `mike-ai-profile-controller` | kein Modell | Startet und stoppt ausschließlich freigegebene Modellprofile und Spezialworker in einer sicheren Reihenfolge. |
|
||||||
|
| `mike-ai-qwen3-tts` | `Qwen/Qwen3-TTS-12Hz-1.7B-Base`, Stimme Serena | Hochwertige deutsche Sprachausgabe auf der RTX 3060 im LLM-Betrieb. |
|
||||||
|
| `mike-ai-router` | kein eigenes Modell | Einzige OpenAI-kompatible Modelladresse; koordiniert Profile, Bildaufträge, Sprache und Betriebsarten. |
|
||||||
|
| `mike-ai-stem-separator` | BS-RoFormer Viperx 1297, `htdemucs_ft`, `htdemucs_6s`, `MossFormer2_SE_48K` | Trennt Gesang, Instrumente oder Sprache/Hintergrundgeräusche im exklusiven Separationsmodus. |
|
||||||
|
| `mike-ai-tts-gateway` | kein eigenes Modell | Normalisiert Text, leitet TTS an Qwen3-TTS weiter und fällt bei Bedarf auf Piper zurück. |
|
||||||
|
| `mike-ai-voice-studio` | `k2-fsa/OmniVoice` 0.2.1 mit Whisper-ASR | Erzeugt Text-to-Speech mit einer Referenzstimme; kein Audio-to-Audio-Voice-Changer. |
|
||||||
|
| `mike-ai-whisper` | Whisper.cpp `ggml-small` | Lokale deutsche Spracherkennung auf der CPU über `/v1/audio/transcriptions`. |
|
||||||
|
| `mike-ai-wireguard-gateway` | kein Modell | Veröffentlicht Dashboard und Fachoberflächen ausschließlich über den privaten WireGuard-Pfad. |
|
||||||
|
| `mike-ai-xvc-studio` | `chenxie95/X-VC` und GLM-4-Voice-Tokenizer | Wandelt eine vorhandene Sprachaufnahme anhand einer Referenzstimme in Audio zu Audio um. |
|
||||||
|
|
||||||
|
## Aufräumregel
|
||||||
|
|
||||||
|
Vor dem Löschen muss ein Kandidat gegen Compose-Dateien, Docker-Labels,
|
||||||
|
Mounts, Router-/Controller-Verweise und `/data` geprüft werden. Entfernt werden
|
||||||
|
nur nachweislich abgelöste Images, Gewichte, Versuchsdaten und Build-Caches.
|
||||||
|
Gewollt gestoppte Profilcontainer, persistente Modell-Caches und die letzte
|
||||||
|
funktionierende Produktionsvariante bleiben erhalten.
|
||||||
|
|
||||||
|
Am 9. September wurden die verworfenen Vevo2-Images und -Daten, das alte
|
||||||
|
Vevo2-Projektverzeichnis, ein leeres Test-Lab sowie der Docker-Build-Cache
|
||||||
|
entfernt. Der Build-Cache allein gab 74,43 GB frei; `/data` besitzt danach rund
|
||||||
|
562 GB freien Speicher. Kein produktiver oder bedarfsgesteuerter Container
|
||||||
|
wurde gelöscht.
|
||||||
@@ -40,7 +40,7 @@ akzeptiert werden. Details stehen in [FLUX_9B_BETA.md](FLUX_9B_BETA.md).
|
|||||||
|
|
||||||
## Lokale Spracherkennung
|
## Lokale Spracherkennung
|
||||||
|
|
||||||
Athena betreibt Whisper.cpp v1.9.1 mit `large-v3-turbo` als CPU-Dienst. Der
|
Athena betreibt Whisper.cpp v1.9.1 mit `ggml-small` als CPU-Dienst. Der
|
||||||
Profile Router veröffentlicht ihn als OpenAI-kompatiblen Endpunkt
|
Profile Router veröffentlicht ihn als OpenAI-kompatiblen Endpunkt
|
||||||
`/v1/audio/transcriptions`; Standardsprache ist Deutsch. Modell und Download
|
`/v1/audio/transcriptions`; Standardsprache ist Deutsch. Modell und Download
|
||||||
bleiben im persistenten Docker-Volume `whisper-data` erhalten.
|
bleiben im persistenten Docker-Volume `whisper-data` erhalten.
|
||||||
|
|||||||
@@ -1,6 +1,6 @@
|
|||||||
# Athena-Betriebsmodi
|
# Athena-Betriebsmodi
|
||||||
|
|
||||||
Athena besitzt sechs gegenseitig exklusive Betriebsmodi:
|
Athena besitzt fünf gegenseitig exklusive Betriebsmodi:
|
||||||
|
|
||||||
- `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt.
|
- `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt.
|
||||||
- `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt.
|
- `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt.
|
||||||
|
|||||||
@@ -55,8 +55,9 @@ Titelgenerierung und Kontextkompression in Hermes.
|
|||||||
|---|---|---|---|
|
|---|---|---|---|
|
||||||
| 05.09.2026 | Coqui XTTS v2 | deutsche Satzzeichen, Enden und Streaming-Chunks erzeugten Halluzinationen und unnatürliche Prosodie | **verworfen und entfernt** |
|
| 05.09.2026 | Coqui XTTS v2 | deutsche Satzzeichen, Enden und Streaming-Chunks erzeugten Halluzinationen und unnatürliche Prosodie | **verworfen und entfernt** |
|
||||||
| 05.09.2026 | `Qwen/Qwen3-TTS-12Hz-1.7B-Base` | deutlich natürlichere deutsche Ausgabe ohne die XTTS-Endhalluzinationen | **produktiv** mit Piper-Fallback |
|
| 05.09.2026 | `Qwen/Qwen3-TTS-12Hz-1.7B-Base` | deutlich natürlichere deutsche Ausgabe ohne die XTTS-Endhalluzinationen | **produktiv** mit Piper-Fallback |
|
||||||
| 06.–08.09.2026 | Whisper.cpp `large-v3-turbo` | lokaler TTS→STT-Rundlauf und OpenClaw-Transkription erfolgreich | **produktiv** |
|
| 06.–08.09.2026 | Whisper.cpp `large-v3-turbo` | lokaler TTS→STT-Rundlauf und OpenClaw-Transkription erfolgreich; später zugunsten des kleineren Laufzeitmodells entfernt | **ersetzt** |
|
||||||
| 09.09.2026 | `RMSnow/Vevo2`, Amphion `26f6883110181f1dbfe95c70a7c7dbaf4de5f42a` | Technik und Geschwindigkeit funktionierten, reale deutsche Sprachwandlung mit kurzer und langer Referenz war jedoch unverständlich, halluzinierend oder musikalisch | **qualitativ verworfen**; Container ersetzt, Image und Daten vorerst nur als Rollback erhalten |
|
| seit 03.09.2026 | Whisper.cpp `ggml-small` | tatsächlich im Compose-Stack und im laufenden Container verwendetes CPU-STT-Modell | **produktiv** |
|
||||||
|
| 09.09.2026 | `RMSnow/Vevo2`, Amphion `26f6883110181f1dbfe95c70a7c7dbaf4de5f42a` | Technik und Geschwindigkeit funktionierten, reale deutsche Sprachwandlung mit kurzer und langer Referenz war jedoch unverständlich, halluzinierend oder musikalisch | **qualitativ verworfen und entfernt**; Ergebnis bleibt hier dokumentiert, Images, Daten und altes Projekt wurden am 09.09. bereinigt |
|
||||||
| 09.09.2026 | `k2-fsa/OmniVoice` 0.2.1 | Offizielle Gradio-UI auf RTX 5080 gestartet; Modell plus Whisper-ASR belegen rund 3,7 GiB VRAM. `omnivoice-triton` 0.1.0 ist kompatibel im Image vorhanden, für den ersten Hörtest aber bewusst noch nicht aktiviert | **technischer Starttest bestanden**, Hörabnahme und Basis-vs.-Triton-Messung offen; Gewichte CC BY-NC und daher nur nichtkommerziell einsetzen |
|
| 09.09.2026 | `k2-fsa/OmniVoice` 0.2.1 | Offizielle Gradio-UI auf RTX 5080 gestartet; Modell plus Whisper-ASR belegen rund 3,7 GiB VRAM. `omnivoice-triton` 0.1.0 ist kompatibel im Image vorhanden, für den ersten Hörtest aber bewusst noch nicht aktiviert | **technischer Starttest bestanden**, Hörabnahme und Basis-vs.-Triton-Messung offen; Gewichte CC BY-NC und daher nur nichtkommerziell einsetzen |
|
||||||
| 09.09.2026 | `chenxie95/X-VC`, Code `49df8c591eafc48b096e466d96f9839f9c0dd739`, UI-Basis `d761cd6421e85376b2656dfefd8471d7f35a42be` | Offizielles Beispiel Ende-zu-Ende gewandelt: 5,20 s Audio in 1,07 s (RTF 0,21), gültiges 16-kHz-Mono-PCM-WAV; Modell belegt rund 2,9 GiB auf der RTX 5080. GLM-4-Voice-Tokenizer dokumentiert Chinesisch und Englisch | **technischer Start- und Konvertierungstest bestanden**; deutsche Hörabnahme offen |
|
| 09.09.2026 | `chenxie95/X-VC`, Code `49df8c591eafc48b096e466d96f9839f9c0dd739`, UI-Basis `d761cd6421e85376b2656dfefd8471d7f35a42be` | Offizielles Beispiel Ende-zu-Ende gewandelt: 5,20 s Audio in 1,07 s (RTF 0,21), gültiges 16-kHz-Mono-PCM-WAV; Modell belegt rund 2,9 GiB auf der RTX 5080. GLM-4-Voice-Tokenizer dokumentiert Chinesisch und Englisch | **technischer Start- und Konvertierungstest bestanden**; deutsche Hörabnahme offen |
|
||||||
|
|
||||||
|
|||||||
+1
-1
@@ -337,7 +337,7 @@ MEDIUM_CONTEXT=${MEDIUM_CONTEXT:-160000}
|
|||||||
MEDIUM_BATCH_SIZE=${MEDIUM_BATCH_SIZE:-2048}
|
MEDIUM_BATCH_SIZE=${MEDIUM_BATCH_SIZE:-2048}
|
||||||
MEDIUM_UBATCH_SIZE=${MEDIUM_UBATCH_SIZE:-128}
|
MEDIUM_UBATCH_SIZE=${MEDIUM_UBATCH_SIZE:-128}
|
||||||
MEDIUM_PARALLEL_SLOTS=${MEDIUM_PARALLEL_SLOTS:-1}
|
MEDIUM_PARALLEL_SLOTS=${MEDIUM_PARALLEL_SLOTS:-1}
|
||||||
BETA1_CONTEXT=${BETA1_CONTEXT:-192000}
|
BETA1_CONTEXT=${BETA1_CONTEXT:-112000}
|
||||||
BETA1_BATCH_SIZE=${BETA1_BATCH_SIZE:-2048}
|
BETA1_BATCH_SIZE=${BETA1_BATCH_SIZE:-2048}
|
||||||
BETA1_UBATCH_SIZE=${BETA1_UBATCH_SIZE:-128}
|
BETA1_UBATCH_SIZE=${BETA1_UBATCH_SIZE:-128}
|
||||||
BETA1_PARALLEL_SLOTS=${BETA1_PARALLEL_SLOTS:-1}
|
BETA1_PARALLEL_SLOTS=${BETA1_PARALLEL_SLOTS:-1}
|
||||||
|
|||||||
@@ -11,17 +11,26 @@ die() { printf 'FEHLER: %s\n' "$*" >&2; exit 1; }
|
|||||||
[[ -s $PROFILE_MATRIX ]] || die "Profilmatrix fehlt: $PROFILE_MATRIX"
|
[[ -s $PROFILE_MATRIX ]] || die "Profilmatrix fehlt: $PROFILE_MATRIX"
|
||||||
|
|
||||||
install_skill() {
|
install_skill() {
|
||||||
local source=$1 root=$2 name target
|
local source=$1 root=$2 name source_dir target_dir target path relative
|
||||||
name=${source%/SKILL.md}
|
name=${source%/SKILL.md}
|
||||||
name=${name##*/}
|
name=${name##*/}
|
||||||
target=$root/platform/$name/SKILL.md
|
source_dir=${source%/SKILL.md}
|
||||||
|
target_dir=$root/platform/$name
|
||||||
|
target=$target_dir/SKILL.md
|
||||||
grep -Fxq -- "name: $name" "$source" || \
|
grep -Fxq -- "name: $name" "$source" || \
|
||||||
die "Skill-Quelle hat kein gültiges Frontmatter: $source"
|
die "Skill-Quelle hat kein gültiges Frontmatter: $source"
|
||||||
install -d -o 10000 -g 10000 -m 0750 "${target%/*}"
|
install -d -o 10000 -g 10000 -m 0750 "$target_dir"
|
||||||
if [[ -s $target ]] && ! cmp -s "$source" "$target"; then
|
if [[ -s $target ]] && ! cmp -s "$source" "$target"; then
|
||||||
cp -a "$target" "$target.before-managed-update-$(date +%Y%m%d-%H%M%S)"
|
cp -a "$target" "$target.before-managed-update-$(date +%Y%m%d-%H%M%S)"
|
||||||
fi
|
fi
|
||||||
install -o 10000 -g 10000 -m 0640 "$source" "$target"
|
while IFS= read -r path; do
|
||||||
|
relative=${path#"$source_dir"/}
|
||||||
|
if [[ -d $path ]]; then
|
||||||
|
install -d -o 10000 -g 10000 -m 0750 "$target_dir/$relative"
|
||||||
|
else
|
||||||
|
install -o 10000 -g 10000 -m 0640 "$path" "$target_dir/$relative"
|
||||||
|
fi
|
||||||
|
done < <(find "$source_dir" -mindepth 1 -print)
|
||||||
cmp -s "$source" "$target" || die "Skill-Synchronisierung fehlgeschlagen: $target"
|
cmp -s "$source" "$target" || die "Skill-Synchronisierung fehlgeschlagen: $target"
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -1,44 +1,53 @@
|
|||||||
---
|
---
|
||||||
name: athena-operator
|
name: athena-operator
|
||||||
description: Understand, operate and extend the Athena AI host.
|
description: Operate and extend Mike's Athena AI host safely. Use for Athena models, profiles, Docker services, GPU allocation, inference modes, benchmarks, cleanup, deployment, backups, documentation, or when testing a new local AI model for Hermes.
|
||||||
license: MIT
|
license: MIT
|
||||||
metadata:
|
metadata:
|
||||||
hermes:
|
hermes:
|
||||||
version: 2.0.0
|
version: 3.0.0
|
||||||
author: Michael Roll
|
author: Michael Roll
|
||||||
platforms: [linux]
|
platforms: [linux]
|
||||||
tags: [athena, docker, mcp, models, backup]
|
tags: [athena, docker, gpu, models, benchmark, cleanup, backup]
|
||||||
---
|
---
|
||||||
|
|
||||||
# Athena Operator
|
# Athena Operator
|
||||||
|
|
||||||
Use this skill for work on Athena itself: Docker, MCPs, models, profiles,
|
Use this skill for work on Athena itself: Docker, models, profiles, Hermes
|
||||||
Hermes, OpenWebUI, TTS, STT, image generation, Git and backups.
|
integration, TTS, STT, image/audio/music workers, Git and backups.
|
||||||
|
|
||||||
## Start
|
## Start
|
||||||
|
|
||||||
1. Call `athena_operator_inspect` with `subject=guide`; it returns the current
|
1. Call `athena_operator_inspect` with `subject=guide`.
|
||||||
`ATHENA.md` as the architectural truth.
|
2. Choose and read the matching reference before acting:
|
||||||
2. Inspect the affected live area only if needed.
|
- architecture, containers or modes: `references/architecture-and-modes.md`
|
||||||
3. Search for the concrete source file, then read only the required lines.
|
- a model download, profile or A/B test: `references/model-evaluation.md`
|
||||||
|
- deployment, cleanup, documentation or publication:
|
||||||
|
`references/change-and-cleanup.md`
|
||||||
|
3. Inspect only the affected live area with `overview`, `containers`, `models`,
|
||||||
|
`jobs` or `git_status`.
|
||||||
|
4. Search for the exact source path before reading or editing it.
|
||||||
|
|
||||||
Do not rediscover the complete platform for every task. Do not read entire
|
Do not rediscover the complete platform for every task. Do not read entire
|
||||||
large files when a bounded section is enough. Do not guess file paths.
|
large files when a bounded section is enough. Do not guess file paths.
|
||||||
|
|
||||||
## Change
|
## Change
|
||||||
|
|
||||||
When the user has clearly requested a change, use `athena_operator_change` to
|
When the user clearly requests a change, use `athena_operator_change` to apply
|
||||||
apply the smallest durable change. The operator owns the Git worktree and
|
the smallest durable change. The operator owns the Git worktree and deployment
|
||||||
deployment access; do not clone another repository or request another SSH key.
|
access; do not clone another repository or request another SSH key.
|
||||||
|
|
||||||
Afterwards run focused checks, verify the affected service, commit and push.
|
Afterwards run focused checks, verify the affected service functionally, update
|
||||||
|
the affected documentation, commit and push.
|
||||||
The scheduled Docker-data backup is automatic. After storage-affecting work,
|
The scheduled Docker-data backup is automatic. After storage-affecting work,
|
||||||
run one manual backup and verify its archive instead of building a special
|
run one manual backup and verify its archive instead of building a special
|
||||||
recovery kit.
|
recovery kit.
|
||||||
|
|
||||||
For a new MCP, normally change only its server code, Dockerfile, MCP Compose
|
For every model candidate, preserve the currently working profile until the
|
||||||
service, env example, client registration and a focused test. Reuse an existing
|
candidate is downloaded and ready. Compare like with like, record the exact
|
||||||
backend instead of installing a duplicate service.
|
artifact and result in `docs/TESTED_MODELS.md`, then either promote it or remove
|
||||||
|
its weights and test-only runtime. Never silently lower a standard profile
|
||||||
|
below Q4; Q3 is allowed only after a documented comparison shows negligible
|
||||||
|
quality loss for the intended work.
|
||||||
|
|
||||||
## Tool discipline
|
## Tool discipline
|
||||||
|
|
||||||
@@ -48,6 +57,8 @@ backend instead of installing a duplicate service.
|
|||||||
- Prefer specialist MCPs for Home Assistant, Unraid, ARR, Navidrome and other
|
- Prefer specialist MCPs for Home Assistant, Unraid, ARR, Navidrome and other
|
||||||
external systems.
|
external systems.
|
||||||
- A healthy container is not proof; perform one bounded functional check.
|
- A healthy container is not proof; perform one bounded functional check.
|
||||||
|
- `Created` or cleanly stopped model workers are normal on-demand services, not
|
||||||
|
proof of garbage.
|
||||||
- Never claim a write, deploy, commit, push or backup succeeded without its
|
- Never claim a write, deploy, commit, push or backup succeeded without its
|
||||||
actual result.
|
actual result.
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,48 @@
|
|||||||
|
# Athena architecture and modes
|
||||||
|
|
||||||
|
Use `ATHENA.md` as the short operational truth and
|
||||||
|
`docs/CONTAINER_INVENTORY.md` for the current mapping of container, model and
|
||||||
|
role. If live state disagrees with documentation, report the discrepancy and
|
||||||
|
correct the durable source when the user requested maintenance.
|
||||||
|
|
||||||
|
## Boundaries
|
||||||
|
|
||||||
|
- Hermes, chats, skills and portable specialist MCPs live on Unraid.
|
||||||
|
- Athena is the inference host. Its canonical checkout is
|
||||||
|
`/opt/mike-ai/stack`; models live under `/data/models`; local secrets and
|
||||||
|
runtime configuration live under `/etc/mike-ai` and never enter Git.
|
||||||
|
- The Profile Router is the single OpenAI-compatible address clients use.
|
||||||
|
- The Profile Controller is the only component allowed to orchestrate approved
|
||||||
|
model and specialist workers.
|
||||||
|
- The Athena Operator is host-bound and is the normal maintenance interface.
|
||||||
|
|
||||||
|
## Exclusive states
|
||||||
|
|
||||||
|
Athena has five mutually exclusive persistent modes: `llm`, `music`,
|
||||||
|
`separation`, `voice`, and `voicechange`. Image generation is a transactional
|
||||||
|
request: it temporarily pauses the active text profile and Qwen3-TTS, runs the
|
||||||
|
image worker, then restores the previous LLM state.
|
||||||
|
|
||||||
|
Only one heavy GPU path may be active. Do not manually start a second GPU
|
||||||
|
worker around the controller. The lightweight dashboard, router, controller,
|
||||||
|
gateway, UI, CPU-STT, Piper, backup and operator containers may remain active.
|
||||||
|
|
||||||
|
## Model profiles
|
||||||
|
|
||||||
|
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
|
||||||
|
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
|
||||||
|
- Large: the same Q4 model, 192,000 tokens, vision.
|
||||||
|
- Ultra: the same Q4 model, 262,144 tokens, no vision projector.
|
||||||
|
- Beta 1: Qwen3.8-27B GSQ-RCO IQ3_S-MTP, 112,000 tokens; experimental only.
|
||||||
|
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
|
||||||
|
|
||||||
|
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
|
||||||
|
infer allocation from model size alone; verify the profile's Compose arguments
|
||||||
|
and live VRAM. RTX 3060 also hosts Qwen3-TTS during normal LLM operation.
|
||||||
|
|
||||||
|
## What is and is not stale
|
||||||
|
|
||||||
|
The GPU workers use `restart: "no"` and are created once, then started on
|
||||||
|
demand. A stopped `mike-ai-llama-*`, image, music, separator, OmniVoice or X-VC
|
||||||
|
container is expected. A candidate is stale only after checking Compose,
|
||||||
|
labels, mounts, router/controller references, model paths and test history.
|
||||||
@@ -0,0 +1,43 @@
|
|||||||
|
# Durable changes, cleanup and publication
|
||||||
|
|
||||||
|
## Change flow
|
||||||
|
|
||||||
|
1. Inspect `guide` and the affected live subject.
|
||||||
|
2. Check `git_status`; preserve unrelated user changes.
|
||||||
|
3. Search the canonical checkout and edit the smallest source of truth.
|
||||||
|
4. Validate syntax and Compose before deployment.
|
||||||
|
5. Deploy only the affected service unless the requested change genuinely
|
||||||
|
spans the core stack.
|
||||||
|
6. Verify container health and one real function, not just process existence.
|
||||||
|
7. Update documentation and the container/model inventory when architecture,
|
||||||
|
models, ports, modes or ownership changed.
|
||||||
|
8. Commit and push only after checks pass. Verify the remote result.
|
||||||
|
|
||||||
|
Use an asynchronous operator job only once and poll it with
|
||||||
|
`athena_operator_job`; never launch a duplicate because a long task is quiet.
|
||||||
|
|
||||||
|
## Cleanup proof
|
||||||
|
|
||||||
|
Before deletion, prove that an item is unused by checking:
|
||||||
|
|
||||||
|
- current Compose projects and Docker labels;
|
||||||
|
- router and controller source references;
|
||||||
|
- mounts, volumes and model manifests;
|
||||||
|
- `docs/TESTED_MODELS.md` and any rollback requirement;
|
||||||
|
- whether a stopped container is an intentional on-demand worker.
|
||||||
|
|
||||||
|
Prefer precise targets. Never perform broad recursive deletion from `/`,
|
||||||
|
`/data`, `/opt` or a variable that was not resolved and printed first. Build
|
||||||
|
cache can be pruned after confirming no build is running. Do not prune named
|
||||||
|
volumes or active images generically.
|
||||||
|
|
||||||
|
After storage-affecting work, report exactly what was removed and free space,
|
||||||
|
run the existing manual backup operation, and verify the produced archive.
|
||||||
|
|
||||||
|
## Safety and secrets
|
||||||
|
|
||||||
|
Never reboot or shut down Athena or alter SSH, networking, WireGuard, firewall,
|
||||||
|
kernel, boot, partitions or mounts without a separate explicit current user
|
||||||
|
instruction. Do not print environment dumps, tokens, API keys, private keys or
|
||||||
|
the contents of `/etc/mike-ai`. Redact accidental secret material from reports
|
||||||
|
and never commit it.
|
||||||
@@ -0,0 +1,41 @@
|
|||||||
|
# Model evaluation
|
||||||
|
|
||||||
|
## Before downloading
|
||||||
|
|
||||||
|
1. Read all of `docs/TESTED_MODELS.md`; it is the no-repeat register.
|
||||||
|
2. Record the exact repository, revision, filename, base model, fine-tune,
|
||||||
|
quantization, license, format and estimated disk/VRAM/RAM requirements.
|
||||||
|
3. Confirm the model adds a genuinely new candidate rather than a renamed
|
||||||
|
artifact already tested.
|
||||||
|
4. Do not stop the active production model merely to download or prepare the
|
||||||
|
candidate. Put temporary scripts under `/tmp`; durable code belongs in Git.
|
||||||
|
|
||||||
|
## Quality rule
|
||||||
|
|
||||||
|
Standard text profiles must not fall below Q4. A smaller Q3 quantization may be
|
||||||
|
kept only when an A/B test documents that its quality loss is negligible for
|
||||||
|
Mike's intended workload. Smaller files are not presumed faster: GPU split,
|
||||||
|
memory bandwidth, kernels, cache formats and cross-GPU traffic must be measured.
|
||||||
|
|
||||||
|
## Fair A/B test
|
||||||
|
|
||||||
|
Hold these equal wherever the models permit it:
|
||||||
|
|
||||||
|
- prompt set and conversation history;
|
||||||
|
- context size and filled-context test point;
|
||||||
|
- KV-cache quantization, slots, batch/uBatch, MTP and sampling;
|
||||||
|
- GPU visibility and tensor split;
|
||||||
|
- warm-up state and output-token limit.
|
||||||
|
|
||||||
|
Measure prompt processing, short decode, long-context decode, peak VRAM and
|
||||||
|
wall time. Test meaning preservation, uncertainty, negation, ordered safety
|
||||||
|
constraints, German language consistency, tool-call schema and long-context
|
||||||
|
recall. Do not replace a model on synthetic benchmark scores alone.
|
||||||
|
|
||||||
|
## Completion
|
||||||
|
|
||||||
|
Write the exact artifact, settings, raw result path, interpretation and decision
|
||||||
|
to `docs/TESTED_MODELS.md` in the same commit. If promoted, update the manifest,
|
||||||
|
Compose/env examples, profile matrix and user documentation. If rejected,
|
||||||
|
remove candidate-only weights, images and containers after preserving the
|
||||||
|
result. Restore and functionally test the previous profile, then publish.
|
||||||
Reference in New Issue
Block a user