Synchronize repository with Athena deployment

This commit is contained in:
Mikei386
2026-09-13 20:01:36 +02:00
parent 040a2df48b
commit fce9900389
60 changed files with 2570 additions and 607 deletions
-70
View File
@@ -1,70 +0,0 @@
#!/bin/sh
set -eu
target=/opt/hermes/cron/scheduler_delivery.py
broken='env.pop("HERMES_HOME", None)'
if [ -f "$target" ] && grep -Fq "$broken" "$target"; then
sed -i '/^[[:space:]]*env\.pop("HERMES_HOME", None)[[:space:]]*$/d' "$target"
echo "[cron-profile-root-fix] preserved HERMES_HOME for named-profile delivery"
else
echo "[cron-profile-root-fix] upstream code already fixed or layout changed; no action"
fi
cli_dir="${HERMES_HOME:-/opt/data}/.local/bin"
mkdir -p "$cli_dir"
ln -sfn /opt/hermes/.venv/bin/hermes "$cli_dir/hermes"
ln -sfn /opt/hermes/.venv/bin/hermes /usr/local/bin/hermes
# Hermes' generic sentence splitter treats periods in German abbreviations as
# sentence ends. It also submits every sentence as a separate generative TTS
# request, which creates long gaps with local Qwen3-TTS. Keep the first sentence
# immediate, then group following complete sentences into modest ~240-character
# chunks so playback remains responsive but substantially more continuous.
tts_target=/opt/hermes/tools/tts_streaming.py
if [ -f "$tts_target" ] && grep -Fq 'self.buf = _THINK_BLOCK_RE.sub("", self.buf + delta)' "$tts_target"; then
TTS_TARGET="$tts_target" python3 <<'PY'
import os
from pathlib import Path
p = Path(os.environ["TTS_TARGET"])
s = p.read_text()
base_feed = ' self.buf = _THINK_BLOCK_RE.sub("", self.buf + delta)'
abbr_lines = [
' self.buf = re.sub(r"\\bca\\.(?=\\s)", "circa", self.buf, flags=re.IGNORECASE)',
' self.buf = re.sub(r"\\bv\\.\\s*a\\.(?=\\s)", "vor allem", self.buf, flags=re.IGNORECASE)',
' self.buf = re.sub(r"\\bmax\\.(?=\\s)", "maximal", self.buf, flags=re.IGNORECASE)',
]
for line in abbr_lines:
s = s.replace("\n" + line, "")
init_old = ' self.buf = ""'
init_extra = (
'\n self.followup_target_len = 240'
'\n self._emitted_first = False'
)
s = s.replace(init_old + init_extra, init_old)
threshold_old = ' if len(head.strip()) < self.min_len:'
threshold_new = (
' threshold = (self.min_len if not self._emitted_first '
'else self.followup_target_len)\n'
' if len(head.strip()) < threshold:'
)
s = s.replace(threshold_new, threshold_old)
append_old = ' out.append(head)'
append_new = append_old + '\n self._emitted_first = True'
s = s.replace(append_new, append_old)
s = s.replace(base_feed, base_feed + "\n" + "\n".join(abbr_lines), 1)
s = s.replace(init_old, init_old + init_extra, 1)
s = s.replace(threshold_old, threshold_new, 1)
s = s.replace(append_old, append_new, 1)
p.write_text(s)
PY
echo "[tts-streaming-fix] enabled German normalization and adaptive sentence grouping"
else
echo "[tts-streaming-fix] upstream code already fixed or layout changed; no action"
fi
+2 -3
View File
@@ -144,13 +144,12 @@ stt:
model: "base"
language: "de"
# Reuse Athena's OpenAI-compatible TTS route. It currently serves XTTS v2 with
# Annmarie Nele and transparently falls back to Piper when XTTS is unavailable.
# Reuse Athena's OpenAI-compatible Qwen3-TTS route.
tts:
provider: "openai"
speed: 1.0
openai:
model: "piper"
model: "qwen3-tts"
voice: "alloy"
speed: 1.0
base_url: "http://router:8081/v1"
@@ -1,120 +0,0 @@
---
name: athena-ai-profile-router
description: Use when operating the AI profile router on Athena.
metadata:
hermes:
editorial_name: "Athena AI-Profile-Router"
editorial_description: "Betrieb und Abfragen der lokalen KI-Plattform auf Athena: Profile, Status, Umschaltung."
---
# Athena AI-Profile-Router
Lokale KI-Plattform auf **Athena** (root@192.168.1.212, Debian 13, GPU-Host).
Repo: `/root/AI-Profile-Router` (Gitea: `michael/AI-Profile-Router`).
SSH: `ssh -i /opt/data/athena_key -o UserKnownHostsFile=/opt/data/.ssh/known_hosts -o IdentitiesOnly=yes root@192.168.1.212`
## Wichtigste Regel
**Der Host steht physisch in einer anderen Stadt. Es gibt keinen schnellen
Zugang.**
- **NIEMALS** herunterfahren (`shutdown`, `poweroff`, `halt`) oder neu
starten (`reboot`, `init 6`).
- Keine Aktion, die die Erreichbarkeit gefährdet: keine Änderungen an `lan0`,
Firewall-Default-Policies, `DOCKER-USER`-Regeln oder Routing ohne explizite
Freigabe und Rückfallplan; keine Reboot-Pflicht auslösenden Aktionen
(Kernel-, NVIDIA-Treiber-Installation, `apt full-upgrade` mit Kernel) ohne
ausdrückliche Freigabe; keine Abschaltung von WireGuard, SSH oder dem
Gateway-Container.
- Installer-Exit-Code 20 (NVIDIA-Treiber, Reboot nötig) wird nicht automatisch
nachgeholt — der Betreiber entscheidet.
- Nach jeder Netz-/Firewall-/WG-Änderung verifizieren: (1) SSH-Roundtrip,
(2) WG-Tunnel up, (3) Router `/health` und `/ready` = 200.
## Architektur (Kurzform)
- **Profile Router** (`mike-ai-router`, OpenAI-kompatible API, Port 8081):
einzige Client-Schnittstelle. Clients nutzen `http://<host>:8081/v1` mit
Bearer-Key aus `/etc/mike-ai/router-api-key`.
- **llama.cpp** (`mike-ai-llama-*`): ein Container pro Profil, immer exakt
einer aktiv; nur Docker-intern (Port 8080, nie direkt von Clients).
- **Profile Controller** (`mike-ai-profile-controller`): einziger Dienst mit
Docker-Socket; begrenzter Containerwechsel.
- **Open WebUI** (`mike-ai-open-webui`, Port 8080): einzige normale Oberfläche.
- **MCP-Tool-Stack** (`platform/mcp/`): getrennte Container für Web, HA, ARR,
Unraid — nur über internes `mike-ai-tools`-Netz.
- WireGuard-Isolation: KI-Dienste nur über WG-Adresse erreichbar,
fail-closed bei Tunnelausfall (Blackhole-Default-Route).
## Profile (Live-Stand 11.09.2026)
| Profil | Virtuelles Modell | Kontext | Zweck |
|---|---|---:|---|
| fast | `qwen-fast` | 76.800 | Alltag, Agenten, hohe Geschwindigkeit |
| medium | `qwen-medium` | 160.000 | mehr Kontext |
| large | `qwen-large` | 192.000 | lange Sitzungen |
| ultra | `qwen-ultra` | 262.144 | maximale Kontextlänge |
| uncensored | `qwen-uncensored` | 80.000 | unzensiert |
Aktives Profil: `GET /status` → `current_profile`.
Profilwechsel: `POST /fast`, `/medium`, `/large`, `/ultra`, `/uncensored`
(authentifiziert) oder manuell `llama-profile <name>` auf dem Host.
## Wichtige Befehle (per SSH auf Athena)
```bash
# Status (Router lebt, Profil, Upstream, Telemetrie)
KEY=$(cat /etc/mike-ai/router-api-key)
curl -s -H "Authorization: Bearer ***" http://172.30.20.3:8081/status
# Health / Ready (ohne Auth)
curl -s -o /dev/null -w '%{http_code}' http://172.30.20.3:8081/health # 200 = Router lebt
curl -s -o /dev/null -w '%{http_code}' http://172.30.20.3:8081/ready # 200 = Modell bereit
# Virtuelle Modelle
curl -s -H "Authorization: Bearer ***" http://172.30.20.3:8081/v1/models
# Profilwechsel
curl -s -X POST -H "Authorization: Bearer ***" http://172.30.20.3:8081/ultra
# Aktive Profile-Registry (Router-Container)
docker exec mike-ai-router cat /app/router_profiles.json
# Container-Status
docker ps --format '{{.Names}}\t{{.Status}}' | grep mike-ai
```
Router-IP auf dem internen `mike-ai_control`-Netz: `172.30.20.3`
(neu ermitteln: `docker inspect mike-ai-router --format '{{.NetworkSettings.Networks.mike-ai_control.IPAddress}}'`).
## Fehler- und Recovery-Verhalten
- `/health=200` + `/ready=503` → Router lebt, Modell lädt/wechselt noch.
- `429` → Parallelitätsgrenze (max. 16) erreicht; Client mit Backoff.
- Profilwechsel bricht ab, wenn laufende Chats den Drain-Timeout (600 s)
überschreiten — er beendet niemals absichtlich einen Chat.
- Nach Routerabsturz: letztes stabiles Profil aus atomarer Zustandsdatei
(`/var/lib/mike-ai-profile-router/state.json`) rekonstruiert.
- Upgrade-Regel: niemals Build, Quantisierung und Profil gleichzeitig ändern.
## Git / Repo
- Remote: `git@192.168.1.2:michael/AI-Profile-Router.git` (Gitea auf Unraid).
- Athena erreicht das Heimnetz **nur über WireGuard** — ist der WG-Tunnel
down, schlägt `git push` mit Connection timed out fehl. Nicht mit
Firewall-/Routing-Änderungen gegensteuern; Tunnel-Status prüfen und
melden. Commits bleiben lokal auf Athena, bis der Tunnel wieder up ist.
- Git-Identität auf Athena ist repo-lokal gesetzt (Mikei386).
## Sicherheitsregeln
- API-Key (`/etc/mike-ai/router-api-key`) nie in Chat, Repo oder URLs.
- Clients nie direkt Port 8080 (llama.cpp) — immer über Router 8081.
- `config/install.env` bleibt lokal (0600), wird nicht committet.
- Systempartition dauerhaft unter 85 % halten; nur ein Textmodell gleichzeitig.
## Backup
Gesichert: Repo, Modellmanifest (Hashes, ohne Secrets), `/etc/mike-ai`
verschlüsselt, systemd-Konfiguration, Benchmarkresultate.
Nicht nötig: Builds, Venvs, Caches, Modelle (Quelle + Prüfsumme dokumentiert).
@@ -0,0 +1,47 @@
# Athena architecture and modes
Use `ATHENA.md` as the short operational truth and
`docs/CONTAINER_INVENTORY.md` for the current mapping of container, model and
role. If live state disagrees with documentation, report the discrepancy and
correct the durable source when the user requested maintenance.
## Boundaries
- Hermes, chats, skills and portable specialist MCPs live on Unraid.
- Athena is the inference host. Its canonical checkout is
`/opt/mike-ai/stack`; models live under `/data/models`; local secrets and
runtime configuration live under `/etc/mike-ai` and never enter Git.
- The Profile Router is the single OpenAI-compatible address clients use.
- The Profile Controller is the only component allowed to orchestrate approved
model and specialist workers.
- The Athena Operator is host-bound and is the normal maintenance interface.
## Exclusive states
Athena has five mutually exclusive persistent modes: `llm`, `music`,
`separation`, `voice`, and `voicechange`. Image generation is a transactional
request: it temporarily pauses the active text profile and Qwen3-TTS, runs the
image worker, then restores the previous LLM state.
Only one heavy GPU path may be active. Do not manually start a second GPU
worker around the controller. The lightweight dashboard, router, controller,
gateway, UI, CPU-STT, backup and operator containers may remain active.
## Model profiles
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
- Large: the same Q4 model, 192,000 tokens, vision.
- Ultra: the same Q4 model, 262,144 tokens, no vision projector.
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
infer allocation from model size alone; verify the profile's Compose arguments
and live VRAM. RTX 3060 also hosts Qwen3-TTS during normal LLM operation.
## What is and is not stale
The GPU workers use `restart: "no"` and are created once, then started on
demand. A stopped `mike-ai-llama-*`, image, music, separator, OmniVoice or X-VC
container is expected. A candidate is stale only after checking Compose,
labels, mounts, router/controller references, model paths and test history.