diff --git a/ATHENA.md b/ATHENA.md index 8a4d9fc..3d980be 100644 --- a/ATHENA.md +++ b/ATHENA.md @@ -13,9 +13,9 @@ Fehlersuche sie benötigt. - Lokale Konfiguration und Secrets: `/etc/mike-ai` (niemals in Git). - Benutzerzugriff auf KI-Dienste: über WireGuard, nicht über das Uni-LAN. - OpenAI-kompatible Modell-API: Profile Router auf Port 8081. -- Oberfläche auf Athena: OpenWebUI. Der offizielle Hermes Agent läuft auf - Unraid unter `/mnt/nvme-storage/appdata/Hermes-Agent` und verwendet Athenas - Router als Modell-Backend. +- Oberfläche und Agent: Hermes auf Unraid unter + `/mnt/nvme-storage/appdata/Hermes-Agent`. Athena stellt dafür nur die + OpenAI-kompatible Router-API bereit. OpenWebUI ist abgeschaltetes Rückfallnetz. - Inferenz: genau ein aktives llama.cpp-Textprofil; der Router wechselt bei Bedarf zwischen Fast, Medium, Large, Ultra und Uncensored. @@ -35,16 +35,11 @@ Migration höchstens als gekennzeichnetes Archiv existieren. ## Container-Prinzip -Ein eigener Container ist sinnvoll, wenn ein Dienst eigene Abhängigkeiten, -Zugangsdaten oder eine eigene Fehlergrenze hat. Deshalb bleiben die fachlichen -MCPs getrennt, beispielsweise Home Assistant, Unraid, ARR, Navidrome, Deemix, -Web und SSH. Es gibt jedoch keine zusätzlichen MCPs für einzelne -Zwischenschritte einer Installation. - -Hermes auf Unraid ist die primäre Oberfläche für längere administrative und -agentische Aufgaben. OpenWebUI auf Athena erhält weiterhin die Werkzeuge, die -für kurze Abfragen sinnvoll sind. Ein neuer MCP muss nicht automatisch in jede -Oberfläche eingebunden werden; das richtet sich nach dem Auftrag. +Athena betreibt nur Inferenz, Router, Sprache/Bild, Backup und den +hostgebundenen Athena-Operator. Portable Fach-MCPs laufen als getrennte +Unterprozesse im **einen MCPHub-Container auf Unraid**. Sie bleiben unter +`/mcp/NAME` einzeln sichtbar und abschaltbar, brauchen aber nicht je einen +Docker-Container. `config/mcp-registry.json` ist die einzige Serverliste. ## Standardablauf für Änderungen @@ -63,18 +58,14 @@ Werkzeugaufruf wird höchstens einmal wiederholt. ## Neuer MCP -Für einen neuen MCP sind gewöhnlich nur diese Teile nötig: +Portable MCPs werden nach dem Skill `mcphub-deployer` in das reproduzierbare +MCPHub-Image eingebaut. Dazu gehören gepinnte Quelle, genau ein Registry-Eintrag, +Secret-Datei nur im Appdata, Build, Handshake und eine read-only-Probe. Danach +wird ausschließlich MCPHub über Unraid DockerMan neu erstellt. Router, Qwen, +WireGuard und andere Container werden nicht neu gestartet. -1. Servercode und Dockerfile unter `platform/mcp/`. -2. Ein Service im eingebundenen `platform/mcp/compose.yaml`. -3. Eine Env-Beispieldatei unter `config/`; echte Werte nach `/etc/mike-ai`. -4. Genau ein Eintrag in `config/mcp-registry.json`; daraus werden Hermes und - OpenWebUI automatisch erzeugt. -5. Ein kleiner Test sowie ein kurzer Eintrag in dieser Datei oder in der - Komponentenübersicht, falls wirklich zusätzliche Erklärung nötig ist. - -Dann wird nur der neue MCP gebaut und gestartet. Router, Qwen, WireGuard, -Hermes und der komplette Stack werden nicht pauschal neu gestartet. +Nur ein Werkzeug, das Athenas Host selbst verwalten muss, gehört in den +Athena-Operator. Es wird kein zweiter allgemeiner Terminal- oder Doku-MCP gebaut. ## Temporär oder dauerhaft diff --git a/README.md b/README.md index 50497e0..297c4de 100644 --- a/README.md +++ b/README.md @@ -1,9 +1,8 @@ # Athena AI -Ein reproduzierbarer Docker-Stack für Athenas lokale KI. Athena stellt Router, -llama.cpp-Profile, OpenWebUI, Sprache und Bildgenerierung bereit. Der offizielle -Hermes Agent und die portablen Werkzeuge laufen auf Unraid und werden dort -zusammen mit dem übrigen Appdata gesichert. +Ein reproduzierbarer Docker-Stack für Athenas lokale Inferenz. Athena stellt +Router, llama.cpp-Profile, Sprache und Bildgenerierung bereit. Der offizielle +Hermes Agent und MCPHub laufen auf Unraid und werden dort mit Appdata gesichert. ## Aufbau @@ -21,11 +20,10 @@ zusammen mit dem übrigen Appdata gesichert. - Home Assistant, ARR, Unraid, Navidrome, Deemix und GitHub laufen gemeinsam im MCPHub-Container auf Unraid, bleiben aber als getrennte MCP-Server unter `/mcp/NAME` sichtbar, abschaltbar und unabhängig für Clients freigebbar. -- SearXNG/TinySearch dürfen als eigene Such-Backends laufen. Der eigentliche - Web-MCP-Adapter zieht ebenfalls in den MCPHub, sobald der Backend-Pfad aus dem - Unraid-Netz verfügbar ist. -- `config/mcp-registry.json` ist die einzige Liste der MCPs für Hermes und - OpenWebUI. `platform/mcp/sync-clients.py` erzeugt beide Registrierungen. +- Hermes verwendet für allgemeine Recherche seine integrierten Webwerkzeuge; + der frühere Athena-Webadapter wird nicht mehr gestartet. +- `config/mcp-registry.json` ist die einzige Liste der MCPHub-Server und ihrer + Client-Registrierungen. `platform/mcp/sync-clients.py` erzeugt Hermes daraus. - Modelle und Athena-Backups liegen auf `/data`. Hermes liegt vollständig unter `/mnt/nvme-storage/appdata/Hermes-Agent`; MCPHub-Zustand, Client-Schlüssel und MCP-Zugänge liegen unter `/mnt/nvme-storage/appdata/MCPHub`. @@ -43,7 +41,7 @@ sudo ./install.sh --config config/install.env ``` Das Skript installiert Docker und NVIDIA-Unterstützung, lädt die konfigurierten -Modelle, baut den Stack und startet die benötigten Profile und MCPs. +Modelle und startet ausschließlich Athenas Inferenz-Kern plus Operator. ## Bedienung @@ -51,14 +49,16 @@ Modelle, baut den Stack und startet die benötigten Profile und MCPs. # Gesamten Stack anzeigen docker compose --env-file /etc/mike-ai/stack.env ps -# Eine gezielte Athena-Komponente ausrollen -docker compose --env-file /etc/mike-ai/stack.env up -d --build router +# Erst anzeigen, dann eine gezielte Komponente ohne Nebenwirkungen ausrollen +./manage.sh --dry-run deploy router +./manage.sh deploy router # Sofortiges Datenbackup zusätzlich zum Fünf-Stunden-Zeitplan docker exec mike-ai-backup backup -``` -OpenWebUI: `http://:8080` +# Kurzer read-only Ende-zu-Ende-Test nach jedem Release +sudo ./smoke-test.sh +``` Router-API: `http://:8081/v1` @@ -81,6 +81,7 @@ Details, Prüfschritte und der exakte Sicherungsumfang stehen in - [`ATHENA.md`](ATHENA.md) – kurze Maschinen- und Operatoranleitung - [`docs/STANDARD_PROFILE_MATRIX.md`](docs/STANDARD_PROFILE_MATRIX.md) – Profile und Messwerte +- [`docs/MCP_SERVERS.md`](docs/MCP_SERVERS.md) – automatisch erzeugte MCP-Liste - [`docs/RECOVERY.md`](docs/RECOVERY.md) – Backup und Neuaufbau Die MCPHub-Installation, Endpunkte und der schrittweise Rückbau der alten diff --git a/compose.yaml b/compose.yaml index ab124b2..e49d1ea 100644 --- a/compose.yaml +++ b/compose.yaml @@ -928,11 +928,9 @@ services: - /data/docker-backups:/archive - /etc/mike-ai:/backup/etc-mike-ai:ro - /opt/mike-ai/stack:/backup/stack:ro - - open-webui-data:/backup/volumes/open-webui-data:ro - piper-data:/backup/volumes/piper-data:ro - router-state:/backup/volumes/router-state:ro - router-images:/backup/volumes/router-images:ro - - tinysearch-models:/backup/volumes/tinysearch-models:ro security_opt: ["no-new-privileges:true"] networks: diff --git a/config/mcp-registry.json b/config/mcp-registry.json index 5a2daca..e864ab9 100644 --- a/config/mcp-registry.json +++ b/config/mcp-registry.json @@ -11,16 +11,13 @@ "env_file": "/etc/mike-ai/mcphub-client.env", "key_env": "MCPHUB_BEARER_TOKEN", "auth_type": "bearer", - "timeout": 900 - }, - { - "id": "web-general-local", - "hermes_id": "web-general", - "name": "Allgemeines Web (TinySearch)", - "description": "Breite Websuche und Seitenabruf für beliebige öffentliche Websites. Kurz und gezielt suchen; keine vollständigen Websites rekursiv einlesen.", - "url": "http://tinysearch:8000/mcp", - "clients": ["hermes", "openwebui"], - "timeout": 180 + "timeout": 900, + "hub": { + "type": "streamable-http", + "url": "http://192.168.1.212:8202/mcp", + "owner": "admin", + "enabled": true + } }, { "id": "github-local", @@ -33,7 +30,13 @@ "key_env": "MCPHUB_BEARER_TOKEN", "auth_type": "bearer", "timeout": 300, - "functions": "github-search_repositories,github-get_file_contents,github-search_code" + "functions": "github-search_repositories,github-get_file_contents,github-search_code", + "hub": { + "type": "stdio", + "command": "/usr/local/bin/run-with-env", + "args": ["/run/secrets/mcphub/github.env", "--", "/usr/local/bin/github-mcp-server", "stdio", "--read-only", "--tools", "search_repositories,get_file_contents,search_code"], + "enabled": true + } }, { "id": "homeassistant-local", @@ -45,7 +48,15 @@ "env_file": "/etc/mike-ai/mcphub-client.env", "key_env": "MCPHUB_BEARER_TOKEN", "auth_type": "bearer", - "timeout": 300 + "timeout": 300, + "hub": { + "type": "streamable-http", + "secret_file": "homeassistant.env", + "url": "${HASS_URL}/api/hass_mcp", + "headers": {"Authorization": "Bearer ${HASS_TOKEN}"}, + "owner": "admin", + "enabled": true + } }, { "id": "arr-local", @@ -57,7 +68,13 @@ "env_file": "/etc/mike-ai/mcphub-client.env", "key_env": "MCPHUB_BEARER_TOKEN", "auth_type": "bearer", - "timeout": 600 + "timeout": 600, + "hub": { + "type": "stdio", + "command": "/usr/local/bin/run-with-env", + "args": ["/run/secrets/mcphub/arr.env", "--", "arr-mcp", "--transport", "stdio", "--auth-type", "none"], + "enabled": true + } }, { "id": "navidrome-local", @@ -69,7 +86,14 @@ "env_file": "/etc/mike-ai/mcphub-client.env", "key_env": "MCPHUB_BEARER_TOKEN", "auth_type": "bearer", - "timeout": 300 + "timeout": 300, + "hub": { + "type": "stdio", + "command": "/usr/local/bin/run-with-env", + "args": ["/run/secrets/mcphub/navidrome.env", "--", "node", "/opt/casaderoll/navidrome/dist/index.js"], + "env": {"MCP_TRANSPORT": "stdio", "MCP_HTTP_EXPOSE": "false", "WEBUI_ENABLED": "false"}, + "enabled": true + } }, { "id": "deemix-local", @@ -81,7 +105,14 @@ "env_file": "/etc/mike-ai/mcphub-client.env", "key_env": "MCPHUB_BEARER_TOKEN", "auth_type": "bearer", - "timeout": 300 + "timeout": 300, + "hub": { + "type": "stdio", + "command": "/usr/local/bin/run-with-env", + "args": ["/run/secrets/mcphub/deemix.env", "--", "python3", "/opt/casaderoll/mcps/deemix_mcp.py"], + "env": {"MCP_TRANSPORT": "stdio"}, + "enabled": true + } }, { "id": "mua", @@ -93,7 +124,39 @@ "env_file": "/etc/mike-ai/mcphub-client.env", "auth_type": "bearer", "clients": ["hermes", "openwebui"], - "timeout": 900 + "timeout": 900, + "hub": { + "type": "streamable-http", + "secret_file": "mua.env", + "url": "${MUA_MCP_URL}", + "headers": {"Authorization": "Bearer ${MUA_MCP_BEARER_TOKEN}"}, + "owner": "admin", + "enabled": true + } + }, + { + "id": "fritzbox-local", + "hermes_id": "fritzbox", + "name": "FRITZ!Box", + "description": "FRITZ!Box-Status, WAN/Glasfaser, Netzwerkgeräte, WLAN und Telefonie. Zuerst die aktive WAN-Verbindung ermitteln; Änderungen nur auf ausdrücklichen Auftrag.", + "url": "http://192.168.1.2:8787/mcp/fritzbox", + "key_env": "MCPHUB_BEARER_TOKEN", + "env_file": "/etc/mike-ai/mcphub-client.env", + "auth_type": "bearer", + "clients": ["hermes", "openwebui"], + "timeout": 300, + "tool_include": [ + "fritzbox-list_services", + "fritzbox-list_actions", + "fritzbox-describe_action", + "fritzbox-call_action" + ], + "hub": { + "type": "stdio", + "command": "/usr/local/bin/run-with-env", + "args": ["/run/secrets/mcphub/fritzbox.env", "--", "/opt/casaderoll/fritz-mcp"], + "enabled": true + } }, { "id": "mua-readonly-local", diff --git a/config/profile-matrix.json b/config/profile-matrix.json new file mode 100644 index 0000000..3d83a5b --- /dev/null +++ b/config/profile-matrix.json @@ -0,0 +1,62 @@ +{ + "version": 1, + "default_profile": "medium", + "max_output_tokens": 8192, + "profiles": [ + { + "id": "fast", + "alias": "qwen-fast", + "context": 76800, + "model_env": "FAST_MODEL_FILE", + "model_family": "Qwen3.8-27B IQ4 Mix", + "gpu_split": "5080 only", + "vision": true, + "mtp": 2, + "description": "Schnelles Profil für kurze Chats und zügige Werkzeugaufgaben." + }, + { + "id": "medium", + "alias": "qwen-medium", + "context": 160000, + "model_env": "MEDIUM_MODEL_FILE", + "model_family": "Qwen3.8-27B IQ4 XS Pure", + "gpu_split": "90:10", + "vision": true, + "mtp": 3, + "description": "Ausgewogenes Standardprofil für Alltag und lange agentische Aufgaben." + }, + { + "id": "large", + "alias": "qwen-large", + "context": 192000, + "model_env": "LARGE_MODEL_FILE", + "model_family": "Qwen3.8-27B IQ4 XS Pure", + "gpu_split": "86:14", + "vision": true, + "mtp": 3, + "description": "Großes Profil für umfangreiche Dokumente und lange technische Arbeiten." + }, + { + "id": "ultra", + "alias": "qwen-ultra", + "context": 262144, + "model_env": "ULTRA_MODEL_FILE", + "model_family": "Qwen3.8-27B IQ4 XS Pure", + "gpu_split": "80:20", + "vision": false, + "mtp": 2, + "description": "Maximaler Textkontext; bewusst ohne Vision-Projektor." + }, + { + "id": "uncensored", + "alias": "qwen-uncensored", + "context": 80000, + "model_env": "UNCENSORED_MODEL_FILE", + "model_family": "Qwen3.8-27B Abliterated Q4_K_M", + "gpu_split": "90:10", + "vision": true, + "mtp": 2, + "description": "Weniger restriktives Spezialprofil; Werkzeugrechte bleiben unverändert." + } + ] +} diff --git a/config/unraid-templates/my-MCPHub.xml b/config/unraid-templates/my-MCPHub.xml index 802427c..e5032ec 100644 --- a/config/unraid-templates/my-MCPHub.xml +++ b/config/unraid-templates/my-MCPHub.xml @@ -1,7 +1,7 @@ MCPHub - casaderoll/mcphub:1.0.0 + casaderoll/mcphub:1.1.1 https://hub.docker.com/r/samanhappy/mcphub bridge @@ -11,7 +11,7 @@ https://github.com/samanhappy/mcphub/issues https://github.com/samanhappy/mcphub https://github.com/samanhappy/mcphub#readme - Zentrale MCP-Verwaltung mit Weboberfläche. Das lokale CasaDeRoll-Image basiert reproduzierbar auf MCPHub 1.0.32 und enthält die versionierten ARR-, Deemix-, Navidrome-, GitHub- und Web-MCP-Laufzeiten. Home Assistant und MUA werden als vorhandene HTTP-MCPs eingebunden. Einzelne Server bleiben unter /mcp/NAME getrennt sichtbar und schaltbar. + Zentrale MCP-Verwaltung mit Weboberfläche. Das lokale CasaDeRoll-Image basiert reproduzierbar auf MCPHub 1.0.32 und enthält die versionierten ARR-, Deemix-, Navidrome-, GitHub- und FritzBox-Laufzeiten. Home Assistant, MUA und der Athena Operator werden als vorhandene HTTP-MCPs eingebunden. Einzelne Server bleiben unter /mcp/NAME getrennt sichtbar und schaltbar. Hermes nutzt für allgemeine Webrecherche seine eingebauten Werkzeuge. Weboberfläche: http://[IP]:[PORT:3000]/ Benutzer beim ersten Start: admin diff --git a/dev/mock_upstream.py b/dev/mock_upstream.py index 2139e6f..1158d7a 100755 --- a/dev/mock_upstream.py +++ b/dev/mock_upstream.py @@ -99,6 +99,9 @@ class Handler(BaseHTTPRequestHandler): self.send_header("Connection", "close") self.end_headers() model = body.get("model") + delay = body.get("mock_stream_delay", 0.05) + if not isinstance(delay, (int, float)) or delay < 0 or delay > 10: + delay = 0.05 for tok in ["Mock-", "Streaming", "-Antwort", f" ({model})"]: chunk = { "id": "chatcmpl-mock", @@ -109,9 +112,12 @@ class Handler(BaseHTTPRequestHandler): } self.wfile.write(f"data: {json.dumps(chunk)}\n\n".encode()) self.wfile.flush() - time.sleep(0.05) - self.wfile.write(b"data: [DONE]\n\n") - self.wfile.flush() + time.sleep(delay) + try: + self.wfile.write(b"data: [DONE]\n\n") + self.wfile.flush() + except (BrokenPipeError, ConnectionResetError): + pass def _json(self, code: int, payload: dict) -> None: body = json.dumps(payload).encode() diff --git a/dev/test_local.sh b/dev/test_local.sh index 77ff8dd..a6d0789 100755 --- a/dev/test_local.sh +++ b/dev/test_local.sh @@ -183,6 +183,34 @@ echo "$RESP" echo "$RESP" | grep -q "data: " && echo "$RESP" | grep -q "\[DONE\]" \ && ok "SSE-Stream mit [DONE] erhalten" || bad "Streaming" +echo "== Test 4b: Client-Abbruch gibt den Router-Slot frei" +command curl -H "Authorization: Bearer $TEST_ROUTER_KEY" -sfN \ + "$BASE/v1/chat/completions" -H "Content-Type: application/json" \ + -d '{"model":"qwen-fast","stream":true,"mock_stream_delay":2,"messages":[{"role":"user","content":"Abbruch"}]}' \ + >/tmp/aborted-stream.txt 2>/dev/null & +ABORT_PID=$! +sleep 0.4 +kill "$ABORT_PID" 2>/dev/null || true +wait "$ABORT_PID" 2>/dev/null || true +for _ in $(seq 1 30); do + ACTIVE=$(curl -sf "$BASE/status" | python3 -c 'import json,sys; print(json.load(sys.stdin)["qwen"]["active_chats"])') + [[ $ACTIVE == 0 ]] && break + sleep 0.1 +done +CODE=$(curl -s -o /tmp/after-abort.json -w "%{http_code}" "$BASE/v1/chat/completions" \ + -H "Content-Type: application/json" \ + -d '{"model":"qwen-fast","messages":[{"role":"user","content":"noch frei?"}]}') +[[ ${ACTIVE:-1} == 0 && $CODE == 200 ]] \ + && ok "abgebrochener Stream gibt Lease frei; Folgerequest erfolgreich" \ + || bad "Stream-Abbruch hinterließ active_chats=${ACTIVE:-?}, HTTP $CODE" + +if [[ ${ROUTER_TEST_QUICK:-0} == 1 ]]; then + echo + echo "== Schnellergebnis: $PASS bestanden, $FAIL fehlgeschlagen ==" + [[ $FAIL -eq 0 ]] + exit +fi + # --- 5. Tool Calls ----------------------------------------------------------------- echo "== Test 5: Tool Calls" RESP=$(curl -sf "$BASE/v1/chat/completions" -H "Content-Type: application/json" \ diff --git a/dev/test_mcp_registry.py b/dev/test_mcp_registry.py index 6e3b3d4..2158753 100644 --- a/dev/test_mcp_registry.py +++ b/dev/test_mcp_registry.py @@ -66,6 +66,41 @@ class RegistryTests(unittest.TestCase): con.close() self.assertEqual(ids, ["unmanaged", "one"]) + def test_production_registry_has_unique_ids_and_fritzbox(self): + document = json.loads((ROOT / "config/mcp-registry.json").read_text()) + ids = [item["id"] for item in document["servers"]] + hermes_ids = [ + item.get("hermes_id", item["id"]) + for item in document["servers"] if "hermes" in item.get("clients", []) + ] + self.assertEqual(len(ids), len(set(ids))) + self.assertEqual(len(hermes_ids), len(set(hermes_ids))) + fritz = next(item for item in document["servers"] if item["id"] == "fritzbox-local") + self.assertEqual(fritz["hub"]["type"], "stdio") + self.assertIn("fritz-mcp", fritz["hub"]["args"][-1]) + self.assertEqual(len(fritz["tool_include"]), 4) + + def test_hermes_tool_filter_is_generated(self): + item = { + "id": "wide", "name": "Wide", "description": "Test", + "url": "http://wide/mcp", "clients": ["hermes"], + "tool_include": ["list", "describe", "call"], + } + block = self.module.hermes_block([item]) + self.assertIn(" tools:\n include:", block) + self.assertIn(' - "describe"', block) + + def test_raw_mcphub_token_can_replace_host_specific_env_file(self): + item = { + "id": "hub", "name": "Hub", "description": "Test", + "url": "http://hub/mcp/test", "clients": ["hermes"], + "env_file": "/missing/client.env", + "key_env": "MCPHUB_BEARER_TOKEN", + } + self.module.CLIENT_TOKEN = "local-token" + self.assertTrue(self.module.enabled(item)) + self.assertEqual(self.module.resolved(item), ("http://hub/mcp/test", "local-token")) + if __name__ == "__main__": unittest.main() diff --git a/dev/test_mcphub_settings.py b/dev/test_mcphub_settings.py new file mode 100644 index 0000000..fd99e6e --- /dev/null +++ b/dev/test_mcphub_settings.py @@ -0,0 +1,81 @@ +#!/usr/bin/env python3 +"""Regression tests for declarative, update-safe MCPHub settings.""" + +from __future__ import annotations + +import importlib.util +import json +import os +import tempfile +import unittest +from pathlib import Path + + +ROOT = Path(__file__).parents[1] +SOURCE = ROOT / "platform/mcphub/configure-settings.py" + + +def load_module(): + spec = importlib.util.spec_from_file_location("configure_settings", SOURCE) + module = importlib.util.module_from_spec(spec) + assert spec.loader + spec.loader.exec_module(module) + return module + + +class MCPHubSettingsTests(unittest.TestCase): + def setUp(self): + self.module = load_module() + self.temp = tempfile.TemporaryDirectory() + self.root = Path(self.temp.name) + self.secrets = self.root / "secrets" + self.secrets.mkdir() + (self.secrets / "remote.env").write_text( + "REMOTE_URL=http://example.test/mcp\nTOKEN=secret-value\n", + encoding="utf-8", + ) + self.registry = self.root / "registry.json" + self.registry.write_text(json.dumps({ + "version": 1, + "servers": [ + { + "id": "remote-local", + "hermes_id": "remote", + "hub": { + "type": "streamable-http", + "secret_file": "remote.env", + "url": "${REMOTE_URL}", + "headers": {"Authorization": "Bearer ${TOKEN}"}, + "enabled": True, + }, + }, + {"id": "client-only", "url": "http://unused/mcp"}, + ], + }), encoding="utf-8") + + def tearDown(self): + self.temp.cleanup() + + def test_registry_renders_only_hub_servers_and_expands_secrets(self): + servers = self.module.registry_servers(self.registry, self.secrets, {}) + self.assertEqual(list(servers), ["remote"]) + self.assertEqual(servers["remote"]["url"], "http://example.test/mcp") + self.assertEqual( + servers["remote"]["headers"]["Authorization"], + "Bearer secret-value", + ) + + def test_existing_enabled_toggle_survives_reconciliation(self): + servers = self.module.registry_servers( + self.registry, self.secrets, {"remote": {"enabled": False}} + ) + self.assertFalse(servers["remote"]["enabled"]) + + def test_missing_secret_fails_closed(self): + os.unlink(self.secrets / "remote.env") + with self.assertRaises(SystemExit): + self.module.registry_servers(self.registry, self.secrets, {}) + + +if __name__ == "__main__": + unittest.main() diff --git a/docs/MCP_SERVERS.md b/docs/MCP_SERVERS.md new file mode 100644 index 0000000..07fb947 --- /dev/null +++ b/docs/MCP_SERVERS.md @@ -0,0 +1,16 @@ +# MCP-Server + +Diese Datei wird aus `config/mcp-registry.json` erzeugt. Änderungen gehören nur in die JSON-Registry. + +| Server | Hermes-ID | Endpunkt | Clients | Werkzeuge | +|---|---|---|---|---:| +| Athena Operator | `athena-operator` | `http://192.168.1.2:8787/mcp/athena-operator` | hermes, openwebui | alle | +| GitHub (offiziell, read-only) | `github` | `http://192.168.1.2:8787/mcp/github` | hermes, openwebui | alle | +| Home Assistant | `homeassistant-admin` | `http://192.168.1.2:8787/mcp/homeassistant` | hermes, openwebui | alle | +| Sonarr und Radarr | `arr` | `http://192.168.1.2:8787/mcp/arr` | hermes, openwebui | alle | +| Navidrome | `navidrome` | `http://192.168.1.2:8787/mcp/navidrome` | hermes, openwebui | alle | +| Deemix | `deemix` | `http://192.168.1.2:8787/mcp/deemix` | hermes, openwebui | alle | +| MUA (Unraid-Verwaltung) | `unraid` | `http://192.168.1.2:8787/mcp/unraid` | hermes, openwebui | alle | +| FRITZ!Box | `fritzbox` | `http://192.168.1.2:8787/mcp/fritzbox` | hermes, openwebui | 4 | + +Allgemeine Webrecherche ist ein eingebautes Hermes-Werkzeug und kein MCPHub-Server. diff --git a/docs/RECOVERY.md b/docs/RECOVERY.md index e14cfa9..f210d2c 100644 --- a/docs/RECOVERY.md +++ b/docs/RECOVERY.md @@ -3,18 +3,14 @@ ## Was automatisch gesichert wird Der Container `mike-ai-backup` erstellt alle fünf Stunden ein komprimiertes -Archiv unter `/data/docker-backups` und behält 14 Tage. Während der kurzen -Sicherung wird nur OpenWebUI angehalten, damit seine SQLite-Datenbank konsistent -ist. Netzwerk, WireGuard, Router und Qwen bleiben erreichbar. Hermes läuft -unabhängig davon auf Unraid weiter. +Archiv unter `/data/docker-backups` und behält 14 Tage. Netzwerk, WireGuard, +Router und Qwen bleiben dabei erreichbar. Hermes läuft unabhängig auf Unraid. Enthalten sind: - `/etc/mike-ai` mit lokalen Konfigurationen und Secrets -- OpenWebUI-Daten - Router-Zustand und Router-Bildablage - Piper-Daten -- TinySearch-Modellcache - ein Quellbaum-Snapshot als zusätzliche Bequemlichkeit Nicht kopiert werden `/data/models`, die nur noch als Rückfall vorhandene Kopie @@ -84,7 +80,7 @@ docker compose --env-file /etc/mike-ai/stack.env ps test -s /data/docker-backups/athena-latest.tar.gz ``` -Danach einen OpenWebUI-Login, einen Router-Request und je einen read-only +Danach einen Router-Request, einen Hermes-Zweiturn-Chat und je einen read-only MCP-Aufruf über `http://UNRAID-IP:8787/mcp/NAME` testen. Alte Recovery-Koffer sind für Neuinstallationen nicht mehr erforderlich; Git, Athenas Datenbackup und das Unraid-Appdata-Backup bilden die Wiederherstellung. diff --git a/docs/STANDARD_PROFILE_MATRIX.md b/docs/STANDARD_PROFILE_MATRIX.md index 6aa7ffc..e2932fb 100644 --- a/docs/STANDARD_PROFILE_MATRIX.md +++ b/docs/STANDARD_PROFILE_MATRIX.md @@ -1,50 +1,21 @@ -# Verbindliche Standard-Profilmatrix +# Profilmatrix -Stand: 22. August 2026. Diese fünf Profile sind die Produktionsmatrix für -Athena. Änderungen gelten erst nach einem vergleichbaren synthetischen Test -und einer bewussten Aktualisierung dieser Datei. +Diese Datei wird aus `config/profile-matrix.json` erzeugt. Änderungen gehören nur in die JSON-Matrix. -| Profil | Virtuelles Modell | GGUF | Kontext | GPUs / Split | MTP | Vision | gemessene kurze Ausgabe | -|---|---|---|---:|---|---:|---|---:| -| Fast | `qwen-fast` | IQ4-MIX | 76.800 | RTX 5080 | 2 | ja, Projektor auf RTX 3060 | 85,5 Tok/s | -| **Medium (Default)** | `qwen-medium` | IQ4_XS Pure | 160.000 | RTX 5080 + RTX 3060, 90:10 | 3 | ja, Projektor auf RTX 3060 | 77,2 Tok/s | -| Large | `qwen-large` | IQ4_XS Pure | 192.000 | RTX 5080 + RTX 3060, 86:14 | 3 | ja, Projektor auf RTX 3060 | 75,3 Tok/s | -| Ultra | `qwen-ultra` | IQ4_XS Pure | 262.144 | RTX 5080 + RTX 3060, 80:20 | 2 | nein, text-only | 68,2 Tok/s | -| Uncensored | `qwen-uncensored` | Abliterated Q4_K_M | 80.000 | RTX 5080 + RTX 3060, 90:10 | 2 | ja, eigener Projektor auf RTX 3060 | 52,2 Tok/s | +Standardprofil: **medium** · globales Ausgabelimit: **8192 Token** -## Standardverhalten +| Profil | API-Alias | Kontext | Modell | GPU-Verteilung | Vision | MTP | +|---|---|---:|---|---|---|---:| +| fast | `qwen-fast` | 76,800 | Qwen3.8-27B IQ4 Mix | 5080 only | ja | 2 | +| medium | `qwen-medium` | 160,000 | Qwen3.8-27B IQ4 XS Pure | 90:10 | ja | 3 | +| large | `qwen-large` | 192,000 | Qwen3.8-27B IQ4 XS Pure | 86:14 | ja | 3 | +| ultra | `qwen-ultra` | 262,144 | Qwen3.8-27B IQ4 XS Pure | 80:20 | nein | 2 | +| uncensored | `qwen-uncensored` | 80,000 | Qwen3.8-27B Abliterated Q4_K_M | 90:10 | ja | 2 | -- Medium ist nach Neuinstallation und bewusstem Plattformstart das aktive - Standardprofil. -- Open WebUI erhält `mikeai-medium` als Standardmodell. -- Im OpenWebUI-Arbeitsbereich ist dieses auf `qwen-medium` basierende Preset - `mikeai-medium` die sichtbare Standardauswahl; die fünf rohen `qwen-*`-Aliase - bleiben ausgeblendet. -- Ein explizit gewähltes anderes Modell löst den zugehörigen Containerwechsel - aus; es ist immer nur ein Inferenzcontainer aktiv. -- Ultra reserviert den verfügbaren Speicher für nativen 256K-Textkontext und - lädt deshalb keinen Vision-Projektor. -- Uncensored ist ein bewusst gewähltes Spezialprofil mit reduzierter - Verweigerungsneigung. Es ändert weder Router-Authentifizierung noch - Tool-Rechte, Bestätigungsregeln oder Secret-Filter und ist nicht Default. -- Das gesonderte Experimentalprofil gehört nicht zur Benutzer-Matrix und wird - in Open WebUI nicht als reguläres Modell angeboten. +## Zweck -## Hermes-Kontextpflege - -Hermes verwendet das jeweils ausgewählte 27B-Profil auch für die -Kontextzusammenfassung. Micro-Compaction übernimmt alle fünf abgeschlossenen -Turns einen älteren Assistenten-/Werkzeugabschnitt in die laufende -Zusammenfassung. Die normale schwellenbasierte Kompression bleibt als -Rückfallebene aktiv. Ein separates kleines Kompressionsmodell wird nicht -installiert, weil 2B/4B-Tests exakte technische Zustände verlieren konnten. - -## Nachweise - -- Fast 76,8K: synthetischer Referenzlauf, 85,50 Tok/s. -- Medium 160K: Pure 90:10 mit MTP3, 77,22 Tok/s. -- Large 192K: Pure 86:14 mit MTP3, 75,28 Tok/s. -- Ultra 256K: Pure 80:20 mit MTP2, 68,19 Tok/s; 220.190 Tokens - erfolgreich verarbeitet und Sentinel korrekt wiedergefunden. -- Uncensored 80K: Abliterated Q4_K_M 90:10 mit MTP2, 52,2 Tok/s; 60K-Fülltest, - Vision und synthetischer Logik-/Sicherheitslauf bestanden. +- **fast**: Schnelles Profil für kurze Chats und zügige Werkzeugaufgaben. +- **medium**: Ausgewogenes Standardprofil für Alltag und lange agentische Aufgaben. +- **large**: Großes Profil für umfangreiche Dokumente und lange technische Arbeiten. +- **ultra**: Maximaler Textkontext; bewusst ohne Vision-Projektor. +- **uncensored**: Weniger restriktives Spezialprofil; Werkzeugrechte bleiben unverändert. diff --git a/install.sh b/install.sh index 8f81b43..2d23869 100755 --- a/install.sh +++ b/install.sh @@ -485,16 +485,14 @@ build_and_start() { "from huggingface_hub import snapshot_download; snapshot_download('black-forest-labs/FLUX.2-klein-4B', revision='303481f0390afb112393f9d77e8f0be72fcefeb7', local_dir='/download')" chmod -R a-w "${FLUX_MODEL_DIR:-/data/models/FLUX.2-klein-4B}" fi - # Creates the shared internal tools network before Open WebUI is created. - # Web search always starts; HA/ARR/Unraid only start when their root-only - # secret files and required local artifacts are present. + # Creates the tools network and deploys the only host-bound MCP: Operator. + # Portable MCPs and Hermes live on Unraid and are restored through Appdata. "$STACK_DIR/platform/mcp/install-tools.sh" - "$STACK_DIR/platform/hermes/install-hermes.sh" docker compose --env-file "$SECRETS_DIR/stack.env" --profile inference create \ llama-fast llama-medium llama-large llama-ultra llama-experimental docker compose --env-file "$SECRETS_DIR/stack.env" --profile image create flux-worker docker compose --env-file "$SECRETS_DIR/stack.env" up -d --build \ - xtts piper tts-gateway profile-controller router open-webui hermes + wireguard-gateway xtts piper tts-gateway profile-controller router backup if [[ ${WIREGUARD_MODE:-container} == container ]]; then systemctl restart mike-ai-container-vpn-guard.service fi @@ -541,16 +539,6 @@ PY grep -Fxq 'INSTALL_READINESS_OK' <<<"$activation_output" || \ die "Medium-Standardprofil lieferte keinen bestätigten Readiness-Marker" - log "Hermes-Profile aus der Standardmatrix anlegen" - "$STACK_DIR/platform/hermes/install-profiles.sh" - - log "OpenWebUI-Filter und MCP-Registry synchronisieren" - "$STACK_DIR/platform/openwebui/install-filters.sh" - - if [[ ${INSTALL_HERMES_WEBUI:-true} == true ]]; then - log "Entfernbare Hermes Community-WebUI installieren" - "$STACK_DIR/platform/hermes/install-webui.sh" - fi } hostnamectl set-hostname "$AI_HOSTNAME" @@ -578,9 +566,8 @@ else vpn_address=$AI_BIND_ADDRESS fi cat <&2; exit 2; } + run python3 "$ROOT_DIR/platform/scripts/sync-profile-matrix.py" --check + run python3 "$ROOT_DIR/platform/scripts/render-mcp-registry.py" --check + if [[ $1 == core ]]; then + shift + [[ $# -eq 0 ]] || { echo "core akzeptiert keine weiteren Services" >&2; exit 2; } + run "$ROOT_DIR/platform/mcp/install-tools.sh" + run "${compose[@]}" up -d --build \ + wireguard-gateway piper xtts tts-gateway profile-controller router backup + else + run "${compose[@]}" up -d --build --no-deps "$@" + fi + ;; + stop-legacy) + # Reversible cleanup: stop only obsolete Athena frontends/tool backends. + for name in mike-ai-open-webui mike-ai-hermes mike-ai-hermes-webui \ + mike-ai-hermes-webui-vpn-proxy mike-ai-mcp-web mike-ai-tools-searxng \ + mike-ai-tools-tinysearch mike-ai-mcp-platform-context; do + if docker inspect "$name" >/dev/null 2>&1; then + run docker stop "$name" + fi + done + ;; + -h|--help|"") usage ;; + *) echo "Unbekannter Befehl: $command" >&2; usage >&2; exit 2 ;; +esac diff --git a/platform/checks/verify-platform.sh b/platform/checks/verify-platform.sh index 02a75cd..5a21de1 100755 --- a/platform/checks/verify-platform.sh +++ b/platform/checks/verify-platform.sh @@ -1,102 +1,6 @@ #!/usr/bin/env bash -set -uo pipefail +# Compatibility entrypoint: there is only one authoritative release check. +set -Eeuo pipefail -PASS=0 -WARN=0 -FAIL=0 - -pass() { printf 'PASS %s\n' "$*"; PASS=$((PASS + 1)); } -warn() { printf 'WARN %s\n' "$*"; WARN=$((WARN + 1)); } -fail() { printf 'FAIL %s\n' "$*"; FAIL=$((FAIL + 1)); } - -echo "== Local AI platform verification ==" - -ROOT_USE="$(df -P / | awk 'NR==2 {gsub(/%/, "", $5); print $5}')" -if [[ -n "$ROOT_USE" && "$ROOT_USE" -lt 85 ]]; then - pass "Systempartition bei ${ROOT_USE}%" -else - fail "Systempartition bei ${ROOT_USE:-unbekannt}% (Ziel: unter 85%)" -fi - -if command -v nvidia-smi >/dev/null 2>&1; then - GPU_LIST="$(nvidia-smi --query-gpu=name,memory.total --format=csv,noheader 2>/dev/null)" - GPU_COUNT="$(printf '%s\n' "$GPU_LIST" | sed '/^[[:space:]]*$/d' | wc -l)" - if [[ "$GPU_COUNT" -ge 2 ]]; then - pass "beide GPUs erkannt: $(printf '%s' "$GPU_LIST" | paste -sd ' | ' -)" - elif [[ "$GPU_COUNT" -eq 1 ]]; then - fail "nur eine von zwei erwarteten GPUs erkannt: $GPU_LIST" - else - fail "nvidia-smi liefert keine GPU" - fi -else - fail "nvidia-smi fehlt" -fi - -if [[ -e /sys/class/net/lan0 ]] && \ - [[ "$(< /sys/class/net/lan0/address)" == "58:11:22:bb:ad:0c" ]] && \ - [[ "$(< /sys/class/net/lan0/operstate)" == "up" ]]; then - pass "stabiles LAN-Interface lan0 aktiv (58:11:22:bb:ad:0c)" -else - fail "stabiles LAN-Interface lan0 fehlt, ist down oder hat die falsche MAC" -fi - -container_healthy() { - local name=$1 state health - state="$(docker inspect --format '{{.State.Status}}' "$name" 2>/dev/null || true)" - health="$(docker inspect --format '{{if .State.Health}}{{.State.Health.Status}}{{end}}' "$name" 2>/dev/null || true)" - [[ $state == running && ( -z $health || $health == healthy ) ]] -} - -for container in mike-ai-profile-controller mike-ai-router mike-ai-xtts \ - mike-ai-piper mike-ai-tts-gateway mike-ai-open-webui; do - if container_healthy "$container"; then - pass "$container gesund" - else - fail "$container fehlt oder ist nicht gesund" - fi -done - -ACTIVE_LLAMA="$(docker ps --format '{{.Names}}' | grep -Ec '^mike-ai-llama-(fast|medium|large|ultra|uncensored|experimental)$' || true)" -if [[ $ACTIVE_LLAMA -eq 1 ]]; then - pass "exakt ein llama.cpp-Profil aktiv" -else - fail "$ACTIVE_LLAMA llama.cpp-Profile aktiv (erwartet: 1)" -fi - -for optional in mike-ai-mcp-web mike-ai-mcp-homeassistant mike-ai-mcp-arr \ - mike-ai-mcp-github mike-ai-mcp-athena-operator mike-ai-backup; do - if container_healthy "$optional"; then - pass "$optional aktiv" - else - warn "$optional nicht aktiv oder nicht installiert" - fi -done - -if docker ps -a --format '{{.Names}}' | grep -qx 'mike-ai-mcp-unraid-official'; then - fail "veralteter GraphQL-basierter Unraid-MCP ist noch vorhanden" -else - pass "kein GraphQL-basierter Unraid-MCP vorhanden" -fi - -if docker exec mike-ai-router python -c \ - "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8081/health', timeout=3)" \ - >/dev/null 2>&1; then - pass "Router-Liveness intern erreichbar" -else - fail "Router-Liveness intern nicht erreichbar" -fi - -if systemctl list-unit-files --no-legend 2>/dev/null | grep -Eq '(vision-rx|whisper-rx|granite-rx).*enabled'; then - warn "Aktivierte RX-Altlast gefunden" -else - pass "Keine aktivierte RX-Altlast" -fi - -if systemctl list-unit-files --no-legend 2>/dev/null | grep -E '(benchmark|race).*enabled' >/dev/null; then - warn "Automatisch aktivierter Benchmark-/Race-Dienst gefunden" -else - pass "Keine automatisch aktivierten Benchmarks" -fi - -printf '\nErgebnis: %d PASS, %d WARN, %d FAIL\n' "$PASS" "$WARN" "$FAIL" -[[ "$FAIL" -eq 0 ]] +ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)" +exec "$ROOT_DIR/smoke-test.sh" "$@" diff --git a/platform/hermes/config.yaml b/platform/hermes/config.yaml index 0dd468c..e225975 100644 --- a/platform/hermes/config.yaml +++ b/platform/hermes/config.yaml @@ -59,7 +59,7 @@ agent: # Keep enough deliberation for tool choice while avoiding the provider's # unbounded `auto` reasoning mode on ordinary turns. Users can still raise it # per session with /reasoning. - reasoning_effort: "minimal" + reasoning_effort: "low" gateway_timeout: 3600 session_stall_timeout: 600 tool_loop_guardrails: @@ -77,7 +77,9 @@ agent: max_web_searches: 8 max_subagents: 8 -# Keep ample room for long agent work. Compression starts at 82% of whichever +# Keep ample room for long agent work. Compression starts late; deterministic +# tool-result pruning and rolling micro-compaction keep it from getting there +# during normal jobs. # router profile is selected. Keep the result compact enough that a local # model does not spend many minutes generating the handoff. compression: @@ -89,19 +91,26 @@ compression: micro_compact_every_n_turns: 5 micro_compact_defrag_threshold_tokens: 2000 progress_notices: true - threshold: 0.82 + threshold: 0.95 target_ratio: 0.15 + max_attempts: 1 tail_mode: "lean" protect_last_n: 20 protect_first_n: 0 - proactive_prune_tokens: 50000 - proactive_prune_min_result_chars: 4000 - proactive_prune_min_reclaim_tokens: 4096 + proactive_prune_tokens: 32000 + proactive_prune_min_result_chars: 2000 + proactive_prune_min_reclaim_tokens: 2048 + context_timeout_seconds: 45 # A failed local summarizer must not block an interactive client for ten # minutes. Continue without dropping messages after two minutes. context_total_ceiling_seconds: 120 auxiliary: + compression: + provider: "main" + reasoning_effort: "none" + timeout: 120 + max_concurrency: 1 # Session names are cosmetic and used to create a second concurrent LLM # request after every first reply. On a single inference slot this blocks the # actual chat, so keep the original timestamp/session id instead. diff --git a/platform/hermes/install-profiles.sh b/platform/hermes/install-profiles.sh index af883e3..d24199e 100755 --- a/platform/hermes/install-profiles.sh +++ b/platform/hermes/install-profiles.sh @@ -1,23 +1,26 @@ #!/usr/bin/env bash set -Eeuo pipefail -HERMES_CONTAINER=${HERMES_CONTAINER:-mike-ai-hermes} +HERMES_CONTAINER=${HERMES_CONTAINER:-Hermes-Agent} SECRETS_DIR=${SECRETS_DIR:-/etc/mike-ai} -HERMES_DATA_DIR=${HERMES_DATA_DIR:-/data/hermes} +HERMES_DATA_DIR=${HERMES_DATA_DIR:-/mnt/nvme-storage/appdata/Hermes-Agent} +STACK_DIR=${STACK_DIR:-$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)} +PROFILE_MATRIX=${PROFILE_MATRIX:-$STACK_DIR/config/profile-matrix.json} die() { printf 'FEHLER: %s\n' "$*" >&2; exit 1; } docker inspect "$HERMES_CONTAINER" >/dev/null 2>&1 || \ die "Hermes-Container fehlt: $HERMES_CONTAINER" +[[ -s $PROFILE_MATRIX ]] || die "Profilmatrix fehlt: $PROFILE_MATRIX" deadline=$((SECONDS + 180)) until [[ $(docker inspect --format '{{if .State.Health}}{{.State.Health.Status}}{{else}}{{.State.Status}}{{end}}' \ - "$HERMES_CONTAINER" 2>/dev/null || true) == healthy ]]; do + "$HERMES_CONTAINER" 2>/dev/null || true) =~ ^(healthy|running)$ ]]; do (( SECONDS < deadline )) || die "Hermes wurde nicht rechtzeitig gesund." sleep 2 done create_profile() { - local name=$1 model=$2 context=$3 description=$4 + local name=$1 model=$2 context=$3 description=$4 max_tokens=$5 if ! docker exec "$HERMES_CONTAINER" hermes profile show "$name" >/dev/null 2>&1; then docker exec "$HERMES_CONTAINER" hermes profile create "$name" \ --clone-from default --description "$description" @@ -26,23 +29,32 @@ create_profile() { docker exec "$HERMES_CONTAINER" hermes -p "$name" config set model.context_length "$context" # These are platform-wide latency and loop safeguards, not model-specific # tuning. Enforce them on old profiles as well as newly cloned profiles. - docker exec "$HERMES_CONTAINER" hermes -p "$name" config set model.max_tokens 8192 + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set model.max_tokens "$max_tokens" docker exec "$HERMES_CONTAINER" hermes -p "$name" config set agent.max_turns 64 - docker exec "$HERMES_CONTAINER" hermes -p "$name" config set agent.reasoning_effort minimal + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set agent.reasoning_effort low docker exec "$HERMES_CONTAINER" hermes -p "$name" config set display.show_reasoning false docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.title_generation.enabled false docker exec "$HERMES_CONTAINER" hermes -p "$name" config unset compression.threshold_tokens || true - docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.threshold 0.82 + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.threshold 0.95 docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.target_ratio 0.15 + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.max_attempts 1 + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.context_timeout_seconds 45 docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.context_total_ceiling_seconds 120 docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.protect_last_n 20 docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.protect_first_n 0 docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.micro_compact true docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.micro_compact_every_n_turns 5 docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.micro_compact_defrag_threshold_tokens 2000 + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.proactive_prune_tokens 32000 + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.proactive_prune_min_result_chars 2000 + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.proactive_prune_min_reclaim_tokens 2048 # The selected 27B profile performs production compaction. A separate small # compressor was removed after it lost exact technical state in benchmarks. docker exec "$HERMES_CONTAINER" hermes -p "$name" config unset auxiliary.compression || true + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.provider main + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.reasoning_effort none + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.timeout 120 + docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.max_concurrency 1 docker exec "$HERMES_CONTAINER" hermes -p "$name" config set platform_toolsets.cli \ '["web","terminal","file","skills","todo","memory","vision","tts"]' [[ $(docker exec "$HERMES_CONTAINER" hermes -p "$name" config get model.default) == "$model" ]] || \ @@ -51,51 +63,32 @@ create_profile() { die "Kontext von Profil $name konnte nicht verifiziert werden." } -create_profile fast qwen-fast 76800 \ - "Schnelles Qwen3.8-27B-Profil mit 76,8K Kontext fuer kurze Chats und schnelle Aufgaben." -create_profile medium qwen-medium 160000 \ - "Ausgewogenes Qwen3.8-27B-Standardprofil mit 160K Kontext fuer Alltag und agentische Aufgaben." -create_profile large qwen-large 192000 \ - "Grosses Qwen3.8-27B-Profil mit 192K Kontext fuer umfangreiche Dokumente und lange Aufgaben." -create_profile ultra qwen-ultra 262144 \ - "Maximales Qwen3.8-27B-Profil mit 262K Kontext fuer sehr grosse Kontexte; langsamer als die Standardprofile." -create_profile uncensored qwen-uncensored 80000 \ - "Unzensiertes Qwen3.8-27B-Profil mit 80K Kontext fuer spezielle Anfragen." +while IFS=$'\t' read -r name model context description max_tokens; do + create_profile "$name" "$model" "$context" "$description" "$max_tokens" +done < <(jq -r --argjson max "$(jq '.max_output_tokens' "$PROFILE_MATRIX")" \ + '.profiles[] | [.id,.alias,(.context|tostring),(.description|gsub("[\\t\\n]";" ")),($max|tostring)] | @tsv' \ + "$PROFILE_MATRIX") -# Existing profiles may predate managed secret rendering and therefore contain -# the literal ${ROUTER_API_KEY}. Repair only that exact placeholder; never log -# or commit the secret itself. -[[ -s $SECRETS_DIR/router-api-key ]] || die "Router-API-Key fehlt." -router_key=$(<"$SECRETS_DIR/router-api-key") -ROUTER_API_KEY="$router_key" HERMES_DATA_DIR="$HERMES_DATA_DIR" python3 <<'PY' -import os -import pathlib - -root = pathlib.Path(os.environ["HERMES_DATA_DIR"]) / "profiles" -placeholder = "${ROUTER_API_KEY}" -paths = [pathlib.Path(os.environ["HERMES_DATA_DIR"]) / "config.yaml"] -paths.extend(sorted(root.glob("*/config.yaml"))) -for path in paths: - if not path.exists(): - continue - text = path.read_text() - if placeholder in text: - text = text.replace(placeholder, os.environ["ROUTER_API_KEY"], 1) - # Avoid collision with Hermes' disabled built-in `homeassistant` toolset. - # The collision filters the healthy external MCP out of agent snapshots. - text = text.replace( - "\n homeassistant:\n url: http://mcp-homeassistant:8000/mcp\n", - "\n homeassistant-admin:\n url: http://mcp-homeassistant:8000/mcp\n", - ) - path.write_text(text) -PY - -sync_args=(--registry "${STACK_DIR:-/opt/mike-ai/stack}/config/mcp-registry.json") +# The Unraid host deliberately has no system Python. Run the small declarative +# client renderer in the already version-pinned MCPHub image instead of adding +# host dependencies. +sync_args=(--registry /stack/config/mcp-registry.json) +mcphub_token=${MCPHUB_TOKEN_FILE:-/mnt/nvme-storage/appdata/MCPHub/client-token} +token_mount=() +if [[ -s $mcphub_token ]]; then + token_mount=(-v "$mcphub_token:/run/input/mcphub-token:ro") + sync_args+=(--mcphub-token-file /run/input/mcphub-token) +fi while IFS= read -r config; do - sync_args+=(--hermes "$config") + sync_args+=(--hermes "/hermes/${config#"$HERMES_DATA_DIR"/}") done < <(find "$HERMES_DATA_DIR" -name config.yaml -type f -print) -python3 "${STACK_DIR:-/opt/mike-ai/stack}/platform/mcp/sync-clients.py" "${sync_args[@]}" +docker run --rm --entrypoint python \ + -v "$STACK_DIR:/stack:ro" \ + -v "$HERMES_DATA_DIR:/hermes:rw" \ + "${token_mount[@]}" \ + casaderoll/mcphub:1.1.0 \ + /stack/platform/mcp/sync-clients.py "${sync_args[@]}" -"${STACK_DIR:-/opt/mike-ai/stack}/platform/hermes/install-skills.sh" +"$STACK_DIR/platform/hermes/install-skills.sh" docker exec "$HERMES_CONTAINER" hermes profile list printf 'HERMES_PROFILES_OK\n' diff --git a/platform/hermes/install-skills.sh b/platform/hermes/install-skills.sh index 829a467..a139694 100755 --- a/platform/hermes/install-skills.sh +++ b/platform/hermes/install-skills.sh @@ -1,31 +1,39 @@ #!/usr/bin/env bash set -Eeuo pipefail -STACK_DIR=${STACK_DIR:-/opt/mike-ai/stack} -HERMES_DATA_DIR=${HERMES_DATA_DIR:-/data/hermes} -SKILL_SOURCE=$STACK_DIR/platform/hermes/skills/athena-operator/SKILL.md -PROFILES=(fast medium large ultra uncensored) +STACK_DIR=${STACK_DIR:-$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)} +HERMES_DATA_DIR=${HERMES_DATA_DIR:-/mnt/nvme-storage/appdata/Hermes-Agent} +PROFILE_MATRIX=${PROFILE_MATRIX:-$STACK_DIR/config/profile-matrix.json} +SKILL_ROOT=$STACK_DIR/platform/hermes/skills die() { printf 'FEHLER: %s\n' "$*" >&2; exit 1; } [[ $EUID -eq 0 ]] || die "Bitte als root ausführen." -[[ -s $SKILL_SOURCE ]] || die "Skill-Quelle fehlt: $SKILL_SOURCE" -grep -Fxq -- 'name: athena-operator' "$SKILL_SOURCE" || \ - die "Skill-Quelle hat kein gültiges Athena-Operator-Frontmatter." +[[ -s $PROFILE_MATRIX ]] || die "Profilmatrix fehlt: $PROFILE_MATRIX" install_skill() { - local root=$1 target=$1/platform/athena-operator/SKILL.md + local source=$1 root=$2 name target + name=${source%/SKILL.md} + name=${name##*/} + target=$root/platform/$name/SKILL.md + grep -Fxq -- "name: $name" "$source" || \ + die "Skill-Quelle hat kein gültiges Frontmatter: $source" install -d -o 10000 -g 10000 -m 0750 "${target%/*}" - if [[ -s $target ]] && ! cmp -s "$SKILL_SOURCE" "$target"; then + if [[ -s $target ]] && ! cmp -s "$source" "$target"; then cp -a "$target" "$target.before-managed-update-$(date +%Y%m%d-%H%M%S)" fi - install -o 10000 -g 10000 -m 0640 "$SKILL_SOURCE" "$target" - cmp -s "$SKILL_SOURCE" "$target" || die "Skill-Synchronisierung fehlgeschlagen: $target" + install -o 10000 -g 10000 -m 0640 "$source" "$target" + cmp -s "$source" "$target" || die "Skill-Synchronisierung fehlgeschlagen: $target" } -install_skill "$HERMES_DATA_DIR/skills" -for profile in "${PROFILES[@]}"; do - [[ -d $HERMES_DATA_DIR/profiles/$profile ]] || continue - install_skill "$HERMES_DATA_DIR/profiles/$profile/skills" +mapfile -t profiles < <(jq -r '.profiles[].id' "$PROFILE_MATRIX") + +for source in "$SKILL_ROOT"/*/SKILL.md; do + [[ -s $source ]] || continue + install_skill "$source" "$HERMES_DATA_DIR/skills" + for profile in "${profiles[@]}"; do + [[ -d $HERMES_DATA_DIR/profiles/$profile ]] || continue + install_skill "$source" "$HERMES_DATA_DIR/profiles/$profile/skills" + done done -printf 'HERMES_ATHENA_OPERATOR_SKILL_OK\n' +printf 'HERMES_SKILLS_OK\n' diff --git a/platform/hermes/skills/mcphub-deployer/SKILL.md b/platform/hermes/skills/mcphub-deployer/SKILL.md index 0147c03..6bef8fb 100644 --- a/platform/hermes/skills/mcphub-deployer/SKILL.md +++ b/platform/hermes/skills/mcphub-deployer/SKILL.md @@ -16,10 +16,10 @@ Use these paths directly. Do not search the filesystem for alternatives. - Host: Unraid `192.168.1.2` - Container: `MCPHub` - UI/base URL: `http://192.168.1.2:8787` -- Operational source/build tree: `/mnt/nvme-storage/appdata/MCPHub/build` +- Operational source/build tree: `/mnt/nvme-storage/appdata/MCPHub/build/repo` - Dockerfile: `platform/mcphub/Dockerfile` below that tree -- Server declaration: `platform/mcphub/configure-settings.py` -- Client registry: `config/mcp-registry.json` +- Single server and client registry: `config/mcp-registry.json` +- Registry renderer: `platform/mcphub/configure-settings.py` (normally unchanged) - Unraid template: `config/unraid-templates/my-MCPHub.xml` - Persistent state: `/mnt/nvme-storage/appdata/MCPHub` - Secrets: `/mnt/nvme-storage/appdata/MCPHub/secrets/.env`, mode `0600` @@ -91,8 +91,8 @@ Modify only the necessary fixed production files. Rules: - Put credentials only in the matching secret file with mode `0600`. - Never print, log, commit, summarize, or return a secret. - Use `/usr/local/bin/run-with-env` for secret-backed stdio servers. -- Declare servers in `configure-settings.py`; do not manually treat - `mcp_settings.json` as the source of truth. +- Declare servers only in `config/mcp-registry.json`; do not hard-code a server + in `configure-settings.py` and do not edit `mcp_settings.json` manually. - Preserve existing users, bearer keys, prompts, resources, enabled states, and per-tool toggles. - Build a new image tag. Never overwrite the tag currently running. @@ -104,7 +104,7 @@ unfiltered large server to clients. ### 5. Deploy without collateral changes -Build from `/mnt/nvme-storage/appdata/MCPHub/build`, then recreate only +Build from `/mnt/nvme-storage/appdata/MCPHub/build/repo`, then recreate only `MCPHub` through Unraid DockerMan so it stays a managed Unraid container. Preserve all Appdata and mounts. diff --git a/platform/mcp/install-tools.sh b/platform/mcp/install-tools.sh index 32fa9d4..5bff534 100755 --- a/platform/mcp/install-tools.sh +++ b/platform/mcp/install-tools.sh @@ -11,13 +11,6 @@ if [[ -s $STACK_ENV ]]; then COMPOSE+=(--env-file "$STACK_ENV") fi COMPOSE+=(-f "$STACK_DIR/compose.yaml") -export SEARXNG_SETTINGS_FILE="${SEARXNG_SETTINGS_FILE:-$MCP_DIR/../web-search/searxng-settings.yml}" - -[[ -s $SEARXNG_SETTINGS_FILE ]] || { - echo "SearXNG-Konfiguration fehlt: $SEARXNG_SETTINGS_FILE" >&2 - exit 1 -} - docker network inspect mike-ai-tools >/dev/null 2>&1 || \ docker network create --internal --subnet 172.30.40.0/24 mike-ai-tools >/dev/null docker network inspect mike-ai-tools-egress >/dev/null 2>&1 || \ @@ -26,36 +19,11 @@ docker network inspect mike-ai-tools-egress >/dev/null 2>&1 || \ # One administrative MCP exposes ATHENA.md plus the bounded host operator. "$MCP_DIR/../operator/install-operator.sh" -services=(searxng tinysearch mcp-athena-operator) +services=(mcp-athena-operator) # Portable Fach-MCPs laufen zentral im MCPHub auf Unraid. Ihre Compose-Blöcke # bleiben vorläufig als explizite Rollback-Profile erhalten, werden bei einer # normalen Athena-Installation aber weder gebaut noch gestartet. echo "ARR, Deemix, GitHub, Home Assistant und Navidrome werden über MCPHub auf Unraid bereitgestellt." -# TinySearch keeps the embedding bundle outside the container. Download it -# once on a fresh host; subsequent rebuilds reuse the named volume. -# -# Do not call `tinysearch setup` here. The pinned container image already -# contains a complete Playwright/Chromium installation, while that command -# unconditionally tries to install Chromium again. On IPv4-only hosts this -# redundant download can hang indefinitely. Prepare only the persistent ONNX -# bundle that is actually absent on a fresh installation. -docker volume create mike-ai-tools_tinysearch-models >/dev/null -tiny_image="marcellm01/tinysearch@sha256:7a7d0585f5000f462e699e42b97409715826a9e2edcd09a166afa93a4b7cba31" -if ! docker run --rm --entrypoint test \ - -v mike-ai-tools_tinysearch-models:/data/models "$tiny_image" \ - -f /data/models/all-minilm-l6-v2-onnx/model.onnx; then - echo "TinySearch-Modell wird einmalig geladen." - docker run --rm --entrypoint python \ - -v mike-ai-tools_tinysearch-models:/data/models "$tiny_image" \ - -c 'from tinysearch.services.onnx_bundle_service import ensure_onnx_bundle_sync; ensure_onnx_bundle_sync("fast")' -fi - "${COMPOSE[@]}" up -d --build "${services[@]}" -for webui in mike-ai-open-webui Open-WebUI; do - if docker container inspect "$webui" >/dev/null 2>&1; then - docker network connect mike-ai-tools "$webui" 2>/dev/null || true - fi -done - "${COMPOSE[@]}" ps diff --git a/platform/mcp/sync-clients.py b/platform/mcp/sync-clients.py index 90e36dc..a583901 100755 --- a/platform/mcp/sync-clients.py +++ b/platform/mcp/sync-clients.py @@ -13,6 +13,7 @@ import time BEGIN = "# BEGIN MANAGED MCP SERVERS" END = "# END MANAGED MCP SERVERS" +CLIENT_TOKEN = "" def env_file(path: str) -> dict[str, str]: @@ -37,7 +38,10 @@ def enabled(item: dict) -> bool: if source: values = env_file(source) url_ready = bool(item.get("url")) or bool(values.get(item.get("url_env", ""))) - key_ready = not item.get("key_env") or bool(values.get(item["key_env"])) + key_ready = (not item.get("key_env") + or bool(values.get(item["key_env"])) + or (item.get("key_env") == "MCPHUB_BEARER_TOKEN" + and bool(CLIENT_TOKEN))) return url_ready and key_ready return True @@ -47,6 +51,8 @@ def resolved(item: dict) -> tuple[str, str]: values = env_file(item["env_file"]) url = item.get("url") or values[item["url_env"]] key = values.get(item.get("key_env", ""), "") + if not key and item.get("key_env") == "MCPHUB_BEARER_TOKEN": + key = CLIENT_TOKEN return url, key return item["url"], "" @@ -72,6 +78,16 @@ def hermes_block(items: list[dict]) -> str: ]) if key: lines.extend([" headers:", f" Authorization: {yaml_quote('Bearer ' + key)}"]) + if "tool_include" in item: + lines.append(" tools:") + lines.append(" include:") + for tool in item["tool_include"]: + lines.append(f" - {yaml_quote(str(tool))}") + elif "tool_exclude" in item: + lines.append(" tools:") + lines.append(" exclude:") + for tool in item["tool_exclude"]: + lines.append(f" - {yaml_quote(str(tool))}") lines.extend([ f" timeout: {int(item.get('timeout', 300))}", " connect_timeout: 30", @@ -103,6 +119,8 @@ def openwebui_connection(item: dict) -> dict: config = {"enable": True, "access_grants": []} if item.get("functions"): config["function_name_filter_list"] = item["functions"] + elif item.get("tool_include"): + config["function_name_filter_list"] = ",".join(item["tool_include"]) return { "url": url, "path": "", "type": "mcp", "auth_type": item.get("auth_type", "none"), "headers": None, @@ -135,11 +153,17 @@ def update_openwebui(db: pathlib.Path, items: list[dict]) -> None: def main() -> None: + global CLIENT_TOKEN parser = argparse.ArgumentParser() parser.add_argument("--registry", type=pathlib.Path, required=True) parser.add_argument("--hermes", type=pathlib.Path, action="append", default=[]) parser.add_argument("--openwebui-db", type=pathlib.Path) + parser.add_argument("--mcphub-token-file", type=pathlib.Path) args = parser.parse_args() + if args.mcphub_token_file: + CLIENT_TOKEN = args.mcphub_token_file.read_text(encoding="utf-8").strip() + if not CLIENT_TOKEN: + raise SystemExit("MCPHub token file is empty") if args.hermes: block = hermes_block(active(args.registry, "hermes")) for path in args.hermes: diff --git a/platform/mcphub/Dockerfile b/platform/mcphub/Dockerfile index 5f36d25..97da1ef 100644 --- a/platform/mcphub/Dockerfile +++ b/platform/mcphub/Dockerfile @@ -2,11 +2,13 @@ FROM ghcr.io/github/github-mcp-server@sha256:1817b57d43916532dc002bdc5f344d639bd FROM ghcr.io/blakeem/navidrome-mcp:2.2.0@sha256:047f911a5a8f7cc8f185bb4d6e7ca6c435542edefff4694a00c2f718ab0ee7f5 AS navidrome -FROM samanhappy/mcphub:1.0.32 +FROM samanhappy/mcphub:1.0.32@sha256:df34df85e639743d0bf4b64182fc2f47586bf544beb08c85a5e5d693891daad3 ARG MCP_VERSION=1.29.0 ARG ARR_MCP_VERSION=1.0.1 ARG YT_DLP_VERSION=2026.7.4 +ARG FRITZ_MCP_VERSION=0.8.0 +ARG FRITZ_MCP_SHA256=47f4e2b5595a522aeda4aed9f5c65a57e500c3589c9f4262b8fbf72bd8c72bd4 USER root @@ -18,13 +20,21 @@ RUN python3 -m pip install --no-cache-dir \ "arr-mcp[mcp]==${ARR_MCP_VERSION}" \ "yt-dlp==${YT_DLP_VERSION}" +# Released, checksum-pinned Go binary. Keeping the download in the reproducible +# image build prevents MCPHub upgrades from silently dropping the Fritz tools. +RUN mkdir -p /opt/casaderoll \ + && python3 -c 'import sys,urllib.request; v=sys.argv[1]; urllib.request.urlretrieve(f"https://github.com/kambriso/fritzbox-mcp-server/releases/download/v{v}/fritz-mcp-linux-amd64", "/opt/casaderoll/fritz-mcp")' "${FRITZ_MCP_VERSION}" \ + && echo "${FRITZ_MCP_SHA256} /opt/casaderoll/fritz-mcp" | sha256sum -c - \ + && chmod 0755 /opt/casaderoll/fritz-mcp + COPY --from=github /server/github-mcp-server /usr/local/bin/github-mcp-server COPY --from=navidrome /app /opt/casaderoll/navidrome COPY platform/mcp/deemix_mcp.py /opt/casaderoll/mcps/deemix_mcp.py -COPY platform/web-search/web_search_mcp.py /opt/casaderoll/mcps/web_search_mcp.py COPY platform/mcp/patches/mcp_sonarr.py /usr/local/lib/python3.13/site-packages/arr_mcp/mcp/mcp_sonarr.py COPY platform/mcp/patches/mcp_radarr.py /usr/local/lib/python3.13/site-packages/arr_mcp/mcp/mcp_radarr.py +COPY config/mcp-registry.json /opt/casaderoll/config/mcp-registry.json +COPY platform/mcphub/configure-settings.py /opt/casaderoll/configure-settings.py COPY platform/mcphub/run-with-env.py /usr/local/bin/run-with-env COPY platform/mcphub/casaderoll-entrypoint.sh /usr/local/bin/casaderoll-mcphub-entrypoint @@ -36,7 +46,8 @@ RUN node -e 'const fs=require("node:fs"); const p="/opt/casaderoll/navidrome/dis ENV MCPHUB_SETTING_PATH=/app/data/ \ REQUEST_TIMEOUT=120000 \ - NODE_ENV=production + NODE_ENV=production \ + FRITZ_MCP_VERSION=${FRITZ_MCP_VERSION} ENTRYPOINT ["/usr/local/bin/casaderoll-mcphub-entrypoint"] CMD ["/usr/local/bin/entrypoint.sh", "pnpm", "start"] diff --git a/platform/mcphub/README.md b/platform/mcphub/README.md index b33c9a8..9a5d2d3 100644 --- a/platform/mcphub/README.md +++ b/platform/mcphub/README.md @@ -9,8 +9,8 @@ sie sich einen Docker-Container und ein Appdata-Backup teilen. - ARR, Deemix, Navidrome und GitHub laufen als lokale stdio-Unterprozesse. - Home Assistant und MUA/Unraid sind vorhandene HTTP-MCP-Endpunkte und werden vom Hub direkt weitergereicht. -- Der Web-Adapter kann hier laufen; SearXNG/TinySearch dürfen getrennte - Backend-Dienste bleiben. +- Allgemeine Webrecherche bleibt ein eingebautes Hermes-Werkzeug. Der alte + Athena-Webadapter sowie SearXNG/TinySearch gehören nicht zum MCPHub-Image. - Athenas administrativer Operator ist hostgebunden und bleibt auf Athena. MCPHub reicht den vorhandenen, nur über WireGuard erreichbaren HTTP-Endpunkt `http://192.168.1.212:8202/mcp` als `/mcp/athena-operator` weiter. Dadurch @@ -64,7 +64,7 @@ read-only-Aufruf prüfen und erst danach Clients auf `http://UNRAID-IP:8787/mcp/{server}` umstellen. Der alte Athena-Container wird erst gestoppt, wenn Hermes und OpenWebUI nachweislich über MCPHub funktionieren. -Aktueller Stand: Athena Operator, ARR, Deemix, Navidrome, GitHub, Home Assistant -und MUA/Unraid sind auf MCPHub registriert. Alte portable Athena-MCP-Container -bleiben vorläufig als ausgeschaltetes Rückfallnetz bestehen. Der Web-Adapter -ist der letzte noch offene Migrationspunkt. +Aktueller Stand: Athena Operator, ARR, Deemix, Navidrome, GitHub, Home Assistant, +MUA/Unraid und FRITZ!Box sind auf MCPHub registriert. Alte portable +Athena-MCP-Container bleiben ausgeschaltet als kurzfristiges Rückfallnetz +bestehen. Der frühere Webadapter ist nicht mehr Bestandteil des Images. diff --git a/platform/mcphub/casaderoll-entrypoint.sh b/platform/mcphub/casaderoll-entrypoint.sh index faa8090..69fca90 100644 --- a/platform/mcphub/casaderoll-entrypoint.sh +++ b/platform/mcphub/casaderoll-entrypoint.sh @@ -18,4 +18,13 @@ fi JWT_SECRET=$(cat "$jwt_file") export JWT_SECRET +# Reconcile the persistent MCPHub state with the versioned registry on every +# start. Existing users, bearer tokens and per-server enabled flags survive. +# This makes image upgrades reproducible instead of relying on manual edits in +# MCPHub's database/UI. +settings_file="$state_dir/mcp_settings.json" +python3 /opt/casaderoll/configure-settings.py \ + "$settings_file" /run/secrets/mcphub \ + --registry /opt/casaderoll/config/mcp-registry.json + exec "$@" diff --git a/platform/mcphub/configure-settings.py b/platform/mcphub/configure-settings.py index d2f096c..0839563 100644 --- a/platform/mcphub/configure-settings.py +++ b/platform/mcphub/configure-settings.py @@ -11,6 +11,7 @@ import argparse import json import os import pathlib +import re import secrets import tempfile import uuid @@ -34,75 +35,70 @@ def env_file(path: pathlib.Path) -> dict[str, str]: return values -def required(values: dict[str, str], key: str, source: pathlib.Path) -> str: - value = values.get(key, "").strip() - if not value: - raise SystemExit(f"{key} is missing in {source}") +PLACEHOLDER = re.compile(r"\$\{([A-Za-z_][A-Za-z0-9_]*)\}") + + +def expand(value: object, values: dict[str, str], source: pathlib.Path) -> object: + """Resolve secret placeholders without ever logging their values.""" + if isinstance(value, str): + def replace(match: re.Match[str]) -> str: + key = match.group(1) + resolved = values.get(key, "").strip() + if not resolved: + raise SystemExit(f"{key} is missing in {source}") + return resolved + return PLACEHOLDER.sub(replace, value) + if isinstance(value, list): + return [expand(item, values, source) for item in value] + if isinstance(value, dict): + return {key: expand(item, values, source) for key, item in value.items()} return value +def registry_servers(registry: pathlib.Path, secrets_dir: pathlib.Path, + existing: dict[str, object]) -> dict[str, object]: + document = json.loads(registry.read_text(encoding="utf-8")) + if document.get("version") != 1: + raise SystemExit("Unsupported MCP registry schema") + result: dict[str, object] = {} + for item in document.get("servers", []): + spec = item.get("hub") + if not isinstance(spec, dict): + continue + server_id = str(item.get("hermes_id") or item["id"]) + spec = dict(spec) + secret_name = str(spec.pop("secret_file", "")) + secret_path = secrets_dir / secret_name if secret_name else secrets_dir + values = env_file(secret_path) if secret_name else {} + rendered = expand(spec, values, secret_path) + if isinstance(rendered, dict) and isinstance(rendered.get("url"), str): + rendered["url"] = re.sub(r"(? None: parser = argparse.ArgumentParser() parser.add_argument("settings", type=pathlib.Path) parser.add_argument("secrets", type=pathlib.Path) + parser.add_argument( + "--registry", type=pathlib.Path, + default=pathlib.Path("/opt/casaderoll/config/mcp-registry.json"), + ) parser.add_argument("--web-backend", default="", help="TinySearch MCP URL; empty keeps web disabled") parser.add_argument("--searxng", default="", help="SearXNG base URL") args = parser.parse_args() + args.settings.parent.mkdir(parents=True, exist_ok=True) settings = json.loads(args.settings.read_text(encoding="utf-8")) if args.settings.exists() else {} - ha_path = args.secrets / "homeassistant.env" - mua_path = args.secrets / "mua.env" - ha = env_file(ha_path) - mua = env_file(mua_path) - - servers: dict[str, object] = { - "athena-operator": { - "type": "streamable-http", - "url": "http://192.168.1.212:8202/mcp", - "owner": "admin", - "enabled": True, - }, - "arr": { - "type": "stdio", - "command": "/usr/local/bin/run-with-env", - "args": ["/run/secrets/mcphub/arr.env", "--", "arr-mcp", "--transport", "stdio", "--auth-type", "none"], - "enabled": True, - }, - "deemix": { - "type": "stdio", - "command": "/usr/local/bin/run-with-env", - "args": ["/run/secrets/mcphub/deemix.env", "--", "python3", "/opt/casaderoll/mcps/deemix_mcp.py"], - "env": {"MCP_TRANSPORT": "stdio"}, - "enabled": True, - }, - "navidrome": { - "type": "stdio", - "command": "/usr/local/bin/run-with-env", - "args": ["/run/secrets/mcphub/navidrome.env", "--", "node", "/opt/casaderoll/navidrome/dist/index.js"], - "env": {"MCP_TRANSPORT": "stdio", "MCP_HTTP_EXPOSE": "false", "WEBUI_ENABLED": "false"}, - "enabled": True, - }, - "github": { - "type": "stdio", - "command": "/usr/local/bin/run-with-env", - "args": ["/run/secrets/mcphub/github.env", "--", "/usr/local/bin/github-mcp-server", "stdio", "--read-only", "--tools", "search_repositories,get_file_contents,search_code"], - "enabled": True, - }, - "homeassistant": { - "type": "streamable-http", - "url": required(ha, "HASS_URL", ha_path).rstrip("/") + "/api/hass_mcp", - "headers": {"Authorization": "Bearer " + required(ha, "HASS_TOKEN", ha_path)}, - "owner": "admin", - "enabled": True, - }, - "unraid": { - "type": "streamable-http", - "url": required(mua, "MUA_MCP_URL", mua_path), - "headers": {"Authorization": "Bearer " + required(mua, "MUA_MCP_BEARER_TOKEN", mua_path)}, - "owner": "admin", - "enabled": True, - }, - } + servers = registry_servers( + args.registry, + args.secrets, + settings.get("mcpServers", {}) if isinstance(settings.get("mcpServers"), dict) else {}, + ) if args.web_backend: servers["web"] = { "type": "stdio", @@ -144,7 +140,6 @@ def main() -> None: system = settings.setdefault("systemConfig", {}) system.setdefault("routing", {})["skipAuth"] = False - args.settings.parent.mkdir(parents=True, exist_ok=True) fd, temporary = tempfile.mkstemp(prefix=".mcp-settings-", dir=args.settings.parent) try: with os.fdopen(fd, "w", encoding="utf-8") as handle: diff --git a/platform/scripts/render-mcp-registry.py b/platform/scripts/render-mcp-registry.py new file mode 100644 index 0000000..b837041 --- /dev/null +++ b/platform/scripts/render-mcp-registry.py @@ -0,0 +1,57 @@ +#!/usr/bin/env python3 +"""Render the human-readable MCP list from the single JSON registry.""" + +from __future__ import annotations + +import argparse +import json +import pathlib + + +ROOT = pathlib.Path(__file__).resolve().parents[2] +REGISTRY = ROOT / "config" / "mcp-registry.json" +OUTPUT = ROOT / "docs" / "MCP_SERVERS.md" + + +def render() -> str: + document = json.loads(REGISTRY.read_text(encoding="utf-8")) + rows = [] + for item in document["servers"]: + if not item.get("hub"): + continue + clients = ", ".join(item.get("clients", [])) or "–" + selection = str(len(item["tool_include"])) if item.get("tool_include") else "alle" + rows.append( + f"| {item['name']} | `{item['hermes_id']}` | `{item['url']}` | " + f"{clients} | {selection} |" + ) + return "\n".join([ + "# MCP-Server", + "", + "Diese Datei wird aus `config/mcp-registry.json` erzeugt. Änderungen gehören nur in die JSON-Registry.", + "", + "| Server | Hermes-ID | Endpunkt | Clients | Werkzeuge |", + "|---|---|---|---|---:|", + *rows, + "", + "Allgemeine Webrecherche ist ein eingebautes Hermes-Werkzeug und kein MCPHub-Server.", + "", + ]) + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("--check", action="store_true") + args = parser.parse_args() + expected = render() + if args.check: + if not OUTPUT.is_file() or OUTPUT.read_text(encoding="utf-8") != expected: + raise SystemExit("MCP registry documentation is out of date") + print("MCP_REGISTRY_OK") + return + OUTPUT.write_text(expected, encoding="utf-8") + print("MCP_REGISTRY_RENDERED") + + +if __name__ == "__main__": + main() diff --git a/platform/scripts/sync-profile-matrix.py b/platform/scripts/sync-profile-matrix.py new file mode 100755 index 0000000..c93e31e --- /dev/null +++ b/platform/scripts/sync-profile-matrix.py @@ -0,0 +1,126 @@ +#!/usr/bin/env python3 +"""Render and verify all client-facing profile data from one matrix.""" + +from __future__ import annotations + +import argparse +import json +import pathlib +import re + + +def load(path: pathlib.Path) -> dict: + data = json.loads(path.read_text(encoding="utf-8")) + if data.get("version") != 1 or not isinstance(data.get("profiles"), list): + raise SystemExit("Unsupported profile matrix schema") + ids: set[str] = set() + aliases: set[str] = set() + for item in data["profiles"]: + required = {"id", "alias", "context", "model_env", "description"} + missing = required - item.keys() + if missing: + raise SystemExit(f"Profile entry missing: {sorted(missing)}") + if item["id"] in ids or item["alias"] in aliases: + raise SystemExit(f"Duplicate profile or alias: {item['id']}") + if not isinstance(item["context"], int) or item["context"] < 8192: + raise SystemExit(f"Invalid context for {item['id']}") + ids.add(item["id"]) + aliases.add(item["alias"]) + if data.get("default_profile") not in ids: + raise SystemExit("default_profile is not defined") + return data + + +def write_if_changed(path: pathlib.Path, content: str) -> None: + if path.exists() and path.read_text(encoding="utf-8") == content: + return + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(content, encoding="utf-8") + + +def render_router(data: dict) -> str: + payload = { + "profiles": { + item["id"]: { + "context": item["context"], + "model_alias": item["alias"], + } + for item in data["profiles"] + } + } + return json.dumps(payload, indent=2, ensure_ascii=False) + "\n" + + +def render_docs(data: dict) -> str: + lines = [ + "# Profilmatrix", + "", + "Diese Datei wird aus `config/profile-matrix.json` erzeugt. Änderungen gehören nur in die JSON-Matrix.", + "", + f"Standardprofil: **{data['default_profile']}** · globales Ausgabelimit: **{data['max_output_tokens']} Token**", + "", + "| Profil | API-Alias | Kontext | Modell | GPU-Verteilung | Vision | MTP |", + "|---|---|---:|---|---|---|---:|", + ] + for item in data["profiles"]: + lines.append( + f"| {item['id']} | `{item['alias']}` | {item['context']:,} | " + f"{item.get('model_family', '')} | {item.get('gpu_split', '')} | " + f"{'ja' if item.get('vision') else 'nein'} | {item.get('mtp', '')} |" + ) + lines.extend(["", "## Zweck", ""]) + for item in data["profiles"]: + lines.append(f"- **{item['id']}**: {item['description']}") + return "\n".join(lines) + "\n" + + +def verify_compose(data: dict, compose: pathlib.Path) -> None: + text = compose.read_text(encoding="utf-8") + expected_cap = f'MAX_GENERATION_TOKENS: "{int(data["max_output_tokens"])}"' + if expected_cap not in text: + raise SystemExit("Router output-token cap drift") + for item in data["profiles"]: + block_match = re.search( + rf"(?ms)^ llama-{re.escape(item['id'])}:\n(?P.*?)(?=^ [a-zA-Z0-9_-]+:|\Z)", + text, + ) + if not block_match: + raise SystemExit(f"Compose service llama-{item['id']} is missing") + block = block_match.group("body") + if f"- {item['alias']}" not in block: + raise SystemExit(f"Compose alias drift for {item['id']}") + env_prefix = item["id"].upper() + if f"${{{env_prefix}_CONTEXT:-{item['context']}}}" not in block: + raise SystemExit(f"Compose context drift for {item['id']}") + + +def main() -> None: + root = pathlib.Path(__file__).resolve().parents[2] + parser = argparse.ArgumentParser() + parser.add_argument("--matrix", type=pathlib.Path, + default=root / "config/profile-matrix.json") + parser.add_argument("--router", type=pathlib.Path, + default=root / "router/router_profiles.json") + parser.add_argument("--docs", type=pathlib.Path, + default=root / "docs/STANDARD_PROFILE_MATRIX.md") + parser.add_argument("--compose", type=pathlib.Path, + default=root / "compose.yaml") + parser.add_argument("--check", action="store_true") + args = parser.parse_args() + data = load(args.matrix) + router = render_router(data) + docs = render_docs(data) + verify_compose(data, args.compose) + if args.check: + if args.router.read_text(encoding="utf-8") != router: + raise SystemExit("router_profiles.json is not generated from the matrix") + if args.docs.read_text(encoding="utf-8") != docs: + raise SystemExit("profile documentation is not generated from the matrix") + else: + write_if_changed(args.router, router) + write_if_changed(args.docs, docs) + print("PROFILE_MATRIX_OK") + + +if __name__ == "__main__": + main() diff --git a/restore.sh b/restore.sh index 1c1f64a..9128008 100755 --- a/restore.sh +++ b/restore.sh @@ -23,8 +23,7 @@ tar -xzf "$ARCHIVE" -C "$work" [[ -d $work/volumes ]] || die "Backup enthält keine Docker-Volumes." # Stop only users of the restored volumes. WireGuard, SSH and networking stay up. -for container in mike-ai-open-webui mike-ai-router mike-ai-profile-controller \ - mike-ai-piper mike-ai-tools-tinysearch; do +for container in mike-ai-router mike-ai-profile-controller mike-ai-piper; do if [[ $(docker inspect -f '{{.State.Running}}' "$container" 2>/dev/null || true) == true ]]; then docker stop "$container" >/dev/null fi @@ -42,19 +41,11 @@ restore_volume() { rsync -a --delete "$source/" "$mountpoint/" } -restore_volume mike-ai_open-webui-data "$work/volumes/open-webui-data" restore_volume mike-ai_piper-data "$work/volumes/piper-data" restore_volume mike-ai_router-state "$work/volumes/router-state" restore_volume mike-ai_router-images "$work/volumes/router-images" -restore_volume mike-ai-tools_tinysearch-models "$work/volumes/tinysearch-models" cd "$ROOT_DIR" -docker compose --env-file /etc/mike-ai/stack.env \ - --profile homeassistant --profile arr --profile navidrome --profile deemix --profile github \ - up -d -python3 platform/mcp/sync-clients.py \ - --registry config/mcp-registry.json \ - --hermes /data/hermes/config.yaml \ - --openwebui-db "$(docker volume inspect -f '{{.Mountpoint}}' mike-ai_open-webui-data)/webui.db" +./manage.sh deploy core printf 'ATHENA_RESTORE_OK %s\n' "$ARCHIVE" diff --git a/router/ai_profile_router.py b/router/ai_profile_router.py index ae3b849..698f093 100755 --- a/router/ai_profile_router.py +++ b/router/ai_profile_router.py @@ -62,6 +62,7 @@ import logging import os import queue import re +import select import socket import subprocess import sys @@ -2098,6 +2099,21 @@ class Handler(BaseHTTPRequestHandler): chunk = resp.read1(16384) if not chunk: break + # A closed downstream socket may still accept one small write + # into the kernel buffer. Detect the peer FIN before writing so + # slow upstream streams are cancelled promptly and do not keep + # an inference lease occupied until generation completes. + try: + readable, _, _ = select.select([self.connection], [], [], 0) + if readable: + flags = socket.MSG_PEEK | getattr(socket, "MSG_DONTWAIT", 0) + if self.connection.recv(1, flags) == b"": + log.info("Client trennte Streaming-Verbindung; Upstream wird abgebrochen") + break + except (BlockingIOError, InterruptedError): + pass + except OSError: + break self.wfile.write(chunk) self.wfile.flush() except (OSError, http.client.HTTPException) as e: diff --git a/smoke-test.sh b/smoke-test.sh new file mode 100755 index 0000000..6d22e8a --- /dev/null +++ b/smoke-test.sh @@ -0,0 +1,82 @@ +#!/usr/bin/env bash +# Short, read-only end-to-end test for every Athena release. +set -Eeuo pipefail + +ROOT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +ENV_FILE=${STACK_ENV:-/etc/mike-ai/stack.env} + +pass() { printf 'PASS %s\n' "$*"; } +fail() { printf 'FAIL %s\n' "$*" >&2; exit 1; } + +[[ $EUID -eq 0 ]] || fail "Bitte als root ausfuehren." +command -v docker >/dev/null || fail "Docker fehlt." +[[ -s $ENV_FILE ]] || fail "Stack-Konfiguration fehlt: $ENV_FILE" + +cd "$ROOT_DIR" +./manage.sh validate >/dev/null +pass "Profilmatrix, MCP-Liste und Compose-Konfiguration stimmen" + +healthy() { + local name=$1 state health + state=$(docker inspect -f '{{.State.Status}}' "$name" 2>/dev/null || true) + health=$(docker inspect -f '{{if .State.Health}}{{.State.Health.Status}}{{end}}' "$name" 2>/dev/null || true) + [[ $state == running && ( -z $health || $health == healthy ) ]] +} + +for name in \ + mike-ai-wireguard-gateway \ + mike-ai-profile-controller \ + mike-ai-router \ + mike-ai-piper \ + mike-ai-xtts \ + mike-ai-tts-gateway \ + mike-ai-mcp-athena-operator \ + mike-ai-backup; do + healthy "$name" || fail "$name fehlt oder ist nicht gesund" +done +pass "Athena-Kerndienste sind gesund" + +active_llama=$(docker ps --format '{{.Names}}' | \ + grep -Ec '^mike-ai-llama-(fast|medium|large|ultra|uncensored)$' || true) +[[ $active_llama -eq 1 ]] || fail "$active_llama aktive llama-Profile (erwartet: 1)" +pass "Genau ein llama.cpp-Profil ist aktiv" + +for legacy in \ + mike-ai-open-webui \ + mike-ai-hermes \ + mike-ai-hermes-webui \ + mike-ai-hermes-webui-vpn-proxy \ + mike-ai-mcp-web \ + mike-ai-tools-searxng \ + mike-ai-tools-tinysearch \ + mike-ai-mcp-platform-context; do + [[ $(docker inspect -f '{{.State.Running}}' "$legacy" 2>/dev/null || true) != true ]] || \ + fail "Altlast laeuft noch: $legacy" +done +pass "OpenWebUI, alte Hermes-Dienste und migrierte MCPs sind aus" + +docker exec mike-ai-router python - <<'PY' >/dev/null +import json +import os +import urllib.request + +key = os.environ["ROUTER_API_KEY"] +for path in ("/health", "/v1/models"): + request = urllib.request.Request( + "http://127.0.0.1:8081" + path, + headers={"Authorization": "Bearer " + key}, + ) + with urllib.request.urlopen(request, timeout=10) as response: + body = response.read() + if response.status != 200: + raise SystemExit(f"{path}: HTTP {response.status}") + if path.endswith("models") and not json.loads(body).get("data"): + raise SystemExit("Router meldet keine Modelle") +PY +pass "Router-Liveness und OpenAI-Modellliste funktionieren" + +latest=/data/docker-backups/athena-latest.tar.gz +[[ -s $latest ]] || fail "Kein gueltiges Athena-Backup: $latest" +pass "Automatisches Athena-Backup ist vorhanden" + +printf '\nATHENA_E2E_OK\n'