--- name: athena-operator description: Operate and extend the Athena AI platform safely. license: MIT metadata: hermes: version: 0.1.0 author: Michael Roll, Hermes Agent platforms: [linux, macos, windows] tags: [athena, operations, docker, mcp, models, recovery] related_skills: [] --- # Athena Operator Skill Operate, diagnose, extend, and recover the Athena AI platform through its platform-context and operator MCPs. Keep durable truth in the versioned Athena repository and use live measurements only as evidence of current state. ## When to Use - Use for Athena, MikeAI, Docker-stack, router, inference-profile, model, benchmark, MCP, Hermes, Open WebUI, TTS, STT, image, Git-deploy, and recovery work on the Athena host. - Use when a user asks how Athena is built or whether an Athena component is running, configured, documented, reproducible, or recoverable. - Do not use as the primary tool for Home Assistant, Unraid, Sonarr, Radarr, Navidrome, or GitHub data when their specialist MCP is available. ## Prerequisites - Require `Athena Plattformwissen` for architecture and versioned knowledge. - Require `Athena Operator` for live host inspection and execution. - Treat missing or failed tools as missing evidence. Never invent state, output, files, logs, or completed actions. - Never request, reveal, copy into chat, or commit secret values. ## Tool Selection 1. Identify the target system before calling a tool. 2. For Athena itself, begin with `athena_operator_inspect`. Use `athena_get_overview` when architectural context is needed. 3. For external services, prefer the narrow specialist MCP. Use `athena_get_external_services` before planning a duplicate service. 4. Use `athena_operator_search_source` and `athena_operator_read_source` for deployed code. Use `athena_search_knowledge` and `athena_read_source` for documentation. Do not guess paths or configuration. 5. Use `athena_operator_terminal` as the broad Athena escape hatch only when a structured tool is too narrow. Keep commands focused and outputs bounded. ## Procedure 1. **Establish evidence.** Inspect only the relevant live subject and read the smallest authoritative source section. Completion: runtime facts and source facts are separately identified. 2. **Check drift.** Compare live state with versioned source and current reference documentation. Completion: any mismatch is named before changes. 3. **Protect the active workload.** Check jobs and active containers. Do not restart, recreate, switch profiles, alter shared configuration, or consume required GPU capacity while an important request or benchmark is active. Completion: work is either proven idle or the change is staged only. 4. **Plan rollback.** Name the files, services, validation, rollback artifact, and expected user-visible effect. Completion: rollback is possible without relying on chat history. 5. **Change the source of truth.** Modify repository sources, not only a live container. Prefer `athena_operator_prepare`; show its full preview and stop for the exact user confirmation before `athena_operator_execute`. Use `file_update` for the exact source paths. Never clone the repository in the Hermes sandbox and never request or copy an SSH key; the Operator owns the canonical worktree and its deploy credentials. Completion: the approved ticket matches the intended content. 6. **Deploy narrowly.** Change only named services. Never restart the entire stack merely to activate one component. Completion: unrelated containers and the active inference request remain undisturbed. 7. **Verify behavior.** Run syntax/config checks, focused tests, service health, and one bounded functional test. A running container alone is not proof. Completion: expected behavior and rollback path are both verified. 8. **Close the maintenance loop.** Update relevant docs, publish only the explicitly selected changed paths with `git_publish`, create a newer recovery bundle, then check maintenance status. Completion: source commit, deployed state, docs, and recovery agree. ## Persistent Versus Temporary Work - "Use" a missing helper for one task: place it in a task-specific temporary location or ephemeral container and remove it afterward. - "Install", "add", "deploy", or "make permanent": implement it in repository source, documentation, installation flow, and recovery. - Do not create a second backend merely because an existing service is stopped, inaccessible, or absent from one tool catalogue. ## Remote-Safety Boundary - Athena has no physical console or KVM. Never attempt power control or changes to Athena SSH, LAN, WireGuard, firewall, boot, kernel, drivers, mounts, or partitions through this workflow. - Do not stop or restart the WireGuard gateway as a side effect of ordinary deployment. Bind user services to the VPN path; keep them unavailable from the university LAN. - Inside the trusted VPN, normal service communication and Internet access are allowed. Do not add extra egress restrictions unless the user requests them. ## Pitfalls - Profile names are not simultaneous models; exactly one text profile is active. - A profile switch can terminate active generation and invalidate prompt cache. - A healthy container can still expose the wrong model, route, or tool set. - `/opt/mike-ai/stack` is deployed source, not automatically the canonical Git worktree. Complete durable changes through Operator operations `file_update`, `run_checks`, `compose_deploy`, `git_publish`, and `recovery`. - New skills are loaded at the next Hermes session; absence in the current session is expected. ## Verification - State the tools that supplied each important live claim. - List every modified source file and every deployed service. - Report focused test and health results, not vague success language. - If Git publication or recovery creation is incomplete, call it unfinished maintenance rather than declaring the task fully complete.