diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..b908d4c --- /dev/null +++ b/.gitignore @@ -0,0 +1,3 @@ +__pycache__/ +*.pyc +.DS_Store diff --git a/README.md b/README.md new file mode 100644 index 0000000..6e0b9a8 --- /dev/null +++ b/README.md @@ -0,0 +1,132 @@ +# AI Profile Router + +Kleiner OpenAI-kompatibler Proxy (Python, nur Standardbibliothek) vor einem +lokalen llama.cpp-Server. Er leitet normale OpenAI-Requests transparent +weiter (Streaming, Tool Calls, JSON) und schaltet zwischen drei festen +llama.cpp-Profilen um. + +## Zielsystem + +| | | +|---|---| +| Host | 192.168.1.196 | +| SSH | `root` mit Key `lmstudio_unraid` | +| llama.cpp | `http://127.0.0.1:8080` (Service `mike-ai-llama-ui.service`) | +| Profil-Skript | `/usr/local/bin/llama-profile {fast\|medium\|long}` | +| Router-Port | **8081** | +| Router-Service | `mike-ai-profile-router.service` | + +## Profile / virtuelle Modelle + +| Profil | Modell | Kontext | +|---|---|---| +| `fast` | `qwen-fast` | 73728 | +| `medium` | `qwen-medium` | 94208 | +| `long` | `qwen-long` | 131072 | + +## Endpunkte + +| Endpunkt | Beschreibung | +|---|---| +| `GET /v1/models` | Die drei virtuellen Modelle inkl. `context_length`/`context_window` | +| `GET /status` | Aktives Profil, Upstream-Zustand, Modell, Kontext, Uptime | +| `POST /fast` `/medium` `/long` | Profilwechsel (auch `GET` möglich) | +| `POST /v1/chat/completions` | Weiterleitung an llama.cpp (Streaming + Tool Calls) | +| alles andere | Transparente Weiterleitung an llama.cpp | + +### Verhalten + +- **Virtuelles Modell** (`qwen-fast`/`qwen-medium`/`qwen-long` in + `chat.completions`): Der Router stellt sicher, dass das passende Profil + aktiv ist (notfalls Wechsel + Warten), ersetzt das Modell durch das echte + llama.cpp-Modell und leitet weiter. +- **Profilwechsel** (`POST /fast` …): Führt `/usr/local/bin/llama-profile + ` aus (ohne Shell, feste Argumente → keine Injection), wartet dann, + bis llama.cpp wieder erreichbar ist, und liefert erst dann `200`. +- **Impliziter Wechsel bei downem llama.cpp**: Ein Chat-Request mit virtuellem + Modell liefert sofort `502`, wenn das Profil bereits aktiv ist, aber + llama.cpp down ist (kein stiller Neustart). Der Neustart wird explizit über + `POST /` angestoßen. +- **Ungültige Profile/Modelle**: `POST /` → `400`; + `qwen-` als Modell → `400`. Nur die drei festen Profile sind + schaltbar. +- **Fehlerformat**: OpenAI-kompatibel (`{"error": {"message", "type", "code"}}`). + +## Repository-Struktur + +``` +router/ai_profile_router.py # der Router (einzige Laufzeit-Datei) +deploy/mike-ai-profile-router.service # systemd-Unit +deploy/install.sh # läuft auf dem Zielsystem (per SSH) +deploy/deploy.sh # läuft lokal: SCP + SSH +dev/ # lokale Tests (Mock-llama.cpp, Fake-Profil-Skript) +``` + +Entwicklungsdateien (`dev/`) und Deployment-Dateien (`router/`, `deploy/`) +sind getrennt. Auf dem Zielsystem landet nur `router/` + `deploy/`. + +## Deployment + +Voraussetzung: SSH-Key `~/.ssh/lmstudio_unraid` (bereits vorhanden). + +```bash +./deploy/deploy.sh +``` + +Das Skript: + +1. Überträgt `ai_profile_router.py`, `install.sh` und die systemd-Unit per + SCP nach `/tmp/ai-profile-router/` auf dem Zielsystem. +2. Führt `install.sh` per SSH aus, das: + - den alten Router (`mike-ai-local-llm-router.service` + + `/opt/mike-ai/local-llm-router`) **mit Backup** entfernt, + - den neuen Router nach `/opt/mike-ai/ai-profile-router/` installiert, + - `mike-ai-profile-router.service` aktiviert (Start beim Boot) und startet, + - `GET /status` verifiziert. + +Die Installation ist idempotent (Update = erneut ausführen). + +## Konfiguration (Umgebungsvariablen in der systemd-Unit) + +| Variable | Default | Bedeutung | +|---|---|---| +| `ROUTER_HOST` | `0.0.0.0` | Bind-Adresse | +| `ROUTER_PORT` | `8081` | Port | +| `UPSTREAM_URL` | `http://127.0.0.1:8080` | llama.cpp | +| `PROFILE_SCRIPT` | `/usr/local/bin/llama-profile` | Profil-Skript | +| `PROFILE_DIR` | `/etc/systemd/system/mike-ai-llama-ui.service.d` | Ort der `override.conf` | +| `SWITCH_TIMEOUT` | `600` | Warten auf llama.cpp nach Wechsel (s) | +| `REQUEST_TIMEOUT` | `600` | Read-Timeout für Upstream-Requests (s) | +| `CONNECT_TIMEOUT` | `10` | Connect-Timeout Upstream (s) | +| `POLL_INTERVAL` | `2` | Polling-Intervall (s) | +| `LOG_LEVEL` | `INFO` | Logging-Level | + +## Lokale Tests + +```bash +./dev/test_local.sh +``` + +Startet einen Mock-llama.cpp und den Router mit einem Fake-Profil-Skript und +prüft: `/v1/models`, `/status`, Forwarding, Streaming, Tool Calls, +Profilwechsel (fast→medium→fast), virtuelles Modell triggert Wechsel, +ungültige Profile, Upstream down → 502, Recovery. + +## Betrieb + +```bash +systemctl status mike-ai-profile-router +journalctl -u mike-ai-profile-router -f +curl -s http://192.168.1.196:8081/status | python3 -m json.tool +curl -s -X POST http://192.168.1.196:8081/medium +``` + +## Sicherheit + +- Keine Shell-Aufrufe: Profil-Skript wird mit `subprocess.run([script, profil])` + aufgerufen, `profil` ist Whitelist-geprüft (`fast|medium|long`). +- Keine Secrets/Tokens im Code oder in der Unit. +- systemd-Hardening: `NoNewPrivileges=true`, `PrivateTmp=true`. +- Die bestehende llama.cpp-/Profil-Konfiguration wird nicht verändert; der + Router nutzt nur das vorhandene `llama-profile`-Skript und liest die + `override.conf`.