Rebuild profile router as secure recoverable v2

This commit is contained in:
Mikei386
2026-08-20 13:41:00 +02:00
parent bdd6c08643
commit fd4a3055e8
18 changed files with 1162 additions and 129 deletions
+36 -6
View File
@@ -52,6 +52,7 @@ llama.cpp :8080 + profilabhängige MCP-Server
- [Saubere Installation](docs/INSTALLATION.md)
- [Betrieb und Profilwechsel](docs/OPERATIONS.md)
- [Sicherheitsmodell](docs/SECURITY.md)
- [Router V2: Migration und Kompatibilität](docs/ROUTER_V2_MIGRATION.md)
- [Migration vom bestehenden Host](docs/MIGRATION.md)
- [MCP-Aufteilung](platform/mcp/README.md)
- [llama.cpp-Build und Profile](platform/llama/README.md)
@@ -92,11 +93,13 @@ Sprachausgabe bereit (XTTS-v2, CPU-only, OpenAI-kompatibel).
| Endpunkt | Beschreibung |
|---|---|
| `GET /v1/models` | Die drei virtuellen Modelle inkl. `context_length`/`context_window` |
| `GET /health` | öffentliche Liveness-Prüfung des Routerprozesses |
| `GET /ready` | öffentliche Readiness-Prüfung von Router + Textmodell |
| `GET /status` | Aktives Profil, Upstream-Zustand, Modell, Kontext, Uptime |
| `POST /fast` `/medium` `/long` | Profilwechsel (auch `GET` möglich) |
| `POST /fast` `/medium` `/long` | expliziter Profilwechsel |
| `POST /v1/chat/completions` | Weiterleitung an llama.cpp (Streaming + Tool Calls) |
| `POST /v1/images/generations` | Bildgenerierung (FLUX.2 [klein] 4B Base, OpenAI-kompatibel) |
| `GET /images` | Liste der gespeicherten Bilder (max. 200) |
| `GET /images` | Liste der aufbewahrten Bilder |
| `GET /images/<datei>` | PNG-Download (nur `images/`-Verzeichnis, validiert) |
| `POST /v1/audio/speech` | Sprachausgabe (XTTS-v2, OpenAI-kompatibel) |
| `POST /v1/audio/transcriptions` | Deutsche Spracherkennung (whisper.cpp, OpenAI-kompatibel) |
@@ -105,6 +108,12 @@ Sprachausgabe bereit (XTTS-v2, CPU-only, OpenAI-kompatibel).
| `POST /vision/test` | Direkter Vision-Test (Bild + Frage → Q3-Analyse) |
| alles andere | Transparente Weiterleitung an llama.cpp |
Bis auf `GET /health` und `GET /ready` benötigen alle Endpunkte einen
Router-API-Key als `Authorization: Bearer …` oder `X-API-Key`. Der Installer
erzeugt ihn einmalig in `/etc/mike-ai/router-api-key` und gibt ihn niemals im
Installationslog aus. Ein fehlender oder zu kurzer Schlüssel verhindert den
Produktivstart (fail closed).
### Verhalten
- **Virtuelles Modell** (`qwen-fast`/`qwen-medium`/`qwen-long` in
@@ -159,6 +168,7 @@ Beispiel:
```bash
curl -s http://AI_HOST:8081/v1/images/generations \
-H "Authorization: Bearer $ROUTER_KEY" \
-H 'Content-Type: application/json' \
-d '{"prompt":"ein roter Würfel auf weißem Grund","size":"1024x1024","quality":"standard"}'
```
@@ -306,11 +316,13 @@ Beispiele:
```bash
# MP3 (Default), Stimme claribel
curl -s http://AI_HOST:8081/v1/audio/speech \
-H "Authorization: Bearer $ROUTER_KEY" \
-H 'Content-Type: application/json' \
-d '{"input":"Hallo, dies ist ein Test.","voice":"claribel"}' -o out.mp3
# WAV, 1.5x Tempo
curl -s http://AI_HOST:8081/v1/audio/speech \
-H "Authorization: Bearer $ROUTER_KEY" \
-H 'Content-Type: application/json' \
-d '{"input":"Guten Tag.","voice":"claribel","speed":1.5,"response_format":"wav"}' -o out.wav
```
@@ -400,17 +412,20 @@ Beispiele:
```bash
# WebM/Opus (z.B. aus Open WebUI-Mikrofon)
curl -s http://AI_HOST:8081/v1/audio/transcriptions \
-H "Authorization: Bearer $ROUTER_KEY" \
-F "file=@aufnahme.webm" \
-F "model=whisper-1"
# WAV mit expliziter Sprache
curl -s http://AI_HOST:8081/v1/audio/transcriptions \
-H "Authorization: Bearer $ROUTER_KEY" \
-F "file=@aufnahme.wav" \
-F "model=whisper-1" \
-F "language=de"
# Verbose-Format
curl -s http://AI_HOST:8081/v1/audio/transcriptions \
-H "Authorization: Bearer $ROUTER_KEY" \
-F "file=@aufnahme.wav" \
-F "model=whisper-1" \
-F "response_format=verbose_json"
@@ -510,6 +525,11 @@ Die Installation ist idempotent (Update = erneut ausführen).
|---|---|---|
| `ROUTER_HOST` | `0.0.0.0` | Bind-Adresse |
| `ROUTER_PORT` | `8081` | Port |
| `ROUTER_AUTH_MODE` | `required` | Authentifizierung; `off` nur für lokale Tests |
| `ROUTER_API_KEY_FILE` | `/etc/mike-ai/router-api-key` | Schlüsseldatei (0600) |
| `ROUTER_STATE_FILE` | `/var/lib/mike-ai-profile-router/state.json` | atomarer Crash-/Recovery-Zustand |
| `ROUTER_PROFILES_FILE` | `/etc/mike-ai/router-profiles.json` | Profile, Kontext und erwarteter Alias |
| `ROUTER_MAX_CONCURRENT_REQUESTS` | `16` | harte Grenze paralleler Requests |
| `UPSTREAM_URL` | `http://127.0.0.1:8080` | llama.cpp |
| `PROFILE_SCRIPT` | `/usr/local/bin/llama-profile` | Profil-Skript |
| `PROFILE_DIR` | `/etc/systemd/system/mike-ai-llama-ui.service.d` | Ort der `override.conf` |
@@ -538,6 +558,8 @@ Die Installation ist idempotent (Update = erneut ausführen).
| `VISION_UNLOAD_TIMEOUT` | `120` | Warten auf VRAM-Freiheit (s) |
| `VISION_MAX_TOKENS` | `4096` | Max. Tokens der Vision-Analyse |
| `VISION_CACHE_MAX` | `64` | Größe des Analyse-Caches (LRU) |
| `VISION_MAX_IMAGE_BYTES` | `20971520` | maximales dekodiertes Bild (20 MiB) |
| `VISION_ALLOW_REMOTE_URLS` | `false` | externe Bild-URLs; standardmäßig SSRF-sicher aus |
| `VISION_LOG` | `/opt/mike-ai/ai-profile-router/vision_server.log` | Vision-Server-Log |
TTS-Worker (`mike-ai-xtts.service`):
@@ -567,7 +589,8 @@ paralleler Chat während Bild-Job wartet statt 502) und **Sprachausgabe**
(`/status` mit tts-Section, `POST /v1/audio/speech` wav/mp3, Validierung,
Worker-Fehler→503, Worker down→503, Worker-Neustart→Recovery).
Aktuell: **43 Tests** (32 bestehende + 11 TTS-Assertions).
Aktuell: **63 Integrationsassertions** plus Chunked-, Multipart- und
Security/Recovery-Unit-Tests.
## Betrieb
@@ -576,10 +599,12 @@ systemctl status mike-ai-profile-router
systemctl status mike-ai-xtts
journalctl -u mike-ai-profile-router -f
journalctl -u mike-ai-xtts -f
curl -s http://AI_HOST:8081/status | python3 -m json.tool
curl -s -X POST http://AI_HOST:8081/medium
ROUTER_KEY='aus lokalem Secret-Store'
curl -s http://AI_HOST:8081/status -H "Authorization: Bearer $ROUTER_KEY" | python3 -m json.tool
curl -s -X POST http://AI_HOST:8081/medium -H "Authorization: Bearer $ROUTER_KEY"
# TTS-Test
curl -s http://AI_HOST:8081/v1/audio/speech \
-H "Authorization: Bearer $ROUTER_KEY" \
-H 'Content-Type: application/json' \
-d '{"input":"Hallo","voice":"claribel"}' -o test.mp3
```
@@ -589,7 +614,12 @@ curl -s http://AI_HOST:8081/v1/audio/speech \
- Keine Shell-Aufrufe: Profil-Skript wird mit `subprocess.run([script, profil])`
aufgerufen, `profil` ist Whitelist-geprüft (`fast|medium|long`).
- Keine Secrets/Tokens im Code oder in der Unit.
- systemd-Hardening: `NoNewPrivileges=true`, `PrivateTmp=true`.
- API-Key-Pflicht, konstante Schlüsselprüfung und keine Weitergabe des
Router-Schlüssels an llama.cpp.
- Remote-Bilder standardmäßig gesperrt; lokale Data-URLs sind typ- und
größenvalidiert.
- systemd-Hardening: `NoNewPrivileges`, `PrivateTmp`, `ProtectHome`,
Kernel-/Control-Group-Schutz und restriktive Dateirechte.
- Die bestehende llama.cpp-/Profil-Konfiguration wird nicht verändert; der
Router nutzt nur das vorhandene `llama-profile`-Skript und liest die
`override.conf`.