docs: README und .gitignore
This commit is contained in:
@@ -0,0 +1,132 @@
|
||||
# AI Profile Router
|
||||
|
||||
Kleiner OpenAI-kompatibler Proxy (Python, nur Standardbibliothek) vor einem
|
||||
lokalen llama.cpp-Server. Er leitet normale OpenAI-Requests transparent
|
||||
weiter (Streaming, Tool Calls, JSON) und schaltet zwischen drei festen
|
||||
llama.cpp-Profilen um.
|
||||
|
||||
## Zielsystem
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| Host | 192.168.1.196 |
|
||||
| SSH | `root` mit Key `lmstudio_unraid` |
|
||||
| llama.cpp | `http://127.0.0.1:8080` (Service `mike-ai-llama-ui.service`) |
|
||||
| Profil-Skript | `/usr/local/bin/llama-profile {fast\|medium\|long}` |
|
||||
| Router-Port | **8081** |
|
||||
| Router-Service | `mike-ai-profile-router.service` |
|
||||
|
||||
## Profile / virtuelle Modelle
|
||||
|
||||
| Profil | Modell | Kontext |
|
||||
|---|---|---|
|
||||
| `fast` | `qwen-fast` | 73728 |
|
||||
| `medium` | `qwen-medium` | 94208 |
|
||||
| `long` | `qwen-long` | 131072 |
|
||||
|
||||
## Endpunkte
|
||||
|
||||
| Endpunkt | Beschreibung |
|
||||
|---|---|
|
||||
| `GET /v1/models` | Die drei virtuellen Modelle inkl. `context_length`/`context_window` |
|
||||
| `GET /status` | Aktives Profil, Upstream-Zustand, Modell, Kontext, Uptime |
|
||||
| `POST /fast` `/medium` `/long` | Profilwechsel (auch `GET` möglich) |
|
||||
| `POST /v1/chat/completions` | Weiterleitung an llama.cpp (Streaming + Tool Calls) |
|
||||
| alles andere | Transparente Weiterleitung an llama.cpp |
|
||||
|
||||
### Verhalten
|
||||
|
||||
- **Virtuelles Modell** (`qwen-fast`/`qwen-medium`/`qwen-long` in
|
||||
`chat.completions`): Der Router stellt sicher, dass das passende Profil
|
||||
aktiv ist (notfalls Wechsel + Warten), ersetzt das Modell durch das echte
|
||||
llama.cpp-Modell und leitet weiter.
|
||||
- **Profilwechsel** (`POST /fast` …): Führt `/usr/local/bin/llama-profile
|
||||
<profil>` aus (ohne Shell, feste Argumente → keine Injection), wartet dann,
|
||||
bis llama.cpp wieder erreichbar ist, und liefert erst dann `200`.
|
||||
- **Impliziter Wechsel bei downem llama.cpp**: Ein Chat-Request mit virtuellem
|
||||
Modell liefert sofort `502`, wenn das Profil bereits aktiv ist, aber
|
||||
llama.cpp down ist (kein stiller Neustart). Der Neustart wird explizit über
|
||||
`POST /<profil>` angestoßen.
|
||||
- **Ungültige Profile/Modelle**: `POST /<anderes>` → `400`;
|
||||
`qwen-<anderes>` als Modell → `400`. Nur die drei festen Profile sind
|
||||
schaltbar.
|
||||
- **Fehlerformat**: OpenAI-kompatibel (`{"error": {"message", "type", "code"}}`).
|
||||
|
||||
## Repository-Struktur
|
||||
|
||||
```
|
||||
router/ai_profile_router.py # der Router (einzige Laufzeit-Datei)
|
||||
deploy/mike-ai-profile-router.service # systemd-Unit
|
||||
deploy/install.sh # läuft auf dem Zielsystem (per SSH)
|
||||
deploy/deploy.sh # läuft lokal: SCP + SSH
|
||||
dev/ # lokale Tests (Mock-llama.cpp, Fake-Profil-Skript)
|
||||
```
|
||||
|
||||
Entwicklungsdateien (`dev/`) und Deployment-Dateien (`router/`, `deploy/`)
|
||||
sind getrennt. Auf dem Zielsystem landet nur `router/` + `deploy/`.
|
||||
|
||||
## Deployment
|
||||
|
||||
Voraussetzung: SSH-Key `~/.ssh/lmstudio_unraid` (bereits vorhanden).
|
||||
|
||||
```bash
|
||||
./deploy/deploy.sh
|
||||
```
|
||||
|
||||
Das Skript:
|
||||
|
||||
1. Überträgt `ai_profile_router.py`, `install.sh` und die systemd-Unit per
|
||||
SCP nach `/tmp/ai-profile-router/` auf dem Zielsystem.
|
||||
2. Führt `install.sh` per SSH aus, das:
|
||||
- den alten Router (`mike-ai-local-llm-router.service` +
|
||||
`/opt/mike-ai/local-llm-router`) **mit Backup** entfernt,
|
||||
- den neuen Router nach `/opt/mike-ai/ai-profile-router/` installiert,
|
||||
- `mike-ai-profile-router.service` aktiviert (Start beim Boot) und startet,
|
||||
- `GET /status` verifiziert.
|
||||
|
||||
Die Installation ist idempotent (Update = erneut ausführen).
|
||||
|
||||
## Konfiguration (Umgebungsvariablen in der systemd-Unit)
|
||||
|
||||
| Variable | Default | Bedeutung |
|
||||
|---|---|---|
|
||||
| `ROUTER_HOST` | `0.0.0.0` | Bind-Adresse |
|
||||
| `ROUTER_PORT` | `8081` | Port |
|
||||
| `UPSTREAM_URL` | `http://127.0.0.1:8080` | llama.cpp |
|
||||
| `PROFILE_SCRIPT` | `/usr/local/bin/llama-profile` | Profil-Skript |
|
||||
| `PROFILE_DIR` | `/etc/systemd/system/mike-ai-llama-ui.service.d` | Ort der `override.conf` |
|
||||
| `SWITCH_TIMEOUT` | `600` | Warten auf llama.cpp nach Wechsel (s) |
|
||||
| `REQUEST_TIMEOUT` | `600` | Read-Timeout für Upstream-Requests (s) |
|
||||
| `CONNECT_TIMEOUT` | `10` | Connect-Timeout Upstream (s) |
|
||||
| `POLL_INTERVAL` | `2` | Polling-Intervall (s) |
|
||||
| `LOG_LEVEL` | `INFO` | Logging-Level |
|
||||
|
||||
## Lokale Tests
|
||||
|
||||
```bash
|
||||
./dev/test_local.sh
|
||||
```
|
||||
|
||||
Startet einen Mock-llama.cpp und den Router mit einem Fake-Profil-Skript und
|
||||
prüft: `/v1/models`, `/status`, Forwarding, Streaming, Tool Calls,
|
||||
Profilwechsel (fast→medium→fast), virtuelles Modell triggert Wechsel,
|
||||
ungültige Profile, Upstream down → 502, Recovery.
|
||||
|
||||
## Betrieb
|
||||
|
||||
```bash
|
||||
systemctl status mike-ai-profile-router
|
||||
journalctl -u mike-ai-profile-router -f
|
||||
curl -s http://192.168.1.196:8081/status | python3 -m json.tool
|
||||
curl -s -X POST http://192.168.1.196:8081/medium
|
||||
```
|
||||
|
||||
## Sicherheit
|
||||
|
||||
- Keine Shell-Aufrufe: Profil-Skript wird mit `subprocess.run([script, profil])`
|
||||
aufgerufen, `profil` ist Whitelist-geprüft (`fast|medium|long`).
|
||||
- Keine Secrets/Tokens im Code oder in der Unit.
|
||||
- systemd-Hardening: `NoNewPrivileges=true`, `PrivateTmp=true`.
|
||||
- Die bestehende llama.cpp-/Profil-Konfiguration wird nicht verändert; der
|
||||
Router nutzt nur das vorhandene `llama-profile`-Skript und liest die
|
||||
`override.conf`.
|
||||
Reference in New Issue
Block a user