Expand router into reproducible local AI platform

This commit is contained in:
Mikei386 committed 2026-08-20 12:56:53 +02:00
1 parent a84220725a
commit 0e4a9de5ba
34 files changed
+2584 -31

No files matched your search

+69 -13
View File
@@ -1,4 +1,59 @@
# AI Profile Router
# Lokale KI-Plattform
Dieses private Repository dokumentiert und installiert die reproduzierbare
KI-Umgebung rund um einen lokalen `llama.cpp`-Host. Der **AI Profile Router**
bleibt der zentrale Bestandteil, ist aber nicht mehr die einzige Komponente.
## Plattform auf einen Blick
```text
Clients (Open WebUI, Hermes, Apps)
|
v
AI Profile Router :8081
|-- qwen-fast (72K, maximale Geschwindigkeit)
|-- qwen-medium (92K, reines IQ4_XS)
|-- qwen-long (128K, CPU-Offload)
|-- Vision-Hotswap
|-- lokale Bildgenerierung
|-- Whisper STT
`-- XTTS TTS
|
v
llama.cpp :8080 + profilabhängige MCP-Server
```
### Enthalten
- produktiver Router samt Tests und Deployment
- fest definierte Fast-/Medium-/Long-Profile
- reproduzierbarer `llama.cpp`-Build über einen festgelegten Commit
- systemd-Vorlagen und Profilumschaltung
- Modellmanifest ohne Modelldateien
- sichere MCP-Beispielkonfiguration ohne Zugangsdaten
- Installations-, Betriebs-, Sicherheits- und Migrationsanleitung
- Prüfskript für einen frisch installierten Host
### Bewusst nicht enthalten
- GGUF-, Whisper-, FLUX- oder XTTS-Modelldateien
- API-Schlüssel, Tokens, SSH-Schlüssel oder Zertifikate
- Chats, Prompts, Logs, Bilder oder Audiodateien
- alte Benchmarks, experimentelle Builds und ausgemusterte RX-Dienste
- hostgebundene Backups und Cache-Verzeichnisse
## Dokumentation
- [Architektur](docs/ARCHITECTURE.md)
- [Komponentenverzeichnis](docs/COMPONENTS.md)
- [Saubere Installation](docs/INSTALLATION.md)
- [Betrieb und Profilwechsel](docs/OPERATIONS.md)
- [Sicherheitsmodell](docs/SECURITY.md)
- [Migration vom bestehenden Host](docs/MIGRATION.md)
- [MCP-Aufteilung](platform/mcp/README.md)
- [llama.cpp-Build und Profile](platform/llama/README.md)
## AI Profile Router
Kleiner OpenAI-kompatibler Proxy (Python, nur Standardbibliothek) vor einem
lokalen llama.cpp-Server. Er leitet normale OpenAI-Requests transparent
@@ -12,8 +67,8 @@ Sprachausgabe bereit (XTTS-v2, CPU-only, OpenAI-kompatibel).
| | |
|---|---|
| Host | 192.168.1.196 |
| SSH | `root` mit Key `lmstudio_unraid` |
| Host | frei wählbarer Linux-KI-Host |
| SSH | administrativer Zugang nur für Installation und Wartung |
| llama.cpp | `http://127.0.0.1:8080` (Service `mike-ai-llama-ui.service`) |
| Profil-Skript | `/usr/local/bin/llama-profile {fast\|medium\|long}` |
| Router-Port | **8081** |
@@ -100,7 +155,7 @@ OpenAI-kompatibel. Unterstützt `prompt`, `size`, `n`, `seed`, `quality`,
Beispiel:
```bash
curl -s http://192.168.1.196:8081/v1/images/generations \
curl -s http://AI_HOST:8081/v1/images/generations \
-H 'Content-Type: application/json' \
-d '{"prompt":"ein roter Würfel auf weißem Grund","size":"1024x1024","quality":"standard"}'
```
@@ -247,12 +302,12 @@ Beispiele:
```bash
# MP3 (Default), Stimme claribel
curl -s http://192.168.1.196:8081/v1/audio/speech \
curl -s http://AI_HOST:8081/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"input":"Hallo, dies ist ein Test.","voice":"claribel"}' -o out.mp3
# WAV, 1.5x Tempo
curl -s http://192.168.1.196:8081/v1/audio/speech \
curl -s http://AI_HOST:8081/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"input":"Guten Tag.","voice":"claribel","speed":1.5,"response_format":"wav"}' -o out.wav
```
@@ -341,18 +396,18 @@ Beispiele:
```bash
# WebM/Opus (z.B. aus Open WebUI-Mikrofon)
curl -s http://192.168.1.196:8081/v1/audio/transcriptions \
curl -s http://AI_HOST:8081/v1/audio/transcriptions \
-F "file=@aufnahme.webm" \
-F "model=whisper-1"
# WAV mit expliziter Sprache
curl -s http://192.168.1.196:8081/v1/audio/transcriptions \
curl -s http://AI_HOST:8081/v1/audio/transcriptions \
-F "file=@aufnahme.wav" \
-F "model=whisper-1" \
-F "language=de"
# Verbose-Format
curl -s http://192.168.1.196:8081/v1/audio/transcriptions \
curl -s http://AI_HOST:8081/v1/audio/transcriptions \
-F "file=@aufnahme.wav" \
-F "model=whisper-1" \
-F "response_format=verbose_json"
@@ -415,7 +470,8 @@ sind getrennt. Auf dem Zielsystem landet nur `router/` + `deploy/`.
## Deployment
Voraussetzung: SSH-Key `~/.ssh/lmstudio_unraid` (bereits vorhanden).
Voraussetzung: administrativer SSH-Key für den Zielhost. Ziel und Key werden
explizit über `TARGET` und `SSH_KEY` übergeben.
```bash
./deploy/deploy.sh
@@ -517,10 +573,10 @@ systemctl status mike-ai-profile-router
systemctl status mike-ai-xtts
journalctl -u mike-ai-profile-router -f
journalctl -u mike-ai-xtts -f
curl -s http://192.168.1.196:8081/status | python3 -m json.tool
curl -s -X POST http://192.168.1.196:8081/medium
curl -s http://AI_HOST:8081/status | python3 -m json.tool
curl -s -X POST http://AI_HOST:8081/medium
# TTS-Test
curl -s http://192.168.1.196:8081/v1/audio/speech \
curl -s http://AI_HOST:8081/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"input":"Hallo","voice":"claribel"}' -o test.mp3
```