Add XTTS primary voice with Piper fallback
This commit is contained in:
+16
-6
@@ -24,7 +24,9 @@ Heimnetz / VPN-Clients
|
||||
+-- llama-uncensored (80K, Abliterated, dual GPU)
|
||||
+-- llama-experimental
|
||||
+-- llama-ultra (256K, text-only, dual GPU)
|
||||
+-- Piper-TTS (CPU, nur intern)
|
||||
+-- TTS-Gateway
|
||||
| +-- XTTS-v2 / Annmarie Nele (RTX 3060, primär)
|
||||
| +-- Piper-TTS (CPU, automatischer Fallback)
|
||||
+-- internes MCP-Netz
|
||||
+-- Web-MCP + TinySearch + SearXNG
|
||||
+-- Home-Assistant-MCP-Relay
|
||||
@@ -41,7 +43,9 @@ Heimnetz / VPN-Clients
|
||||
| Profile Router | nur Docker-intern | OpenAI-API und Profilwahl |
|
||||
| Profile Controller | nein | startet ausschließlich fest erlaubte Profile |
|
||||
| llama.cpp Profile | nein | Inferenz, Tool Calling, integrierte Vision |
|
||||
| Piper | nein | lokale deutsche Text-to-Speech-Ausgabe |
|
||||
| XTTS-v2 | nein | primäre deutsche/englische Text-to-Speech-Ausgabe auf RTX 3060 |
|
||||
| TTS-Gateway | nein | serialisiert XTTS, segmentiert Sprachwechsel und fällt auf Piper zurück |
|
||||
| Piper | nein | CPU-basierte Text-to-Speech-Rückfallebene |
|
||||
| MCP-Tool-Stack | nein | voneinander getrennte Werkzeugbereiche |
|
||||
| SearXNG/TinySearch | nein | Suchbackend des Web-MCP |
|
||||
|
||||
@@ -99,10 +103,16 @@ Ultra bleibt für maximalen Kontext bewusst text-only.
|
||||
|
||||
## Optionale Erweiterungen
|
||||
|
||||
Piper läuft als eigener CPU-Container und wird von Open WebUI über den Router
|
||||
angesprochen. Sein Port wird nicht veröffentlicht. Das Stimmenmodell liegt im
|
||||
persistenten Volume `piper-data` und wird beim ersten Start reproduzierbar
|
||||
nachgeladen. Bildgenerierung und Whisper bleiben im Basissystem deaktiviert.
|
||||
Open WebUI spricht ausschließlich den Router an. Dieser reicht TTS intern an
|
||||
das TTS-Gateway weiter. Das Gateway nutzt primär XTTS-v2 mit der Stimme
|
||||
`Annmarie Nele` auf der RTX 3060. Deutsche Texte werden an bekannten
|
||||
englischen IT-Begriffen segmentiert; reine englische Texte laufen vollständig
|
||||
mit `language=en`. Da der offizielle XTTS-Streamingserver nur einen Auftrag
|
||||
gleichzeitig unterstützt, serialisiert das Gateway die Aufträge. Bei Fehler,
|
||||
Timeout oder belegter Queue übernimmt automatisch Piper auf der CPU. Kein
|
||||
TTS-Port wird veröffentlicht. Der äußere Kompatibilitätsname bleibt bewusst
|
||||
`piper/alloy`, damit persistente Open-WebUI-Einstellungen nach Updates und
|
||||
Restores gültig bleiben. Bildgenerierung und Whisper bleiben im Basissystem deaktiviert.
|
||||
Home Assistant, ARR und Unraid sind vorbereitete
|
||||
MCP-Profile: Sie werden erst gestartet, wenn die jeweilige root-only
|
||||
Secret-Datei vorhanden ist. Multimodale Bildanalyse erfolgt direkt über Qwen
|
||||
|
||||
+3
-1
@@ -12,7 +12,9 @@
|
||||
| ARR-MCP | `arr-mcp` 1.0.1 plus dokumentierter Sonarr-Patch | eigener optionaler Container | optional |
|
||||
| Unraid-MCP | lokales `runraid`-Binary | eigener optionaler Container | optional |
|
||||
| Whisper | ggml-org/whisper.cpp | Service im Router-Deploy | optional |
|
||||
| Piper TTS | Open Home Foundation, `piper-tts` 1.6.0 | eigener interner CPU-Container, Stimme `de_DE-thorsten-high` | Kern |
|
||||
| XTTS-v2 | Coqui, offizielles CUDA-12.1-Image per Digest | RTX-3060-Container, Stimme `Annmarie Nele`, CPML | Kern |
|
||||
| TTS-Gateway | `platform/docker/tts-gateway/` | interne Queue, Deutsch/Englisch-Segmentierung und Piper-Fallback | Kern |
|
||||
| Piper TTS | Open Home Foundation, `piper-tts` 1.6.0 | interner CPU-Fallback, Stimme `de_DE-thorsten-high` | Kern |
|
||||
| FLUX.2 klein | Black Forest Labs | Worker und Modellmanifest | optional |
|
||||
| LLama-GUI | separates Upstream-Projekt | nur Betriebsrolle dokumentiert | optional |
|
||||
| Glances | Distribution | nur Betriebsrolle dokumentiert | optional |
|
||||
|
||||
@@ -126,15 +126,20 @@ leitet das Bild dann direkt weiter und führt keinen Modellwechsel mehr aus.
|
||||
|
||||
### XTTS
|
||||
|
||||
Dieser Abschnitt beschreibt ausschließlich den alten Referenzhost. Im neuen
|
||||
Docker-Zielsystem ersetzt Piper (`de_DE-thorsten-high`) diesen Dienst.
|
||||
Der aktuelle Docker-Stack nutzt Coqui XTTS-v2 als primäre Sprachausgabe.
|
||||
Der isolierte Eignungs- und Ausfalltest ist in
|
||||
[`XTTS_EVALUATION_2026-08-23.md`](XTTS_EVALUATION_2026-08-23.md) dokumentiert.
|
||||
|
||||
- Modell: Coqui XTTS-v2
|
||||
- CPU-only
|
||||
- Stimme: `claribel`
|
||||
- Deutsch und Englisch
|
||||
- Port 8085, auf dem alten Host noch im LAN gebunden
|
||||
- eigenes Python-3.11-Venv
|
||||
- Modell: Coqui XTTS-v2, offizielles CUDA-12.1-Image per Digest gepinnt
|
||||
- GPU: ausschließlich RTX 3060 über ihre stabile GPU-UUID
|
||||
- Stimme: `Annmarie Nele`
|
||||
- Deutsch und Englisch; bekannte englische IT-Begriffe werden segmentiert
|
||||
- kein veröffentlichter Port, nur Docker-intern erreichbar
|
||||
- serielles TTS-Gateway vor XTTS, weil der Server nur einen Auftrag zugleich
|
||||
zuverlässig verarbeitet
|
||||
- Piper mit `de_DE-thorsten-high` bleibt als automatischer CPU-Fallback aktiv
|
||||
- OpenWebUI behält aus Kompatibilitätsgründen `model=piper` und `voice=alloy`;
|
||||
der Router leitet diese Werte an das Gateway weiter
|
||||
|
||||
## Websuche
|
||||
|
||||
|
||||
@@ -76,8 +76,11 @@ laufen, sondern alle fachlichen Funktionen geprüft wurden.
|
||||
- [ ] FLUX erzeugt Standard- und High-Bild
|
||||
- [ ] Qwen-Profil wird nach FLUX wiederhergestellt
|
||||
- [ ] Whisper transkribiert deutsche und englische Testdatei
|
||||
- [ ] Piper ist gesund und erzeugt über den Router deutsche WAV- und MP3-Ausgabe
|
||||
- [ ] Stimme und `piper-tts`-Version entsprechen der Installationskonfiguration
|
||||
- [ ] XTTS-v2 läuft ausschließlich auf der RTX 3060 und meldet `Annmarie Nele`
|
||||
- [ ] TTS-Gateway erzeugt über den Router deutsche und englische WAV-/MP3-Ausgabe
|
||||
- [ ] englische IT-Begriffe im deutschen Satz werden sprachlich segmentiert
|
||||
- [ ] gestopptes XTTS fällt ohne Router-/OpenWebUI-Neustart auf Piper zurück
|
||||
- [ ] Piper-Fallback und `piper-tts`-Version entsprechen der Installationskonfiguration
|
||||
- [ ] STT/TTS blockieren das Textmodell nicht unzulässig
|
||||
|
||||
## Phase F – Sicherheitsprüfung
|
||||
|
||||
+12
-2
@@ -84,7 +84,9 @@ OpenSSH-Dienst des Hosts.
|
||||
|
||||
Die TTS-Verbindung wird für eine frische Open-WebUI-Datenbank automatisch als
|
||||
OpenAI-kompatibler Audio-Endpunkt des Routers vorbelegt. Der Router reicht sie
|
||||
intern an Piper weiter; Port 8085 wird nicht am Host veröffentlicht. Ein
|
||||
intern an das TTS-Gateway weiter. Primär spricht XTTS-v2 mit `Annmarie Nele`
|
||||
auf der RTX 3060; bei Fehlern oder Queue-Timeout übernimmt Piper auf der CPU.
|
||||
Der Port 8085 wird nicht am Host veröffentlicht. Ein
|
||||
Restore setzt zusätzlich die vier persistenten Audiofelder gezielt neu, damit
|
||||
alte Werte wie `tts-1` oder `coral` die Compose-Vorgaben nicht überstimmen. Ein
|
||||
Ende-zu-Ende-Test ohne Ausgabe des API-Schlüssels:
|
||||
@@ -95,9 +97,17 @@ curl -fsS http://127.0.0.1:8081/v1/audio/speech \
|
||||
-H "Authorization: Bearer $ROUTER_API_KEY" \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"model":"piper","voice":"alloy","input":"Hallo von Athena.","response_format":"mp3"}' \
|
||||
-o /tmp/athena-piper-test.mp3
|
||||
-o /tmp/athena-tts-test.mp3
|
||||
```
|
||||
|
||||
Der beibehaltene API-Name `piper/alloy` ist eine Kompatibilitätsschnittstelle;
|
||||
bei gesundem XTTS stammt die Ausgabe von `Annmarie Nele`. Der interne Status
|
||||
des TTS-Gateways nennt `last_backend`, `primary_ready`, `fallback_ready` und
|
||||
die Zahl der Piper-Rückfälle. Ein Fallback-Test stoppt ausschließlich XTTS,
|
||||
erzeugt einen synthetischen Satz über denselben Router-Endpunkt und startet
|
||||
XTTS anschließend wieder. OpenWebUI und Router müssen dafür nicht geändert
|
||||
oder neu gestartet werden.
|
||||
|
||||
Zusätzlich prüfen: Standort-LAN sieht keine KI-Ports; Heimnetz erreicht beide;
|
||||
gestopptes VPN-Gateway lässt KI-Container nicht ins Internet; jeder Profilwechsel
|
||||
startet exakt einen llama-Container; Text, Tool Call, Bild und Sprachausgabe funktionieren.
|
||||
|
||||
@@ -35,7 +35,10 @@ Pflichtrollen:
|
||||
- BF16 Vision-Projektor
|
||||
- Whisper large-v3-turbo
|
||||
- FLUX.2 klein
|
||||
- Piper `piper-tts` 1.6.0 und Stimme `de_DE-thorsten-high`
|
||||
- XTTS-v2, per Digest gepinntes CUDA-12.1-Image und CPML-Akzeptanz
|
||||
- XTTS-Stimme `Annmarie Nele`, RTX-3060-UUID und persistenter Modellcache
|
||||
- internes TTS-Gateway mit Queue, Sprachsegmentierung und Piper-Fallback
|
||||
- Piper `piper-tts` 1.6.0 und Stimme `de_DE-thorsten-high` als CPU-Fallback
|
||||
|
||||
## 2. Externe Komponenten und Commits – teilweise gesichert
|
||||
|
||||
|
||||
@@ -0,0 +1,113 @@
|
||||
# XTTS-v2 GPU evaluation on Athena (2026-08-23)
|
||||
|
||||
## Purpose and safety boundary
|
||||
|
||||
This was an isolated, reversible evaluation of Coqui XTTS-v2 as a possible
|
||||
replacement for Piper. The user accepted the Coqui Public Model License for
|
||||
this private test.
|
||||
|
||||
- Official image: `ghcr.io/coqui-ai/xtts-streaming-server:latest-cuda121`
|
||||
- Pulled digest: `sha256:f7fb3b1f9d4bc88af94da1b5959d8002f1e0b003c97557164034eb8a29f01b90`
|
||||
- Test container: `mike-ai-xtts-test`
|
||||
- GPU visibility: RTX 3060 only
|
||||
- Host binding: `127.0.0.1:18105` only
|
||||
- Restart policy: `no`
|
||||
- Model cache: `/data/xtts-test/cache`
|
||||
- Piper, Open WebUI and the router were not reconfigured.
|
||||
|
||||
The official server describes itself as a demo server. In particular, it does
|
||||
not support concurrent streaming requests and is not an OpenAI-compatible
|
||||
production endpoint. A queueing/OpenAI compatibility proxy is therefore
|
||||
required before integration with Open WebUI.
|
||||
|
||||
## XTTS resource use
|
||||
|
||||
With the Medium profile already running, XTTS increased RTX 3060 use from
|
||||
about 4,471 MiB to about 6,419 MiB. XTTS therefore occupied approximately
|
||||
1,948 MiB and left about 5,492 MiB free. It did not use the RTX 5080.
|
||||
|
||||
With Ultra (256K) and XTTS loaded together:
|
||||
|
||||
| GPU | Used | Free |
|
||||
|---|---:|---:|
|
||||
| RTX 3060 12 GB | 8,669 MiB | 3,242 MiB |
|
||||
| RTX 5080 16 GB | 15,770 MiB | 89 MiB |
|
||||
|
||||
The combination loaded successfully without OOM. This confirms that XTTS fits
|
||||
even beside the largest standard text profile. The RTX 5080 must remain
|
||||
unavailable to XTTS because Ultra already fills it almost completely.
|
||||
|
||||
## Streaming measurements
|
||||
|
||||
The initial measurements used built-in female speaker `Ana Florence`. A
|
||||
subsequent five-voice German comparison selected **`Annmarie Nele`** as the
|
||||
production voice. Tests used harmless synthetic text.
|
||||
|
||||
| Test | First audio | Generation time | Produced audio | RTF |
|
||||
|---|---:|---:|---:|---:|
|
||||
| German | 0.701 s | 2.889 s | 6.965 s | 0.415 |
|
||||
| English | 0.305 s | 1.330 s | 3.989 s | 0.333 |
|
||||
| German sentence with English IT terms | 0.309 s | 2.496 s | 7.339 s | 0.340 |
|
||||
|
||||
After warm-up, audio starts after roughly 0.3 seconds and synthesis is around
|
||||
2.4 to 3 times faster than real time. Perceived Open WebUI latency also
|
||||
includes Qwen's time to finish the first sentence and proxy buffering.
|
||||
|
||||
## Effect on Qwen throughput
|
||||
|
||||
| Profile | XTTS state | Generation speed |
|
||||
|---|---|---:|
|
||||
| Medium 160K | loaded but idle | 71.92 token/s |
|
||||
| Medium 160K | actively speaking | 59.90 token/s |
|
||||
| Ultra 256K | loaded but idle | 66.89 token/s |
|
||||
| Ultra 256K | actively speaking | 55.27 token/s |
|
||||
|
||||
Active synthesis costs roughly 17% of Qwen generation speed because Qwen also
|
||||
uses the RTX 3060. The slowdown ends with the speech request. Merely keeping
|
||||
XTTS resident did not cause instability.
|
||||
|
||||
## Result and recommendation
|
||||
|
||||
XTTS-v2 is technically viable on the RTX 3060 and fits alongside every current
|
||||
profile, including Ultra 256K. It provides early streaming and substantially
|
||||
more natural multilingual speech than the current German-only Piper voice.
|
||||
|
||||
The production design keeps Piper and adds a small internal proxy that provides:
|
||||
|
||||
1. OpenAI-compatible `/v1/audio/speech` input and output.
|
||||
2. A one-request queue because the official XTTS server has no concurrency.
|
||||
3. German/English text segmentation so English product names are synthesized
|
||||
with `language=en` while surrounding German remains `language=de`.
|
||||
4. Cached speaker conditioning and a fixed allowlist of voices.
|
||||
5. Health checks, bounded timeouts and automatic fallback to Piper.
|
||||
|
||||
This gateway now lives under `platform/docker/tts-gateway/`. The externally
|
||||
visible compatibility values remain `model=piper` and `voice=alloy`; internally
|
||||
that alias selects `Annmarie Nele` whenever XTTS is healthy.
|
||||
|
||||
## Production result and rollback
|
||||
|
||||
After the isolated evaluation, the compatibility gateway was tested in three
|
||||
stages and then deployed to production:
|
||||
|
||||
1. Healthy XTTS produced valid WAV through the router-compatible endpoint.
|
||||
2. XTTS was deliberately stopped; the same endpoint returned valid Piper WAV.
|
||||
3. XTTS was restarted and automatically became the active backend again.
|
||||
|
||||
The production services are `mike-ai-xtts` and `mike-ai-tts-gateway`, both
|
||||
Docker-internal. Piper remained healthy throughout. OpenWebUI required no
|
||||
configuration or database change. The router's public compatibility values
|
||||
remain `model=piper` and `voice=alloy`.
|
||||
|
||||
The initial Compose GPU declaration exposed both NVIDIA cards and caused XTTS
|
||||
to select the nearly full RTX 5080. This was caught before the router switch.
|
||||
The final declaration uses a Docker device reservation with the stable RTX
|
||||
3060 UUID; inspecting the container must show exactly that UUID in
|
||||
`DeviceRequests`.
|
||||
|
||||
The verified pre-deployment state is backed up below
|
||||
`/data/backups/mike-ai/20260823-xtts-production`. The reusable rollback helper
|
||||
is `platform/scripts/rollback-tts-production.sh`; it restores the saved Compose
|
||||
and environment files, recreates the old Piper-connected router and removes
|
||||
only XTTS and its gateway. Both VPN and university-network SSH paths were
|
||||
verified after deployment.
|
||||
Reference in New Issue
Block a user