Add XTTS primary voice with Piper fallback

This commit is contained in:
Mikei386
2026-08-23 12:22:59 +02:00
parent 16a6288a4d
commit f90fc93f9c
18 changed files with 715 additions and 28 deletions
+113
View File
@@ -0,0 +1,113 @@
# XTTS-v2 GPU evaluation on Athena (2026-08-23)
## Purpose and safety boundary
This was an isolated, reversible evaluation of Coqui XTTS-v2 as a possible
replacement for Piper. The user accepted the Coqui Public Model License for
this private test.
- Official image: `ghcr.io/coqui-ai/xtts-streaming-server:latest-cuda121`
- Pulled digest: `sha256:f7fb3b1f9d4bc88af94da1b5959d8002f1e0b003c97557164034eb8a29f01b90`
- Test container: `mike-ai-xtts-test`
- GPU visibility: RTX 3060 only
- Host binding: `127.0.0.1:18105` only
- Restart policy: `no`
- Model cache: `/data/xtts-test/cache`
- Piper, Open WebUI and the router were not reconfigured.
The official server describes itself as a demo server. In particular, it does
not support concurrent streaming requests and is not an OpenAI-compatible
production endpoint. A queueing/OpenAI compatibility proxy is therefore
required before integration with Open WebUI.
## XTTS resource use
With the Medium profile already running, XTTS increased RTX 3060 use from
about 4,471 MiB to about 6,419 MiB. XTTS therefore occupied approximately
1,948 MiB and left about 5,492 MiB free. It did not use the RTX 5080.
With Ultra (256K) and XTTS loaded together:
| GPU | Used | Free |
|---|---:|---:|
| RTX 3060 12 GB | 8,669 MiB | 3,242 MiB |
| RTX 5080 16 GB | 15,770 MiB | 89 MiB |
The combination loaded successfully without OOM. This confirms that XTTS fits
even beside the largest standard text profile. The RTX 5080 must remain
unavailable to XTTS because Ultra already fills it almost completely.
## Streaming measurements
The initial measurements used built-in female speaker `Ana Florence`. A
subsequent five-voice German comparison selected **`Annmarie Nele`** as the
production voice. Tests used harmless synthetic text.
| Test | First audio | Generation time | Produced audio | RTF |
|---|---:|---:|---:|---:|
| German | 0.701 s | 2.889 s | 6.965 s | 0.415 |
| English | 0.305 s | 1.330 s | 3.989 s | 0.333 |
| German sentence with English IT terms | 0.309 s | 2.496 s | 7.339 s | 0.340 |
After warm-up, audio starts after roughly 0.3 seconds and synthesis is around
2.4 to 3 times faster than real time. Perceived Open WebUI latency also
includes Qwen's time to finish the first sentence and proxy buffering.
## Effect on Qwen throughput
| Profile | XTTS state | Generation speed |
|---|---|---:|
| Medium 160K | loaded but idle | 71.92 token/s |
| Medium 160K | actively speaking | 59.90 token/s |
| Ultra 256K | loaded but idle | 66.89 token/s |
| Ultra 256K | actively speaking | 55.27 token/s |
Active synthesis costs roughly 17% of Qwen generation speed because Qwen also
uses the RTX 3060. The slowdown ends with the speech request. Merely keeping
XTTS resident did not cause instability.
## Result and recommendation
XTTS-v2 is technically viable on the RTX 3060 and fits alongside every current
profile, including Ultra 256K. It provides early streaming and substantially
more natural multilingual speech than the current German-only Piper voice.
The production design keeps Piper and adds a small internal proxy that provides:
1. OpenAI-compatible `/v1/audio/speech` input and output.
2. A one-request queue because the official XTTS server has no concurrency.
3. German/English text segmentation so English product names are synthesized
with `language=en` while surrounding German remains `language=de`.
4. Cached speaker conditioning and a fixed allowlist of voices.
5. Health checks, bounded timeouts and automatic fallback to Piper.
This gateway now lives under `platform/docker/tts-gateway/`. The externally
visible compatibility values remain `model=piper` and `voice=alloy`; internally
that alias selects `Annmarie Nele` whenever XTTS is healthy.
## Production result and rollback
After the isolated evaluation, the compatibility gateway was tested in three
stages and then deployed to production:
1. Healthy XTTS produced valid WAV through the router-compatible endpoint.
2. XTTS was deliberately stopped; the same endpoint returned valid Piper WAV.
3. XTTS was restarted and automatically became the active backend again.
The production services are `mike-ai-xtts` and `mike-ai-tts-gateway`, both
Docker-internal. Piper remained healthy throughout. OpenWebUI required no
configuration or database change. The router's public compatibility values
remain `model=piper` and `voice=alloy`.
The initial Compose GPU declaration exposed both NVIDIA cards and caused XTTS
to select the nearly full RTX 5080. This was caught before the router switch.
The final declaration uses a Docker device reservation with the stable RTX
3060 UUID; inspecting the container must show exactly that UUID in
`DeviceRequests`.
The verified pre-deployment state is backed up below
`/data/backups/mike-ai/20260823-xtts-production`. The reusable rollback helper
is `platform/scripts/rollback-tts-production.sh`; it restores the saved Compose
and environment files, recreates the old Piper-connected router and removes
only XTTS and its gateway. Both VPN and university-network SSH paths were
verified after deployment.