Add XTTS primary voice with Piper fallback
This commit is contained in:
@@ -0,0 +1,113 @@
|
||||
# XTTS-v2 GPU evaluation on Athena (2026-08-23)
|
||||
|
||||
## Purpose and safety boundary
|
||||
|
||||
This was an isolated, reversible evaluation of Coqui XTTS-v2 as a possible
|
||||
replacement for Piper. The user accepted the Coqui Public Model License for
|
||||
this private test.
|
||||
|
||||
- Official image: `ghcr.io/coqui-ai/xtts-streaming-server:latest-cuda121`
|
||||
- Pulled digest: `sha256:f7fb3b1f9d4bc88af94da1b5959d8002f1e0b003c97557164034eb8a29f01b90`
|
||||
- Test container: `mike-ai-xtts-test`
|
||||
- GPU visibility: RTX 3060 only
|
||||
- Host binding: `127.0.0.1:18105` only
|
||||
- Restart policy: `no`
|
||||
- Model cache: `/data/xtts-test/cache`
|
||||
- Piper, Open WebUI and the router were not reconfigured.
|
||||
|
||||
The official server describes itself as a demo server. In particular, it does
|
||||
not support concurrent streaming requests and is not an OpenAI-compatible
|
||||
production endpoint. A queueing/OpenAI compatibility proxy is therefore
|
||||
required before integration with Open WebUI.
|
||||
|
||||
## XTTS resource use
|
||||
|
||||
With the Medium profile already running, XTTS increased RTX 3060 use from
|
||||
about 4,471 MiB to about 6,419 MiB. XTTS therefore occupied approximately
|
||||
1,948 MiB and left about 5,492 MiB free. It did not use the RTX 5080.
|
||||
|
||||
With Ultra (256K) and XTTS loaded together:
|
||||
|
||||
| GPU | Used | Free |
|
||||
|---|---:|---:|
|
||||
| RTX 3060 12 GB | 8,669 MiB | 3,242 MiB |
|
||||
| RTX 5080 16 GB | 15,770 MiB | 89 MiB |
|
||||
|
||||
The combination loaded successfully without OOM. This confirms that XTTS fits
|
||||
even beside the largest standard text profile. The RTX 5080 must remain
|
||||
unavailable to XTTS because Ultra already fills it almost completely.
|
||||
|
||||
## Streaming measurements
|
||||
|
||||
The initial measurements used built-in female speaker `Ana Florence`. A
|
||||
subsequent five-voice German comparison selected **`Annmarie Nele`** as the
|
||||
production voice. Tests used harmless synthetic text.
|
||||
|
||||
| Test | First audio | Generation time | Produced audio | RTF |
|
||||
|---|---:|---:|---:|---:|
|
||||
| German | 0.701 s | 2.889 s | 6.965 s | 0.415 |
|
||||
| English | 0.305 s | 1.330 s | 3.989 s | 0.333 |
|
||||
| German sentence with English IT terms | 0.309 s | 2.496 s | 7.339 s | 0.340 |
|
||||
|
||||
After warm-up, audio starts after roughly 0.3 seconds and synthesis is around
|
||||
2.4 to 3 times faster than real time. Perceived Open WebUI latency also
|
||||
includes Qwen's time to finish the first sentence and proxy buffering.
|
||||
|
||||
## Effect on Qwen throughput
|
||||
|
||||
| Profile | XTTS state | Generation speed |
|
||||
|---|---|---:|
|
||||
| Medium 160K | loaded but idle | 71.92 token/s |
|
||||
| Medium 160K | actively speaking | 59.90 token/s |
|
||||
| Ultra 256K | loaded but idle | 66.89 token/s |
|
||||
| Ultra 256K | actively speaking | 55.27 token/s |
|
||||
|
||||
Active synthesis costs roughly 17% of Qwen generation speed because Qwen also
|
||||
uses the RTX 3060. The slowdown ends with the speech request. Merely keeping
|
||||
XTTS resident did not cause instability.
|
||||
|
||||
## Result and recommendation
|
||||
|
||||
XTTS-v2 is technically viable on the RTX 3060 and fits alongside every current
|
||||
profile, including Ultra 256K. It provides early streaming and substantially
|
||||
more natural multilingual speech than the current German-only Piper voice.
|
||||
|
||||
The production design keeps Piper and adds a small internal proxy that provides:
|
||||
|
||||
1. OpenAI-compatible `/v1/audio/speech` input and output.
|
||||
2. A one-request queue because the official XTTS server has no concurrency.
|
||||
3. German/English text segmentation so English product names are synthesized
|
||||
with `language=en` while surrounding German remains `language=de`.
|
||||
4. Cached speaker conditioning and a fixed allowlist of voices.
|
||||
5. Health checks, bounded timeouts and automatic fallback to Piper.
|
||||
|
||||
This gateway now lives under `platform/docker/tts-gateway/`. The externally
|
||||
visible compatibility values remain `model=piper` and `voice=alloy`; internally
|
||||
that alias selects `Annmarie Nele` whenever XTTS is healthy.
|
||||
|
||||
## Production result and rollback
|
||||
|
||||
After the isolated evaluation, the compatibility gateway was tested in three
|
||||
stages and then deployed to production:
|
||||
|
||||
1. Healthy XTTS produced valid WAV through the router-compatible endpoint.
|
||||
2. XTTS was deliberately stopped; the same endpoint returned valid Piper WAV.
|
||||
3. XTTS was restarted and automatically became the active backend again.
|
||||
|
||||
The production services are `mike-ai-xtts` and `mike-ai-tts-gateway`, both
|
||||
Docker-internal. Piper remained healthy throughout. OpenWebUI required no
|
||||
configuration or database change. The router's public compatibility values
|
||||
remain `model=piper` and `voice=alloy`.
|
||||
|
||||
The initial Compose GPU declaration exposed both NVIDIA cards and caused XTTS
|
||||
to select the nearly full RTX 5080. This was caught before the router switch.
|
||||
The final declaration uses a Docker device reservation with the stable RTX
|
||||
3060 UUID; inspecting the container must show exactly that UUID in
|
||||
`DeviceRequests`.
|
||||
|
||||
The verified pre-deployment state is backed up below
|
||||
`/data/backups/mike-ai/20260823-xtts-production`. The reusable rollback helper
|
||||
is `platform/scripts/rollback-tts-production.sh`; it restores the saved Compose
|
||||
and environment files, recreates the old Piper-connected router and removes
|
||||
only XTTS and its gateway. Both VPN and university-network SSH paths were
|
||||
verified after deployment.
|
||||
Reference in New Issue
Block a user