114 lines
5.0 KiB
Markdown
114 lines
5.0 KiB
Markdown
# XTTS-v2 GPU evaluation on Athena (2026-08-23)
|
|
|
|
## Purpose and safety boundary
|
|
|
|
This was an isolated, reversible evaluation of Coqui XTTS-v2 as a possible
|
|
replacement for Piper. The user accepted the Coqui Public Model License for
|
|
this private test.
|
|
|
|
- Official image: `ghcr.io/coqui-ai/xtts-streaming-server:latest-cuda121`
|
|
- Pulled digest: `sha256:f7fb3b1f9d4bc88af94da1b5959d8002f1e0b003c97557164034eb8a29f01b90`
|
|
- Test container: `mike-ai-xtts-test`
|
|
- GPU visibility: RTX 3060 only
|
|
- Host binding: `127.0.0.1:18105` only
|
|
- Restart policy: `no`
|
|
- Model cache: `/data/xtts-test/cache`
|
|
- Piper, Open WebUI and the router were not reconfigured.
|
|
|
|
The official server describes itself as a demo server. In particular, it does
|
|
not support concurrent streaming requests and is not an OpenAI-compatible
|
|
production endpoint. A queueing/OpenAI compatibility proxy is therefore
|
|
required before integration with Open WebUI.
|
|
|
|
## XTTS resource use
|
|
|
|
With the Medium profile already running, XTTS increased RTX 3060 use from
|
|
about 4,471 MiB to about 6,419 MiB. XTTS therefore occupied approximately
|
|
1,948 MiB and left about 5,492 MiB free. It did not use the RTX 5080.
|
|
|
|
With Ultra (256K) and XTTS loaded together:
|
|
|
|
| GPU | Used | Free |
|
|
|---|---:|---:|
|
|
| RTX 3060 12 GB | 8,669 MiB | 3,242 MiB |
|
|
| RTX 5080 16 GB | 15,770 MiB | 89 MiB |
|
|
|
|
The combination loaded successfully without OOM. This confirms that XTTS fits
|
|
even beside the largest standard text profile. The RTX 5080 must remain
|
|
unavailable to XTTS because Ultra already fills it almost completely.
|
|
|
|
## Streaming measurements
|
|
|
|
The initial measurements used built-in female speaker `Ana Florence`. A
|
|
subsequent five-voice German comparison selected **`Annmarie Nele`** as the
|
|
production voice. Tests used harmless synthetic text.
|
|
|
|
| Test | First audio | Generation time | Produced audio | RTF |
|
|
|---|---:|---:|---:|---:|
|
|
| German | 0.701 s | 2.889 s | 6.965 s | 0.415 |
|
|
| English | 0.305 s | 1.330 s | 3.989 s | 0.333 |
|
|
| German sentence with English IT terms | 0.309 s | 2.496 s | 7.339 s | 0.340 |
|
|
|
|
After warm-up, audio starts after roughly 0.3 seconds and synthesis is around
|
|
2.4 to 3 times faster than real time. Perceived Open WebUI latency also
|
|
includes Qwen's time to finish the first sentence and proxy buffering.
|
|
|
|
## Effect on Qwen throughput
|
|
|
|
| Profile | XTTS state | Generation speed |
|
|
|---|---|---:|
|
|
| Medium 160K | loaded but idle | 71.92 token/s |
|
|
| Medium 160K | actively speaking | 59.90 token/s |
|
|
| Ultra 256K | loaded but idle | 66.89 token/s |
|
|
| Ultra 256K | actively speaking | 55.27 token/s |
|
|
|
|
Active synthesis costs roughly 17% of Qwen generation speed because Qwen also
|
|
uses the RTX 3060. The slowdown ends with the speech request. Merely keeping
|
|
XTTS resident did not cause instability.
|
|
|
|
## Result and recommendation
|
|
|
|
XTTS-v2 is technically viable on the RTX 3060 and fits alongside every current
|
|
profile, including Ultra 256K. It provides early streaming and substantially
|
|
more natural multilingual speech than the current German-only Piper voice.
|
|
|
|
The production design keeps Piper and adds a small internal proxy that provides:
|
|
|
|
1. OpenAI-compatible `/v1/audio/speech` input and output.
|
|
2. A one-request queue because the official XTTS server has no concurrency.
|
|
3. German/English text segmentation so English product names are synthesized
|
|
with `language=en` while surrounding German remains `language=de`.
|
|
4. Cached speaker conditioning and a fixed allowlist of voices.
|
|
5. Health checks, bounded timeouts and automatic fallback to Piper.
|
|
|
|
This gateway now lives under `platform/docker/tts-gateway/`. The externally
|
|
visible compatibility values remain `model=piper` and `voice=alloy`; internally
|
|
that alias selects `Annmarie Nele` whenever XTTS is healthy.
|
|
|
|
## Production result and rollback
|
|
|
|
After the isolated evaluation, the compatibility gateway was tested in three
|
|
stages and then deployed to production:
|
|
|
|
1. Healthy XTTS produced valid WAV through the router-compatible endpoint.
|
|
2. XTTS was deliberately stopped; the same endpoint returned valid Piper WAV.
|
|
3. XTTS was restarted and automatically became the active backend again.
|
|
|
|
The production services are `mike-ai-xtts` and `mike-ai-tts-gateway`, both
|
|
Docker-internal. Piper remained healthy throughout. OpenWebUI required no
|
|
configuration or database change. The router's public compatibility values
|
|
remain `model=piper` and `voice=alloy`.
|
|
|
|
The initial Compose GPU declaration exposed both NVIDIA cards and caused XTTS
|
|
to select the nearly full RTX 5080. This was caught before the router switch.
|
|
The final declaration uses a Docker device reservation with the stable RTX
|
|
3060 UUID; inspecting the container must show exactly that UUID in
|
|
`DeviceRequests`.
|
|
|
|
The verified pre-deployment state is backed up below
|
|
`/data/backups/mike-ai/20260823-xtts-production`. The reusable rollback helper
|
|
is `platform/scripts/rollback-tts-production.sh`; it restores the saved Compose
|
|
and environment files, recreates the old Piper-connected router and removes
|
|
only XTTS and its gateway. Both VPN and university-network SSH paths were
|
|
verified after deployment.
|