5.0 KiB
XTTS-v2 GPU evaluation on Athena (2026-08-23)
Purpose and safety boundary
This was an isolated, reversible evaluation of Coqui XTTS-v2 as a possible replacement for Piper. The user accepted the Coqui Public Model License for this private test.
- Official image:
ghcr.io/coqui-ai/xtts-streaming-server:latest-cuda121 - Pulled digest:
sha256:f7fb3b1f9d4bc88af94da1b5959d8002f1e0b003c97557164034eb8a29f01b90 - Test container:
mike-ai-xtts-test - GPU visibility: RTX 3060 only
- Host binding:
127.0.0.1:18105only - Restart policy:
no - Model cache:
/data/xtts-test/cache - Piper, Open WebUI and the router were not reconfigured.
The official server describes itself as a demo server. In particular, it does not support concurrent streaming requests and is not an OpenAI-compatible production endpoint. A queueing/OpenAI compatibility proxy is therefore required before integration with Open WebUI.
XTTS resource use
With the Medium profile already running, XTTS increased RTX 3060 use from about 4,471 MiB to about 6,419 MiB. XTTS therefore occupied approximately 1,948 MiB and left about 5,492 MiB free. It did not use the RTX 5080.
With Ultra (256K) and XTTS loaded together:
| GPU | Used | Free |
|---|---|---|
| RTX 3060 12 GB | 8,669 MiB | 3,242 MiB |
| RTX 5080 16 GB | 15,770 MiB | 89 MiB |
The combination loaded successfully without OOM. This confirms that XTTS fits even beside the largest standard text profile. The RTX 5080 must remain unavailable to XTTS because Ultra already fills it almost completely.
Streaming measurements
The initial measurements used built-in female speaker Ana Florence. A
subsequent five-voice German comparison selected Annmarie Nele as the
production voice. Tests used harmless synthetic text.
| Test | First audio | Generation time | Produced audio | RTF |
|---|---|---|---|---|
| German | 0.701 s | 2.889 s | 6.965 s | 0.415 |
| English | 0.305 s | 1.330 s | 3.989 s | 0.333 |
| German sentence with English IT terms | 0.309 s | 2.496 s | 7.339 s | 0.340 |
After warm-up, audio starts after roughly 0.3 seconds and synthesis is around 2.4 to 3 times faster than real time. Perceived Open WebUI latency also includes Qwen's time to finish the first sentence and proxy buffering.
Effect on Qwen throughput
| Profile | XTTS state | Generation speed |
|---|---|---|
| Medium 160K | loaded but idle | 71.92 token/s |
| Medium 160K | actively speaking | 59.90 token/s |
| Ultra 256K | loaded but idle | 66.89 token/s |
| Ultra 256K | actively speaking | 55.27 token/s |
Active synthesis costs roughly 17% of Qwen generation speed because Qwen also uses the RTX 3060. The slowdown ends with the speech request. Merely keeping XTTS resident did not cause instability.
Result and recommendation
XTTS-v2 is technically viable on the RTX 3060 and fits alongside every current profile, including Ultra 256K. It provides early streaming and substantially more natural multilingual speech than the current German-only Piper voice.
The production design keeps Piper and adds a small internal proxy that provides:
- OpenAI-compatible
/v1/audio/speechinput and output. - A one-request queue because the official XTTS server has no concurrency.
- German/English text segmentation so English product names are synthesized
with
language=enwhile surrounding German remainslanguage=de. - Cached speaker conditioning and a fixed allowlist of voices.
- Health checks, bounded timeouts and automatic fallback to Piper.
This gateway now lives under platform/docker/tts-gateway/. The externally
visible compatibility values remain model=piper and voice=alloy; internally
that alias selects Annmarie Nele whenever XTTS is healthy.
Production result and rollback
After the isolated evaluation, the compatibility gateway was tested in three stages and then deployed to production:
- Healthy XTTS produced valid WAV through the router-compatible endpoint.
- XTTS was deliberately stopped; the same endpoint returned valid Piper WAV.
- XTTS was restarted and automatically became the active backend again.
The production services are mike-ai-xtts and mike-ai-tts-gateway, both
Docker-internal. Piper remained healthy throughout. OpenWebUI required no
configuration or database change. The router's public compatibility values
remain model=piper and voice=alloy.
The initial Compose GPU declaration exposed both NVIDIA cards and caused XTTS
to select the nearly full RTX 5080. This was caught before the router switch.
The final declaration uses a Docker device reservation with the stable RTX
3060 UUID; inspecting the container must show exactly that UUID in
DeviceRequests.
The verified pre-deployment state is backed up below
/data/backups/mike-ai/20260823-xtts-production. The reusable rollback helper
is platform/scripts/rollback-tts-production.sh; it restores the saved Compose
and environment files, recreates the old Piper-connected router and removes
only XTTS and its gateway. Both VPN and university-network SSH paths were
verified after deployment.