Add X-VC voice conversion mode

This commit is contained in:
Mikei386
2026-09-09 16:14:10 +02:00
parent 17f1a08d7d
commit 87a2ae5704
15 changed files with 544 additions and 28 deletions
@@ -0,0 +1,24 @@
# X-VC voice conversion on Athena
Isolated quality gate for `chenxie95/X-VC`, using the official inference path
from commit `49df8c591eafc48b096e466d96f9839f9c0dd739`. The UI is adapted from the
public Hugging Face Space at commit
`d761cd6421e85376b2656dfefd8471d7f35a42be` and runs locally without
ZeroGPU.
- Private URL: `http://192.168.1.212:8009`
- Source clip: speech content and timing to preserve
- Reference clip: target speaker identity
- Output: 16 kHz PCM WAV
- GPU: RTX 5080 only
- Persistent cache: `/data/voice/xvc/huggingface`
- Code and model license: MIT
The semantic tokenizer documents Chinese and English. German is therefore a
quality gate, not an assumed supported language. Keep OmniVoice installed: it
does text-to-speech cloning, while X-VC tests true audio-to-audio conversion.
Technical acceptance on 9 September 2026 used the repository's source and
target examples: 5.20 seconds were converted in 1.07 seconds (RTF 0.21). The
result was valid mono PCM WAV at 16 kHz, and the loaded process occupied about
2.9 GiB on the RTX 5080. German listening quality remains open.