Files
AI-Profile-Router/experiments/xvc-voice-conversion

X-VC voice conversion on Athena

Isolated quality gate for chenxie95/X-VC, using the official inference path from commit 49df8c591eafc48b096e466d96f9839f9c0dd739. The UI is adapted from the public Hugging Face Space at commit d761cd6421e85376b2656dfefd8471d7f35a42be and runs locally without ZeroGPU.

  • Private URL: http://192.168.1.212:8009
  • Source clip: speech content and timing to preserve
  • Reference clip: target speaker identity
  • Output: native 16 kHz PCM WAV plus optional Resemble-Enhance restoration at 44.1 kHz
  • GPU: RTX 5080 only
  • Persistent cache: /data/voice/xvc/huggingface
  • Code and model license: MIT

The semantic tokenizer documents Chinese and English. German is therefore a quality gate, not an assumed supported language. Keep OmniVoice installed: it does text-to-speech cloning, while X-VC tests true audio-to-audio conversion.

Technical acceptance on 9 September 2026 used the repository's source and target examples: 5.20 seconds were converted in 1.07 seconds (RTF 0.21). The result was valid mono PCM WAV at 16 kHz, and the loaded process occupied about 2.9 GiB on the RTX 5080. German listening quality remains open.

The optional high-quality path uses resemble-enhance 0.0.1 with model revision 4e3510ce4a8391159f665903544c5150bee7b2cb. It does not change X-VC's native 16-kHz architecture. Instead, it reconstructs missing speech bandwidth after conversion and writes a second 44.1-kHz WAV. The UI always retains the native output for an honest A/B comparison.