31 lines
1.5 KiB
Markdown
31 lines
1.5 KiB
Markdown
# X-VC voice conversion on Athena
|
|
|
|
Isolated quality gate for `chenxie95/X-VC`, using the official inference path
|
|
from commit `49df8c591eafc48b096e466d96f9839f9c0dd739`. The UI is adapted from the
|
|
public Hugging Face Space at commit
|
|
`d761cd6421e85376b2656dfefd8471d7f35a42be` and runs locally without
|
|
ZeroGPU.
|
|
|
|
- Private URL: `http://192.168.1.212:8009`
|
|
- Source clip: speech content and timing to preserve
|
|
- Reference clip: target speaker identity
|
|
- Output: native 16 kHz PCM WAV plus optional Resemble-Enhance restoration at 44.1 kHz
|
|
- GPU: RTX 5080 only
|
|
- Persistent cache: `/data/voice/xvc/huggingface`
|
|
- Code and model license: MIT
|
|
|
|
The semantic tokenizer documents Chinese and English. German is therefore a
|
|
quality gate, not an assumed supported language. Keep OmniVoice installed: it
|
|
does text-to-speech cloning, while X-VC tests true audio-to-audio conversion.
|
|
|
|
Technical acceptance on 9 September 2026 used the repository's source and
|
|
target examples: 5.20 seconds were converted in 1.07 seconds (RTF 0.21). The
|
|
result was valid mono PCM WAV at 16 kHz, and the loaded process occupied about
|
|
2.9 GiB on the RTX 5080. German listening quality remains open.
|
|
|
|
The optional high-quality path uses `resemble-enhance` 0.0.1 with model
|
|
revision `4e3510ce4a8391159f665903544c5150bee7b2cb`. It does not change X-VC's
|
|
native 16-kHz architecture. Instead, it reconstructs missing speech bandwidth
|
|
after conversion and writes a second 44.1-kHz WAV. The UI always retains the
|
|
native output for an honest A/B comparison.
|