25 lines
1.1 KiB
Markdown
25 lines
1.1 KiB
Markdown
# X-VC voice conversion on Athena
|
|
|
|
Isolated quality gate for `chenxie95/X-VC`, using the official inference path
|
|
from commit `49df8c591eafc48b096e466d96f9839f9c0dd739`. The UI is adapted from the
|
|
public Hugging Face Space at commit
|
|
`d761cd6421e85376b2656dfefd8471d7f35a42be` and runs locally without
|
|
ZeroGPU.
|
|
|
|
- Private URL: `http://192.168.1.212:8009`
|
|
- Source clip: speech content and timing to preserve
|
|
- Reference clip: target speaker identity
|
|
- Output: 16 kHz PCM WAV
|
|
- GPU: RTX 5080 only
|
|
- Persistent cache: `/data/voice/xvc/huggingface`
|
|
- Code and model license: MIT
|
|
|
|
The semantic tokenizer documents Chinese and English. German is therefore a
|
|
quality gate, not an assumed supported language. Keep OmniVoice installed: it
|
|
does text-to-speech cloning, while X-VC tests true audio-to-audio conversion.
|
|
|
|
Technical acceptance on 9 September 2026 used the repository's source and
|
|
target examples: 5.20 seconds were converted in 1.07 seconds (RTF 0.21). The
|
|
result was valid mono PCM WAV at 16 kHz, and the loaded process occupied about
|
|
2.9 GiB on the RTX 5080. German listening quality remains open.
|