Add X-VC voice conversion mode
This commit is contained in:
@@ -0,0 +1,24 @@
|
||||
# X-VC voice conversion on Athena
|
||||
|
||||
Isolated quality gate for `chenxie95/X-VC`, using the official inference path
|
||||
from commit `49df8c591eafc48b096e466d96f9839f9c0dd739`. The UI is adapted from the
|
||||
public Hugging Face Space at commit
|
||||
`d761cd6421e85376b2656dfefd8471d7f35a42be` and runs locally without
|
||||
ZeroGPU.
|
||||
|
||||
- Private URL: `http://192.168.1.212:8009`
|
||||
- Source clip: speech content and timing to preserve
|
||||
- Reference clip: target speaker identity
|
||||
- Output: 16 kHz PCM WAV
|
||||
- GPU: RTX 5080 only
|
||||
- Persistent cache: `/data/voice/xvc/huggingface`
|
||||
- Code and model license: MIT
|
||||
|
||||
The semantic tokenizer documents Chinese and English. German is therefore a
|
||||
quality gate, not an assumed supported language. Keep OmniVoice installed: it
|
||||
does text-to-speech cloning, while X-VC tests true audio-to-audio conversion.
|
||||
|
||||
Technical acceptance on 9 September 2026 used the repository's source and
|
||||
target examples: 5.20 seconds were converted in 1.07 seconds (RTF 0.21). The
|
||||
result was valid mono PCM WAV at 16 kHz, and the loaded process occupied about
|
||||
2.9 GiB on the RTX 5080. German listening quality remains open.
|
||||
Reference in New Issue
Block a user