Files
AI-Profile-Router/experiments/vevo2-voice-cloning

Vevo2 Voice Conversion on Athena

This directory contains Athena's private Vevo2 voice-conversion studio.

  • Code: open-mmlab/Amphion commit 26f6883110181f1dbfe95c70a7c7dbaf4de5f42a (MIT)
  • Weights: RMSnow/Vevo2 (CC BY-NC-ND 4.0)
  • Runtime: PyTorch 2.7.1 + CUDA 12.8 inherited from Athena's tested audio separator image; the historical Amphion torch 2.0/cu118 pins are not used.
  • Storage: /data/voice/vevo2; removing that directory and the test image removes all downloaded artifacts.
  • Private UI: http://192.168.1.212:8008 through WireGuard only.
  • Output: uncompressed mono WAV, 24 kHz.

The 9 September technical gate converted the official 8.6-second speech sample through the production HTTP API in 2.342 seconds. A warm service start loaded the model in 12.216 seconds, and peak CUDA allocation during conversion was 5605.6 MiB. The returned 24-kHz WAV was 412878 bytes and 8.6 seconds long. The service stores named reference voices, accepts a source clip, and returns a transient WAV download. Jobs and generated outputs are removed after delivery. It is deliberately not exposed on the university interface.

The installed Compose project is named mike-ai-voice. Rebuild or recreate it with an explicit VOICE_GPU_UUID so Docker keeps the service attached to the RTX 5080. The profile controller starts and stops the existing container; it does not rebuild it during a mode switch.

The weights are licensed CC BY-NC-ND 4.0. This deployment is for Mike's private, non-commercial use only. Do not use or expose it as a public or commercial voice-cloning service.