YuE2 3B isolated quality test
Prepared, non-starting evaluation of m-a-p/YuE2-3B with the standard
m-a-p/YuE2-Vae listening decoder. The source is pinned to the official
yue2-v0.1.6 commit 9c6c4b349be978b06a9d0d958471a07a6cdeff4d.
Preparation on Athena is complete. The model and VAE files were checked
against their published weights_manifest.json SHA-256 values. The Docker
image is built, but no YuE2 container has been created or started.
Safety and isolation
- This experiment is not part of the profile controller or dashboard.
- The Compose service uses the
manualprofile, has no restart policy and cannot start through an ordinarydocker compose up. - Only the RTX 5080 is exposed to the container.
- Building and downloading do not load the model or use a GPU.
- Do not start it while another Athena GPU job is active.
Persistent files
/data/models/yue2/
├── YuE2-3B/
└── YuE2-Vae/
/data/music/yue2/
The initial control request is a true empty-lyrics instrumental request. No
invented [Instrumental] lyrics marker is used.
Community WebUI
The community-ui profile runs Ladypoly's YuE2 WebUI at the pinned commit
8fc05609bde5dcd345d7c6d57fff3da0839164c5. Only the WebUI layer is copied
from that repository. Athena continues to use the verified official YuE2
0.1.6 runtime and the existing model cache; the WebUI fork's bundled model
code and Windows installers are not used.
The UI is bound only to Athena's localhost on port 8014 and stores complete
takes below /data/music/yue2. Optional Windows-only installers for llama.cpp
and stable-diffusion.cpp are not part of the Athena setup. Manual composition,
generation, result playback, score editing and the take library work without
those optional components.
Uploaded-audio remix (SheetSage2)
The Remix a take drawer also accepts WAV, FLAC, MP3, M4A, OGG, Opus and AAC uploads. SheetSage2 transcribes the recording into an editable ABC melody and chord plan, then YuE2 renders that structure in a newly selected style. It does not preserve the original samples, singer or production verbatim.
SheetSage2 is kept in /opt/yue2/.venv-sheetsage2, while its persistent model
files live below /data/models/yue2. The environment intentionally reuses the
image's PyTorch 2.10/CUDA 12.8 runtime: the upstream cu126 recipe is not
Blackwell-capable. Required persistent directories are:
/data/models/yue2/
├── SheetSage2/
└── MERT-v2-FullSong/
The local SheetSage2/config.json must point base_model_name_or_path to
/opt/yue2/models/MERT-v2-FullSong, allowing the complete transcription path
to run offline. Analysis and generation share the RTX 5080 and therefore run
sequentially; the UI parks YuE2 before starting SheetSage2.
The small integration patch under community-webui-patches/ fixes the
community release's missing refreshArt() function on Linux and permits the
native YuE2 empty-lyrics request for true instrumentals. It deliberately does
not modify the model runtime.
docker compose --profile community-ui up -d yue2-ui
The original small German playground remains available as a stopped fallback
on localhost port 8016 through the playground-fallback profile. Do not run
both frontends concurrently because both can submit work to the same GPU.
Manual test (only after GPU availability was checked)
From /opt/mike-ai/yue2-3b on Athena:
docker compose --profile manual run --rm yue2-test generate \
--offline \
--device cuda:0 \
--budget 16 \
--request /workspace/requests/instrumental-synthwave.json \
--output /workspace/runs
Start with the official unquantized BF16 path. If and only if this fails from
VRAM pressure, repeat with --quantization fp8 --offload-ar; keep the outputs
separate because that is a different inference configuration.
YuE2 is newly released and officially specifies a 24-GB BF16 GPU. Readiness of this image and the downloaded weights is not evidence that the 16-GB RTX 5080 run will fit or that its audio quality is acceptable.