# YuE2 3B isolated quality test Prepared, non-starting evaluation of `m-a-p/YuE2-3B` with the standard `m-a-p/YuE2-Vae` listening decoder. The source is pinned to the official `yue2-v0.1.6` commit `9c6c4b349be978b06a9d0d958471a07a6cdeff4d`. Preparation on Athena is complete. The model and VAE files were checked against their published `weights_manifest.json` SHA-256 values. The Docker image is built, but no YuE2 container has been created or started. ## Safety and isolation - This experiment is not part of the profile controller or dashboard. - The Compose service uses the `manual` profile, has no restart policy and cannot start through an ordinary `docker compose up`. - Only the RTX 5080 is exposed to the container. - Building and downloading do not load the model or use a GPU. - Do not start it while another Athena GPU job is active. ## Persistent files ```text /data/models/yue2/ ├── YuE2-3B/ └── YuE2-Vae/ /data/music/yue2/ ``` The initial control request is a true empty-lyrics instrumental request. No invented `[Instrumental]` lyrics marker is used. ## Community WebUI The `community-ui` profile runs Ladypoly's YuE2 WebUI at the pinned commit `8fc05609bde5dcd345d7c6d57fff3da0839164c5`. Only the WebUI layer is copied from that repository. Athena continues to use the verified official YuE2 `0.1.6` runtime and the existing model cache; the WebUI fork's bundled model code and Windows installers are not used. The UI is bound only to Athena's localhost on port 8014 and stores complete takes below `/data/music/yue2`. Optional Windows-only installers for llama.cpp and stable-diffusion.cpp are not part of the Athena setup. Manual composition, generation, result playback, score editing and the take library work without those optional components. ### Uploaded-audio remix (SheetSage2) The **Remix a take** drawer also accepts WAV, FLAC, MP3, M4A, OGG, Opus and AAC uploads. SheetSage2 transcribes the recording into an editable ABC melody and chord plan, then YuE2 renders that structure in a newly selected style. It does not preserve the original samples, singer or production verbatim. SheetSage2 is kept in `/opt/yue2/.venv-sheetsage2`, while its persistent model files live below `/data/models/yue2`. The environment intentionally reuses the image's PyTorch 2.10/CUDA 12.8 runtime: the upstream cu126 recipe is not Blackwell-capable. Required persistent directories are: ```text /data/models/yue2/ ├── SheetSage2/ └── MERT-v2-FullSong/ ``` The local `SheetSage2/config.json` must point `base_model_name_or_path` to `/opt/yue2/models/MERT-v2-FullSong`, allowing the complete transcription path to run offline. Analysis and generation share the RTX 5080 and therefore run sequentially; the UI parks YuE2 before starting SheetSage2. The small integration patch under `community-webui-patches/` fixes the community release's missing `refreshArt()` function on Linux and permits the native YuE2 empty-lyrics request for true instrumentals. It deliberately does not modify the model runtime. ```sh docker compose --profile community-ui up -d yue2-ui ``` The original small German playground remains available as a stopped fallback on localhost port 8016 through the `playground-fallback` profile. Do not run both frontends concurrently because both can submit work to the same GPU. ## Manual test (only after GPU availability was checked) From `/opt/mike-ai/yue2-3b` on Athena: ```sh docker compose --profile manual run --rm yue2-test generate \ --offline \ --device cuda:0 \ --budget 16 \ --request /workspace/requests/instrumental-synthwave.json \ --output /workspace/runs ``` Start with the official unquantized BF16 path. If and only if this fails from VRAM pressure, repeat with `--quantization fp8 --offload-ar`; keep the outputs separate because that is a different inference configuration. YuE2 is newly released and officially specifies a 24-GB BF16 GPU. Readiness of this image and the downloaded weights is not evidence that the 16-GB RTX 5080 run will fit or that its audio quality is acceptable.