feat: add multi-stem audio separation modes

This commit is contained in:
Mikei386 committed 2026-09-09 00:30:08 +02:00
1 parent 30203bf13b
commit e2f35517f8
6 files changed
+76 -40

No files matched your search

+1 -1
View File
@@ -124,7 +124,7 @@ Details, Installation, Prüfung und Rollback stehen in
- Athena-Dashboard: `http://192.168.1.212:8099` - Athena-Dashboard: `http://192.168.1.212:8099`
- Musikstudio, Original UI (stabil): `http://192.168.1.212:7862` - Musikstudio, Original UI (stabil): `http://192.168.1.212:7862`
- Musikstudio, Community UI (experimentell): `http://192.168.1.212:7861` - Musikstudio, Community UI (experimentell): `http://192.168.1.212:7861`
- Stimmen trennen (BS-RoFormer): `http://192.168.1.212:8007` - Spuren trennen (BS-RoFormer + Demucs, 2/4/6 Stems): `http://192.168.1.212:8007`
Der Betriebsmodus lässt sich dort direkt umschalten. In Hermes funktionieren Der Betriebsmodus lässt sich dort direkt umschalten. In Hermes funktionieren
außerdem `/athena music`, `/athena stems`, `/athena llm` und `/athena status`; Details stehen in außerdem `/athena music`, `/athena stems`, `/athena llm` und `/athena status`; Details stehen in
+8 -5
View File
@@ -4,8 +4,8 @@ Athena besitzt drei gegenseitig exklusive Betriebsmodi:
- `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt. - `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt.
- `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt. - `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt.
- `separation`: BS-RoFormer Viperx 1297 trennt Gesang und Instrumental; - `separation`: wahlweise BS-RoFormer für Gesang/Instrumental oder Demucs für
LLM, Bild, TTS und ACE-Step sind gestoppt. vier beziehungsweise sechs Spuren; LLM, Bild, TTS und ACE-Step sind gestoppt.
Die Zustandsmaschine lebt im Athena-Router. Das Dashboard und Chat-Clients wie Die Zustandsmaschine lebt im Athena-Router. Das Dashboard und Chat-Clients wie
Hermes sind nur Bedienoberflächen derselben API. Der zuletzt aktive LLM-Modus Hermes sind nur Bedienoberflächen derselben API. Der zuletzt aktive LLM-Modus
@@ -36,9 +36,12 @@ Docker-Netz `mike-ai-music` auf `http://music-worker:7860` zu.
Im Trennmodus öffnet das Dashboard die private Athena-Oberfläche unter Im Trennmodus öffnet das Dashboard die private Athena-Oberfläche unter
`http://192.168.1.212:8007`. Sie nimmt WAV, FLAC, MP3, M4A und weitere `http://192.168.1.212:8007`. Sie nimmt WAV, FLAC, MP3, M4A und weitere
übliche Formate an und liefert ein ZIP mit verlustfreien `vocals.flac` und übliche Formate an und liefert verlustfreie FLAC-Spuren als ZIP. Verfügbar sind
`instrumental.flac`. Grundlage ist `audio-separator` 0.47.0 mit die spezialisierte Gesangstrennung (`vocals`, `instrumental`), vier Spuren
`model_bs_roformer_ep_317_sdr_12.9755.ckpt`. (`vocals`, `drums`, `bass`, `other`) und experimentell sechs Spuren
(`vocals`, `drums`, `bass`, `guitar`, `piano`, `other`). Grundlage ist
`audio-separator` 0.47.0 mit BS-RoFormer Viperx 1297, `htdemucs_ft` und
`htdemucs_6s`.
Hermes benötigt dafür kein Plugin. Exakt eingegebene Steuerbefehle werden vom Hermes benötigt dafür kein Plugin. Exakt eingegebene Steuerbefehle werden vom
Router lokal beantwortet, auch wenn gerade kein LLM geladen ist: Router lokal beantwortet, auch wenn gerade kein LLM geladen ist:
@@ -19,4 +19,4 @@ ENV MODEL_FILENAME=model_bs_roformer_ep_317_sdr_12.9755.ckpt \
JOB_DIR=/data/jobs JOB_DIR=/data/jobs
EXPOSE 8080 EXPOSE 8080
CMD ["sh", "-c", "audio-separator --model_filename \"$MODEL_FILENAME\" --model_file_dir \"$MODEL_DIR\" --download_model_only && exec uvicorn app:app --host 0.0.0.0 --port 8080 --workers 1"] CMD ["sh", "-c", "for model in \"$MODEL_FILENAME\" htdemucs_ft.yaml htdemucs_6s.yaml; do audio-separator --model_filename \"$model\" --model_file_dir \"$MODEL_DIR\" --download_model_only || exit 1; done; exec uvicorn app:app --host 0.0.0.0 --port 8080 --workers 1"]
@@ -1,16 +1,23 @@
# Athena Vocal Separator # Athena Stem Separator
Exklusiver dritter Athena-Betriebsmodus für lokale Zwei-Spur-Trennung in Exklusiver dritter Athena-Betriebsmodus für lokale Zwei- bis Sechs-Spur-Trennung.
`vocals.flac` und `instrumental.flac`.
- Engine: `audio-separator` 0.47.0 (MIT) - Engine: `audio-separator` 0.47.0 (MIT)
- Modell: BS-RoFormer Viperx 1297, - **Gesang / Instrumental:** BS-RoFormer Viperx 1297,
`model_bs_roformer_ep_317_sdr_12.9755.ckpt` `model_bs_roformer_ep_317_sdr_12.9755.ckpt`; Vocal SDR 12,9,
- Modellbewertung im audio-separator-Katalog: Vocal SDR 12,9, Instrumental SDR 17,0.
Instrumental SDR 17,0 - **4 Spuren:** `htdemucs_ft.yaml` mit Gesang, Schlagzeug, Bass und Rest.
- **6 Spuren (experimentell):** `htdemucs_6s.yaml` ergänzt gemeinsame Gitarre
und Piano. Die Instrumentqualität liegt unter der spezialisierten
Zwei-Spur-Trennung.
- GPU: RTX 5080; LLM, Bildmodelle, TTS und ACE-Step sind dabei verriegelt. - GPU: RTX 5080; LLM, Bildmodelle, TTS und ACE-Step sind dabei verriegelt.
- Privat erreichbar: `http://192.168.1.212:8007/` - Privat erreichbar: `http://192.168.1.212:8007/`
Das Modell wird beim ersten Start nach `/data/models/audio-separator` Die Modelle werden beim ersten Start nach `/data/models/audio-separator`
heruntergeladen. Temporäre Jobs liegen unter `/data/audio/separation` und heruntergeladen. Temporäre Jobs liegen unter `/data/audio/separation` und
werden nach dem ZIP-Download entfernt. werden nach dem ZIP-Download entfernt. Eigene Spuren für E-/Akustikgitarre,
Synthesizer und Streicher sind bewusst noch nicht angeboten: Dafür braucht es
weitere Zielmodelle; sie aus `other` umzubenennen wäre fachlich falsch.
Die API erwartet `multipart/form-data` mit `file` und optional `mode`:
`vocals` (Standard), `four_stem` oder `six_stem`.
+42 -21
View File
@@ -9,7 +9,7 @@ import time
import zipfile import zipfile
from pathlib import Path from pathlib import Path
from fastapi import FastAPI, File, HTTPException, UploadFile from fastapi import FastAPI, File, Form, HTTPException, UploadFile
from fastapi.responses import FileResponse, HTMLResponse from fastapi.responses import FileResponse, HTMLResponse
from starlette.background import BackgroundTask from starlette.background import BackgroundTask
@@ -22,7 +22,13 @@ ALLOWED = {".wav", ".flac", ".mp3", ".m4a", ".aac", ".ogg", ".opus", ".wma"}
SEPARATION_LOCK = asyncio.Lock() SEPARATION_LOCK = asyncio.Lock()
STARTED = time.time() STARTED = time.time()
app = FastAPI(title="Athena Vocal Separator", version="1.0") MODES = {
"vocals": {"model": MODEL, "stems": ("vocals", "instrumental"), "archive": "athena-vocals-instrumental.zip", "engine": "mdxc"},
"four_stem": {"model": "htdemucs_ft.yaml", "stems": ("vocals", "drums", "bass", "other"), "archive": "athena-4-stems.zip", "engine": "demucs"},
"six_stem": {"model": "htdemucs_6s.yaml", "stems": ("vocals", "drums", "bass", "guitar", "piano", "other"), "archive": "athena-6-stems-experimental.zip", "engine": "demucs"},
}
app = FastAPI(title="Athena Stem Separator", version="2.0")
@app.get("/", response_class=HTMLResponse) @app.get("/", response_class=HTMLResponse)
@@ -32,11 +38,11 @@ def index() -> str:
@app.get("/health") @app.get("/health")
def health() -> dict: def health() -> dict:
checkpoint = MODEL_DIR / MODEL available = {name: (MODEL_DIR / mode["model"]).exists() for name, mode in MODES.items()}
return { return {
"status": "ok" if checkpoint.exists() else "starting", "status": "ok" if all(available.values()) else "starting",
"model": MODEL, "models": {name: mode["model"] for name, mode in MODES.items()},
"model_ready": checkpoint.exists(), "models_ready": available,
"busy": SEPARATION_LOCK.locked(), "busy": SEPARATION_LOCK.locked(),
"uptime_seconds": round(time.time() - STARTED, 1), "uptime_seconds": round(time.time() - STARTED, 1),
} }
@@ -46,28 +52,40 @@ def _cleanup(path: Path) -> None:
shutil.rmtree(path, ignore_errors=True) shutil.rmtree(path, ignore_errors=True)
def _run_separator(input_path: Path, output_dir: Path) -> None: def _run_separator(input_path: Path, output_dir: Path, mode: dict) -> None:
args = [ args = [
"audio-separator", str(input_path), "audio-separator", str(input_path),
"--model_filename", MODEL, "--model_filename", mode["model"],
"--model_file_dir", str(MODEL_DIR), "--model_file_dir", str(MODEL_DIR),
"--output_dir", str(output_dir), "--output_dir", str(output_dir),
"--output_format", "FLAC", "--output_format", "FLAC",
"--sample_rate", "44100", "--sample_rate", "44100",
"--use_soundfile", "--use_soundfile",
"--use_autocast", "--use_autocast",
"--mdxc_segment_size", "256",
"--mdxc_overlap", "8",
"--mdxc_batch_size", "1",
] ]
if mode["engine"] == "mdxc":
args.extend(["--mdxc_segment_size", "256", "--mdxc_overlap", "8", "--mdxc_batch_size", "1"])
else:
args.extend(["--demucs_segment_size", "40", "--demucs_shifts", "2", "--demucs_overlap", "0.25"])
completed = subprocess.run(args, capture_output=True, text=True, timeout=7200) completed = subprocess.run(args, capture_output=True, text=True, timeout=7200)
if completed.returncode: if completed.returncode:
detail = (completed.stderr or completed.stdout or "unknown error")[-4000:] detail = (completed.stderr or completed.stdout or "unknown error")[-4000:]
raise RuntimeError(detail) raise RuntimeError(detail)
def _stem_name(path: Path, expected: tuple[str, ...]) -> str | None:
lower = path.stem.lower()
for stem in sorted(expected, key=len, reverse=True):
if stem in lower:
return stem
return None
@app.post("/v1/separate") @app.post("/v1/separate")
async def separate(file: UploadFile = File(...)) -> FileResponse: async def separate(file: UploadFile = File(...), mode: str = Form("vocals")) -> FileResponse:
selected_mode = MODES.get(mode)
if selected_mode is None:
raise HTTPException(422, f"Unbekannter Trennmodus: {mode}")
suffix = Path(file.filename or "upload.wav").suffix.lower() suffix = Path(file.filename or "upload.wav").suffix.lower()
if suffix not in ALLOWED: if suffix not in ALLOWED:
raise HTTPException(415, "Dieses Audioformat wird nicht unterstützt.") raise HTTPException(415, "Dieses Audioformat wird nicht unterstützt.")
@@ -87,21 +105,24 @@ async def separate(file: UploadFile = File(...)) -> FileResponse:
raise HTTPException(413, "Datei ist größer als 1 GiB.") raise HTTPException(413, "Datei ist größer als 1 GiB.")
handle.write(chunk) handle.write(chunk)
async with SEPARATION_LOCK: async with SEPARATION_LOCK:
await asyncio.to_thread(_run_separator, input_path, output_dir) await asyncio.to_thread(_run_separator, input_path, output_dir, selected_mode)
stems = sorted(output_dir.glob("*.flac")) stems = sorted(output_dir.glob("*.flac"))
if len(stems) != 2: expected = selected_mode["stems"]
raise RuntimeError(f"Erwartet wurden zwei FLAC-Dateien, gefunden: {len(stems)}") recognized = {_stem_name(stem, expected): stem for stem in stems}
archive = job / "athena-vocals-instrumental.zip" recognized.pop(None, None)
missing = [stem for stem in expected if stem not in recognized]
if missing:
found = ", ".join(stem.name for stem in stems) or "keine"
raise RuntimeError(f"Fehlende Spuren: {', '.join(missing)}; gefunden: {found}")
archive = job / selected_mode["archive"]
with zipfile.ZipFile(archive, "w", compression=zipfile.ZIP_STORED) as bundle: with zipfile.ZipFile(archive, "w", compression=zipfile.ZIP_STORED) as bundle:
for stem in stems: for stem in expected:
lower = stem.name.lower() bundle.write(recognized[stem], f"{stem}.flac")
target = "vocals.flac" if "vocal" in lower else "instrumental.flac"
bundle.write(stem, target)
return FileResponse( return FileResponse(
archive, archive,
media_type="application/zip", media_type="application/zip",
filename="athena-vocals-instrumental.zip", filename=selected_mode["archive"],
background=BackgroundTask(_cleanup, job), background=BackgroundTask(_cleanup, job),
) )
except HTTPException: except HTTPException:
@@ -1,4 +1,9 @@
<!doctype html><html lang="de"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>Athena · Stimmen trennen</title><style> <!doctype html>
:root{color-scheme:dark;--bg:#07111c;--card:#101d2b;--line:#26384b;--cyan:#48d7f5;--mint:#63e6be;--text:#ecf5ff;--muted:#91a4b7}*{box-sizing:border-box}body{margin:0;background:radial-gradient(circle at 20% 0,#142a42 0,#07111c 42%);font:16px system-ui,sans-serif;color:var(--text);min-height:100vh;display:grid;place-items:center;padding:24px}.card{width:min(760px,100%);padding:32px;border:1px solid var(--line);border-radius:22px;background:rgba(16,29,43,.96);box-shadow:0 25px 70px #0008}.eyebrow{color:var(--cyan);font-weight:800;letter-spacing:.14em;text-transform:uppercase;font-size:12px}h1{font-size:clamp(30px,5vw,52px);margin:.3em 0 .15em}p{color:var(--muted);line-height:1.6}.drop{display:block;margin:28px 0;padding:44px 24px;border:2px dashed #3f5a72;border-radius:18px;text-align:center;cursor:pointer;transition:.2s}.drop:hover,.drop.drag{border-color:var(--cyan);background:#48d7f50b}input{display:none}.file{color:var(--mint);font-weight:700;margin-top:8px}button{width:100%;border:0;border-radius:13px;padding:15px;font-weight:800;font-size:16px;background:linear-gradient(90deg,var(--cyan),var(--mint));color:#05202a;cursor:pointer}button:disabled{opacity:.45;cursor:not-allowed}.status{min-height:28px;margin-top:18px;color:var(--muted)}.bar{height:7px;background:#07111c;border-radius:9px;overflow:hidden;margin-top:12px}.fill{height:100%;width:0;background:linear-gradient(90deg,var(--cyan),var(--mint));transition:.4s}.run .fill{width:85%;animation:pulse 1.5s infinite alternate}@keyframes pulse{to{opacity:.45}}small{display:block;color:#71879a;margin-top:20px}</style></head><body><main class="card"><div class="eyebrow">Athena Audio Lab</div><h1>Stimmen sauber trennen</h1><p>BS‑RoFormer Viperx 1297 zerlegt deinen Titel in zwei verlustfreie FLAC-Spuren: Gesang und Instrumental.</p><label class="drop" id="drop">Audio auswählen oder hier ablegen<input id="file" type="file" accept="audio/*"><div class="file" id="name">Noch keine Datei gewählt</div></label><button id="start" disabled>Vocals und Instrumental erzeugen</button><div class="status" id="status">Bereit.</div><div class="bar" id="bar"><div class="fill"></div></div><small>Die Verarbeitung läuft lokal auf Athena. Nichts wird in eine Cloud hochgeladen.</small></main><script> <html lang="de"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1">
const file=document.querySelector('#file'),drop=document.querySelector('#drop'),name=document.querySelector('#name'),start=document.querySelector('#start'),status=document.querySelector('#status'),bar=document.querySelector('#bar');let selected;function choose(f){selected=f;name.textContent=f?`${f.name} · ${(f.size/1048576).toFixed(1)} MiB`:'Noch keine Datei gewählt';start.disabled=!f}file.onchange=()=>choose(file.files[0]);drop.ondragover=e=>{e.preventDefault();drop.classList.add('drag')};drop.ondragleave=()=>drop.classList.remove('drag');drop.ondrop=e=>{e.preventDefault();drop.classList.remove('drag');choose(e.dataTransfer.files[0])};start.onclick=async()=>{start.disabled=true;bar.classList.add('run');status.textContent='Modell trennt den Titel – das kann einige Minuten dauern …';let body=new FormData();body.append('file',selected);try{let r=await fetch('/v1/separate',{method:'POST',body});if(!r.ok)throw Error((await r.json()).detail||`HTTP ${r.status}`);let blob=await r.blob(),a=document.createElement('a');a.href=URL.createObjectURL(blob);a.download='athena-vocals-instrumental.zip';a.click();setTimeout(()=>URL.revokeObjectURL(a.href),5000);status.textContent='Fertig – ZIP mit vocals.flac und instrumental.flac wurde geladen.'}catch(e){status.textContent=`Fehler: ${e.message}`}finally{bar.classList.remove('run');start.disabled=false}}; <title>Athena · Spuren trennen</title><style>
:root{color-scheme:dark;--bg:#07111c;--card:#101d2b;--line:#26384b;--cyan:#48d7f5;--mint:#63e6be;--text:#ecf5ff;--muted:#91a4b7}*{box-sizing:border-box}body{margin:0;background:radial-gradient(circle at 20% 0,#142a42 0,#07111c 42%);font:16px system-ui,sans-serif;color:var(--text);min-height:100vh;display:grid;place-items:center;padding:24px}.card{width:min(820px,100%);padding:32px;border:1px solid var(--line);border-radius:22px;background:rgba(16,29,43,.96);box-shadow:0 25px 70px #0008}.eyebrow{color:var(--cyan);font-weight:800;letter-spacing:.14em;text-transform:uppercase;font-size:12px}h1{font-size:clamp(30px,5vw,52px);margin:.3em 0 .15em}p{color:var(--muted);line-height:1.6}.modes{display:grid;grid-template-columns:repeat(3,1fr);gap:10px;margin:24px 0}.mode{display:block;border:1px solid var(--line);border-radius:14px;padding:14px;cursor:pointer}.mode:has(input:checked){border-color:var(--cyan);background:#48d7f510}.mode input{display:none}.mode b,.mode span{display:block}.mode span{color:var(--muted);font-size:13px;margin-top:5px;line-height:1.35}.drop{display:block;margin:20px 0;padding:36px 24px;border:2px dashed #3f5a72;border-radius:18px;text-align:center;cursor:pointer;transition:.2s}.drop:hover,.drop.drag{border-color:var(--cyan);background:#48d7f50b}.drop input{display:none}.file{color:var(--mint);font-weight:700;margin-top:8px}button{width:100%;border:0;border-radius:13px;padding:15px;font-weight:800;font-size:16px;background:linear-gradient(90deg,var(--cyan),var(--mint));color:#05202a;cursor:pointer}button:disabled{opacity:.45;cursor:not-allowed}.status{min-height:28px;margin-top:18px;color:var(--muted)}.bar{height:7px;background:#07111c;border-radius:9px;overflow:hidden;margin-top:12px}.fill{height:100%;width:0;background:linear-gradient(90deg,var(--cyan),var(--mint));transition:.4s}.run .fill{width:85%;animation:pulse 1.5s infinite alternate}@keyframes pulse{to{opacity:.45}}small{display:block;color:#71879a;margin-top:20px}@media(max-width:680px){.modes{grid-template-columns:1fr}}
</style></head><body><main class="card"><div class="eyebrow">Athena Audio Lab</div><h1>Instrumente und Stimmen trennen</h1><p>Wähle zwischen der besonders sauberen Gesangstrennung und zusammengehörigen Mehrspur-Modellen.</p>
<div class="modes"><label class="mode"><input type="radio" name="mode" value="vocals" checked><b>Gesang · beste Qualität</b><span>Vocals + Instrumental<br>BS‑RoFormer</span></label><label class="mode"><input type="radio" name="mode" value="four_stem"><b>4 Spuren</b><span>Gesang, Schlagzeug, Bass, Rest<br>HTDemucs FT</span></label><label class="mode"><input type="radio" name="mode" value="six_stem"><b>6 Spuren · experimentell</b><span>Zusätzlich Gitarre und Piano<br>HTDemucs 6s</span></label></div>
<label class="drop" id="drop">Audio auswählen oder hier ablegen<input id="file" type="file" accept="audio/*"><div class="file" id="name">Noch keine Datei gewählt</div></label><button id="start" disabled>Spuren erzeugen</button><div class="status" id="status">Bereit.</div><div class="bar" id="bar"><div class="fill"></div></div><small>Alles läuft lokal auf Athena. Synthesizer, Streicher sowie elektrische und akustische Gitarre separat benötigen zusätzliche Spezialmodelle und sind noch nicht freigeschaltet.</small></main><script>
const file=document.querySelector('#file'),drop=document.querySelector('#drop'),name=document.querySelector('#name'),start=document.querySelector('#start'),status=document.querySelector('#status'),bar=document.querySelector('#bar');let selected;const info={vocals:['athena-vocals-instrumental.zip','vocals.flac und instrumental.flac'],four_stem:['athena-4-stems.zip','vier Spuren'],six_stem:['athena-6-stems-experimental.zip','sechs Spuren']};function choose(f){selected=f;name.textContent=f?`${f.name} · ${(f.size/1048576).toFixed(1)} MiB`:'Noch keine Datei gewählt';start.disabled=!f}file.onchange=()=>choose(file.files[0]);drop.ondragover=e=>{e.preventDefault();drop.classList.add('drag')};drop.ondragleave=()=>drop.classList.remove('drag');drop.ondrop=e=>{e.preventDefault();drop.classList.remove('drag');choose(e.dataTransfer.files[0])};start.onclick=async()=>{const mode=document.querySelector('input[name=mode]:checked').value;start.disabled=true;bar.classList.add('run');status.textContent='Modell trennt den Titel – das kann einige Minuten dauern …';let body=new FormData();body.append('file',selected);body.append('mode',mode);try{let r=await fetch('/v1/separate',{method:'POST',body});if(!r.ok)throw Error((await r.json()).detail||`HTTP ${r.status}`);let blob=await r.blob(),a=document.createElement('a');a.href=URL.createObjectURL(blob);a.download=info[mode][0];a.click();setTimeout(()=>URL.revokeObjectURL(a.href),5000);status.textContent=`Fertig – ZIP mit ${info[mode][1]} wurde geladen.`}catch(e){status.textContent=`Fehler: ${e.message}`}finally{bar.classList.remove('run');start.disabled=false}};
</script></body></html> </script></body></html>