Stimmen sauber trennen
BS‑RoFormer Viperx 1297 zerlegt deinen Titel in zwei verlustfreie FLAC-Spuren: Gesang und Instrumental.
diff --git a/README.md b/README.md index 245a40e..480189e 100644 --- a/README.md +++ b/README.md @@ -124,7 +124,7 @@ Details, Installation, Prüfung und Rollback stehen in - Athena-Dashboard: `http://192.168.1.212:8099` - Musikstudio, Original UI (stabil): `http://192.168.1.212:7862` - Musikstudio, Community UI (experimentell): `http://192.168.1.212:7861` -- Stimmen trennen (BS-RoFormer): `http://192.168.1.212:8007` +- Spuren trennen (BS-RoFormer + Demucs, 2/4/6 Stems): `http://192.168.1.212:8007` Der Betriebsmodus lässt sich dort direkt umschalten. In Hermes funktionieren außerdem `/athena music`, `/athena stems`, `/athena llm` und `/athena status`; Details stehen in diff --git a/docs/OPERATING_MODES.md b/docs/OPERATING_MODES.md index 5e0fc5a..337dca5 100644 --- a/docs/OPERATING_MODES.md +++ b/docs/OPERATING_MODES.md @@ -4,8 +4,8 @@ Athena besitzt drei gegenseitig exklusive Betriebsmodi: - `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt. - `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt. -- `separation`: BS-RoFormer Viperx 1297 trennt Gesang und Instrumental; - LLM, Bild, TTS und ACE-Step sind gestoppt. +- `separation`: wahlweise BS-RoFormer für Gesang/Instrumental oder Demucs für + vier beziehungsweise sechs Spuren; LLM, Bild, TTS und ACE-Step sind gestoppt. Die Zustandsmaschine lebt im Athena-Router. Das Dashboard und Chat-Clients wie Hermes sind nur Bedienoberflächen derselben API. Der zuletzt aktive LLM-Modus @@ -36,9 +36,12 @@ Docker-Netz `mike-ai-music` auf `http://music-worker:7860` zu. Im Trennmodus öffnet das Dashboard die private Athena-Oberfläche unter `http://192.168.1.212:8007`. Sie nimmt WAV, FLAC, MP3, M4A und weitere -übliche Formate an und liefert ein ZIP mit verlustfreien `vocals.flac` und -`instrumental.flac`. Grundlage ist `audio-separator` 0.47.0 mit -`model_bs_roformer_ep_317_sdr_12.9755.ckpt`. +übliche Formate an und liefert verlustfreie FLAC-Spuren als ZIP. Verfügbar sind +die spezialisierte Gesangstrennung (`vocals`, `instrumental`), vier Spuren +(`vocals`, `drums`, `bass`, `other`) und experimentell sechs Spuren +(`vocals`, `drums`, `bass`, `guitar`, `piano`, `other`). Grundlage ist +`audio-separator` 0.47.0 mit BS-RoFormer Viperx 1297, `htdemucs_ft` und +`htdemucs_6s`. Hermes benötigt dafür kein Plugin. Exakt eingegebene Steuerbefehle werden vom Router lokal beantwortet, auch wenn gerade kein LLM geladen ist: diff --git a/experiments/bs-roformer-vocal-separation/Dockerfile b/experiments/bs-roformer-vocal-separation/Dockerfile index 442aa7b..4ac6e57 100644 --- a/experiments/bs-roformer-vocal-separation/Dockerfile +++ b/experiments/bs-roformer-vocal-separation/Dockerfile @@ -19,4 +19,4 @@ ENV MODEL_FILENAME=model_bs_roformer_ep_317_sdr_12.9755.ckpt \ JOB_DIR=/data/jobs EXPOSE 8080 -CMD ["sh", "-c", "audio-separator --model_filename \"$MODEL_FILENAME\" --model_file_dir \"$MODEL_DIR\" --download_model_only && exec uvicorn app:app --host 0.0.0.0 --port 8080 --workers 1"] +CMD ["sh", "-c", "for model in \"$MODEL_FILENAME\" htdemucs_ft.yaml htdemucs_6s.yaml; do audio-separator --model_filename \"$model\" --model_file_dir \"$MODEL_DIR\" --download_model_only || exit 1; done; exec uvicorn app:app --host 0.0.0.0 --port 8080 --workers 1"] diff --git a/experiments/bs-roformer-vocal-separation/README.md b/experiments/bs-roformer-vocal-separation/README.md index 6dc0dca..10036b5 100644 --- a/experiments/bs-roformer-vocal-separation/README.md +++ b/experiments/bs-roformer-vocal-separation/README.md @@ -1,16 +1,23 @@ -# Athena Vocal Separator +# Athena Stem Separator -Exklusiver dritter Athena-Betriebsmodus für lokale Zwei-Spur-Trennung in -`vocals.flac` und `instrumental.flac`. +Exklusiver dritter Athena-Betriebsmodus für lokale Zwei- bis Sechs-Spur-Trennung. - Engine: `audio-separator` 0.47.0 (MIT) -- Modell: BS-RoFormer Viperx 1297, - `model_bs_roformer_ep_317_sdr_12.9755.ckpt` -- Modellbewertung im audio-separator-Katalog: Vocal SDR 12,9, - Instrumental SDR 17,0 +- **Gesang / Instrumental:** BS-RoFormer Viperx 1297, + `model_bs_roformer_ep_317_sdr_12.9755.ckpt`; Vocal SDR 12,9, + Instrumental SDR 17,0. +- **4 Spuren:** `htdemucs_ft.yaml` mit Gesang, Schlagzeug, Bass und Rest. +- **6 Spuren (experimentell):** `htdemucs_6s.yaml` ergänzt gemeinsame Gitarre + und Piano. Die Instrumentqualität liegt unter der spezialisierten + Zwei-Spur-Trennung. - GPU: RTX 5080; LLM, Bildmodelle, TTS und ACE-Step sind dabei verriegelt. - Privat erreichbar: `http://192.168.1.212:8007/` -Das Modell wird beim ersten Start nach `/data/models/audio-separator` +Die Modelle werden beim ersten Start nach `/data/models/audio-separator` heruntergeladen. Temporäre Jobs liegen unter `/data/audio/separation` und -werden nach dem ZIP-Download entfernt. +werden nach dem ZIP-Download entfernt. Eigene Spuren für E-/Akustikgitarre, +Synthesizer und Streicher sind bewusst noch nicht angeboten: Dafür braucht es +weitere Zielmodelle; sie aus `other` umzubenennen wäre fachlich falsch. + +Die API erwartet `multipart/form-data` mit `file` und optional `mode`: +`vocals` (Standard), `four_stem` oder `six_stem`. diff --git a/experiments/bs-roformer-vocal-separation/app.py b/experiments/bs-roformer-vocal-separation/app.py index c0a7579..23aae6a 100644 --- a/experiments/bs-roformer-vocal-separation/app.py +++ b/experiments/bs-roformer-vocal-separation/app.py @@ -9,7 +9,7 @@ import time import zipfile from pathlib import Path -from fastapi import FastAPI, File, HTTPException, UploadFile +from fastapi import FastAPI, File, Form, HTTPException, UploadFile from fastapi.responses import FileResponse, HTMLResponse from starlette.background import BackgroundTask @@ -22,7 +22,13 @@ ALLOWED = {".wav", ".flac", ".mp3", ".m4a", ".aac", ".ogg", ".opus", ".wma"} SEPARATION_LOCK = asyncio.Lock() STARTED = time.time() -app = FastAPI(title="Athena Vocal Separator", version="1.0") +MODES = { + "vocals": {"model": MODEL, "stems": ("vocals", "instrumental"), "archive": "athena-vocals-instrumental.zip", "engine": "mdxc"}, + "four_stem": {"model": "htdemucs_ft.yaml", "stems": ("vocals", "drums", "bass", "other"), "archive": "athena-4-stems.zip", "engine": "demucs"}, + "six_stem": {"model": "htdemucs_6s.yaml", "stems": ("vocals", "drums", "bass", "guitar", "piano", "other"), "archive": "athena-6-stems-experimental.zip", "engine": "demucs"}, +} + +app = FastAPI(title="Athena Stem Separator", version="2.0") @app.get("/", response_class=HTMLResponse) @@ -32,11 +38,11 @@ def index() -> str: @app.get("/health") def health() -> dict: - checkpoint = MODEL_DIR / MODEL + available = {name: (MODEL_DIR / mode["model"]).exists() for name, mode in MODES.items()} return { - "status": "ok" if checkpoint.exists() else "starting", - "model": MODEL, - "model_ready": checkpoint.exists(), + "status": "ok" if all(available.values()) else "starting", + "models": {name: mode["model"] for name, mode in MODES.items()}, + "models_ready": available, "busy": SEPARATION_LOCK.locked(), "uptime_seconds": round(time.time() - STARTED, 1), } @@ -46,28 +52,40 @@ def _cleanup(path: Path) -> None: shutil.rmtree(path, ignore_errors=True) -def _run_separator(input_path: Path, output_dir: Path) -> None: +def _run_separator(input_path: Path, output_dir: Path, mode: dict) -> None: args = [ "audio-separator", str(input_path), - "--model_filename", MODEL, + "--model_filename", mode["model"], "--model_file_dir", str(MODEL_DIR), "--output_dir", str(output_dir), "--output_format", "FLAC", "--sample_rate", "44100", "--use_soundfile", "--use_autocast", - "--mdxc_segment_size", "256", - "--mdxc_overlap", "8", - "--mdxc_batch_size", "1", ] + if mode["engine"] == "mdxc": + args.extend(["--mdxc_segment_size", "256", "--mdxc_overlap", "8", "--mdxc_batch_size", "1"]) + else: + args.extend(["--demucs_segment_size", "40", "--demucs_shifts", "2", "--demucs_overlap", "0.25"]) completed = subprocess.run(args, capture_output=True, text=True, timeout=7200) if completed.returncode: detail = (completed.stderr or completed.stdout or "unknown error")[-4000:] raise RuntimeError(detail) +def _stem_name(path: Path, expected: tuple[str, ...]) -> str | None: + lower = path.stem.lower() + for stem in sorted(expected, key=len, reverse=True): + if stem in lower: + return stem + return None + + @app.post("/v1/separate") -async def separate(file: UploadFile = File(...)) -> FileResponse: +async def separate(file: UploadFile = File(...), mode: str = Form("vocals")) -> FileResponse: + selected_mode = MODES.get(mode) + if selected_mode is None: + raise HTTPException(422, f"Unbekannter Trennmodus: {mode}") suffix = Path(file.filename or "upload.wav").suffix.lower() if suffix not in ALLOWED: raise HTTPException(415, "Dieses Audioformat wird nicht unterstützt.") @@ -87,21 +105,24 @@ async def separate(file: UploadFile = File(...)) -> FileResponse: raise HTTPException(413, "Datei ist größer als 1 GiB.") handle.write(chunk) async with SEPARATION_LOCK: - await asyncio.to_thread(_run_separator, input_path, output_dir) + await asyncio.to_thread(_run_separator, input_path, output_dir, selected_mode) stems = sorted(output_dir.glob("*.flac")) - if len(stems) != 2: - raise RuntimeError(f"Erwartet wurden zwei FLAC-Dateien, gefunden: {len(stems)}") - archive = job / "athena-vocals-instrumental.zip" + expected = selected_mode["stems"] + recognized = {_stem_name(stem, expected): stem for stem in stems} + recognized.pop(None, None) + missing = [stem for stem in expected if stem not in recognized] + if missing: + found = ", ".join(stem.name for stem in stems) or "keine" + raise RuntimeError(f"Fehlende Spuren: {', '.join(missing)}; gefunden: {found}") + archive = job / selected_mode["archive"] with zipfile.ZipFile(archive, "w", compression=zipfile.ZIP_STORED) as bundle: - for stem in stems: - lower = stem.name.lower() - target = "vocals.flac" if "vocal" in lower else "instrumental.flac" - bundle.write(stem, target) + for stem in expected: + bundle.write(recognized[stem], f"{stem}.flac") return FileResponse( archive, media_type="application/zip", - filename="athena-vocals-instrumental.zip", + filename=selected_mode["archive"], background=BackgroundTask(_cleanup, job), ) except HTTPException: diff --git a/experiments/bs-roformer-vocal-separation/index.html b/experiments/bs-roformer-vocal-separation/index.html index ef43742..f972bad 100644 --- a/experiments/bs-roformer-vocal-separation/index.html +++ b/experiments/bs-roformer-vocal-separation/index.html @@ -1,4 +1,9 @@ -
BS‑RoFormer Viperx 1297 zerlegt deinen Titel in zwei verlustfreie FLAC-Spuren: Gesang und Instrumental.