Stimmen sauber trennen
BS‑RoFormer Viperx 1297 zerlegt deinen Titel in zwei verlustfreie FLAC-Spuren: Gesang und Instrumental.
diff --git a/README.md b/README.md index 0792a40..245a40e 100644 --- a/README.md +++ b/README.md @@ -15,8 +15,8 @@ Bild- und Sprachausgabe. **Hermes und die Fach-MCPs laufen auf Unraid.** - Qwen3-TTS 1.7B auf der RTX 3060 mit Piper als CPU-Fallback - Whisper.cpp `large-v3-turbo` auf der CPU für lokale deutsche Spracherkennung - Live-Dashboard mit 21 Tagen Detailhistorie auf Port 8099 -- Dashboard-Umschaltung zwischen LLM-Betrieb und ACE-Step-Musikstudio mit - persistenter `fspecii/ace-step-ui`-Bibliothek +- Dashboard-Umschaltung zwischen LLM-Betrieb, ACE-Step-Musikstudio und + BS-RoFormer-Stimmtrennung - Portainer CE als optionale Container-Ansicht auf Port 9443 - WireGuard-Gateway, Datenbackup und Athena-Operator - keine produktive Hermes-, OpenWebUI- oder portable Fach-MCP-Instanz @@ -124,9 +124,10 @@ Details, Installation, Prüfung und Rollback stehen in - Athena-Dashboard: `http://192.168.1.212:8099` - Musikstudio, Original UI (stabil): `http://192.168.1.212:7862` - Musikstudio, Community UI (experimentell): `http://192.168.1.212:7861` +- Stimmen trennen (BS-RoFormer): `http://192.168.1.212:8007` Der Betriebsmodus lässt sich dort direkt umschalten. In Hermes funktionieren -außerdem `/athena music`, `/athena llm` und `/athena status`; Details stehen in +außerdem `/athena music`, `/athena stems`, `/athena llm` und `/athena status`; Details stehen in [docs/OPERATING_MODES.md](docs/OPERATING_MODES.md). Der Router stellt Sprache OpenAI-kompatibel bereit: Sprachausgabe über diff --git a/compose.yaml b/compose.yaml index 3eabdf0..c22ce74 100644 --- a/compose.yaml +++ b/compose.yaml @@ -570,6 +570,7 @@ services: RESTORE_WORKER: restore TTS_WORKER: qwen3 MUSIC_WORKER: acestep + SEPARATOR_WORKER: bs-roformer networks: [control] security_opt: ["no-new-privileges:true"] healthcheck: @@ -858,6 +859,7 @@ services: ROUTER_API_KEY: "${ROUTER_API_KEY:?ROUTER_API_KEY is required}" MUSIC_COMMUNITY_UI_URL: "${MUSIC_COMMUNITY_UI_URL:-http://192.168.1.212:7861/}" MUSIC_ORIGINAL_UI_URL: "${MUSIC_ORIGINAL_UI_URL:-http://192.168.1.212:7862/}" + SEPARATOR_UI_URL: "${SEPARATOR_UI_URL:-http://192.168.1.212:8007/}" HOST_PROC: /host/proc HOST_DATA: /host/data HOST_MODELS: /host/models diff --git a/dev/test_profile_controller.py b/dev/test_profile_controller.py index c84106d..95d0f06 100644 --- a/dev/test_profile_controller.py +++ b/dev/test_profile_controller.py @@ -41,7 +41,40 @@ def music_item(state="exited"): "Labels": {controller.MUSIC_LABEL_KEY: "acestep"}} +def separator_item(state="exited"): + return {"Id": "id-separator", "State": state, + "Labels": {controller.SEPARATOR_LABEL_KEY: "bs-roformer"}} + + class ProfileControllerTests(unittest.TestCase): + def test_separator_start_exclusively_stops_gpu_workers(self): + profiles = {name: item(name) for name in controller.ALLOWED} + profiles["large"] = item("large", "running") + calls = [] + + def request(method, path): + calls.append((method, path)) + return 204, b"" + + with patch.object(controller, "SEPARATOR_WORKER", "bs-roformer"), \ + patch.object(controller, "MUSIC_WORKER", "acestep"), \ + patch.object(controller, "containers", return_value=profiles), \ + patch.object(controller, "separator_container", return_value=separator_item()), \ + patch.object(controller, "music_container", return_value=music_item("running")), \ + patch.object(controller, "image_containers", return_value=[image_item("running")]), \ + patch.object(controller, "tts_container", return_value=tts_item()), \ + patch.object(controller, "docker_request", side_effect=request): + result = controller.set_separator_worker(True) + + self.assertEqual(result, {"separator_worker": "bs-roformer", "state": "running"}) + self.assertEqual(calls, [ + ("POST", "/containers/id-large/stop?t=120"), + ("POST", "/containers/id-flux/stop?t=20"), + ("POST", "/containers/id-tts/stop?t=30"), + ("POST", "/containers/id-music/stop?t=30"), + ("POST", "/containers/id-separator/start"), + ]) + def test_music_start_exclusively_stops_gpu_workers(self): profiles = {name: item(name) for name in controller.ALLOWED} profiles["ultra"] = item("ultra", "running") diff --git a/docs/OPERATING_MODES.md b/docs/OPERATING_MODES.md index 7de8904..5e0fc5a 100644 --- a/docs/OPERATING_MODES.md +++ b/docs/OPERATING_MODES.md @@ -1,18 +1,20 @@ # Athena-Betriebsmodi -Athena besitzt zwei gegenseitig exklusive Betriebsmodi: +Athena besitzt drei gegenseitig exklusive Betriebsmodi: -- `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; ACE-Step ist gestoppt. -- `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild- und TTS-Worker sind gestoppt. +- `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt. +- `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt. +- `separation`: BS-RoFormer Viperx 1297 trennt Gesang und Instrumental; + LLM, Bild, TTS und ACE-Step sind gestoppt. Die Zustandsmaschine lebt im Athena-Router. Das Dashboard und Chat-Clients wie Hermes sind nur Bedienoberflächen derselben API. Der zuletzt aktive LLM-Modus -wird persistent gespeichert und beim Verlassen des Musikmodus wieder geladen. +wird persistent gespeichert und beim Verlassen eines Spezialmodus wieder geladen. ## Bedienung -Im Athena-Dashboard stehen die Schaltflächen **LLM-Betrieb** und -**Musikstudio starten** bereit. Im Musikmodus werden zwei Oberflächen angeboten: +Im Athena-Dashboard stehen **LLM-Betrieb**, **Musikstudio** und +**Stimmen trennen** bereit. Im Musikmodus werden zwei Oberflächen angeboten: - **Original UI · stabil** öffnet die zum laufenden ACE-Step-Image gehörende Gradio-Oberfläche. Sie ist für Cover, Remix und erweiterte Workflows der @@ -32,11 +34,18 @@ veröffentlicht. `ace-step-ui` ist reproduzierbar auf Commit `a1fdf91829ec6f7b98844f80e323529cd155dbf2` fixiert und greift intern über das Docker-Netz `mike-ai-music` auf `http://music-worker:7860` zu. +Im Trennmodus öffnet das Dashboard die private Athena-Oberfläche unter +`http://192.168.1.212:8007`. Sie nimmt WAV, FLAC, MP3, M4A und weitere +übliche Formate an und liefert ein ZIP mit verlustfreien `vocals.flac` und +`instrumental.flac`. Grundlage ist `audio-separator` 0.47.0 mit +`model_bs_roformer_ep_317_sdr_12.9755.ckpt`. + Hermes benötigt dafür kein Plugin. Exakt eingegebene Steuerbefehle werden vom Router lokal beantwortet, auch wenn gerade kein LLM geladen ist: ```text /athena music +/athena stems /athena llm /athena status ``` @@ -46,17 +55,19 @@ Die HTTP-Schnittstelle verwendet authentifizierte Requests: ```text GET /mode POST /mode {"mode":"music"} +POST /mode {"mode":"separation"} POST /mode {"mode":"llm"} ``` Der Wechsel läuft asynchron. Fortschritt und Fehler stehen unter `mode` in `GET /status`. Der Profile-Controller akzeptiert ausschließlich den mit -`com.mike-ai.music-worker=acestep` markierten Container; freie Container- oder -Docker-Befehle werden nicht entgegengenommen. +`com.mike-ai.music-worker=acestep` beziehungsweise +`com.mike-ai.stem-separator=bs-roformer` markierten Container; freie +Container- oder Docker-Befehle werden nicht entgegengenommen. ## Wiederanlauf Der Router speichert `mode`, `last_profile` und `return_profile` atomar. War -beim Router-Neustart der Musikmodus aktiv, startet er ACE-Step erneut. Beim +beim Router-Neustart ein Spezialmodus aktiv, startet er den passenden Worker erneut. Beim Wechsel zurück wird das gespeicherte LLM-Profil semantisch auf Alias und Kontextfenster geprüft, bevor Chat-Anfragen wieder freigegeben werden. diff --git a/docs/SPECIALIZED_MODEL_ROADMAP.md b/docs/SPECIALIZED_MODEL_ROADMAP.md index 82d0c8e..ea8a46b 100644 --- a/docs/SPECIALIZED_MODEL_ROADMAP.md +++ b/docs/SPECIALIZED_MODEL_ROADMAP.md @@ -9,7 +9,7 @@ im zentralen Register `TESTED_MODELS.md` dokumentiert wurde. | Prioritaet | Aufgabe | Kandidat | Geplanter Betrieb | Status | |---:|---|---|---|---| | 1 | Musik erzeugen und bearbeiten | `ACE-Step 1.5 XL SFT` mit `acestep-5Hz-lm-1.7B` | exklusives On-Demand-Profil auf der RTX 5080; CPU-Offload; Qwen, Vision und TTS werden waehrenddessen entladen | **integriert; Klangabnahme laeuft** | -| 2 | Gesang und Instrumente trennen | BS-RoFormer Viperx 1297; alternativ MelBand-RoFormer, fuer Mehrspur `htdemucs_ft` | eigener Audio-Worker; GPU bevorzugt, CPU als langsamer Fallback | offen | +| 2 | Gesang und Instrumente trennen | BS-RoFormer Viperx 1297, `ep_317` | exklusiver Audio-Worker auf der RTX 5080; FLAC-Ausgabe; eigener Dashboard-Modus | **integriert; Qualitätstest läuft** | | 3 | Voice Cloning | vorhandenes `Qwen3-TTS-12Hz-1.7B-Base` | bestehender TTS-Worker auf der RTX 3060; zunaechst den eingebauten 3-Sekunden-Klonpfad freilegen | offen | | 4 | Objekte lokalisieren und zaehlen | Grounding DINO oder RF-DETR | optionaler Vision-Worker; normales Erkennen bleibt beim vorhandenen Qwen-Vision-Projektor | offen | | 5 | Bildort schaetzen | GeoAgent 8B | exklusives Vision-Profil; Ergebnis nur als Wahrscheinlichkeitsrangliste | offen | diff --git a/docs/TESTED_MODELS.md b/docs/TESTED_MODELS.md index 4847c1d..1d1ef28 100644 --- a/docs/TESTED_MODELS.md +++ b/docs/TESTED_MODELS.md @@ -63,6 +63,12 @@ Titelgenerierung und Kontextkompression in Hermes. |---|---|---|---|---|---| | 08.09.2026 | `ACE-Step/acestep-v15-xl-sft` mit `acestep-5Hz-lm-1.7B`, offizielles ACE-Step-1.5-Image `sha256:95652cd780c78a1b1a7f6f0335530430f0ae53d96c7c12d59f9f39fa23d38567` | 30 s Instrumental, Thinking/LM aktiv, Batch 1, RTX 5080 16 GiB, automatischer CPU-Offload und INT8 Weight-only DiT | erfolgreich in 15,39 s: LM 8,00 s, DiT 7,39 s, MP3 0,82 s; PyTorch meldete maximal 9,38 GiB CUDA-Allokation; kein OOM/CUDA-Fehler | **Beta-Test bestanden**; Klangabnahme und Hermes-/Router-Integration noch offen | Athena: `/data/music/acestep/batch_1788873234/`; [SPECIALIZED_MODEL_ROADMAP.md](SPECIALIZED_MODEL_ROADMAP.md) | +## Audio-Trennung + +| Datum | Modell | Test | Ergebnis | Status / Entscheidung | Beleg | +|---|---|---|---|---|---| +| 08.09.2026 | BS-RoFormer Viperx 1297, `model_bs_roformer_ep_317_sdr_12.9755.ckpt`, `audio-separator` 0.47.0 | 20-s-FLAC eines vorhandenen ACE-Step-Titels, RTX 5080, CUDA 12.8, ONNX Runtime GPU 1.22.0 | zwei gültige FLAC-Spuren mit jeweils exakt 20,0 s; Verarbeitung 19 s; Vocal-Datei 1,45 MB, Instrumental-Datei 3,84 MB | **technischer Ende-zu-Ende-Test bestanden**; Hörabnahme durch Nutzer offen | [bs-roformer-vocal-separation](../experiments/bs-roformer-vocal-separation/README.md) | + ## Ablauf für zukünftige Kandidaten 1. Exakten Hugging-Face-/Ollama-Namen und Dateinamen in diesem Dokument suchen. diff --git a/experiments/bs-roformer-vocal-separation/Dockerfile b/experiments/bs-roformer-vocal-separation/Dockerfile new file mode 100644 index 0000000..442aa7b --- /dev/null +++ b/experiments/bs-roformer-vocal-separation/Dockerfile @@ -0,0 +1,22 @@ +FROM pytorch/pytorch:2.7.1-cuda12.8-cudnn9-runtime@sha256:c16f4c749e2d9e96878875cdf6cc45cddda1d1a36fddd371dd6f2360f1b6e2a2 + +RUN apt-get update \ + && DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends build-essential curl ffmpeg libsndfile1 \ + && rm -rf /var/lib/apt/lists/* + +RUN python -m pip install --no-cache-dir \ + "audio-separator[gpu]==0.47.0" \ + "onnxruntime-gpu==1.22.0" \ + "fastapi==0.116.1" \ + "python-multipart==0.0.20" \ + "uvicorn[standard]==0.35.0" + +WORKDIR /app +COPY app.py index.html ./ + +ENV MODEL_FILENAME=model_bs_roformer_ep_317_sdr_12.9755.ckpt \ + MODEL_DIR=/models \ + JOB_DIR=/data/jobs + +EXPOSE 8080 +CMD ["sh", "-c", "audio-separator --model_filename \"$MODEL_FILENAME\" --model_file_dir \"$MODEL_DIR\" --download_model_only && exec uvicorn app:app --host 0.0.0.0 --port 8080 --workers 1"] diff --git a/experiments/bs-roformer-vocal-separation/README.md b/experiments/bs-roformer-vocal-separation/README.md new file mode 100644 index 0000000..6dc0dca --- /dev/null +++ b/experiments/bs-roformer-vocal-separation/README.md @@ -0,0 +1,16 @@ +# Athena Vocal Separator + +Exklusiver dritter Athena-Betriebsmodus für lokale Zwei-Spur-Trennung in +`vocals.flac` und `instrumental.flac`. + +- Engine: `audio-separator` 0.47.0 (MIT) +- Modell: BS-RoFormer Viperx 1297, + `model_bs_roformer_ep_317_sdr_12.9755.ckpt` +- Modellbewertung im audio-separator-Katalog: Vocal SDR 12,9, + Instrumental SDR 17,0 +- GPU: RTX 5080; LLM, Bildmodelle, TTS und ACE-Step sind dabei verriegelt. +- Privat erreichbar: `http://192.168.1.212:8007/` + +Das Modell wird beim ersten Start nach `/data/models/audio-separator` +heruntergeladen. Temporäre Jobs liegen unter `/data/audio/separation` und +werden nach dem ZIP-Download entfernt. diff --git a/experiments/bs-roformer-vocal-separation/app.py b/experiments/bs-roformer-vocal-separation/app.py new file mode 100644 index 0000000..c0a7579 --- /dev/null +++ b/experiments/bs-roformer-vocal-separation/app.py @@ -0,0 +1,124 @@ +from __future__ import annotations + +import asyncio +import os +import shutil +import subprocess +import tempfile +import time +import zipfile +from pathlib import Path + +from fastapi import FastAPI, File, HTTPException, UploadFile +from fastapi.responses import FileResponse, HTMLResponse +from starlette.background import BackgroundTask + + +MODEL = os.getenv("MODEL_FILENAME", "model_bs_roformer_ep_317_sdr_12.9755.ckpt") +MODEL_DIR = Path(os.getenv("MODEL_DIR", "/models")) +JOB_DIR = Path(os.getenv("JOB_DIR", "/data/jobs")) +MAX_UPLOAD = int(os.getenv("MAX_UPLOAD_BYTES", str(1024 ** 3))) +ALLOWED = {".wav", ".flac", ".mp3", ".m4a", ".aac", ".ogg", ".opus", ".wma"} +SEPARATION_LOCK = asyncio.Lock() +STARTED = time.time() + +app = FastAPI(title="Athena Vocal Separator", version="1.0") + + +@app.get("/", response_class=HTMLResponse) +def index() -> str: + return Path("/app/index.html").read_text(encoding="utf-8") + + +@app.get("/health") +def health() -> dict: + checkpoint = MODEL_DIR / MODEL + return { + "status": "ok" if checkpoint.exists() else "starting", + "model": MODEL, + "model_ready": checkpoint.exists(), + "busy": SEPARATION_LOCK.locked(), + "uptime_seconds": round(time.time() - STARTED, 1), + } + + +def _cleanup(path: Path) -> None: + shutil.rmtree(path, ignore_errors=True) + + +def _run_separator(input_path: Path, output_dir: Path) -> None: + args = [ + "audio-separator", str(input_path), + "--model_filename", MODEL, + "--model_file_dir", str(MODEL_DIR), + "--output_dir", str(output_dir), + "--output_format", "FLAC", + "--sample_rate", "44100", + "--use_soundfile", + "--use_autocast", + "--mdxc_segment_size", "256", + "--mdxc_overlap", "8", + "--mdxc_batch_size", "1", + ] + completed = subprocess.run(args, capture_output=True, text=True, timeout=7200) + if completed.returncode: + detail = (completed.stderr or completed.stdout or "unknown error")[-4000:] + raise RuntimeError(detail) + + +@app.post("/v1/separate") +async def separate(file: UploadFile = File(...)) -> FileResponse: + suffix = Path(file.filename or "upload.wav").suffix.lower() + if suffix not in ALLOWED: + raise HTTPException(415, "Dieses Audioformat wird nicht unterstützt.") + if SEPARATION_LOCK.locked(): + raise HTTPException(409, "Eine Trennung läuft bereits.") + + job = Path(tempfile.mkdtemp(prefix="separate-", dir=JOB_DIR)) + input_path = job / f"input{suffix}" + output_dir = job / "output" + output_dir.mkdir() + size = 0 + try: + with input_path.open("wb") as handle: + while chunk := await file.read(1024 * 1024): + size += len(chunk) + if size > MAX_UPLOAD: + raise HTTPException(413, "Datei ist größer als 1 GiB.") + handle.write(chunk) + async with SEPARATION_LOCK: + await asyncio.to_thread(_run_separator, input_path, output_dir) + + stems = sorted(output_dir.glob("*.flac")) + if len(stems) != 2: + raise RuntimeError(f"Erwartet wurden zwei FLAC-Dateien, gefunden: {len(stems)}") + archive = job / "athena-vocals-instrumental.zip" + with zipfile.ZipFile(archive, "w", compression=zipfile.ZIP_STORED) as bundle: + for stem in stems: + lower = stem.name.lower() + target = "vocals.flac" if "vocal" in lower else "instrumental.flac" + bundle.write(stem, target) + return FileResponse( + archive, + media_type="application/zip", + filename="athena-vocals-instrumental.zip", + background=BackgroundTask(_cleanup, job), + ) + except HTTPException: + _cleanup(job) + raise + except subprocess.TimeoutExpired: + _cleanup(job) + raise HTTPException(504, "Die Trennung hat das Zeitlimit überschritten.") + except Exception as exc: + _cleanup(job) + raise HTTPException(500, f"Trennung fehlgeschlagen: {exc}") + + +@app.on_event("startup") +def prepare() -> None: + JOB_DIR.mkdir(parents=True, exist_ok=True) + MODEL_DIR.mkdir(parents=True, exist_ok=True) + for old in JOB_DIR.glob("separate-*"): + if old.is_dir() and time.time() - old.stat().st_mtime > 86400: + _cleanup(old) diff --git a/experiments/bs-roformer-vocal-separation/compose.yaml b/experiments/bs-roformer-vocal-separation/compose.yaml new file mode 100644 index 0000000..4196737 --- /dev/null +++ b/experiments/bs-roformer-vocal-separation/compose.yaml @@ -0,0 +1,38 @@ +services: + stem-separator: + build: . + image: mike-ai/bs-roformer-separator:0.47.0 + container_name: mike-ai-stem-separator + labels: + com.mike-ai.stem-separator: "bs-roformer" + environment: + NVIDIA_VISIBLE_DEVICES: ${SEPARATOR_GPU_UUID:?set SEPARATOR_GPU_UUID to the RTX 5080 UUID} + MODEL_FILENAME: model_bs_roformer_ep_317_sdr_12.9755.ckpt + MODEL_DIR: /models + JOB_DIR: /data/jobs + deploy: + resources: + reservations: + devices: + - driver: nvidia + device_ids: ["${SEPARATOR_GPU_UUID:?set SEPARATOR_GPU_UUID to the RTX 5080 UUID}"] + capabilities: [gpu] + ports: + - "127.0.0.1:${SEPARATOR_PORT:-8007}:8080" + volumes: + - ${SEPARATOR_MODEL_DIR:-/data/models/audio-separator}:/models + - ${SEPARATOR_DATA_DIR:-/data/audio/separation}:/data + shm_size: "2gb" + restart: "no" + healthcheck: + test: ["CMD-SHELL", "curl -fsS http://127.0.0.1:8080/health | grep -q '\"status\":\"ok\"'"] + interval: 15s + timeout: 5s + start_period: 600s + retries: 3 + networks: [frontend] + +networks: + frontend: + external: true + name: mike-ai_frontend diff --git a/experiments/bs-roformer-vocal-separation/index.html b/experiments/bs-roformer-vocal-separation/index.html new file mode 100644 index 0000000..ef43742 --- /dev/null +++ b/experiments/bs-roformer-vocal-separation/index.html @@ -0,0 +1,4 @@ +
BS‑RoFormer Viperx 1297 zerlegt deinen Titel in zwei verlustfreie FLAC-Spuren: Gesang und Instrumental.