diff --git a/docs/OPERATING_MODES.md b/docs/OPERATING_MODES.md index 536a84b..0bc9fd6 100644 --- a/docs/OPERATING_MODES.md +++ b/docs/OPERATING_MODES.md @@ -4,8 +4,8 @@ Athena besitzt drei gegenseitig exklusive Betriebsmodi: - `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt. - `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt. -- `separation`: wahlweise BS-RoFormer für Gesang/Instrumental oder Demucs für - vier beziehungsweise sechs Spuren; LLM, Bild, TTS und ACE-Step sind gestoppt. +- `separation`: BS-RoFormer und Demucs trennen Musikspuren; ClearVoice trennt + Sprache von Hintergrundgeräuschen. LLM, Bild, TTS und ACE-Step sind gestoppt. Die Zustandsmaschine lebt im Athena-Router. Das Dashboard und Chat-Clients wie Hermes sind nur Bedienoberflächen derselben API. Der zuletzt aktive LLM-Modus @@ -14,7 +14,7 @@ wird persistent gespeichert und beim Verlassen eines Spezialmodus wieder geladen ## Bedienung Im Athena-Dashboard stehen **LLM-Betrieb**, **Musikstudio** und -**Stimmen trennen** bereit. Im Musikmodus werden zwei Oberflächen angeboten: +**Audio trennen** bereit. Im Musikmodus werden zwei Oberflächen angeboten: - **Original UI · stabil** öffnet die zum laufenden ACE-Step-Image gehörende Gradio-Oberfläche. Sie ist für Cover, Remix und erweiterte Workflows der @@ -37,14 +37,16 @@ Docker-Netz `mike-ai-music` auf `http://music-worker:7860` zu. Im Trennmodus öffnet das Dashboard die private Athena-Oberfläche unter `http://192.168.1.212:8007`. Sie nimmt WAV, FLAC, MP3, M4A und weitere übliche Formate an. Gewählt wird die herauszulösende Quelle: Gesang, -Schlagzeug, Bass, Gitarre, Piano oder Sonstiges. Das ZIP enthält genau diese Zielspur und +Schlagzeug, Bass, Gitarre, Piano, Sonstiges oder gereinigte Sprache. Das ZIP enthält genau diese Zielspur und eine zweite FLAC-Datei mit dem vollständigen Rest ohne die Zielspur. Gesang nutzt BS-RoFormer Viperx 1297, Schlagzeug/Bass `htdemucs_ft` und Gitarre/Piano/Sonstiges experimentell `htdemucs_6s`. „Sonstiges“ ist dessen gemischter `other`-Stem (unter anderem Synthesizer, Streicher, Bläser und Effekte), -nicht eine reine Synthesizer-Spur. Grundlage ist `audio-separator` -0.47.0. Die ältere API-Auswahl kompletter 2-/4-/6-Stem-Sätze bleibt -rückwärtskompatibel. +nicht eine reine Synthesizer-Spur. Sprache nutzt das 48-kHz-Modell +`MossFormer2_SE_48K`; der Download enthält `speech.flac` und +`hintergrund-ohne-sprache.flac`. Die Musiktrennung basiert auf +`audio-separator` 0.47.0. Die ältere API-Auswahl kompletter 2-/4-/6-Stem-Sätze +bleibt rückwärtskompatibel. Hermes benötigt dafür kein Plugin. Exakt eingegebene Steuerbefehle werden vom Router lokal beantwortet, auch wenn gerade kein LLM geladen ist: diff --git a/experiments/bs-roformer-vocal-separation/Dockerfile b/experiments/bs-roformer-vocal-separation/Dockerfile index 4ac6e57..d722e4a 100644 --- a/experiments/bs-roformer-vocal-separation/Dockerfile +++ b/experiments/bs-roformer-vocal-separation/Dockerfile @@ -11,12 +11,20 @@ RUN python -m pip install --no-cache-dir \ "python-multipart==0.0.20" \ "uvicorn[standard]==0.35.0" +# audio-separator 0.47 requires NumPy 2 while ClearVoice 0.1.2 still pins +# NumPy 1.x. Keep ClearVoice in a small overlay venv but share the image's +# CUDA-enabled PyTorch installation instead of duplicating it. +RUN python -m venv --system-site-packages /opt/clearvoice-venv \ + && /opt/clearvoice-venv/bin/python -m pip install --no-cache-dir \ + "clearvoice==0.1.2" \ + "numpy>=1.24.3,<2.0" + WORKDIR /app -COPY app.py index.html ./ +COPY app.py index.html speech_enhance.py ./ ENV MODEL_FILENAME=model_bs_roformer_ep_317_sdr_12.9755.ckpt \ MODEL_DIR=/models \ JOB_DIR=/data/jobs EXPOSE 8080 -CMD ["sh", "-c", "for model in \"$MODEL_FILENAME\" htdemucs_ft.yaml htdemucs_6s.yaml; do audio-separator --model_filename \"$model\" --model_file_dir \"$MODEL_DIR\" --download_model_only || exit 1; done; exec uvicorn app:app --host 0.0.0.0 --port 8080 --workers 1"] +CMD ["sh", "-c", "mkdir -p \"$MODEL_DIR/clearvoice\" && ln -sfn \"$MODEL_DIR/clearvoice\" /app/checkpoints && for model in \"$MODEL_FILENAME\" htdemucs_ft.yaml htdemucs_6s.yaml; do audio-separator --model_filename \"$model\" --model_file_dir \"$MODEL_DIR\" --download_model_only || exit 1; done; /opt/clearvoice-venv/bin/python /app/speech_enhance.py --download-only || exit 1; exec uvicorn app:app --host 0.0.0.0 --port 8080 --workers 1"] diff --git a/experiments/bs-roformer-vocal-separation/README.md b/experiments/bs-roformer-vocal-separation/README.md index c5da84f..1deefa7 100644 --- a/experiments/bs-roformer-vocal-separation/README.md +++ b/experiments/bs-roformer-vocal-separation/README.md @@ -15,6 +15,10 @@ Rest ohne dieses Ziel. liegt unter der spezialisierten Gesangstrennung. - **Sonstiges:** der `other`-Stem von `htdemucs_6s.yaml`. Er bündelt unter anderem Synthesizer, Streicher, Bläser und Effekte und ist keine reine Synthesizer-Spur. +- **Sprache / Hintergrund:** ClearVoice `MossFormer2_SE_48K` (Apache-2.0) + verbessert Sprache bei 48 kHz. Die zweite Spur ist das vom Originalsignal + abgezogene Sprachsignal und enthält den verbleibenden Hintergrund. Stereo wird + kanalweise verarbeitet und anschließend wieder zusammengesetzt. - GPU: RTX 5080; LLM, Bildmodelle, TTS und ACE-Step sind dabei verriegelt. - Privat erreichbar: `http://192.168.1.212:8007/` @@ -26,6 +30,6 @@ weitere Zielmodelle. Die Oberfläche bietet stattdessen den ehrlich benannten, gemischten `other`-Stem als **Sonstiges** an. Die API erwartet `multipart/form-data` mit `file` und optional `target`: -`vocals` (Standard), `drums`, `bass`, `guitar`, `piano` oder `other`. Das ältere Feld +`vocals` (Standard), `drums`, `bass`, `guitar`, `piano`, `other` oder `speech`. Das ältere Feld `mode` mit `vocals`, `four_stem` oder `six_stem` bleibt für vorhandene Clients erhalten und liefert weiterhin alle Modell-Stems. diff --git a/experiments/bs-roformer-vocal-separation/app.py b/experiments/bs-roformer-vocal-separation/app.py index 3d7cecf..456cc17 100644 --- a/experiments/bs-roformer-vocal-separation/app.py +++ b/experiments/bs-roformer-vocal-separation/app.py @@ -26,6 +26,7 @@ MODES = { "vocals": {"model": MODEL, "stems": ("vocals", "instrumental"), "archive": "athena-vocals-instrumental.zip", "engine": "mdxc"}, "four_stem": {"model": "htdemucs_ft.yaml", "stems": ("vocals", "drums", "bass", "other"), "archive": "athena-4-stems.zip", "engine": "demucs"}, "six_stem": {"model": "htdemucs_6s.yaml", "stems": ("vocals", "drums", "bass", "guitar", "piano", "other"), "archive": "athena-6-stems-experimental.zip", "engine": "demucs"}, + "speech": {"model": "MossFormer2_SE_48K", "stems": ("speech", "noise"), "archive": "athena-sprache-und-hintergrund.zip", "engine": "clearvoice"}, } TARGETS = { "vocals": {"mode": "vocals", "stem": "vocals", "remainder": "instrumental", "archive": "athena-gesang-und-rest.zip", "rest_file": "instrumental.flac"}, @@ -34,6 +35,7 @@ TARGETS = { "guitar": {"mode": "six_stem", "stem": "guitar", "archive": "athena-gitarre-und-rest.zip", "rest_file": "rest-ohne-gitarre.flac"}, "piano": {"mode": "six_stem", "stem": "piano", "archive": "athena-piano-und-rest.zip", "rest_file": "rest-ohne-piano.flac"}, "other": {"mode": "six_stem", "stem": "other", "archive": "athena-sonstiges-und-rest.zip", "rest_file": "rest-ohne-sonstiges.flac"}, + "speech": {"mode": "speech", "stem": "speech", "remainder": "noise", "archive": "athena-sprache-und-hintergrund.zip", "rest_file": "hintergrund-ohne-sprache.flac"}, } app = FastAPI(title="Athena Stem Separator", version="2.0") @@ -46,7 +48,14 @@ def index() -> str: @app.get("/health") def health() -> dict: - available = {name: (MODEL_DIR / mode["model"]).exists() for name, mode in MODES.items()} + available = { + name: ( + (MODEL_DIR / "clearvoice" / mode["model"] / "last_best_checkpoint").exists() + if mode["engine"] == "clearvoice" + else (MODEL_DIR / mode["model"]).exists() + ) + for name, mode in MODES.items() + } return { "status": "ok" if all(available.values()) else "starting", "models": {name: mode["model"] for name, mode in MODES.items()}, @@ -62,6 +71,20 @@ def _cleanup(path: Path) -> None: def _run_separator(input_path: Path, output_dir: Path, mode: dict) -> None: + if mode["engine"] == "clearvoice": + completed = subprocess.run( + [ + "/opt/clearvoice-venv/bin/python", "/app/speech_enhance.py", str(input_path), + str(output_dir / "speech.flac"), str(output_dir / "noise.flac"), + ], + capture_output=True, + text=True, + timeout=7200, + ) + if completed.returncode: + detail = (completed.stderr or completed.stdout or "unknown ClearVoice error")[-4000:] + raise RuntimeError(detail) + return args = [ "audio-separator", str(input_path), "--model_filename", mode["model"], diff --git a/experiments/bs-roformer-vocal-separation/index.html b/experiments/bs-roformer-vocal-separation/index.html index c5129f2..94a2112 100644 --- a/experiments/bs-roformer-vocal-separation/index.html +++ b/experiments/bs-roformer-vocal-separation/index.html @@ -1,16 +1,19 @@ Athena · Spuren herauslösen
Athena Audio Lab

Was möchtest du herauslösen?

Der Download enthält immer die gewählte Spur separat und zusätzlich den vollständigen Rest ohne diese Spur.

+
Musik
+
Sprache und Geräusche
+
Bereit.
Alles läuft lokal auf Athena. Synthesizer, Streicher sowie elektrische und akustische Gitarre separat benötigen zusätzliche Spezialmodelle.
diff --git a/experiments/bs-roformer-vocal-separation/speech_enhance.py b/experiments/bs-roformer-vocal-separation/speech_enhance.py new file mode 100644 index 0000000..584db51 --- /dev/null +++ b/experiments/bs-roformer-vocal-separation/speech_enhance.py @@ -0,0 +1,81 @@ +from __future__ import annotations + +import argparse +import subprocess +import tempfile +from pathlib import Path + +import numpy as np +import soundfile as sf +from clearvoice import ClearVoice + + +MODEL = "MossFormer2_SE_48K" +SAMPLE_RATE = 48_000 + + +def convert_input(source: Path, target: Path) -> None: + completed = subprocess.run( + [ + "ffmpeg", "-hide_banner", "-loglevel", "error", "-y", + "-i", str(source), "-vn", "-ar", str(SAMPLE_RATE), + "-c:a", "pcm_f32le", str(target), + ], + capture_output=True, + text=True, + timeout=1800, + ) + if completed.returncode: + raise RuntimeError(completed.stderr[-4000:] or "ffmpeg input conversion failed") + + +def enhance(source: Path, speech_path: Path, noise_path: Path) -> None: + with tempfile.TemporaryDirectory(prefix="clearvoice-") as temp_dir: + converted = Path(temp_dir) / "input-48k.wav" + convert_input(source, converted) + audio, sample_rate = sf.read(converted, dtype="float32", always_2d=True) + if sample_rate != SAMPLE_RATE: + raise RuntimeError(f"unexpected sample rate: {sample_rate}") + + model = ClearVoice(task="speech_enhancement", model_names=[MODEL]) + # Use ClearVoice's file-I/O path so recordings longer than its 20-second + # one-pass window are segmented correctly. Run each channel separately + # because the enhancement network itself is mono, then restore stereo. + channels = [] + for channel_index in range(audio.shape[1]): + channel_path = Path(temp_dir) / f"channel-{channel_index}.wav" + sf.write(channel_path, audio[:, channel_index], SAMPLE_RATE, subtype="FLOAT") + result = np.asarray(model(str(channel_path), False), dtype=np.float32).squeeze() + if result.ndim != 1: + raise RuntimeError(f"unexpected ClearVoice output shape: {result.shape}") + channels.append(result) + enhanced = np.column_stack(channels) + + length = min(len(audio), len(enhanced)) + original = audio[:length] + speech = enhanced[:length] + noise = original - speech + + # FLAC does not support floating-point samples. PCM_24 retains ample + # headroom and avoids the invalid FLOAT/FLAC combination in libsndfile. + sf.write(speech_path, np.clip(speech, -1.0, 1.0), SAMPLE_RATE, format="FLAC", subtype="PCM_24") + sf.write(noise_path, np.clip(noise, -1.0, 1.0), SAMPLE_RATE, format="FLAC", subtype="PCM_24") + + +def main() -> None: + parser = argparse.ArgumentParser() + parser.add_argument("input", nargs="?", type=Path) + parser.add_argument("speech", nargs="?", type=Path) + parser.add_argument("noise", nargs="?", type=Path) + parser.add_argument("--download-only", action="store_true") + args = parser.parse_args() + if args.download_only: + ClearVoice(task="speech_enhancement", model_names=[MODEL]) + return + if not all((args.input, args.speech, args.noise)): + parser.error("input, speech and noise output paths are required") + enhance(args.input, args.speech, args.noise) + + +if __name__ == "__main__": + main() diff --git a/platform/llama-dashboard/app.py b/platform/llama-dashboard/app.py index eb5f61e..2581b4c 100644 --- a/platform/llama-dashboard/app.py +++ b/platform/llama-dashboard/app.py @@ -635,7 +635,7 @@ HTML = r'''
Mike AI · Live Telemetry

Athena llama.cpp Dashboard

verbinde …
- +
Aktives Profil
–
Router wird abgefragt
Modell
–
–
CPU
–
–