feat: add speech noise separation
This commit is contained in:
1 parent
0fd1966ae3
commit
805228abfd
7 files changed
+135
-14
No files matched your search
@@ -4,8 +4,8 @@ Athena besitzt drei gegenseitig exklusive Betriebsmodi:
|
|||||||
|
|
||||||
- `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt.
|
- `llm`: ein llama.cpp-Profil und Qwen3-TTS laufen; Spezialdienste sind gestoppt.
|
||||||
- `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt.
|
- `music`: ACE-Step 1.5 XL-SFT läuft; alle LLM-, Bild-, TTS- und Separator-Worker sind gestoppt.
|
||||||
- `separation`: wahlweise BS-RoFormer für Gesang/Instrumental oder Demucs für
|
- `separation`: BS-RoFormer und Demucs trennen Musikspuren; ClearVoice trennt
|
||||||
vier beziehungsweise sechs Spuren; LLM, Bild, TTS und ACE-Step sind gestoppt.
|
Sprache von Hintergrundgeräuschen. LLM, Bild, TTS und ACE-Step sind gestoppt.
|
||||||
|
|
||||||
Die Zustandsmaschine lebt im Athena-Router. Das Dashboard und Chat-Clients wie
|
Die Zustandsmaschine lebt im Athena-Router. Das Dashboard und Chat-Clients wie
|
||||||
Hermes sind nur Bedienoberflächen derselben API. Der zuletzt aktive LLM-Modus
|
Hermes sind nur Bedienoberflächen derselben API. Der zuletzt aktive LLM-Modus
|
||||||
@@ -14,7 +14,7 @@ wird persistent gespeichert und beim Verlassen eines Spezialmodus wieder geladen
|
|||||||
## Bedienung
|
## Bedienung
|
||||||
|
|
||||||
Im Athena-Dashboard stehen **LLM-Betrieb**, **Musikstudio** und
|
Im Athena-Dashboard stehen **LLM-Betrieb**, **Musikstudio** und
|
||||||
**Stimmen trennen** bereit. Im Musikmodus werden zwei Oberflächen angeboten:
|
**Audio trennen** bereit. Im Musikmodus werden zwei Oberflächen angeboten:
|
||||||
|
|
||||||
- **Original UI · stabil** öffnet die zum laufenden ACE-Step-Image gehörende
|
- **Original UI · stabil** öffnet die zum laufenden ACE-Step-Image gehörende
|
||||||
Gradio-Oberfläche. Sie ist für Cover, Remix und erweiterte Workflows der
|
Gradio-Oberfläche. Sie ist für Cover, Remix und erweiterte Workflows der
|
||||||
@@ -37,14 +37,16 @@ Docker-Netz `mike-ai-music` auf `http://music-worker:7860` zu.
|
|||||||
Im Trennmodus öffnet das Dashboard die private Athena-Oberfläche unter
|
Im Trennmodus öffnet das Dashboard die private Athena-Oberfläche unter
|
||||||
`http://192.168.1.212:8007`. Sie nimmt WAV, FLAC, MP3, M4A und weitere
|
`http://192.168.1.212:8007`. Sie nimmt WAV, FLAC, MP3, M4A und weitere
|
||||||
übliche Formate an. Gewählt wird die herauszulösende Quelle: Gesang,
|
übliche Formate an. Gewählt wird die herauszulösende Quelle: Gesang,
|
||||||
Schlagzeug, Bass, Gitarre, Piano oder Sonstiges. Das ZIP enthält genau diese Zielspur und
|
Schlagzeug, Bass, Gitarre, Piano, Sonstiges oder gereinigte Sprache. Das ZIP enthält genau diese Zielspur und
|
||||||
eine zweite FLAC-Datei mit dem vollständigen Rest ohne die Zielspur. Gesang
|
eine zweite FLAC-Datei mit dem vollständigen Rest ohne die Zielspur. Gesang
|
||||||
nutzt BS-RoFormer Viperx 1297, Schlagzeug/Bass `htdemucs_ft` und
|
nutzt BS-RoFormer Viperx 1297, Schlagzeug/Bass `htdemucs_ft` und
|
||||||
Gitarre/Piano/Sonstiges experimentell `htdemucs_6s`. „Sonstiges“ ist dessen
|
Gitarre/Piano/Sonstiges experimentell `htdemucs_6s`. „Sonstiges“ ist dessen
|
||||||
gemischter `other`-Stem (unter anderem Synthesizer, Streicher, Bläser und Effekte),
|
gemischter `other`-Stem (unter anderem Synthesizer, Streicher, Bläser und Effekte),
|
||||||
nicht eine reine Synthesizer-Spur. Grundlage ist `audio-separator`
|
nicht eine reine Synthesizer-Spur. Sprache nutzt das 48-kHz-Modell
|
||||||
0.47.0. Die ältere API-Auswahl kompletter 2-/4-/6-Stem-Sätze bleibt
|
`MossFormer2_SE_48K`; der Download enthält `speech.flac` und
|
||||||
rückwärtskompatibel.
|
`hintergrund-ohne-sprache.flac`. Die Musiktrennung basiert auf
|
||||||
|
`audio-separator` 0.47.0. Die ältere API-Auswahl kompletter 2-/4-/6-Stem-Sätze
|
||||||
|
bleibt rückwärtskompatibel.
|
||||||
|
|
||||||
Hermes benötigt dafür kein Plugin. Exakt eingegebene Steuerbefehle werden vom
|
Hermes benötigt dafür kein Plugin. Exakt eingegebene Steuerbefehle werden vom
|
||||||
Router lokal beantwortet, auch wenn gerade kein LLM geladen ist:
|
Router lokal beantwortet, auch wenn gerade kein LLM geladen ist:
|
||||||
|
|||||||
@@ -11,12 +11,20 @@ RUN python -m pip install --no-cache-dir \
|
|||||||
"python-multipart==0.0.20" \
|
"python-multipart==0.0.20" \
|
||||||
"uvicorn[standard]==0.35.0"
|
"uvicorn[standard]==0.35.0"
|
||||||
|
|
||||||
|
# audio-separator 0.47 requires NumPy 2 while ClearVoice 0.1.2 still pins
|
||||||
|
# NumPy 1.x. Keep ClearVoice in a small overlay venv but share the image's
|
||||||
|
# CUDA-enabled PyTorch installation instead of duplicating it.
|
||||||
|
RUN python -m venv --system-site-packages /opt/clearvoice-venv \
|
||||||
|
&& /opt/clearvoice-venv/bin/python -m pip install --no-cache-dir \
|
||||||
|
"clearvoice==0.1.2" \
|
||||||
|
"numpy>=1.24.3,<2.0"
|
||||||
|
|
||||||
WORKDIR /app
|
WORKDIR /app
|
||||||
COPY app.py index.html ./
|
COPY app.py index.html speech_enhance.py ./
|
||||||
|
|
||||||
ENV MODEL_FILENAME=model_bs_roformer_ep_317_sdr_12.9755.ckpt \
|
ENV MODEL_FILENAME=model_bs_roformer_ep_317_sdr_12.9755.ckpt \
|
||||||
MODEL_DIR=/models \
|
MODEL_DIR=/models \
|
||||||
JOB_DIR=/data/jobs
|
JOB_DIR=/data/jobs
|
||||||
|
|
||||||
EXPOSE 8080
|
EXPOSE 8080
|
||||||
CMD ["sh", "-c", "for model in \"$MODEL_FILENAME\" htdemucs_ft.yaml htdemucs_6s.yaml; do audio-separator --model_filename \"$model\" --model_file_dir \"$MODEL_DIR\" --download_model_only || exit 1; done; exec uvicorn app:app --host 0.0.0.0 --port 8080 --workers 1"]
|
CMD ["sh", "-c", "mkdir -p \"$MODEL_DIR/clearvoice\" && ln -sfn \"$MODEL_DIR/clearvoice\" /app/checkpoints && for model in \"$MODEL_FILENAME\" htdemucs_ft.yaml htdemucs_6s.yaml; do audio-separator --model_filename \"$model\" --model_file_dir \"$MODEL_DIR\" --download_model_only || exit 1; done; /opt/clearvoice-venv/bin/python /app/speech_enhance.py --download-only || exit 1; exec uvicorn app:app --host 0.0.0.0 --port 8080 --workers 1"]
|
||||||
@@ -15,6 +15,10 @@ Rest ohne dieses Ziel.
|
|||||||
liegt unter der spezialisierten Gesangstrennung.
|
liegt unter der spezialisierten Gesangstrennung.
|
||||||
- **Sonstiges:** der `other`-Stem von `htdemucs_6s.yaml`. Er bündelt unter anderem
|
- **Sonstiges:** der `other`-Stem von `htdemucs_6s.yaml`. Er bündelt unter anderem
|
||||||
Synthesizer, Streicher, Bläser und Effekte und ist keine reine Synthesizer-Spur.
|
Synthesizer, Streicher, Bläser und Effekte und ist keine reine Synthesizer-Spur.
|
||||||
|
- **Sprache / Hintergrund:** ClearVoice `MossFormer2_SE_48K` (Apache-2.0)
|
||||||
|
verbessert Sprache bei 48 kHz. Die zweite Spur ist das vom Originalsignal
|
||||||
|
abgezogene Sprachsignal und enthält den verbleibenden Hintergrund. Stereo wird
|
||||||
|
kanalweise verarbeitet und anschließend wieder zusammengesetzt.
|
||||||
- GPU: RTX 5080; LLM, Bildmodelle, TTS und ACE-Step sind dabei verriegelt.
|
- GPU: RTX 5080; LLM, Bildmodelle, TTS und ACE-Step sind dabei verriegelt.
|
||||||
- Privat erreichbar: `http://192.168.1.212:8007/`
|
- Privat erreichbar: `http://192.168.1.212:8007/`
|
||||||
|
|
||||||
@@ -26,6 +30,6 @@ weitere Zielmodelle. Die Oberfläche bietet stattdessen den ehrlich benannten,
|
|||||||
gemischten `other`-Stem als **Sonstiges** an.
|
gemischten `other`-Stem als **Sonstiges** an.
|
||||||
|
|
||||||
Die API erwartet `multipart/form-data` mit `file` und optional `target`:
|
Die API erwartet `multipart/form-data` mit `file` und optional `target`:
|
||||||
`vocals` (Standard), `drums`, `bass`, `guitar`, `piano` oder `other`. Das ältere Feld
|
`vocals` (Standard), `drums`, `bass`, `guitar`, `piano`, `other` oder `speech`. Das ältere Feld
|
||||||
`mode` mit `vocals`, `four_stem` oder `six_stem` bleibt für vorhandene Clients
|
`mode` mit `vocals`, `four_stem` oder `six_stem` bleibt für vorhandene Clients
|
||||||
erhalten und liefert weiterhin alle Modell-Stems.
|
erhalten und liefert weiterhin alle Modell-Stems.
|
||||||
@@ -26,6 +26,7 @@ MODES = {
|
|||||||
"vocals": {"model": MODEL, "stems": ("vocals", "instrumental"), "archive": "athena-vocals-instrumental.zip", "engine": "mdxc"},
|
"vocals": {"model": MODEL, "stems": ("vocals", "instrumental"), "archive": "athena-vocals-instrumental.zip", "engine": "mdxc"},
|
||||||
"four_stem": {"model": "htdemucs_ft.yaml", "stems": ("vocals", "drums", "bass", "other"), "archive": "athena-4-stems.zip", "engine": "demucs"},
|
"four_stem": {"model": "htdemucs_ft.yaml", "stems": ("vocals", "drums", "bass", "other"), "archive": "athena-4-stems.zip", "engine": "demucs"},
|
||||||
"six_stem": {"model": "htdemucs_6s.yaml", "stems": ("vocals", "drums", "bass", "guitar", "piano", "other"), "archive": "athena-6-stems-experimental.zip", "engine": "demucs"},
|
"six_stem": {"model": "htdemucs_6s.yaml", "stems": ("vocals", "drums", "bass", "guitar", "piano", "other"), "archive": "athena-6-stems-experimental.zip", "engine": "demucs"},
|
||||||
|
"speech": {"model": "MossFormer2_SE_48K", "stems": ("speech", "noise"), "archive": "athena-sprache-und-hintergrund.zip", "engine": "clearvoice"},
|
||||||
}
|
}
|
||||||
TARGETS = {
|
TARGETS = {
|
||||||
"vocals": {"mode": "vocals", "stem": "vocals", "remainder": "instrumental", "archive": "athena-gesang-und-rest.zip", "rest_file": "instrumental.flac"},
|
"vocals": {"mode": "vocals", "stem": "vocals", "remainder": "instrumental", "archive": "athena-gesang-und-rest.zip", "rest_file": "instrumental.flac"},
|
||||||
@@ -34,6 +35,7 @@ TARGETS = {
|
|||||||
"guitar": {"mode": "six_stem", "stem": "guitar", "archive": "athena-gitarre-und-rest.zip", "rest_file": "rest-ohne-gitarre.flac"},
|
"guitar": {"mode": "six_stem", "stem": "guitar", "archive": "athena-gitarre-und-rest.zip", "rest_file": "rest-ohne-gitarre.flac"},
|
||||||
"piano": {"mode": "six_stem", "stem": "piano", "archive": "athena-piano-und-rest.zip", "rest_file": "rest-ohne-piano.flac"},
|
"piano": {"mode": "six_stem", "stem": "piano", "archive": "athena-piano-und-rest.zip", "rest_file": "rest-ohne-piano.flac"},
|
||||||
"other": {"mode": "six_stem", "stem": "other", "archive": "athena-sonstiges-und-rest.zip", "rest_file": "rest-ohne-sonstiges.flac"},
|
"other": {"mode": "six_stem", "stem": "other", "archive": "athena-sonstiges-und-rest.zip", "rest_file": "rest-ohne-sonstiges.flac"},
|
||||||
|
"speech": {"mode": "speech", "stem": "speech", "remainder": "noise", "archive": "athena-sprache-und-hintergrund.zip", "rest_file": "hintergrund-ohne-sprache.flac"},
|
||||||
}
|
}
|
||||||
|
|
||||||
app = FastAPI(title="Athena Stem Separator", version="2.0")
|
app = FastAPI(title="Athena Stem Separator", version="2.0")
|
||||||
@@ -46,7 +48,14 @@ def index() -> str:
|
|||||||
|
|
||||||
@app.get("/health")
|
@app.get("/health")
|
||||||
def health() -> dict:
|
def health() -> dict:
|
||||||
available = {name: (MODEL_DIR / mode["model"]).exists() for name, mode in MODES.items()}
|
available = {
|
||||||
|
name: (
|
||||||
|
(MODEL_DIR / "clearvoice" / mode["model"] / "last_best_checkpoint").exists()
|
||||||
|
if mode["engine"] == "clearvoice"
|
||||||
|
else (MODEL_DIR / mode["model"]).exists()
|
||||||
|
)
|
||||||
|
for name, mode in MODES.items()
|
||||||
|
}
|
||||||
return {
|
return {
|
||||||
"status": "ok" if all(available.values()) else "starting",
|
"status": "ok" if all(available.values()) else "starting",
|
||||||
"models": {name: mode["model"] for name, mode in MODES.items()},
|
"models": {name: mode["model"] for name, mode in MODES.items()},
|
||||||
@@ -62,6 +71,20 @@ def _cleanup(path: Path) -> None:
|
|||||||
|
|
||||||
|
|
||||||
def _run_separator(input_path: Path, output_dir: Path, mode: dict) -> None:
|
def _run_separator(input_path: Path, output_dir: Path, mode: dict) -> None:
|
||||||
|
if mode["engine"] == "clearvoice":
|
||||||
|
completed = subprocess.run(
|
||||||
|
[
|
||||||
|
"/opt/clearvoice-venv/bin/python", "/app/speech_enhance.py", str(input_path),
|
||||||
|
str(output_dir / "speech.flac"), str(output_dir / "noise.flac"),
|
||||||
|
],
|
||||||
|
capture_output=True,
|
||||||
|
text=True,
|
||||||
|
timeout=7200,
|
||||||
|
)
|
||||||
|
if completed.returncode:
|
||||||
|
detail = (completed.stderr or completed.stdout or "unknown ClearVoice error")[-4000:]
|
||||||
|
raise RuntimeError(detail)
|
||||||
|
return
|
||||||
args = [
|
args = [
|
||||||
"audio-separator", str(input_path),
|
"audio-separator", str(input_path),
|
||||||
"--model_filename", mode["model"],
|
"--model_filename", mode["model"],
|
||||||
|
|||||||
@@ -1,16 +1,19 @@
|
|||||||
<!doctype html>
|
<!doctype html>
|
||||||
<html lang="de"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1">
|
<html lang="de"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1">
|
||||||
<title>Athena · Spuren herauslösen</title><style>
|
<title>Athena · Spuren herauslösen</title><style>
|
||||||
:root{color-scheme:dark;--bg:#07111c;--card:#101d2b;--line:#26384b;--cyan:#48d7f5;--mint:#63e6be;--text:#ecf5ff;--muted:#91a4b7}*{box-sizing:border-box}body{margin:0;background:radial-gradient(circle at 20% 0,#142a42 0,#07111c 42%);font:16px system-ui,sans-serif;color:var(--text);min-height:100vh;display:grid;place-items:center;padding:24px}.card{width:min(880px,100%);padding:32px;border:1px solid var(--line);border-radius:22px;background:rgba(16,29,43,.96);box-shadow:0 25px 70px #0008}.eyebrow{color:var(--cyan);font-weight:800;letter-spacing:.14em;text-transform:uppercase;font-size:12px}h1{font-size:clamp(30px,5vw,52px);margin:.3em 0 .15em}p{color:var(--muted);line-height:1.6}.targets{display:grid;grid-template-columns:repeat(3,1fr);gap:9px;margin:24px 0}.target{display:block;border:1px solid var(--line);border-radius:14px;padding:14px 10px;text-align:center;cursor:pointer}.target:has(input:checked){border-color:var(--cyan);background:#48d7f510;box-shadow:0 0 0 1px #48d7f528}.target input{display:none}.target b,.target span{display:block}.target span{color:var(--muted);font-size:12px;margin-top:5px;line-height:1.35}.drop{display:block;margin:20px 0;padding:36px 24px;border:2px dashed #3f5a72;border-radius:18px;text-align:center;cursor:pointer;transition:.2s}.drop:hover,.drop.drag{border-color:var(--cyan);background:#48d7f50b}.drop input{display:none}.file{color:var(--mint);font-weight:700;margin-top:8px}button{width:100%;border:0;border-radius:13px;padding:15px;font-weight:800;font-size:16px;background:linear-gradient(90deg,var(--cyan),var(--mint));color:#05202a;cursor:pointer}button:disabled{opacity:.45;cursor:not-allowed}.status{min-height:28px;margin-top:18px;color:var(--muted)}.bar{height:7px;background:#07111c;border-radius:9px;overflow:hidden;margin-top:12px}.fill{height:100%;width:0;background:linear-gradient(90deg,var(--cyan),var(--mint));transition:.4s}.run .fill{width:85%;animation:pulse 1.5s infinite alternate}@keyframes pulse{to{opacity:.45}}small{display:block;color:#71879a;margin-top:20px}@media(max-width:760px){.targets{grid-template-columns:repeat(2,1fr)}}
|
:root{color-scheme:dark;--bg:#07111c;--card:#101d2b;--line:#26384b;--cyan:#48d7f5;--mint:#63e6be;--text:#ecf5ff;--muted:#91a4b7}*{box-sizing:border-box}body{margin:0;background:radial-gradient(circle at 20% 0,#142a42 0,#07111c 42%);font:16px system-ui,sans-serif;color:var(--text);min-height:100vh;display:grid;place-items:center;padding:24px}.card{width:min(880px,100%);padding:32px;border:1px solid var(--line);border-radius:22px;background:rgba(16,29,43,.96);box-shadow:0 25px 70px #0008}.eyebrow{color:var(--cyan);font-weight:800;letter-spacing:.14em;text-transform:uppercase;font-size:12px}h1{font-size:clamp(30px,5vw,52px);margin:.3em 0 .15em}p{color:var(--muted);line-height:1.6}.targets{display:grid;grid-template-columns:repeat(3,1fr);gap:9px;margin:24px 0}.target{display:block;border:1px solid var(--line);border-radius:14px;padding:14px 10px;text-align:center;cursor:pointer}.target:has(input:checked){border-color:var(--cyan);background:#48d7f510;box-shadow:0 0 0 1px #48d7f528}.target input{display:none}.target b,.target span{display:block}.target span{color:var(--muted);font-size:12px;margin-top:5px;line-height:1.35}.section{grid-column:1/-1;color:var(--cyan);font-size:12px;font-weight:800;letter-spacing:.12em;text-transform:uppercase;margin-top:8px}.drop{display:block;margin:20px 0;padding:36px 24px;border:2px dashed #3f5a72;border-radius:18px;text-align:center;cursor:pointer;transition:.2s}.drop:hover,.drop.drag{border-color:var(--cyan);background:#48d7f50b}.drop input{display:none}.file{color:var(--mint);font-weight:700;margin-top:8px}button{width:100%;border:0;border-radius:13px;padding:15px;font-weight:800;font-size:16px;background:linear-gradient(90deg,var(--cyan),var(--mint));color:#05202a;cursor:pointer}button:disabled{opacity:.45;cursor:not-allowed}.status{min-height:28px;margin-top:18px;color:var(--muted)}.bar{height:7px;background:#07111c;border-radius:9px;overflow:hidden;margin-top:12px}.fill{height:100%;width:0;background:linear-gradient(90deg,var(--cyan),var(--mint));transition:.4s}.run .fill{width:85%;animation:pulse 1.5s infinite alternate}@keyframes pulse{to{opacity:.45}}small{display:block;color:#71879a;margin-top:20px}@media(max-width:760px){.targets{grid-template-columns:repeat(2,1fr)}}
|
||||||
</style></head><body><main class="card"><div class="eyebrow">Athena Audio Lab</div><h1>Was möchtest du herauslösen?</h1><p>Der Download enthält immer die gewählte Spur separat und zusätzlich den vollständigen Rest ohne diese Spur.</p>
|
</style></head><body><main class="card"><div class="eyebrow">Athena Audio Lab</div><h1>Was möchtest du herauslösen?</h1><p>Der Download enthält immer die gewählte Spur separat und zusätzlich den vollständigen Rest ohne diese Spur.</p>
|
||||||
<div class="targets">
|
<div class="targets">
|
||||||
|
<div class="section">Musik</div>
|
||||||
<label class="target"><input type="radio" name="target" value="vocals" checked><b>Gesang</b><span>BS‑RoFormer<br>beste Qualität</span></label>
|
<label class="target"><input type="radio" name="target" value="vocals" checked><b>Gesang</b><span>BS‑RoFormer<br>beste Qualität</span></label>
|
||||||
<label class="target"><input type="radio" name="target" value="drums"><b>Schlagzeug</b><span>HTDemucs FT</span></label>
|
<label class="target"><input type="radio" name="target" value="drums"><b>Schlagzeug</b><span>HTDemucs FT</span></label>
|
||||||
<label class="target"><input type="radio" name="target" value="bass"><b>Bass</b><span>HTDemucs FT</span></label>
|
<label class="target"><input type="radio" name="target" value="bass"><b>Bass</b><span>HTDemucs FT</span></label>
|
||||||
<label class="target"><input type="radio" name="target" value="guitar"><b>Gitarre</b><span>HTDemucs 6s<br>experimentell</span></label>
|
<label class="target"><input type="radio" name="target" value="guitar"><b>Gitarre</b><span>HTDemucs 6s<br>experimentell</span></label>
|
||||||
<label class="target"><input type="radio" name="target" value="piano"><b>Piano</b><span>HTDemucs 6s<br>experimentell</span></label>
|
<label class="target"><input type="radio" name="target" value="piano"><b>Piano</b><span>HTDemucs 6s<br>experimentell</span></label>
|
||||||
<label class="target"><input type="radio" name="target" value="other"><b>Sonstiges</b><span>Synths, Streicher etc.<br>gemischte Spur</span></label>
|
<label class="target"><input type="radio" name="target" value="other"><b>Sonstiges</b><span>Synths, Streicher etc.<br>gemischte Spur</span></label>
|
||||||
|
<div class="section">Sprache und Geräusche</div>
|
||||||
|
<label class="target"><input type="radio" name="target" value="speech"><b>Sprache reinigen</b><span>MossFormer2 · 48 kHz<br>Sprache + Hintergrund</span></label>
|
||||||
</div>
|
</div>
|
||||||
<label class="drop" id="drop">Audio auswählen oder hier ablegen<input id="file" type="file" accept="audio/*"><div class="file" id="name">Noch keine Datei gewählt</div></label><button id="start" disabled>Ausgewählte Spur und Rest erzeugen</button><div class="status" id="status">Bereit.</div><div class="bar" id="bar"><div class="fill"></div></div><small>Alles läuft lokal auf Athena. Synthesizer, Streicher sowie elektrische und akustische Gitarre separat benötigen zusätzliche Spezialmodelle.</small></main><script>
|
<label class="drop" id="drop">Audio auswählen oder hier ablegen<input id="file" type="file" accept="audio/*"><div class="file" id="name">Noch keine Datei gewählt</div></label><button id="start" disabled>Ausgewählte Spur und Rest erzeugen</button><div class="status" id="status">Bereit.</div><div class="bar" id="bar"><div class="fill"></div></div><small>Alles läuft lokal auf Athena. Synthesizer, Streicher sowie elektrische und akustische Gitarre separat benötigen zusätzliche Spezialmodelle.</small></main><script>
|
||||||
const file=document.querySelector('#file'),drop=document.querySelector('#drop'),name=document.querySelector('#name'),start=document.querySelector('#start'),status=document.querySelector('#status'),bar=document.querySelector('#bar');let selected;const names={vocals:'gesang',drums:'schlagzeug',bass:'bass',guitar:'gitarre',piano:'piano',other:'sonstiges'};function choose(f){selected=f;name.textContent=f?`${f.name} · ${(f.size/1048576).toFixed(1)} MiB`:'Noch keine Datei gewählt';start.disabled=!f}file.onchange=()=>choose(file.files[0]);drop.ondragover=e=>{e.preventDefault();drop.classList.add('drag')};drop.ondragleave=()=>drop.classList.remove('drag');drop.ondrop=e=>{e.preventDefault();drop.classList.remove('drag');choose(e.dataTransfer.files[0])};start.onclick=async()=>{const target=document.querySelector('input[name=target]:checked').value;start.disabled=true;bar.classList.add('run');status.textContent='Modell löst die gewählte Spur heraus – das kann einige Minuten dauern …';let body=new FormData();body.append('file',selected);body.append('target',target);try{let r=await fetch('/v1/separate',{method:'POST',body});if(!r.ok)throw Error((await r.json()).detail||`HTTP ${r.status}`);let blob=await r.blob(),a=document.createElement('a');a.href=URL.createObjectURL(blob);a.download=`athena-${names[target]}-und-rest.zip`;a.click();setTimeout(()=>URL.revokeObjectURL(a.href),5000);status.textContent='Fertig – ZIP mit der ausgewählten Spur und dem Rest wurde geladen.'}catch(e){status.textContent=`Fehler: ${e.message}`}finally{bar.classList.remove('run');start.disabled=false}};
|
const file=document.querySelector('#file'),drop=document.querySelector('#drop'),name=document.querySelector('#name'),start=document.querySelector('#start'),status=document.querySelector('#status'),bar=document.querySelector('#bar');let selected;const names={vocals:'gesang',drums:'schlagzeug',bass:'bass',guitar:'gitarre',piano:'piano',other:'sonstiges',speech:'sprache-und-hintergrund'};function choose(f){selected=f;name.textContent=f?`${f.name} · ${(f.size/1048576).toFixed(1)} MiB`:'Noch keine Datei gewählt';start.disabled=!f}file.onchange=()=>choose(file.files[0]);drop.ondragover=e=>{e.preventDefault();drop.classList.add('drag')};drop.ondragleave=()=>drop.classList.remove('drag');drop.ondrop=e=>{e.preventDefault();drop.classList.remove('drag');choose(e.dataTransfer.files[0])};start.onclick=async()=>{const target=document.querySelector('input[name=target]:checked').value;start.disabled=true;bar.classList.add('run');status.textContent=target==='speech'?'MossFormer2 trennt Sprache und Hintergrund – das kann einige Minuten dauern …':'Modell löst die gewählte Spur heraus – das kann einige Minuten dauern …';let body=new FormData();body.append('file',selected);body.append('target',target);try{let r=await fetch('/v1/separate',{method:'POST',body});if(!r.ok)throw Error((await r.json()).detail||`HTTP ${r.status}`);let blob=await r.blob(),a=document.createElement('a');a.href=URL.createObjectURL(blob);a.download=`athena-${names[target]}-und-rest.zip`;a.click();setTimeout(()=>URL.revokeObjectURL(a.href),5000);status.textContent='Fertig – ZIP mit der ausgewählten Spur und dem Rest wurde geladen.'}catch(e){status.textContent=`Fehler: ${e.message}`}finally{bar.classList.remove('run');start.disabled=false}};
|
||||||
</script></body></html>
|
</script></body></html>
|
||||||
@@ -0,0 +1,81 @@
|
|||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import subprocess
|
||||||
|
import tempfile
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import soundfile as sf
|
||||||
|
from clearvoice import ClearVoice
|
||||||
|
|
||||||
|
|
||||||
|
MODEL = "MossFormer2_SE_48K"
|
||||||
|
SAMPLE_RATE = 48_000
|
||||||
|
|
||||||
|
|
||||||
|
def convert_input(source: Path, target: Path) -> None:
|
||||||
|
completed = subprocess.run(
|
||||||
|
[
|
||||||
|
"ffmpeg", "-hide_banner", "-loglevel", "error", "-y",
|
||||||
|
"-i", str(source), "-vn", "-ar", str(SAMPLE_RATE),
|
||||||
|
"-c:a", "pcm_f32le", str(target),
|
||||||
|
],
|
||||||
|
capture_output=True,
|
||||||
|
text=True,
|
||||||
|
timeout=1800,
|
||||||
|
)
|
||||||
|
if completed.returncode:
|
||||||
|
raise RuntimeError(completed.stderr[-4000:] or "ffmpeg input conversion failed")
|
||||||
|
|
||||||
|
|
||||||
|
def enhance(source: Path, speech_path: Path, noise_path: Path) -> None:
|
||||||
|
with tempfile.TemporaryDirectory(prefix="clearvoice-") as temp_dir:
|
||||||
|
converted = Path(temp_dir) / "input-48k.wav"
|
||||||
|
convert_input(source, converted)
|
||||||
|
audio, sample_rate = sf.read(converted, dtype="float32", always_2d=True)
|
||||||
|
if sample_rate != SAMPLE_RATE:
|
||||||
|
raise RuntimeError(f"unexpected sample rate: {sample_rate}")
|
||||||
|
|
||||||
|
model = ClearVoice(task="speech_enhancement", model_names=[MODEL])
|
||||||
|
# Use ClearVoice's file-I/O path so recordings longer than its 20-second
|
||||||
|
# one-pass window are segmented correctly. Run each channel separately
|
||||||
|
# because the enhancement network itself is mono, then restore stereo.
|
||||||
|
channels = []
|
||||||
|
for channel_index in range(audio.shape[1]):
|
||||||
|
channel_path = Path(temp_dir) / f"channel-{channel_index}.wav"
|
||||||
|
sf.write(channel_path, audio[:, channel_index], SAMPLE_RATE, subtype="FLOAT")
|
||||||
|
result = np.asarray(model(str(channel_path), False), dtype=np.float32).squeeze()
|
||||||
|
if result.ndim != 1:
|
||||||
|
raise RuntimeError(f"unexpected ClearVoice output shape: {result.shape}")
|
||||||
|
channels.append(result)
|
||||||
|
enhanced = np.column_stack(channels)
|
||||||
|
|
||||||
|
length = min(len(audio), len(enhanced))
|
||||||
|
original = audio[:length]
|
||||||
|
speech = enhanced[:length]
|
||||||
|
noise = original - speech
|
||||||
|
|
||||||
|
# FLAC does not support floating-point samples. PCM_24 retains ample
|
||||||
|
# headroom and avoids the invalid FLOAT/FLAC combination in libsndfile.
|
||||||
|
sf.write(speech_path, np.clip(speech, -1.0, 1.0), SAMPLE_RATE, format="FLAC", subtype="PCM_24")
|
||||||
|
sf.write(noise_path, np.clip(noise, -1.0, 1.0), SAMPLE_RATE, format="FLAC", subtype="PCM_24")
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> None:
|
||||||
|
parser = argparse.ArgumentParser()
|
||||||
|
parser.add_argument("input", nargs="?", type=Path)
|
||||||
|
parser.add_argument("speech", nargs="?", type=Path)
|
||||||
|
parser.add_argument("noise", nargs="?", type=Path)
|
||||||
|
parser.add_argument("--download-only", action="store_true")
|
||||||
|
args = parser.parse_args()
|
||||||
|
if args.download_only:
|
||||||
|
ClearVoice(task="speech_enhancement", model_names=[MODEL])
|
||||||
|
return
|
||||||
|
if not all((args.input, args.speech, args.noise)):
|
||||||
|
parser.error("input, speech and noise output paths are required")
|
||||||
|
enhance(args.input, args.speech, args.noise)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@@ -635,7 +635,7 @@ HTML = r'''<!doctype html>
|
|||||||
</style></head><body><main>
|
</style></head><body><main>
|
||||||
<div class="top"><div><div class="eyebrow">Mike AI · Live Telemetry</div><h1>Athena llama.cpp Dashboard</h1></div><div class="live"><span class="dot" id="dot"></span><span id="updated">verbinde …</span></div></div>
|
<div class="top"><div><div class="eyebrow">Mike AI · Live Telemetry</div><h1>Athena llama.cpp Dashboard</h1></div><div class="live"><span class="dot" id="dot"></span><span id="updated">verbinde …</span></div></div>
|
||||||
<section class="grid">
|
<section class="grid">
|
||||||
<article class="card span12"><div class="mode-row"><div><div class="label">Athena Betriebsmodus</div><div class="value" id="operatingMode">–</div><div class="sub" id="modeStatus">Status wird geladen …</div></div><div class="mode-buttons"><button id="llmMode" onclick="setMode('llm')">LLM-Betrieb</button><button id="musicMode" onclick="setMode('music')">Musikstudio</button><button id="separationMode" onclick="setMode('separation')">Stimmen trennen</button><span id="musicOpen" hidden><a class="stable" href="__MUSIC_ORIGINAL_UI_URL__" target="_blank" rel="noopener">Original UI · stabil</a><a class="experimental" href="__MUSIC_COMMUNITY_UI_URL__" target="_blank" rel="noopener">Community UI · experimentell</a></span><span id="separatorOpen" hidden><a class="stable" href="__SEPARATOR_UI_URL__" target="_blank" rel="noopener">Separator öffnen</a></span></div></div></article>
|
<article class="card span12"><div class="mode-row"><div><div class="label">Athena Betriebsmodus</div><div class="value" id="operatingMode">–</div><div class="sub" id="modeStatus">Status wird geladen …</div></div><div class="mode-buttons"><button id="llmMode" onclick="setMode('llm')">LLM-Betrieb</button><button id="musicMode" onclick="setMode('music')">Musikstudio</button><button id="separationMode" onclick="setMode('separation')">Audio trennen</button><span id="musicOpen" hidden><a class="stable" href="__MUSIC_ORIGINAL_UI_URL__" target="_blank" rel="noopener">Original UI · stabil</a><a class="experimental" href="__MUSIC_COMMUNITY_UI_URL__" target="_blank" rel="noopener">Community UI · experimentell</a></span><span id="separatorOpen" hidden><a class="stable" href="__SEPARATOR_UI_URL__" target="_blank" rel="noopener">Separator öffnen</a></span></div></div></article>
|
||||||
<article class="card span3"><div class="label">Aktives Profil</div><div class="value" id="profile">–</div><div class="sub" id="profileSub">Router wird abgefragt</div></article>
|
<article class="card span3"><div class="label">Aktives Profil</div><div class="value" id="profile">–</div><div class="sub" id="profileSub">Router wird abgefragt</div></article>
|
||||||
<article class="card span3"><div class="label">Modell</div><div class="value" id="model">–</div><div class="sub" id="modelSub">–</div></article>
|
<article class="card span3"><div class="label">Modell</div><div class="value" id="model">–</div><div class="sub" id="modelSub">–</div></article>
|
||||||
<article class="card span3"><div class="label">CPU</div><div class="value" id="cpu">–</div><div class="bar"><div class="fill" id="cpuBar"></div></div><div class="sub" id="load">–</div></article>
|
<article class="card span3"><div class="label">CPU</div><div class="value" id="cpu">–</div><div class="bar"><div class="fill" id="cpuBar"></div></div><div class="sub" id="load">–</div></article>
|
||||||
|
|||||||
Reference in new issue
Block a user