Enable llama.cpp builds, release management and reference model auto-fit
This commit is contained in:
+63
-1
@@ -7,7 +7,7 @@ Editor/SSH-Client und Browser. Die Anwendung und ihre Tests laufen auf Athena.
|
|||||||
- Eigener Container: `athena-deck-dev`
|
- Eigener Container: `athena-deck-dev`
|
||||||
- Zustand und Installationsmanifest: `/opt/athena-deck-dev/runtime`
|
- Zustand und Installationsmanifest: `/opt/athena-deck-dev/runtime`
|
||||||
- Serverbindung: ausschließlich `127.0.0.1:8108`
|
- Serverbindung: ausschließlich `127.0.0.1:8108`
|
||||||
- GPU-Zugriff: NVIDIA `utility` für Telemetrie; keine Modellstarts.
|
- GPU-Zugriff: NVIDIA `compute,utility` für Telemetrie und Fit-Prüfung; keine Modellstarts.
|
||||||
|
|
||||||
Vom Arbeitsplatz:
|
Vom Arbeitsplatz:
|
||||||
|
|
||||||
@@ -72,3 +72,65 @@ noch keine vollständige Installation eines mehrteiligen Modells dar.
|
|||||||
|
|
||||||
Update-Sicherungen enthalten den Konfigurationszustand, nicht die heruntergeladenen
|
Update-Sicherungen enthalten den Konfigurationszustand, nicht die heruntergeladenen
|
||||||
Modelldateien. Diese bleiben im persistenten Zustandsverzeichnis erhalten.
|
Modelldateien. Diese bleiben im persistenten Zustandsverzeichnis erhalten.
|
||||||
|
|
||||||
|
## Native Laufzeitverwaltung
|
||||||
|
|
||||||
|
`runtime.py` verwendet ausschließlich native Prozesse (git, cmake, llama-server),
|
||||||
|
keine Docker-API. Die Testverpackung enthält CUDA 12.8.1 und Build-Werkzeuge;
|
||||||
|
Debian-Host und Treiber werden nicht verändert. Build-Last ist auf 2 CPUs,
|
||||||
|
8 GiB RAM und maximal zwei Compiler-Jobs begrenzt. Builds liegen unter
|
||||||
|
`state/runtime`, werden nicht in Konfigurations-Backups dupliziert und bleiben
|
||||||
|
bei Updates erhalten. Geprüfte Versionen können als Standard ausgewählt und
|
||||||
|
zurückgeschaltet werden; das startet kein Modell.
|
||||||
|
|
||||||
|
GET `/api/v1/runtime`, `/prerequisites`, `/releases`, `/log`; POST `/build`
|
||||||
|
(revision, backend, jobs), `/cancel` ({}), `/activate` (build_id), `/rollback` ({}),
|
||||||
|
jeweils unter `/api/v1/runtime`. Nur Administratorsitzungen sind zugelassen.
|
||||||
|
Release-Quellen: https://github.com/ggml-org/llama.cpp/releases .
|
||||||
|
|
||||||
|
Die feste VRAM-Reserve entfällt. Die Buildprüfung ermittelt die tatsächliche
|
||||||
|
Unterstützung von `--fit`; modellbezogene Layer-/Kontextprognose wird über das unten beschriebene
|
||||||
|
Fit-Werkzeug bereitgestellt. Modellstarts sind noch nicht angebunden. Keine erfundenen Schätzungen,
|
||||||
|
keine absichtlichen OOM-Versuche und keine Eingriffe in produktive Worker.
|
||||||
|
|
||||||
|
### Auto-Einpassung mit den bisherigen Routermodellen
|
||||||
|
|
||||||
|
`llama-fit-params` wird mitgebaut. GET `/api/v1/runtime/references` liefert die
|
||||||
|
am 28.09.2026 aus dem alten Router gelesenen Profile Fast/Medium/Large/Ultra/
|
||||||
|
Uncensored. POST `/api/v1/runtime/fit` nimmt profile, context und slots entgegen.
|
||||||
|
Die drei GGUF-Dateien werden einzeln und nur lesend in den Testcontainer
|
||||||
|
eingebunden (`reference_models` im Entwicklungsmanifest). Keine Kopie der
|
||||||
|
Gewichte, kein Zugriff auf Prompts oder Anwendungslogs.
|
||||||
|
|
||||||
|
Das offizielle Fit-Werkzeug verwendet `no_alloc` für Modellgewichte und liefert
|
||||||
|
eine Prognose anhand des aktuellen freien Speichers, q4_0-KV-Cache und explizitem
|
||||||
|
Kontext/Slots. Die automatische Reserve beträgt als Sicherheitsregel 5 % des jeweiligen
|
||||||
|
Gesamt-VRAM, mindestens 512 MiB. Karten mit weniger als 10 % freiem VRAM
|
||||||
|
(mindestens 1 GiB) werden vor CUDA-Initialisierung aus der Prüfung ausgeschlossen.
|
||||||
|
Diese Regeln sind Sicherheitsmargen, keine gemessene exakte Modellkapazität. Es findet kein OOM-Stresstest und kein Modellstart statt.
|
||||||
|
Eine Prognose ist keine Garantie unter wechselnder Parallelbelegung. Der native
|
||||||
|
Betrieb kann dieselben Dateien über DECK_REFERENCE_MODELS bereitstellen.
|
||||||
|
|
||||||
|
Fit-Adapter: Das Upstream-Fit-CLI bietet die Serveroption `--kv-unified` nicht an.
|
||||||
|
Deck setzt deshalb beim Build in `tools/fit-params/fit-params.cpp` vor der
|
||||||
|
Backend-Initialisierung explizit `params.kv_unified = true`, entsprechend den
|
||||||
|
bestehenden Profilen. Die Anpassung ist als `shared-kv-pool-v1` im Buildstand
|
||||||
|
vermerkt; die eigentliche Fit-Berechnung bleibt upstream. Fehlt der erwartete
|
||||||
|
Quellanker, bricht der Build ab, statt eine andere Semantik zu verwenden.
|
||||||
|
|
||||||
|
Die Fit-Prognose gilt zunächst für das Textmodell. Vision-Projektoren und MTP
|
||||||
|
der bisherigen Routerprofile sind nicht eingerechnet. Außerdem ersetzt sie keine
|
||||||
|
RAM-/Lastprüfung eines späteren Worker-Starts im begrenzten Testcontainer.
|
||||||
|
|
||||||
|
Verifikation: CUDA-Release b11229 (c2a9e1606807970f4ee3167bacd951699c89caea)
|
||||||
|
für 86/120 erfolgreich gebaut. Medium, 160000 Kontext, zwei Slots: echte
|
||||||
|
Fit-Prognose erfolgreich. Bei damaliger Parallelbelegung 0 GPU-Modelllayer;
|
||||||
|
GPU-Rechenpuffer 1044 MiB, Host Modell 13635 MiB, Kontext 3111 MiB, Compute
|
||||||
|
93 MiB. Das überschreitet das 8-GiB-RAM-Limit der Testinstanz und wird als
|
||||||
|
nicht startbar markiert. Keine Modellgewichte allokiert, alle zuvor vorhandenen
|
||||||
|
Container-Startzeiten unverändert. 46 Tests erfolgreich.
|
||||||
|
|
||||||
|
Der verifizierte Erstbuild liegt in `state/runtime-verification`;
|
||||||
|
`state/runtime/<Build-ID>` verweist intern darauf, damit seine Build-RPATHs
|
||||||
|
erhalten bleiben. Beide Verzeichnisse gehören ausschließlich Deck und bleiben
|
||||||
|
bei Updates bestehen. Künftige Builds entstehen direkt unter `state/runtime`.
|
||||||
|
|||||||
+5
-3
@@ -2,7 +2,9 @@
|
|||||||
|
|
||||||
Eigenständige Installation ohne WireGuard-Modul und ohne Änderungen am bestehenden
|
Eigenständige Installation ohne WireGuard-Modul und ohne Änderungen am bestehenden
|
||||||
Athena-Router. Der Installer installiert den aktuellen Deck-Anwendungsstand;
|
Athena-Router. Der Installer installiert den aktuellen Deck-Anwendungsstand;
|
||||||
Modellkatalog, Downloads und llama.cpp-Builds bleiben weiterhin GUI-Vorschau.
|
Modellkatalog, Downloads und llama.cpp-Buildverwaltung sind aktiv. Das Docker-Image
|
||||||
|
enthält das CUDA-Build-Toolkit; der Debian-Host wird nicht verändert. Zielprodukt
|
||||||
|
bleibt ein nativer systemd-Dienst. Modellstarts sind noch nicht angebunden.
|
||||||
|
|
||||||
## Voraussetzungen
|
## Voraussetzungen
|
||||||
|
|
||||||
@@ -21,7 +23,7 @@ auf einem bereits genutzten Server. Offizielle Anleitung:
|
|||||||
|
|
||||||
Optional für GPU-Messwerte: vorhandener NVIDIA-Treiber und vorhandenes NVIDIA
|
Optional für GPU-Messwerte: vorhandener NVIDIA-Treiber und vorhandenes NVIDIA
|
||||||
Container Toolkit. `--gpu-telemetry` bindet ausschließlich Treiber-Capability
|
Container Toolkit. `--gpu-telemetry` bindet ausschließlich Treiber-Capability
|
||||||
`utility` ein; es installiert keinen Treiber und startet kein Modell. Ohne diese
|
`compute,utility` für Telemetrie und Fit-Prüfung ein; es installiert keinen Treiber und startet kein Modell. Ohne diese
|
||||||
Option bleiben GPU-Werte in der Server-GUI ausdrücklich nicht verfügbar.
|
Option bleiben GPU-Werte in der Server-GUI ausdrücklich nicht verfügbar.
|
||||||
|
|
||||||
## Installation
|
## Installation
|
||||||
@@ -158,7 +160,7 @@ sudo ./install.sh --status --directory /opt/athena-deck-test
|
|||||||
|
|
||||||
Die Webanwendung läuft als UID/GID 65534, mit schreibgeschütztem Root-Dateisystem,
|
Die Webanwendung läuft als UID/GID 65534, mit schreibgeschütztem Root-Dateisystem,
|
||||||
allen Linux-Capabilities entzogen, ohne Docker-Socket, ohne privilegierten Modus
|
allen Linux-Capabilities entzogen, ohne Docker-Socket, ohne privilegierten Modus
|
||||||
und ohne Host-Netzwerk. Ressourcenlimit: eine CPU, 512 MiB RAM, 128 Prozesse.
|
und ohne Host-Netzwerk. Ressourcenlimit: zwei CPUs, 8 GiB RAM, 128 Prozesse.
|
||||||
GPU-Zugriff ist standardmäßig aus. Nur `state/` und ein temporäres tmpfs sind für
|
GPU-Zugriff ist standardmäßig aus. Nur `state/` und ein temporäres tmpfs sind für
|
||||||
die Anwendung beschreibbar. Docker erstellt seine normalen Regeln für den neuen
|
die Anwendung beschreibbar. Docker erstellt seine normalen Regeln für den neuen
|
||||||
Container; vorhandene Gateways und ihre Konfiguration werden nicht bearbeitet.
|
Container; vorhandene Gateways und ihre Konfiguration werden nicht bearbeitet.
|
||||||
|
|||||||
@@ -15,7 +15,7 @@ Vorhandenes Docker wird vorausgesetzt; produktive Dienste werden nicht veränder
|
|||||||
|
|
||||||
Katalog, Bibliothek, Profilentwürfe und eine Laufzeit-/Update-Ansicht sind als
|
Katalog, Bibliothek, Profilentwürfe und eine Laufzeit-/Update-Ansicht sind als
|
||||||
bedienbarer GUI-Prototyp vorhanden. Entwürfe bleiben im Browser; Modellstarts,
|
bedienbarer GUI-Prototyp vorhanden. Entwürfe bleiben im Browser; Modellstarts,
|
||||||
Builds und Modellstarts sind deaktiviert. Katalogsuche und Datei-Downloads sind live; siehe DEVELOPMENT.md. Details: [GUI-Prototyp](STUDIO.md).
|
Katalog, Datei-Downloads und llama.cpp-Buildverwaltung sind live. Modellstarts sind noch nicht angebunden; siehe DEVELOPMENT.md. Details: [GUI-Prototyp](STUDIO.md).
|
||||||
|
|
||||||
## Zugang und API-Token
|
## Zugang und API-Token
|
||||||
|
|
||||||
|
|||||||
@@ -1,3 +1,10 @@
|
|||||||
|
# Stand Laufzeitverwaltung
|
||||||
|
|
||||||
|
llama.cpp-Prüfung, Releases, Builds, Abbruch und Standard-/Rückfallauswahl sind
|
||||||
|
jetzt live. Kein manuelles VRAM-Reservefeld mehr. Modellbezogene Prognosen verwenden die lesend eingebundenen alten Routermodelle
|
||||||
|
und llama-fit-params; ein Modellstart erfolgt nicht. Die folgenden
|
||||||
|
0.4/0.5-Beschreibungen der Laufzeitansicht sind historisch.
|
||||||
|
|
||||||
# Stand 0.5
|
# Stand 0.5
|
||||||
|
|
||||||
Katalog und Datei-Downloads sind jetzt live. Die folgende Beschreibung der
|
Katalog und Datei-Downloads sind jetzt live. Die folgende Beschreibung der
|
||||||
|
|||||||
+4
-3
@@ -1,7 +1,8 @@
|
|||||||
FROM python:3.13-slim-bookworm
|
FROM nvidia/cuda:12.8.1-devel-ubuntu24.04
|
||||||
|
RUN apt-get update && DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends python3 git cmake g++ make ca-certificates libssl-dev && rm -rf /var/lib/apt/lists/*
|
||||||
WORKDIR /app
|
WORKDIR /app
|
||||||
COPY server.py catalog.py auth.py collect_hardware.py demo.py /app/
|
COPY server.py runtime.py catalog.py auth.py collect_hardware.py demo.py /app/
|
||||||
COPY index.html app.js studio.js catalog-ui.js style.css login.html login.js access-ui.js network-ui.js /app/
|
COPY index.html app.js studio.js runtime-ui.js catalog-ui.js style.css login.html login.js access-ui.js network-ui.js /app/
|
||||||
COPY network/__init__.py network/client.py network/config.py network/rpc.py /app/network/
|
COPY network/__init__.py network/client.py network/config.py network/rpc.py /app/network/
|
||||||
ENV PYTHONDONTWRITEBYTECODE=1 PYTHONUNBUFFERED=1 HOME=/tmp \
|
ENV PYTHONDONTWRITEBYTECODE=1 PYTHONUNBUFFERED=1 HOME=/tmp \
|
||||||
DECK_BIND_HOST=0.0.0.0 DECK_STATE_DIR=/var/lib/deck \
|
DECK_BIND_HOST=0.0.0.0 DECK_STATE_DIR=/var/lib/deck \
|
||||||
|
|||||||
+10
-5
@@ -14,7 +14,7 @@ import urllib.request
|
|||||||
|
|
||||||
ROOT = Path(__file__).resolve().parent.parent
|
ROOT = Path(__file__).resolve().parent.parent
|
||||||
LABEL = 'de.casaderoll.athena-deck.standalone'
|
LABEL = 'de.casaderoll.athena-deck.standalone'
|
||||||
FILES = ['catalog.py','catalog-ui.js','server.py','auth.py','collect_hardware.py','demo.py','index.html','app.js','studio.js','style.css','login.html','login.js','access-ui.js','network-ui.js','network/__init__.py','network/client.py','network/config.py','network/rpc.py','deploy/Dockerfile']
|
FILES = ['runtime.py','runtime-ui.js','catalog.py','catalog-ui.js','server.py','auth.py','collect_hardware.py','demo.py','index.html','app.js','studio.js','style.css','login.html','login.js','access-ui.js','network-ui.js','network/__init__.py','network/client.py','network/config.py','network/rpc.py','deploy/Dockerfile']
|
||||||
|
|
||||||
|
|
||||||
def run(*args, check=True, interactive=False):
|
def run(*args, check=True, interactive=False):
|
||||||
@@ -93,7 +93,7 @@ def launch(config):
|
|||||||
base=Path(config['base'])
|
base=Path(config['base'])
|
||||||
args=['docker','run','-d','--name',config['name'],'--label',LABEL+'='+str(base),
|
args=['docker','run','-d','--name',config['name'],'--label',LABEL+'='+str(base),
|
||||||
'--restart','unless-stopped','--read-only','--cap-drop','ALL','--security-opt','no-new-privileges:true',
|
'--restart','unless-stopped','--read-only','--cap-drop','ALL','--security-opt','no-new-privileges:true',
|
||||||
'--pids-limit','128','--memory','512m','--cpus','1', '--tmpfs','/tmp:rw,nosuid,nodev,size=16m',
|
'--pids-limit','128','--memory','8g','--cpus','2', '--tmpfs','/tmp:rw,nosuid,nodev,size=256m',
|
||||||
'-v',str(base/'state')+':/var/lib/deck:rw','-e','DECK_ALLOWED_HOSTS=127.0.0.1:'+str(config['port'])+',localhost:'+str(config['port']),
|
'-v',str(base/'state')+':/var/lib/deck:rw','-e','DECK_ALLOWED_HOSTS=127.0.0.1:'+str(config['port'])+',localhost:'+str(config['port']),
|
||||||
'-p','127.0.0.1:'+str(config['port'])+':8108',
|
'-p','127.0.0.1:'+str(config['port'])+':8108',
|
||||||
'--health-cmd', 'python3 -c "import urllib.request,json; assert json.load(urllib.request.urlopen(\'http://127.0.0.1:8108/api/v1/auth/status\',timeout=3))[\'initialized\']"',
|
'--health-cmd', 'python3 -c "import urllib.request,json; assert json.load(urllib.request.urlopen(\'http://127.0.0.1:8108/api/v1/auth/status\',timeout=3))[\'initialized\']"',
|
||||||
@@ -103,7 +103,12 @@ def launch(config):
|
|||||||
# Explicit isolated development bootstrap, reachable only through host loopback/SSH.
|
# Explicit isolated development bootstrap, reachable only through host loopback/SSH.
|
||||||
args += ['-e','DECK_REQUIRE_SETUP=0']
|
args += ['-e','DECK_REQUIRE_SETUP=0']
|
||||||
if config['gpu_telemetry']:
|
if config['gpu_telemetry']:
|
||||||
args += ['--gpus','all','-e','NVIDIA_DRIVER_CAPABILITIES=utility']
|
args += ['--gpus','all','-e','NVIDIA_DRIVER_CAPABILITIES=compute,utility']
|
||||||
|
if config.get('reference_models'):
|
||||||
|
for folder,filename in [('qwen3.8-27b-iq4-mix','Qwen3.8-27B-IQ4-MIX.gguf'),('qwen3.8-27b-iq4-xs-pure','qwen3.8-27b-IQ4_XS-pure.gguf'),('qwen3.8-27b-abliterated','Qwen3.8-27B-ABLITERATED-Q4_K_M.gguf')]:
|
||||||
|
source=Path('/data/models')/folder/filename
|
||||||
|
if not source.is_file():raise RuntimeError('Referenzmodell fehlt; keine Ersatzpfade.')
|
||||||
|
args+=['--mount','type=bind,src='+str(source)+',dst=/reference-models/'+filename+',readonly']
|
||||||
args.append(config['image'])
|
args.append(config['image'])
|
||||||
run(*args)
|
run(*args)
|
||||||
|
|
||||||
@@ -181,7 +186,7 @@ def update(base, rollback=False):
|
|||||||
backup_name=config['name']+'-previous'
|
backup_name=config['name']+'-previous'
|
||||||
if inspect(backup_name):raise RuntimeError('Rückfall-Containername belegt. Keine Änderung ausgeführt.')
|
if inspect(backup_name):raise RuntimeError('Rückfall-Containername belegt. Keine Änderung ausgeführt.')
|
||||||
stamp=str(time.time_ns())
|
stamp=str(time.time_ns())
|
||||||
shutil.copytree(base/'state',base/'backups'/stamp,ignore=shutil.ignore_patterns('models'))
|
shutil.copytree(base/'state',base/'backups'/stamp,ignore=shutil.ignore_patterns('models','runtime','runtime-verification'))
|
||||||
was_running=previous['State']['Running']
|
was_running=previous['State']['Running']
|
||||||
if was_running:run('docker','stop','--time','15',config['name'])
|
if was_running:run('docker','stop','--time','15',config['name'])
|
||||||
run('docker','rename',config['name'],backup_name)
|
run('docker','rename',config['name'],backup_name)
|
||||||
@@ -206,7 +211,7 @@ def main():
|
|||||||
parser.add_argument('--directory',type=Path,default=Path('/opt/athena-deck-standalone'))
|
parser.add_argument('--directory',type=Path,default=Path('/opt/athena-deck-standalone'))
|
||||||
parser.add_argument('--name',default='athena-deck-standalone')
|
parser.add_argument('--name',default='athena-deck-standalone')
|
||||||
parser.add_argument('--port',type=int,default=8110)
|
parser.add_argument('--port',type=int,default=8110)
|
||||||
parser.add_argument('--gpu-telemetry',action='store_true',help='Vorhandenes NVIDIA Container Toolkit nur für Messwerte nutzen; installiert keine Treiber.')
|
parser.add_argument('--gpu-telemetry',action='store_true',help='Vorhandenes NVIDIA Container Toolkit für Messwerte und Fit-Prüfung nutzen; installiert keine Treiber.')
|
||||||
args=parser.parse_args()
|
args=parser.parse_args()
|
||||||
import re
|
import re
|
||||||
if not re.fullmatch(r'athena-deck-[a-z0-9-]{1,40}',args.name):
|
if not re.fullmatch(r'athena-deck-[a-z0-9-]{1,40}',args.name):
|
||||||
|
|||||||
+1
-1
@@ -1 +1 @@
|
|||||||
<!doctype html><html lang="de"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>Athena Deck</title><link rel="stylesheet" href="/style.css"><script src="/network-ui.js" defer></script><script src="/access-ui.js" defer></script><script src="/catalog-ui.js" defer></script><script src="/studio.js" defer></script><script src="/app.js" defer></script></head><body><aside><a class="brand" href="#home"><span>Α</span> ATHENA <b>DECK</b></a><p class="eyebrow">WORKSPACE / PROTOTYP 01</p><nav aria-label="Hauptnavigation"><a href="#home">Übersicht <span>01</span></a><a href="#chat">Sprachmodelle / Chat <span>02</span></a><a href="#image">Bildgenerierung <span>03</span></a><a href="#audio">Audio <span>04</span></a><a href="#video">Video <span>05</span></a><a href="#services">Weitere Dienste <span>06</span></a><a href="#hardware">Hardware <span>07</span></a><p class="nav-heading">EINSTELLUNGEN</p><a href="#runtime">llama.cpp & Updates <span>↗</span></a><a href="#network">Netzwerk <span>↗</span></a><a href="#access">Zugang & API <span>↗</span></a></nav><footer><i></i> Isolierte Entwicklungsumgebung<p>Getrennte Steuerungsinstanz<br>Hardware · nur lesend</p></footer></aside><main><header><span>ATHENA CONTROL SURFACE</span><span class="pill">v0.5 · Entwicklung</span></header><div id="view"></div><p id="error" role="alert"></p></main></body></html>
|
<!doctype html><html lang="de"><head><meta charset="utf-8"><meta name="viewport" content="width=device-width,initial-scale=1"><title>Athena Deck</title><link rel="stylesheet" href="/style.css"><script src="/network-ui.js" defer></script><script src="/access-ui.js" defer></script><script src="/catalog-ui.js" defer></script><script src="/runtime-ui.js" defer></script><script src="/studio.js" defer></script><script src="/app.js" defer></script></head><body><aside><a class="brand" href="#home"><span>Α</span> ATHENA <b>DECK</b></a><p class="eyebrow">WORKSPACE / PROTOTYP 01</p><nav aria-label="Hauptnavigation"><a href="#home">Übersicht <span>01</span></a><a href="#chat">Sprachmodelle / Chat <span>02</span></a><a href="#image">Bildgenerierung <span>03</span></a><a href="#audio">Audio <span>04</span></a><a href="#video">Video <span>05</span></a><a href="#services">Weitere Dienste <span>06</span></a><a href="#hardware">Hardware <span>07</span></a><p class="nav-heading">EINSTELLUNGEN</p><a href="#runtime">llama.cpp & Updates <span>↗</span></a><a href="#network">Netzwerk <span>↗</span></a><a href="#access">Zugang & API <span>↗</span></a></nav><footer><i></i> Isolierte Entwicklungsumgebung<p>Getrennte Steuerungsinstanz<br>Hardware · nur lesend</p></footer></aside><main><header><span>ATHENA CONTROL SURFACE</span><span class="pill">v0.6 · Entwicklung</span></header><div id="view"></div><p id="error" role="alert"></p></main></body></html>
|
||||||
|
|||||||
@@ -12,7 +12,7 @@ import sys
|
|||||||
|
|
||||||
NAME = 'athena-deck-network'
|
NAME = 'athena-deck-network'
|
||||||
BASE = Path('/opt/athena-deck')
|
BASE = Path('/opt/athena-deck')
|
||||||
FILES = {'catalog.py', 'catalog-ui.js', 'studio.js', 'auth.py', 'access-ui.js', 'server.py', 'demo.py', 'collect_hardware.py', 'index.html', 'app.js', 'style.css', 'network-ui.js', 'login.html', 'login.js', 'network/__init__.py', 'network/config.py', 'network/policy.py', 'network/rpc.py', 'network/agent.py', 'network/client.py', 'network/Dockerfile', '.dockerignore'}
|
FILES = {'runtime.py', 'runtime-ui.js', 'catalog.py', 'catalog-ui.js', 'studio.js', 'auth.py', 'access-ui.js', 'server.py', 'demo.py', 'collect_hardware.py', 'index.html', 'app.js', 'style.css', 'network-ui.js', 'login.html', 'login.js', 'network/__init__.py', 'network/config.py', 'network/policy.py', 'network/rpc.py', 'network/agent.py', 'network/client.py', 'network/Dockerfile', '.dockerignore'}
|
||||||
|
|
||||||
def run(*args, **kwargs):
|
def run(*args, **kwargs):
|
||||||
return subprocess.run(args, capture_output=True, timeout=600, **kwargs)
|
return subprocess.run(args, capture_output=True, timeout=600, **kwargs)
|
||||||
|
|||||||
@@ -0,0 +1,19 @@
|
|||||||
|
const RuntimeUI=(()=>{
|
||||||
|
const e=v=>String(v??'').replace(/[&<>"']/g,c=>({'&':'&','<':'<','>':'>','"':'"',"'":'''}[c]));
|
||||||
|
async function api(path='',data){const r=await fetch('/api/v1/runtime'+path,data?{method:'POST',headers:{'Content-Type':'application/json','X-Athena-Deck':'1'},body:JSON.stringify(data)}:{});const v=await r.json();if(!r.ok)throw Error(v.error||'Anfrage fehlgeschlagen');return v;}
|
||||||
|
function html(){return `<div id="runtime-live"><div class="kicker">ATHENA / LLAMA.CPP</div><h1>llama.cpp & Updates</h1><p>Offizielle Versionen getrennt bauen und prüfen. Ein Build lädt kein Modell und ändert keine produktiven Laufzeiten.</p><section class="card"><h2>Voraussetzungen</h2><button id="rt-check">Werkzeuge & GPUs prüfen</button><div id="rt-prereq"></div></section><section class="card"><h2>Installation & Build</h2><form id="rt-build"><label>Release oder vollständiger Commit<input name="revision" required pattern="b[0-9]{3,8}|[a-f0-9]{40}" placeholder="Unten Releases abfragen und auswählen"></label><label>Backend<select name="backend"><option>CUDA</option><option>CPU</option></select></label><label>Parallele Build-Jobs<select name="jobs"><option>1</option><option>2</option></select></label><p>GPU-Architekturen werden automatisch aus beiden Karten übernommen. Begrenzte Build-Last für den Parallelbetrieb.</p><button>llama.cpp herunterladen & bauen</button></form><button id="rt-cancel" class="secondary">Laufenden Build abbrechen</button><p id="rt-message" role="status"></p><div id="rt-job"></div><details><summary>Build-Ausgabe (eigener Build)</summary><pre id="rt-log" style="max-height:320px;overflow:auto;white-space:pre-wrap"></pre></details></section><section class="card"><h2>Versionen & Änderungen</h2><button id="rt-releases">Nach Updates suchen</button><div id="rt-releases-list"></div></section><section class="card"><h2>Geprüfte Builds</h2><div id="rt-builds"></div><button id="rt-rollback" class="secondary">Vorherigen Standard wiederherstellen</button><p>Als Standard auswählen startet keinen Prozess. Bestehende Router bleiben unverändert.</p></section><section class="card"><h2>Automatische Speicheranpassung</h2><p>Kein pauschales Reservefeld: Die Laufzeit soll Modell, gewünschten Kontext, Slots und den aktuell freien Speicher gemeinsam berücksichtigen. Unterstützte Builds verwenden llama.cpp <code>--fit on</code>; der Kontext wird ausdrücklich vorgegeben. Keine absichtlichen OOM-Tests auf den produktiv genutzten GPUs.</p><form id="rt-fit"><label>Vorhandenes Routerprofil<select name="profile" id="rt-profiles"></select></label><label>Gewünschter Gesamtkontext<input name="context" type="number" min="512" max="2097152" value="160000" required></label><label>Slots<input name="slots" type="number" min="1" max="16" value="2" required></label><button>Auto-Einpassung berechnen</button></form><pre id="rt-fit-result" style="white-space:pre-wrap"></pre><p>Ein Build allein kann keine Modell-Layerzahl bestimmen. Die Berechnung verwendet llama-fit-params ohne Gewichtsallokation. Sie ist eine Prognose für das Textmodell und kein Lasttest. Vision-Projektor und MTP sind noch nicht eingerechnet. Ein Modellstart erfolgt nicht.</p></section></div>`;}
|
||||||
|
function bind(){const root=document.querySelector('#runtime-live');if(!root)return;let running=false;
|
||||||
|
const msg=t=>{if(root.isConnected)root.querySelector('#rt-message').textContent=t;};
|
||||||
|
async function status(){try{const s=await api();if(!root.isConnected)return;running=s.job?.state==='running';root.querySelector('#rt-job').textContent=s.job?`${s.job.revision} · ${s.job.state} · ${s.job.phase} ${s.job.error||''}`:'Noch kein Build gestartet.';root.querySelector('#rt-cancel').disabled=!running;root.querySelector('#rt-build button').disabled=running;root.querySelector('#rt-rollback').disabled=!s.previous;root.querySelector('#rt-builds').innerHTML=s.builds.map(b=>`<article><h3>${e(b.revision)} · ${e(b.backend)} ${s.active===b.id?'· STANDARD':''}</h3><p>Commit ${e(b.commit)} · GPU-Ziele ${e(b.architectures.join(', ')||'CPU')} · Auto-Fit ${b.fit_supported?'unterstützt':'nicht unterstützt'}</p><button data-build="${e(b.id)}" ${s.active===b.id?'disabled':''}>Als Standard auswählen</button></article>`).join('')||'<p>Noch kein geprüfter Build vorhanden.</p>';root.querySelectorAll('[data-build]').forEach(b=>b.onclick=()=>action('/activate',{build_id:b.dataset.build}));if(s.job){const log=await api('/log');if(root.isConnected)root.querySelector('#rt-log').textContent=log.text;}if(running)setTimeout(()=>{if(root.isConnected)status();},3000);}catch(err){msg(err.message);}}
|
||||||
|
async function action(path,data){try{await api(path,data);msg('Aktion angenommen.');await status();}catch(err){msg(err.message);}}
|
||||||
|
async function prerequisites(){try{const p=await api('/prerequisites');if(!root.isConnected)return;root.querySelector('#rt-prereq').innerHTML=`<p>CPU-Build: ${p.cpu_ready?'bereit':'Werkzeuge fehlen'} · CUDA-Build: ${p.cuda_ready?'bereit':'Werkzeuge oder GPUs fehlen'}</p><p>${['git','cmake','g++','nvcc'].map(k=>e(k)+': '+e(p[k]||'fehlt')).join(' · ')}</p>`+p.gpus.map(g=>`<p>${e(g.name)} · Architektur ${e(g.architecture)} · aktuell ${Math.round(g.free_mib)} / ${Math.round(g.total_mib)} MiB frei</p>`).join('');}catch(err){msg(err.message);}}
|
||||||
|
root.querySelector('#rt-check').onclick=prerequisites;
|
||||||
|
root.querySelector('#rt-build').onsubmit=event=>{event.preventDefault();const d=Object.fromEntries(new FormData(event.target));d.jobs=Number(d.jobs);action('/build',d);};
|
||||||
|
root.querySelector('#rt-cancel').onclick=()=>action('/cancel',{});root.querySelector('#rt-rollback').onclick=()=>action('/rollback',{});
|
||||||
|
root.querySelector('#rt-releases').onclick=async()=>{msg('Offizielle Releases werden abgefragt …');try{const v=await api('/releases');if(!root.isConnected)return;root.querySelector('#rt-releases-list').innerHTML=v.releases.map(r=>`<article><h3>${e(r.tag)} ${r.prerelease?'· Vorabversion':''} · ${e(r.published)}</h3><button data-release="${e(r.tag)}">Für Build auswählen</button><a href="${e(r.url)}" target="_blank" rel="noopener noreferrer">Original auf GitHub</a><details><summary>Änderungen / Release-Notes</summary><pre style="white-space:pre-wrap">${e(r.notes)}</pre></details></article>`).join('');root.querySelectorAll('[data-release]').forEach(b=>b.onclick=()=>{root.querySelector('[name=revision]').value=b.dataset.release;msg('Version ausgewählt. Build kann gestartet werden.');});msg(v.releases.length+' offizielle Releases geladen.');}catch(err){msg(err.message);}};
|
||||||
|
api('/references').then(v=>{if(!root.isConnected)return;const select=root.querySelector('#rt-profiles');select.innerHTML=v.profiles.map(p=>`<option value="${e(p.id)}" ${p.id==='medium'?'selected':''}>${e(p.id)} · ${e(p.file)} ${p.available?'':'(nicht eingebunden)'}</option>`).join('');select.onchange=()=>{const p=v.profiles.find(p=>p.id===select.value);root.querySelector('#rt-fit [name=context]').value=p.context;root.querySelector('#rt-fit [name=slots]').value=p.slots;};}).catch(err=>msg(err.message));
|
||||||
|
root.querySelector('#rt-fit').onsubmit=async event=>{event.preventDefault();const b=event.target.querySelector('button');b.disabled=true;msg('Modellmetadaten und Speicherbedarf werden berechnet …');const d=Object.fromEntries(new FormData(event.target));d.context=Number(d.context);d.slots=Number(d.slots);try{const v=await api('/fit',d);if(root.isConnected)root.querySelector('#rt-fit-result').textContent=(v.success?'Prognose erfolgreich. Kein Modell geladen.':'Aktuell keine passende Einpassung ermittelt.')+(v.within_ram_limit===false?'\nNICHT STARTBAR in der Testinstanz: berechneter Host-RAM überschreitet deren Limit von '+v.ram_limit_mib+' MiB.':'')+'\nGPU-Layer (Prognose): '+(v.fitted?.gpu_layers??'nicht ermittelt')+'\nGPU-Verteilung: '+(v.fitted?.tensor_split??'ein Gerät / nicht ermittelt')+'\nAusgelassene belegte GPUs: '+v.excluded_gpus.join(', ')+'\nAutomatische Reserve (MiB): '+v.reserve_mib.join(', ')+'\n'+v.memory.map(m=>`${m.device}: Modell ${m.model_mib} MiB · Kontext ${m.context_mib} MiB · Rechenspeicher ${m.compute_mib} MiB`).join('\n')+'\n'+v.arguments+'\n'+v.details;msg('Berechnung abgeschlossen. Freier Speicher kann sich durch andere Dienste ändern.');}catch(err){msg(err.message);}finally{b.disabled=false;}};
|
||||||
|
prerequisites();status();
|
||||||
|
}
|
||||||
|
return {html,bind};
|
||||||
|
})();
|
||||||
+192
@@ -0,0 +1,192 @@
|
|||||||
|
"""Unprivileged native build manager. No Docker calls, shell commands or host changes."""
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
from pathlib import Path
|
||||||
|
import re
|
||||||
|
import shutil
|
||||||
|
import shlex
|
||||||
|
import signal
|
||||||
|
import subprocess
|
||||||
|
import threading
|
||||||
|
import time
|
||||||
|
import urllib.request
|
||||||
|
import uuid
|
||||||
|
|
||||||
|
ORIGIN='https://github.com/ggml-org/llama.cpp.git'
|
||||||
|
|
||||||
|
def github(path):
|
||||||
|
req=urllib.request.Request('https://api.github.com/repos/ggml-org/llama.cpp/'+path,headers={'Accept':'application/vnd.github+json','User-Agent':'Athena-Deck'})
|
||||||
|
with urllib.request.urlopen(req,timeout=20) as r:return json.load(r)
|
||||||
|
|
||||||
|
def command(args,timeout=10):
|
||||||
|
return subprocess.run(args,capture_output=True,text=True,timeout=timeout,check=True).stdout.strip()
|
||||||
|
|
||||||
|
def prepare_fit_source(directory):
|
||||||
|
# The reference profiles share a KV pool; upstream fit CLI omits its switch.
|
||||||
|
path=Path(directory)/'tools/fit-params/fit-params.cpp'
|
||||||
|
if not path.is_file():raise ValueError('Diese Version enthält kein unterstütztes Fit-Werkzeug.')
|
||||||
|
source=path.read_text();marker=' params.kv_unified = true; // Athena Deck reference profiles'
|
||||||
|
if marker in source:return
|
||||||
|
anchor=' llama_backend_init();'
|
||||||
|
if source.count(anchor)!=1:raise ValueError('Fit-Adapter passt nicht zu dieser Version; Build abgebrochen.')
|
||||||
|
path.write_text(source.replace(anchor,marker+'\n'+anchor))
|
||||||
|
|
||||||
|
class Runtime:
|
||||||
|
def __init__(self,root):
|
||||||
|
self.root=Path(root);self.lock=threading.RLock();self.process=None;self.cancelled=threading.Event();self.busy=False
|
||||||
|
self.state={'job':None,'builds':[],'active':None,'previous':None}
|
||||||
|
if (self.root/'state.json').exists():
|
||||||
|
self.state=json.loads((self.root/'state.json').read_text())
|
||||||
|
if (self.state.get('job') or {}).get('state')=='running':
|
||||||
|
self.state['job'].update(state='interrupted',phase='Deck wurde neu gestartet; Build erneut starten.')
|
||||||
|
def save(self):
|
||||||
|
self.root.mkdir(parents=True,exist_ok=True,mode=0o700)
|
||||||
|
p=self.root/'state.tmp';p.write_text(json.dumps(self.state));p.replace(self.root/'state.json')
|
||||||
|
def prerequisites(self):
|
||||||
|
result={x:shutil.which(x) for x in ('git','cmake','g++','nvcc')};gpus=[]
|
||||||
|
try:
|
||||||
|
for line in command(['nvidia-smi','--query-gpu=index,uuid,name,compute_cap,memory.total,memory.free','--format=csv,noheader,nounits']).splitlines():
|
||||||
|
idx,gpu_uuid,name,cap,total,free=[x.strip() for x in line.split(',')]
|
||||||
|
gpus.append(dict(index=int(idx),uuid=gpu_uuid,name=name,architecture=cap.replace('.',''),total_mib=float(total),free_mib=float(free)))
|
||||||
|
except (OSError,ValueError,subprocess.SubprocessError):pass
|
||||||
|
result.update(gpus=gpus,cpu_ready=all(result[x] for x in ('git','cmake','g++')),cuda_ready=all(result[x] for x in ('git','cmake','g++','nvcc')) and bool(gpus))
|
||||||
|
return result
|
||||||
|
def status(self):
|
||||||
|
with self.lock:return json.loads(json.dumps(self.state))
|
||||||
|
def releases(self):
|
||||||
|
rows=github('releases?per_page=10')
|
||||||
|
return {'releases':[dict(tag=r['tag_name'],name=r['name'],published=r['published_at'],notes=r.get('body',''),url=r['html_url'],prerelease=r['prerelease']) for r in rows if not r['draft']]}
|
||||||
|
def run(self,args,cwd,phase,timeout=7200):
|
||||||
|
with self.lock:
|
||||||
|
if self.cancelled.is_set():raise InterruptedError()
|
||||||
|
self.state['job']['phase']=phase;self.save()
|
||||||
|
log=self.root/'build.log'
|
||||||
|
stream=log.open('ab')
|
||||||
|
try:
|
||||||
|
self.process=subprocess.Popen(args,cwd=cwd,stdout=stream,stderr=subprocess.STDOUT,start_new_session=True)
|
||||||
|
except Exception:
|
||||||
|
stream.close();raise
|
||||||
|
try:
|
||||||
|
code=self.process.wait(timeout=timeout)
|
||||||
|
if self.cancelled.is_set():raise InterruptedError()
|
||||||
|
if code:raise ValueError('Build-Schritt fehlgeschlagen: '+phase)
|
||||||
|
except subprocess.TimeoutExpired:
|
||||||
|
self.stop();raise ValueError('Build-Zeitlimit erreicht.') from None
|
||||||
|
finally:
|
||||||
|
stream.close()
|
||||||
|
with self.lock:self.process=None
|
||||||
|
def start(self,revision,backend='CUDA',jobs=1):
|
||||||
|
if not isinstance(revision,str) or not re.fullmatch(r'(?:b[0-9]{3,8}|[a-f0-9]{40})',revision):raise ValueError('Offizielles b-Release oder vollständigen Commit-SHA wählen.')
|
||||||
|
if backend not in ('CUDA','CPU') or type(jobs)!=int or not 1<=jobs<=2:raise ValueError('CPU/CUDA und 1–2 Build-Jobs unterstützt.')
|
||||||
|
with self.lock:
|
||||||
|
if self.busy:raise ValueError('Ein Build läuft bereits.')
|
||||||
|
p=self.prerequisites()
|
||||||
|
if not p['cuda_ready' if backend=='CUDA' else 'cpu_ready']:raise ValueError('Build-Werkzeuge fehlen. Voraussetzungen prüfen.')
|
||||||
|
self.root.mkdir(parents=True,exist_ok=True)
|
||||||
|
if shutil.disk_usage(self.root).free<15*1024**3:raise ValueError('Mindestens 15 GiB freier Speicher erforderlich.')
|
||||||
|
ident=uuid.uuid4().hex
|
||||||
|
self.state['job']=dict(id=ident,state='running',phase='Vorbereitung',revision=revision,backend=backend,started=time.time(),error=None)
|
||||||
|
self.cancelled.clear();self.busy=True;self.save()
|
||||||
|
threading.Thread(target=self.build,args=(ident,revision,backend,jobs,p),daemon=True).start()
|
||||||
|
return self.status()
|
||||||
|
def build(self,ident,revision,backend,jobs,prereq):
|
||||||
|
directory=self.root/ident
|
||||||
|
try:
|
||||||
|
directory.mkdir();(self.root/'build.log').write_text('')
|
||||||
|
self.run(['git','init',str(directory)],self.root,'Quellverzeichnis anlegen',30)
|
||||||
|
self.run(['git','-C',str(directory),'fetch','--depth','1',ORIGIN,revision],self.root,'Offizielle Quellen herunterladen',300)
|
||||||
|
self.run(['git','checkout','--detach','FETCH_HEAD'],directory,'Feste Version auswählen',30)
|
||||||
|
commit=command(['git','-C',str(directory),'rev-parse','HEAD'])
|
||||||
|
prepare_fit_source(directory)
|
||||||
|
arch=sorted(set(g['architecture'] for g in prereq['gpus']))
|
||||||
|
args=['cmake','-S','.', '-B','build','-DCMAKE_BUILD_TYPE=Release','-DGGML_NATIVE=OFF','-DLLAMA_BUILD_TESTS=OFF','-DLLAMA_BUILD_EXAMPLES=OFF','-DLLAMA_BUILD_SERVER=ON','-DGGML_CUDA='+('ON' if backend=='CUDA' else 'OFF')]
|
||||||
|
if backend=='CUDA':args+=['-DCMAKE_CUDA_ARCHITECTURES='+';'.join(arch)]
|
||||||
|
self.run(args,directory,'Build konfigurieren',300)
|
||||||
|
self.run(['cmake','--build','build','--target','llama-server','llama-fit-params','-j',str(jobs)],directory,'llama.cpp-Werkzeuge kompilieren')
|
||||||
|
binary=directory/'build/bin/llama-server'
|
||||||
|
self.run([str(binary),'--version'],directory,'Binärdatei prüfen',30)
|
||||||
|
help_text=command([str(binary),'--help'],30)
|
||||||
|
item=dict(id=ident,revision=revision,commit=commit,backend=backend,architectures=arch if backend=='CUDA' else [],created=time.time(),fit_supported='--fit ' in help_text,fit_tool=(directory/'build/bin/llama-fit-params').exists(),fit_adapter='shared-kv-pool-v1')
|
||||||
|
with self.lock:
|
||||||
|
self.state['builds'].append(item);self.state['job'].update(state='complete',phase='Build geprüft; kann als Standard ausgewählt werden.');self.save()
|
||||||
|
except Exception as exc:
|
||||||
|
with self.lock:
|
||||||
|
self.state['job'].update(state='cancelled' if isinstance(exc,InterruptedError) else 'failed',error=str(exc) if isinstance(exc,ValueError) else 'Build unterbrochen oder Werkzeug nicht erreichbar.');self.save()
|
||||||
|
finally:
|
||||||
|
with self.lock:self.busy=False
|
||||||
|
def stop(self):
|
||||||
|
with self.lock:
|
||||||
|
self.cancelled.set()
|
||||||
|
if self.process and self.process.poll() is None:
|
||||||
|
try:os.killpg(self.process.pid,signal.SIGTERM)
|
||||||
|
except ProcessLookupError:pass
|
||||||
|
try:self.process.wait(timeout=3)
|
||||||
|
except subprocess.TimeoutExpired:
|
||||||
|
try:os.killpg(self.process.pid,signal.SIGKILL)
|
||||||
|
except ProcessLookupError:pass
|
||||||
|
return {'cancellation_requested':True}
|
||||||
|
def activate(self,build_id):
|
||||||
|
with self.lock:
|
||||||
|
if not any(b['id']==build_id for b in self.state['builds']):raise ValueError('Kein erfolgreich geprüfter Build.')
|
||||||
|
if self.state['active']!=build_id:
|
||||||
|
self.state['previous']=self.state['active'];self.state['active']=build_id;self.save()
|
||||||
|
return self.status()
|
||||||
|
def rollback(self):
|
||||||
|
with self.lock:
|
||||||
|
if not self.state['previous']:raise ValueError('Kein vorheriger Build vorhanden.')
|
||||||
|
return self.activate(self.state['previous'])
|
||||||
|
def log(self):
|
||||||
|
p=self.root/'build.log'
|
||||||
|
if not p.exists():return {'text':''}
|
||||||
|
with p.open('rb') as f:
|
||||||
|
f.seek(max(0,p.stat().st_size-24000));return {'text':f.read().decode(errors='replace')}
|
||||||
|
|
||||||
|
def references(self):
|
||||||
|
base=Path(os.environ.get('DECK_REFERENCE_MODELS','/reference-models'))
|
||||||
|
specs=[('fast','Qwen3.8-27B-IQ4-MIX.gguf',76800,1,64),('medium','qwen3.8-27b-IQ4_XS-pure.gguf',160000,2,256),('large','qwen3.8-27b-IQ4_XS-pure.gguf',192000,1,256),('ultra','qwen3.8-27b-IQ4_XS-pure.gguf',262144,1,128),('uncensored','Qwen3.8-27B-ABLITERATED-Q4_K_M.gguf',80000,1,256)]
|
||||||
|
return {'profiles':[dict(id=i,file=f,context=c,slots=n,ubatch=u,cache='q4_0',available=(base/f).is_file()) for i,f,c,n,u in specs]}
|
||||||
|
def fit(self,profile,context,slots):
|
||||||
|
if type(context)!=int or not 512<=context<=2097152 or type(slots)!=int or not 1<=slots<=16:raise ValueError('Kontext oder Slots außerhalb des erlaubten Bereichs.')
|
||||||
|
ref=next((p for p in self.references()['profiles'] if p['id']==profile),None)
|
||||||
|
if not ref or not ref['available']:raise ValueError('Referenzmodell nicht lesend eingebunden.')
|
||||||
|
with self.lock:
|
||||||
|
build=next((b for b in self.state['builds'] if b['id']==self.state['active']),None)
|
||||||
|
if not build or not build.get('fit_tool'):raise ValueError('Zuerst einen CUDA-Build mit Fit-Werkzeug erstellen und auswählen.')
|
||||||
|
if build['backend']!='CUDA':raise ValueError('Für GPU-Einpassung einen CUDA-Build auswählen.')
|
||||||
|
if self.busy:raise ValueError('Bitte Ende des Builds abwarten.')
|
||||||
|
if getattr(self,'fitting',False):raise ValueError('Eine Einpassung läuft bereits.')
|
||||||
|
self.fitting=True
|
||||||
|
try:
|
||||||
|
binary=self.root/build['id']/'build/bin/llama-fit-params'
|
||||||
|
model=Path(os.environ.get('DECK_REFERENCE_MODELS','/reference-models'))/ref['file']
|
||||||
|
args=[str(binary),'--model',str(model),'--ctx-size',str(context),'--parallel',str(slots),'--cache-type-k','q4_0','--cache-type-v','q4_0','--flash-attn','on','--batch-size','2048','--ubatch-size',str(ref['ubatch'])]
|
||||||
|
snapshot=self.prerequisites()['gpus']
|
||||||
|
usable=[g for g in snapshot if g['free_mib']>=max(1024,g['total_mib']*.10)]
|
||||||
|
if not usable:raise ValueError('GPUs derzeit belegt. Keine sichere Auto-Prognose; produktive Dienste bleiben unverändert.')
|
||||||
|
margins=[max(512,round(g['total_mib']*.05)) for g in usable]
|
||||||
|
args+=['--fit-target',','.join(map(str,margins))]
|
||||||
|
env=dict(os.environ,CUDA_VISIBLE_DEVICES=','.join(g['uuid'] for g in usable))
|
||||||
|
result=subprocess.run(args,capture_output=True,text=True,timeout=120,env=env)
|
||||||
|
text=result.stdout.strip()
|
||||||
|
fitted={};memory=[]
|
||||||
|
if result.returncode==0:
|
||||||
|
values=shlex.split(text)
|
||||||
|
for flag,key in [('-c','context'),('-ngl','gpu_layers'),('-ts','tensor_split')]:
|
||||||
|
if flag in values and values.index(flag)+1<len(values):fitted[key]=values[values.index(flag)+1]
|
||||||
|
if fitted.get('context')!=str(context):raise ValueError('Fit-Werkzeug hat den gewünschten Kontext nicht beibehalten; Ergebnis wird nicht übernommen.')
|
||||||
|
if len(values)%2==0 and all(values[i] in ('-c','-ngl','-ts') for i in range(0,len(values),2)):
|
||||||
|
probe=subprocess.run(args+['--fit-print','on']+values,capture_output=True,text=True,timeout=120,env=env)
|
||||||
|
if probe.returncode==0:
|
||||||
|
for line in probe.stdout.splitlines():
|
||||||
|
match=re.fullmatch(r'(\S+) +(\d+) +(\d+) +(\d+) *',line)
|
||||||
|
if match:memory.append(dict(device=match[1],model_mib=int(match[2]),context_mib=int(match[3]),compute_mib=int(match[4])))
|
||||||
|
ram_limit=None
|
||||||
|
try:
|
||||||
|
raw=Path('/sys/fs/cgroup/memory.max').read_text().strip()
|
||||||
|
if raw!='max':ram_limit=int(raw)//(1024*1024)
|
||||||
|
except (OSError,ValueError):pass
|
||||||
|
host=next((m for m in memory if m['device']=='Host'),None)
|
||||||
|
within_limit=None if host is None or ram_limit is None else sum(host[k] for k in ('model_mib','context_mib','compute_mib'))<=ram_limit
|
||||||
|
return dict(success=result.returncode==0,ram_limit_mib=ram_limit,within_ram_limit=within_limit,memory=memory,fitted=fitted,profile=profile,context=context,slots=slots,arguments=text[-16000:],details=result.stderr[-16000:],estimated=True,scope='text_model_without_vision_or_mtp',model_loaded=False,gpus=snapshot,used_gpu_uuids=[g['uuid'] for g in usable],reserve_mib=margins,excluded_gpus=[g['name'] for g in snapshot if g not in usable])
|
||||||
|
finally:
|
||||||
|
with self.lock:self.fitting=False
|
||||||
@@ -7,6 +7,7 @@ import secrets
|
|||||||
from http.cookies import SimpleCookie, CookieError
|
from http.cookies import SimpleCookie, CookieError
|
||||||
from auth import CredentialStore, verify_password
|
from auth import CredentialStore, verify_password
|
||||||
from catalog import Catalog
|
from catalog import Catalog
|
||||||
|
from runtime import Runtime
|
||||||
from urllib.parse import urlsplit, parse_qs
|
from urllib.parse import urlsplit, parse_qs
|
||||||
from network.client import NetworkClient
|
from network.client import NetworkClient
|
||||||
from network.config import parse_config, ConfigError
|
from network.config import parse_config, ConfigError
|
||||||
@@ -97,6 +98,7 @@ class Server(ThreadingHTTPServer):
|
|||||||
def __init__(self, port, state_dir=None):
|
def __init__(self, port, state_dir=None):
|
||||||
super().__init__((os.environ.get('DECK_BIND_HOST','127.0.0.1'), port), Handler)
|
super().__init__((os.environ.get('DECK_BIND_HOST','127.0.0.1'), port), Handler)
|
||||||
self.catalog = Catalog(Path(state_dir or os.environ.get("DECK_STATE_DIR", ROOT/".state"))/"models")
|
self.catalog = Catalog(Path(state_dir or os.environ.get("DECK_STATE_DIR", ROOT/".state"))/"models")
|
||||||
|
self.runtime = Runtime(Path(state_dir or os.environ.get("DECK_STATE_DIR", ROOT/".state"))/"runtime")
|
||||||
self.demo = DemoService()
|
self.demo = DemoService()
|
||||||
self.hardware = HardwareProvider()
|
self.hardware = HardwareProvider()
|
||||||
self.started = time.time()
|
self.started = time.time()
|
||||||
@@ -260,14 +262,19 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
return self.respond({'error':'Anmeldung erforderlich.'},401)
|
return self.respond({'error':'Anmeldung erforderlich.'},401)
|
||||||
if not self.authenticated() and self.path == '/':
|
if not self.authenticated() and self.path == '/':
|
||||||
return self.respond((ROOT/'login.html').read_bytes(), mime='text/html; charset=utf-8')
|
return self.respond((ROOT/'login.html').read_bytes(), mime='text/html; charset=utf-8')
|
||||||
routes = {'/': ('index.html', 'text/html; charset=utf-8'), '/app.js': ('app.js', 'text/javascript'), '/style.css': ('style.css', 'text/css'), '/network-ui.js': ('network-ui.js', 'text/javascript'), '/access-ui.js': ('access-ui.js', 'text/javascript'), '/studio.js': ('studio.js', 'text/javascript'), '/catalog-ui.js': ('catalog-ui.js','text/javascript')}
|
routes = {'/': ('index.html', 'text/html; charset=utf-8'), '/app.js': ('app.js', 'text/javascript'), '/style.css': ('style.css', 'text/css'), '/network-ui.js': ('network-ui.js', 'text/javascript'), '/access-ui.js': ('access-ui.js', 'text/javascript'), '/studio.js': ('studio.js', 'text/javascript'), '/catalog-ui.js': ('catalog-ui.js','text/javascript'), '/runtime-ui.js': ('runtime-ui.js','text/javascript')}
|
||||||
if self.path in routes:
|
if self.path in routes:
|
||||||
name, mime = routes[self.path]
|
name, mime = routes[self.path]
|
||||||
return self.respond((ROOT/name).read_bytes(), mime=mime)
|
return self.respond((ROOT/name).read_bytes(), mime=mime)
|
||||||
if self.path == '/api/v1/status':
|
if self.path == '/api/v1/status':
|
||||||
return self.respond(dict(name='Athena Deck', version='0.5.0', state='ready', uptime_seconds=round(time.time()-self.server.started), mode='isolated', location=os.environ.get('DECK_LOCATION', 'Athena · Debian-Server'), demo=self.server.demo.status()))
|
return self.respond(dict(name='Athena Deck', version='0.6.0', state='ready', uptime_seconds=round(time.time()-self.server.started), mode='isolated', location=os.environ.get('DECK_LOCATION', 'Athena · Debian-Server'), demo=self.server.demo.status()))
|
||||||
if self.path == '/api/v1/hardware':
|
if self.path == '/api/v1/hardware':
|
||||||
return self.respond(self.server.hardware.snapshot())
|
return self.respond(self.server.hardware.snapshot())
|
||||||
|
if self.path.startswith('/api/v1/runtime'):
|
||||||
|
actions={'/api/v1/runtime':self.server.runtime.status,'/api/v1/runtime/prerequisites':self.server.runtime.prerequisites,'/api/v1/runtime/releases':self.server.runtime.releases,'/api/v1/runtime/log':self.server.runtime.log,'/api/v1/runtime/references':self.server.runtime.references}
|
||||||
|
if self.path in actions:
|
||||||
|
try:return self.respond(actions[self.path]())
|
||||||
|
except Exception:return self.respond({'error':'Laufzeitabfrage fehlgeschlagen; Verbindung oder Werkzeuge prüfen.'},503)
|
||||||
if self.path.startswith('/api/v1/catalog'):
|
if self.path.startswith('/api/v1/catalog'):
|
||||||
try:
|
try:
|
||||||
path=urlsplit(self.path);q=parse_qs(path.query)
|
path=urlsplit(self.path);q=parse_qs(path.query)
|
||||||
@@ -311,6 +318,17 @@ class Handler(BaseHTTPRequestHandler):
|
|||||||
with self.server.auth_lock:
|
with self.server.auth_lock:
|
||||||
self.server.sessions.pop(self.session_token(),None)
|
self.server.sessions.pop(self.session_token(),None)
|
||||||
return self.respond({'logged_out':True})
|
return self.respond({'logged_out':True})
|
||||||
|
if self.path.startswith('/api/v1/runtime/'):
|
||||||
|
try:
|
||||||
|
data=self.read_json();action=self.path.rsplit('/',1)[1]
|
||||||
|
if action=='build' and set(data)=={'revision','backend','jobs'}:return self.respond(self.server.runtime.start(**data))
|
||||||
|
if action=='fit' and set(data)=={'profile','context','slots'}:return self.respond(self.server.runtime.fit(**data))
|
||||||
|
if action=='cancel' and not data:return self.respond(self.server.runtime.stop())
|
||||||
|
if action=='activate' and set(data)=={'build_id'}:return self.respond(self.server.runtime.activate(**data))
|
||||||
|
if action=='rollback' and not data:return self.respond(self.server.runtime.rollback())
|
||||||
|
raise ValueError('Ungültige Laufzeitaktion.')
|
||||||
|
except ValueError as exc:return self.respond({'error':str(exc)},400)
|
||||||
|
except (OSError, subprocess.SubprocessError):return self.respond({'error':'Laufzeitaktion fehlgeschlagen; Speicher und Werkzeuge prüfen.'},503)
|
||||||
if self.path in ('/api/v1/catalog/download','/api/v1/catalog/cancel'):
|
if self.path in ('/api/v1/catalog/download','/api/v1/catalog/cancel'):
|
||||||
try:
|
try:
|
||||||
data=self.read_json()
|
data=self.read_json()
|
||||||
@@ -359,6 +377,7 @@ def main():
|
|||||||
try:
|
try:
|
||||||
server.serve_forever()
|
server.serve_forever()
|
||||||
finally:
|
finally:
|
||||||
|
server.runtime.stop()
|
||||||
server.catalog.stop()
|
server.catalog.stop()
|
||||||
server.demo.stop()
|
server.demo.stop()
|
||||||
server.server_close()
|
server.server_close()
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
/* Live catalogue/downloads; runtime and profile drafts remain previews. */
|
/* Live catalogue/downloads; runtime and profile drafts remain previews. */
|
||||||
const Studio=(()=>{
|
const Studio=(()=>{
|
||||||
const key='athena-deck-studio-v1';
|
const key='athena-deck-studio-v1';
|
||||||
const seed={profiles:[{id:'alltag',name:'qwen-alltag',model:'qwen-iq4',context:76800,slots:1,gpu:'both',split:'85,15',cache:'q8_0',batch:2048,ubatch:128,threads:8,unified:true,flash:true,vision:false},{id:'parallel',name:'qwen-parallel',model:'qwen-iq4',context:160000,slots:2,gpu:'both',split:'85,15',cache:'q8_0',batch:2048,ubatch:128,threads:8,unified:true,flash:true,vision:false},{id:'lang',name:'qwen-langkontext',model:'qwen-iq4',context:262144,slots:1,gpu:'both',split:'80,20',cache:'q8_0',batch:2048,ubatch:128,threads:8,unified:true,flash:true,vision:false}],runtime:{backend:'CUDA',revision:'',architectures:'auto',custom:'',jobs:4,toolkit:'auto',reserve:2048}};
|
const seed={profiles:[{id:'alltag',name:'qwen-alltag',model:'qwen-iq4',context:76800,slots:1,gpu:'both',split:'85,15',cache:'q8_0',batch:2048,ubatch:128,threads:8,unified:true,flash:true,vision:false},{id:'parallel',name:'qwen-parallel',model:'qwen-iq4',context:160000,slots:2,gpu:'both',split:'85,15',cache:'q8_0',batch:2048,ubatch:128,threads:8,unified:true,flash:true,vision:false},{id:'lang',name:'qwen-langkontext',model:'qwen-iq4',context:262144,slots:1,gpu:'both',split:'80,20',cache:'q8_0',batch:2048,ubatch:128,threads:8,unified:true,flash:true,vision:false}],runtime:{backend:'CUDA',revision:'',architectures:'auto',custom:'',jobs:4,toolkit:'auto'}};
|
||||||
let state=structuredClone(seed);try{const saved=JSON.parse(localStorage.getItem(key));if(saved?.profiles&&Array.isArray(saved.profiles)&&saved.runtime)state=saved;}catch{}
|
let state=structuredClone(seed);try{const saved=JSON.parse(localStorage.getItem(key));if(saved?.profiles&&Array.isArray(saved.profiles)&&saved.runtime)state=saved;}catch{}
|
||||||
let section='discover',category='',editing=null;
|
let section='discover',category='',editing=null;
|
||||||
const models=[{id:'qwen-iq4',kind:'chat',name:'Qwen 27B',variant:'IQ4_XS · GGUF',source:'Hugging Face',description:'Sprachmodell · Beispiel für mehrere Profile mit einer Modelldatei.'},{id:'qwen-q4',kind:'chat',name:'Qwen 27B',variant:'Q4_K_M · GGUF',source:'Hugging Face',description:'Alternative Quantisierung · eigener Bibliothekseintrag.'},{id:'image-example',kind:'image',name:'Bildmodell',variant:'Modellpaket · Beispiel',source:'Modellkatalog',description:'Textencoder, Diffusionsmodell und VAE als zusammengehöriges Paket.'},{id:'tts-example',kind:'audio',name:'Sprachausgabe',variant:'TTS · Beispiel',source:'Modellkatalog',description:'Sprachmodell und Stimmen mit passender Audio-Laufzeit.'},{id:'asr-example',kind:'audio',name:'Spracherkennung',variant:'ASR · Beispiel',source:'Modellkatalog',description:'Audio in Text umwandeln · CPU-/GPU-Eignung später prüfen.'},{id:'video-example',kind:'video',name:'Videomodell',variant:'Modellpaket · Beispiel',source:'Modellkatalog',description:'Auflösung, Bildanzahl und Zusatzkomponenten gemeinsam planen.'}];
|
const models=[{id:'qwen-iq4',kind:'chat',name:'Qwen 27B',variant:'IQ4_XS · GGUF',source:'Hugging Face',description:'Sprachmodell · Beispiel für mehrere Profile mit einer Modelldatei.'},{id:'qwen-q4',kind:'chat',name:'Qwen 27B',variant:'Q4_K_M · GGUF',source:'Hugging Face',description:'Alternative Quantisierung · eigener Bibliothekseintrag.'},{id:'image-example',kind:'image',name:'Bildmodell',variant:'Modellpaket · Beispiel',source:'Modellkatalog',description:'Textencoder, Diffusionsmodell und VAE als zusammengehöriges Paket.'},{id:'tts-example',kind:'audio',name:'Sprachausgabe',variant:'TTS · Beispiel',source:'Modellkatalog',description:'Sprachmodell und Stimmen mit passender Audio-Laufzeit.'},{id:'asr-example',kind:'audio',name:'Spracherkennung',variant:'ASR · Beispiel',source:'Modellkatalog',description:'Audio in Text umwandeln · CPU-/GPU-Eignung später prüfen.'},{id:'video-example',kind:'video',name:'Videomodell',variant:'Modellpaket · Beispiel',source:'Modellkatalog',description:'Auflösung, Bildanzahl und Zusatzkomponenten gemeinsam planen.'}];
|
||||||
@@ -12,7 +12,7 @@ const Studio=(()=>{
|
|||||||
const check=v=>v?'checked':'';
|
const check=v=>v?'checked':'';
|
||||||
function save(){try{localStorage.setItem(key,JSON.stringify(state));return true;}catch{notice('Speichern im Browser nicht möglich. Entwurf bleibt nur für diese Ansicht erhalten.');return false;}}
|
function save(){try{localStorage.setItem(key,JSON.stringify(state));return true;}catch{notice('Speichern im Browser nicht möglich. Entwurf bleibt nur für diese Ansicht erhalten.');return false;}}
|
||||||
function notice(message){const node=document.querySelector('#studio-notice');if(node)node.textContent=message;}
|
function notice(message){const node=document.querySelector('#studio-notice');if(node)node.textContent=message;}
|
||||||
const banner=()=>'<div class="prototype-banner"><span class="pill">GUI-VORSCHAU</span><span>Katalog und Datei-Downloads sind live. Profile und Laufzeiten bleiben Entwürfe; keine Modellstarts.</span></div>';
|
const banner=()=>'<div class="prototype-banner"><span class="pill">GUI-VORSCHAU</span><span>Katalog und Datei-Downloads sind live. Profile bleiben Entwürfe; Builds sind unter Einstellungen aktiv. Keine Modellstarts.</span></div>';
|
||||||
function header(title,description){return `<div class="kicker">ATHENA / MODELLVERWALTUNG</div><h1>${title}</h1><p>${description}</p>${banner()}`;}
|
function header(title,description){return `<div class="kicker">ATHENA / MODELLVERWALTUNG</div><h1>${title}</h1><p>${description}</p>${banner()}`;}
|
||||||
function render(page,force=false){
|
function render(page,force=false){
|
||||||
if(!force&&document.querySelector('#studio')?.dataset.page===page)return;
|
if(!force&&document.querySelector('#studio')?.dataset.page===page)return;
|
||||||
@@ -41,8 +41,9 @@ const Studio=(()=>{
|
|||||||
if(id)document.querySelector('#delete-profile').onclick=()=>{state.profiles=state.profiles.filter(q=>q.id!==id);save();render(category,true);notice('Lokalen Profilentwurf entfernt. Modelldateien wurden nicht verändert.');};
|
if(id)document.querySelector('#delete-profile').onclick=()=>{state.profiles=state.profiles.filter(q=>q.id!==id);save();render(category,true);notice('Lokalen Profilentwurf entfernt. Modelldateien wurden nicht verändert.');};
|
||||||
document.querySelector('#profile-editor').scrollIntoView({behavior:'smooth',block:'start'});
|
document.querySelector('#profile-editor').scrollIntoView({behavior:'smooth',block:'start'});
|
||||||
}
|
}
|
||||||
function runtime(){const r=state.runtime;return header('llama.cpp-Laufzeit','Installation, GPU-Unterstützung und Versionen an einem Ort vorbereiten.')+`<div class="runtime-stats"><section class="card"><span class="label">INSTALLIERTE VERSION</span><h2>Nicht angebunden</h2><p>Kein Build erkannt oder ausgeführt.</p></section><section class="card"><span class="label">UPDATE-STATUS</span><h2>Noch nicht geprüft</h2><p>Keine Live-Abfrage von Releases.</p></section><section class="card"><span class="label">ZIELSYSTEM</span><h2>Athena</h2><p>Geplant: RTX 5080 + RTX 3060</p></section></div><section class="card"><div class="card-top"><h2>Installation & Build</h2><span class="pill">LOKALER ENTWURF</span></div><form id="build-form"><div class="form-grid"><label>Rechen-Backend<select name="backend">${['CUDA','CPU','Vulkan'].map(v=>`<option ${selected(v,r.backend)}>${v}</option>`).join('')}</select></label>${field('Gewünschtes Release / Commit','revision',r.revision,'text','placeholder="Noch keine Version ausgewählt" maxlength="80"')}<label>GPU-Architekturen<select name="architectures"><option value="auto" ${selected('auto',r.architectures)}>Auf dem Zielserver automatisch erkennen</option><option value="custom" ${selected('custom',r.architectures)}>Explizite Build-Ziele vorgeben</option></select></label>${field('Explizite Architekturziele (optional)','custom',r.custom,'text','placeholder="Später anhand beider GPUs prüfen" maxlength="80"')}<label>CUDA-Toolkit<select name="toolkit"><option value="auto">Verfügbarkeit und Kompatibilität prüfen</option></select></label>${field('Parallele Build-Jobs','jobs',r.jobs,'number','required min="1" max="64"')}${field('Geplante VRAM-Reserve je GPU (MiB)','reserve',r.reserve,'number','required min="256" max="16384" step="256"')}</div><p class="note">Bei verschiedenen GPUs muss der CUDA-Build die Architektur beider Karten unterstützen. Die konkreten Build-Ziele und die passende Toolkit-Version werden später auf Athena geprüft. Diese Vorschau installiert keine Treiber oder Host-Pakete.</p><button>Build-Entwurf speichern</button><button disabled>llama.cpp herunterladen & bauen</button></form></section><section class="card"><div class="card-top"><h2>Updates & Änderungen</h2><button id="release-preview" class="secondary">Update-Ansicht öffnen</button></div><p>Vorgesehen: installierte Version mit verfügbaren Releases vergleichen, Fixes ansehen und einen neuen Build getrennt vorbereiten.</p><div id="release-panel" hidden><div class="release-comparison"><div><span class="label">AKTUELL</span><h3>Nicht ermittelt</h3></div><span>→</span><div><span class="label">VERFÜGBAR</span><h3>Nicht abgefragt</h3></div></div><h3>Release-Notes & Fixes</h3><p>Hier erscheinen später die tatsächlichen Änderungen mit Quelle und Veröffentlichungsdatum. Keine Beispieldaten werden als neue Version oder behobener Fehler ausgegeben.</p><a class="link" href="https://github.com/ggml-org/llama.cpp/releases" target="_blank" rel="noopener noreferrer">Offizielle llama.cpp-Releases öffnen ↗</a><div class="card-actions"><button disabled>Nach Updates suchen</button><button disabled>Neue Version bauen</button></div></div></section><section class="card"><h2>Geplanter Build-Ablauf</h2><ol class="flow"><li><b>Voraussetzungen prüfen</b><span>Compiler, CUDA, beide GPUs und freier Speicher.</span></li><li><b>Version beziehen</b><span>Ausgewähltes Release oder festen Commit herunterladen.</span></li><li><b>Separat bauen & prüfen</b><span>Neue Binärdatei prüfen; laufende Version beibehalten.</span></li><li><b>Gezielt aktivieren</b><span>Nach laufenden Anfragen wechseln; vorigen Build behalten.</span></li></ol><button disabled>Vorherigen Build wiederherstellen</button><p class="note">Kein Build aktiv · keine Warteschlange · keine Wiederherstellungsversion vorhanden.</p></section>`;}
|
function runtime(){return RuntimeUI.html();}
|
||||||
function bind(page){
|
function bind(page){
|
||||||
|
if(page==='runtime'){RuntimeUI.bind();return;}
|
||||||
if(section==='discover'||section==='library')CatalogUI.bind(page,section==='library');
|
if(section==='discover'||section==='library')CatalogUI.bind(page,section==='library');
|
||||||
document.querySelectorAll('[data-section]').forEach(b=>b.onclick=()=>{section=b.dataset.section;render(page,true);});
|
document.querySelectorAll('[data-section]').forEach(b=>b.onclick=()=>{section=b.dataset.section;render(page,true);});
|
||||||
document.querySelectorAll('[data-detail]').forEach(b=>b.onclick=()=>detail(b.dataset.detail));
|
document.querySelectorAll('[data-detail]').forEach(b=>b.onclick=()=>detail(b.dataset.detail));
|
||||||
|
|||||||
@@ -0,0 +1,83 @@
|
|||||||
|
import tempfile
|
||||||
|
import unittest
|
||||||
|
from pathlib import Path
|
||||||
|
from unittest.mock import patch
|
||||||
|
from runtime import Runtime
|
||||||
|
|
||||||
|
class RuntimeTests(unittest.TestCase):
|
||||||
|
def test_untrusted_build_arguments(self):
|
||||||
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
r=Runtime(d)
|
||||||
|
for ref in ('main; reboot','--upload-pack=x','../main','https://evil/repo'):
|
||||||
|
with self.assertRaises(ValueError):r.start(ref)
|
||||||
|
with self.assertRaises(ValueError):r.start('b9000',jobs=32)
|
||||||
|
def test_missing_tools_does_not_spawn(self):
|
||||||
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
r=Runtime(d)
|
||||||
|
with patch.object(r,'prerequisites',return_value={'cuda_ready':False}),patch('runtime.subprocess.Popen') as spawn:
|
||||||
|
with self.assertRaises(ValueError):r.start('b9000')
|
||||||
|
spawn.assert_not_called()
|
||||||
|
def test_selection_rollback_persistence(self):
|
||||||
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
r=Runtime(d);r.state['builds']=[{'id':'one'},{'id':'two'}]
|
||||||
|
with self.assertRaises(ValueError):r.activate('../unknown')
|
||||||
|
r.activate('one');r.activate('two');r.rollback()
|
||||||
|
self.assertEqual(Runtime(d).status()['active'],'one')
|
||||||
|
def test_restart_marks_interrupted(self):
|
||||||
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
r=Runtime(d);r.state['job']={'state':'running'};r.save()
|
||||||
|
self.assertEqual(Runtime(d).status()['job']['state'],'interrupted')
|
||||||
|
def test_cancel_before_command(self):
|
||||||
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
r=Runtime(d);r.stop()
|
||||||
|
with patch('runtime.subprocess.Popen') as spawn:
|
||||||
|
with self.assertRaises(InterruptedError):r.run(['git'],Path(d),'test')
|
||||||
|
spawn.assert_not_called()
|
||||||
|
def test_fit_requires_selected_build_and_bounded_context(self):
|
||||||
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
r=Runtime(d)
|
||||||
|
for context,slots in [(0,1),(160000,99),(True,1)]:
|
||||||
|
with self.assertRaises(ValueError):r.fit('medium',context,slots)
|
||||||
|
with patch.object(r,'references',return_value={'profiles':[dict(id='medium',available=True)]}):
|
||||||
|
with self.assertRaises(ValueError):r.fit('medium',160000,2)
|
||||||
|
def test_official_prereleases_are_not_hidden(self):
|
||||||
|
row=dict(tag_name='b11229',name='build',published_at='date',body='notes',html_url='https://github.com/ggml-org/llama.cpp/releases',prerelease=True,draft=False)
|
||||||
|
with tempfile.TemporaryDirectory() as d,patch('runtime.github',return_value=[row]):
|
||||||
|
self.assertTrue(Runtime(d).releases()['releases'][0]['prerelease'])
|
||||||
|
def test_fit_excludes_busy_gpu_and_preserves_context(self):
|
||||||
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
r=Runtime(d);r.state.update(active='build',builds=[dict(id='build',backend='CUDA',fit_tool=True)])
|
||||||
|
gpus=[dict(uuid='GPU-free',name='3060',total_mib=12000,free_mib=2600),dict(uuid='GPU-busy',name='5080',total_mib=16000,free_mib=173)]
|
||||||
|
ref=dict(id='medium',available=True,file='model.gguf',ubatch=256)
|
||||||
|
result=type('Result',(),dict(returncode=0,stdout='-c 160000 -ngl 4',stderr=''))()
|
||||||
|
with patch.object(r,'references',return_value={'profiles':[ref]}),patch.object(r,'prerequisites',return_value={'gpus':gpus}),patch('runtime.subprocess.run',return_value=result) as run:
|
||||||
|
value=r.fit('medium',160000,2)
|
||||||
|
self.assertTrue(value['success']);self.assertFalse(value['model_loaded'])
|
||||||
|
self.assertEqual(value['excluded_gpus'],['5080'])
|
||||||
|
self.assertEqual(run.call_args.kwargs['env']['CUDA_VISIBLE_DEVICES'],'GPU-free')
|
||||||
|
args=run.call_args.args[0];self.assertEqual(args[args.index('--ctx-size')+1],'160000');self.assertEqual(args[args.index('--parallel')+1],'2')
|
||||||
|
def test_fit_source_adapter_is_explicit_and_idempotent(self):
|
||||||
|
from runtime import prepare_fit_source
|
||||||
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
path=Path(d)/'tools/fit-params/fit-params.cpp';path.parent.mkdir(parents=True)
|
||||||
|
path.write_text(' llama_backend_init();\n')
|
||||||
|
prepare_fit_source(d);first=path.read_text();prepare_fit_source(d)
|
||||||
|
self.assertEqual(first,path.read_text());self.assertIn('params.kv_unified = true',first)
|
||||||
|
path.write_text('unknown upstream layout')
|
||||||
|
with self.assertRaises(ValueError):prepare_fit_source(d)
|
||||||
|
def test_cancels_only_owned_real_process(self):
|
||||||
|
import sys,threading,time
|
||||||
|
with tempfile.TemporaryDirectory() as d:
|
||||||
|
r=Runtime(d);r.state['job']={'state':'running'};errors=[]
|
||||||
|
def work():
|
||||||
|
try:r.run([sys.executable,'-c','import time; time.sleep(60)'],Path(d),'test',70)
|
||||||
|
except InterruptedError:pass
|
||||||
|
except Exception as exc:errors.append(exc)
|
||||||
|
t=threading.Thread(target=work);t.start()
|
||||||
|
for _ in range(100):
|
||||||
|
if r.process is not None:break
|
||||||
|
time.sleep(.01)
|
||||||
|
child=r.process
|
||||||
|
self.assertIsNotNone(child)
|
||||||
|
r.stop();t.join(3)
|
||||||
|
self.assertFalse(t.is_alive());self.assertIsNotNone(child.poll());self.assertEqual(errors,[])
|
||||||
Reference in New Issue
Block a user