Harden and demote experimental Hermes compressor

This commit is contained in:
Mikei386
2026-08-26 01:40:38 +02:00
parent 99285fe81a
commit aaadc57c07
7 changed files with 52 additions and 40 deletions
+2 -2
View File
@@ -12,8 +12,8 @@ XTTS_CACHE_DIR=/data/models/xtts-v2-cache
XTTS_GPU_DEVICE=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
AI_DNS=192.168.1.1
# Dedicated Hermes context compressor on the RTX 3060.
COMPRESSION_MODEL_FILE=qwen3.5-2b-compression/Qwen_Qwen3.5-2B-Q4_K_M.gguf
# Experimental context-compression benchmark service on the RTX 3060.
COMPRESSION_MODEL_FILE=qwen3.5-4b-compression/Qwen_Qwen3.5-4B-Q4_K_M.gguf
COMPRESSION_CONTEXT=65536
COMPRESSION_GPU_DEVICE=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
+6 -3
View File
@@ -10,9 +10,12 @@ zusammen mit dem übrigen Appdata gesichert.
- `compose.yaml` ist der einzige Einstieg für die KI-Dienste auf Athena.
- Genau ein llama.cpp-Profil ist aktiv. Der Router schaltet zwischen Fast,
Medium, Large, Ultra und Uncensored.
- Ein unabhängiges Qwen3.5-2B-Modell auf der RTX 3060 komprimiert Hermes-Chats
über `http://192.168.1.212:8099/v1`. Es läuft ohne Thinking mit 65K Kontext
und blockiert dadurch weder Router noch das aktive 27B-Hauptmodell.
- Ein unabhängiges Qwen3.5-4B-Testmodell läuft auf der RTX 3060 über
`http://192.168.1.212:8099/v1`. Es ist ein optionaler Benchmark-Endpunkt und
ausdrücklich **nicht** global für Hermes-Kompression aktiviert: technische
End-to-End-Tests zeigten trotz hoher Gesamttreue ausgelassene ältere
`KEY=value`-Zustände. Produktive Kompression übernimmt das jeweils gewählte
27B-Hauptprofil.
- Der **Athena Operator** bleibt als einziger hostgebundener administrativer
MCP direkt auf Athena. MCPHub veröffentlicht seinen vorhandenen
WireGuard-HTTP-Endpunkt zentral unter `/mcp/athena-operator`; es gibt keinen
+16 -1
View File
@@ -97,7 +97,7 @@ services:
cap_drop: [ALL]
command:
- --model
- "/models/${COMPRESSION_MODEL_FILE:-qwen3.5-2b-compression/Qwen_Qwen3.5-2B-Q4_K_M.gguf}"
- "/models/${COMPRESSION_MODEL_FILE:-qwen3.5-4b-compression/Qwen_Qwen3.5-4B-Q4_K_M.gguf}"
- --alias
- qwen-compression
- --ctx-size
@@ -117,6 +117,21 @@ services:
- --port
- "8099"
- --metrics
# Official Qwen3.5 non-thinking sampler. The presence penalty is
# essential here: without it one technical YAML benchmark repeated the
# summary structure for more than 40K tokens.
- --temperature
- "1.0"
- --top-p
- "1.0"
- --top-k
- "20"
- --presence-penalty
- "2.0"
# Hermes' own summary ceiling is 10K. Keep a little completion margin,
# but never allow an accidental repetition loop to fill the 65K slot.
- --n-predict
- "12000"
- --n-gpu-layers
- all
- --device
+3 -3
View File
@@ -72,9 +72,9 @@ EXPERIMENTAL_MODEL_SHA256=40fac4050e940397dbf13087afd50f4734a11805bf9d65ef8ddd74
VISION_PROJECTOR_FILE=qwen/mmproj-BF16.gguf
VISION_PROJECTOR_URL=https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/resolve/main/mmproj-BF16.gguf
VISION_PROJECTOR_SHA256=83ee4f4f205fa514161778c41df1ea14144faa0f713510893b63c2395f5c2d53
COMPRESSION_MODEL_FILE=qwen3.5-2b-compression/Qwen_Qwen3.5-2B-Q4_K_M.gguf
COMPRESSION_MODEL_URL=https://huggingface.co/bartowski/Qwen_Qwen3.5-2B-GGUF/resolve/main/Qwen_Qwen3.5-2B-Q4_K_M.gguf
COMPRESSION_MODEL_SHA256=57a1085840f497d764a7fc5d346922dbde961efb54cc792ea81d694fd846a1d8
COMPRESSION_MODEL_FILE=qwen3.5-4b-compression/Qwen_Qwen3.5-4B-Q4_K_M.gguf
COMPRESSION_MODEL_URL=https://huggingface.co/bartowski/Qwen_Qwen3.5-4B-GGUF/resolve/main/Qwen_Qwen3.5-4B-Q4_K_M.gguf
COMPRESSION_MODEL_SHA256=13c16f426047e2de38cd075bdade4a7bcbc8c774384876f677740cda65f8a983
COMPRESSION_CONTEXT=65536
COMPRESSION_GPU_DEVICE=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
# All standard profiles use the MTP tensor embedded in their GGUF. A separate
+18 -11
View File
@@ -30,19 +30,26 @@ und einer bewussten Aktualisierung dieser Datei.
- Das gesonderte Experimentalprofil gehört nicht zur Benutzer-Matrix und wird
in Open WebUI nicht als reguläres Modell angeboten.
## Hermes-Kompressionsmodell
## Experimentelles Hermes-Kompressionsmodell
Das Hilfsmodell `qwen-compression` ist kein Chatprofil. Qwen3.5-2B Q4_K_M
läuft unabhängig mit 65.536 Token Kontext auf der RTX 3060 und ist nur über
WireGuard auf Port 8099 erreichbar. Hermes erzwingt dafür Non-Thinking.
Der Alias `qwen-compression` ist kein Chatprofil. Qwen3.5-4B Q4_K_M läuft
unabhängig mit 65.536 Token Kontext auf der RTX 3060 und ist nur über
WireGuard auf Port 8099 erreichbar. Non-Thinking und ein 12K-Notausgangslimit
verhindern bekannte Wiederholungsschleifen. Das Modell ist absichtlich **nicht
global in Hermes aktiviert**; produktive Kompression nutzt das gewählte
27B-Hauptprofil.
- 30.245 Token, kalter Prompt: 4.519,7 Prefill-Tok/s.
- Ausgabe im 30K-Qualitätstest: 117,1 Tok/s; vier von fünf ungefilterten
Ankern bewahrt.
- Echter Hermes-Hilfsmodellaufruf: 158 Token in 4,6 Sekunden; alle geforderten
Pfade, Ports, Verbote und der Commit-Anker wurden bewahrt.
- Zusätzlicher Speicherbedarf auf der RTX 3060: rund 1,8 GiB; nach dem Laden
blieben im gemessenen Produktionszustand rund 4,5 GiB frei.
- 12 technische Fälle: 79,27 % exakte Anker im reinen Modellteil und 98,78 %
im vollständigen Hermes-Lean-Ergebnis; 10 von 12 formal bestanden.
- Echter 205K-Token-End-to-End-Lauf: Kompression auf 22,8K in 199 Sekunden,
aber ältere `BUILD_ID`- und kurze Git-HEAD-Zustände gingen verloren.
- Lange technische Einzeltests benötigten 120 bis 185 Sekunden; Ausgabe bei
großem Kontext etwa 50 Tok/s.
- Speicherbelegung der RTX 3060 im gemessenen Gesamtzustand: 9.450 MiB,
verbleibend rund 2.460 MiB.
- Sicherheitsentscheidung: als reproduzierbarer Testdienst behalten, aber
erst nach zuverlässig vollständiger technischer Zustandsbewahrung als
globalen Kompressor freigeben.
## Nachweise
+1 -11
View File
@@ -86,7 +86,7 @@ compression:
threshold: 0.82
target_ratio: 0.35
tail_mode: "lean"
protect_last_n: 12
protect_last_n: 20
protect_first_n: 0
proactive_prune_tokens: 50000
proactive_prune_min_result_chars: 4000
@@ -99,16 +99,6 @@ auxiliary:
# actual chat, so keep the original timestamp/session id instead.
title_generation:
enabled: false
compression:
provider: "openai-api"
model: "qwen-compression"
base_url: "http://192.168.1.212:8099/v1"
api_key: "local"
reasoning_effort: "none"
extra_body:
chat_template_kwargs:
enable_thinking: false
timeout: 600
# The full hermes-cli preset injects several large, overlapping schemas on
# every turn. Athena already exposes browsing, orchestration and host services
+6 -9
View File
@@ -33,16 +33,13 @@ create_profile() {
docker exec "$HERMES_CONTAINER" hermes -p "$name" config unset compression.threshold_tokens || true
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.threshold 0.82
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.target_ratio 0.35
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.protect_last_n 12
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.protect_last_n 20
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.protect_first_n 0
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.provider openai-api
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.model qwen-compression
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.base_url http://192.168.1.212:8099/v1
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.api_key local
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.reasoning_effort none
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.extra_body \
'{"chat_template_kwargs":{"enable_thinking":false}}'
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.timeout 600
# The small compressor remains an opt-in benchmark endpoint. It is not safe
# as a global technical-session compressor: end-to-end tests showed that it
# can omit older exact KEY=value state. Remove stale opt-in settings so the
# selected 27B profile performs production compaction.
docker exec "$HERMES_CONTAINER" hermes -p "$name" config unset auxiliary.compression || true
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set platform_toolsets.cli \
'["web","terminal","file","skills","todo","memory","vision","tts"]'
[[ $(docker exec "$HERMES_CONTAINER" hermes -p "$name" config get model.default) == "$model" ]] || \