Harden and demote experimental Hermes compressor
This commit is contained in:
+2
-2
@@ -12,8 +12,8 @@ XTTS_CACHE_DIR=/data/models/xtts-v2-cache
|
|||||||
XTTS_GPU_DEVICE=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
|
XTTS_GPU_DEVICE=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
|
||||||
AI_DNS=192.168.1.1
|
AI_DNS=192.168.1.1
|
||||||
|
|
||||||
# Dedicated Hermes context compressor on the RTX 3060.
|
# Experimental context-compression benchmark service on the RTX 3060.
|
||||||
COMPRESSION_MODEL_FILE=qwen3.5-2b-compression/Qwen_Qwen3.5-2B-Q4_K_M.gguf
|
COMPRESSION_MODEL_FILE=qwen3.5-4b-compression/Qwen_Qwen3.5-4B-Q4_K_M.gguf
|
||||||
COMPRESSION_CONTEXT=65536
|
COMPRESSION_CONTEXT=65536
|
||||||
COMPRESSION_GPU_DEVICE=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
|
COMPRESSION_GPU_DEVICE=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
|
||||||
|
|
||||||
|
|||||||
@@ -10,9 +10,12 @@ zusammen mit dem übrigen Appdata gesichert.
|
|||||||
- `compose.yaml` ist der einzige Einstieg für die KI-Dienste auf Athena.
|
- `compose.yaml` ist der einzige Einstieg für die KI-Dienste auf Athena.
|
||||||
- Genau ein llama.cpp-Profil ist aktiv. Der Router schaltet zwischen Fast,
|
- Genau ein llama.cpp-Profil ist aktiv. Der Router schaltet zwischen Fast,
|
||||||
Medium, Large, Ultra und Uncensored.
|
Medium, Large, Ultra und Uncensored.
|
||||||
- Ein unabhängiges Qwen3.5-2B-Modell auf der RTX 3060 komprimiert Hermes-Chats
|
- Ein unabhängiges Qwen3.5-4B-Testmodell läuft auf der RTX 3060 über
|
||||||
über `http://192.168.1.212:8099/v1`. Es läuft ohne Thinking mit 65K Kontext
|
`http://192.168.1.212:8099/v1`. Es ist ein optionaler Benchmark-Endpunkt und
|
||||||
und blockiert dadurch weder Router noch das aktive 27B-Hauptmodell.
|
ausdrücklich **nicht** global für Hermes-Kompression aktiviert: technische
|
||||||
|
End-to-End-Tests zeigten trotz hoher Gesamttreue ausgelassene ältere
|
||||||
|
`KEY=value`-Zustände. Produktive Kompression übernimmt das jeweils gewählte
|
||||||
|
27B-Hauptprofil.
|
||||||
- Der **Athena Operator** bleibt als einziger hostgebundener administrativer
|
- Der **Athena Operator** bleibt als einziger hostgebundener administrativer
|
||||||
MCP direkt auf Athena. MCPHub veröffentlicht seinen vorhandenen
|
MCP direkt auf Athena. MCPHub veröffentlicht seinen vorhandenen
|
||||||
WireGuard-HTTP-Endpunkt zentral unter `/mcp/athena-operator`; es gibt keinen
|
WireGuard-HTTP-Endpunkt zentral unter `/mcp/athena-operator`; es gibt keinen
|
||||||
|
|||||||
+16
-1
@@ -97,7 +97,7 @@ services:
|
|||||||
cap_drop: [ALL]
|
cap_drop: [ALL]
|
||||||
command:
|
command:
|
||||||
- --model
|
- --model
|
||||||
- "/models/${COMPRESSION_MODEL_FILE:-qwen3.5-2b-compression/Qwen_Qwen3.5-2B-Q4_K_M.gguf}"
|
- "/models/${COMPRESSION_MODEL_FILE:-qwen3.5-4b-compression/Qwen_Qwen3.5-4B-Q4_K_M.gguf}"
|
||||||
- --alias
|
- --alias
|
||||||
- qwen-compression
|
- qwen-compression
|
||||||
- --ctx-size
|
- --ctx-size
|
||||||
@@ -117,6 +117,21 @@ services:
|
|||||||
- --port
|
- --port
|
||||||
- "8099"
|
- "8099"
|
||||||
- --metrics
|
- --metrics
|
||||||
|
# Official Qwen3.5 non-thinking sampler. The presence penalty is
|
||||||
|
# essential here: without it one technical YAML benchmark repeated the
|
||||||
|
# summary structure for more than 40K tokens.
|
||||||
|
- --temperature
|
||||||
|
- "1.0"
|
||||||
|
- --top-p
|
||||||
|
- "1.0"
|
||||||
|
- --top-k
|
||||||
|
- "20"
|
||||||
|
- --presence-penalty
|
||||||
|
- "2.0"
|
||||||
|
# Hermes' own summary ceiling is 10K. Keep a little completion margin,
|
||||||
|
# but never allow an accidental repetition loop to fill the 65K slot.
|
||||||
|
- --n-predict
|
||||||
|
- "12000"
|
||||||
- --n-gpu-layers
|
- --n-gpu-layers
|
||||||
- all
|
- all
|
||||||
- --device
|
- --device
|
||||||
|
|||||||
@@ -72,9 +72,9 @@ EXPERIMENTAL_MODEL_SHA256=40fac4050e940397dbf13087afd50f4734a11805bf9d65ef8ddd74
|
|||||||
VISION_PROJECTOR_FILE=qwen/mmproj-BF16.gguf
|
VISION_PROJECTOR_FILE=qwen/mmproj-BF16.gguf
|
||||||
VISION_PROJECTOR_URL=https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/resolve/main/mmproj-BF16.gguf
|
VISION_PROJECTOR_URL=https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/resolve/main/mmproj-BF16.gguf
|
||||||
VISION_PROJECTOR_SHA256=83ee4f4f205fa514161778c41df1ea14144faa0f713510893b63c2395f5c2d53
|
VISION_PROJECTOR_SHA256=83ee4f4f205fa514161778c41df1ea14144faa0f713510893b63c2395f5c2d53
|
||||||
COMPRESSION_MODEL_FILE=qwen3.5-2b-compression/Qwen_Qwen3.5-2B-Q4_K_M.gguf
|
COMPRESSION_MODEL_FILE=qwen3.5-4b-compression/Qwen_Qwen3.5-4B-Q4_K_M.gguf
|
||||||
COMPRESSION_MODEL_URL=https://huggingface.co/bartowski/Qwen_Qwen3.5-2B-GGUF/resolve/main/Qwen_Qwen3.5-2B-Q4_K_M.gguf
|
COMPRESSION_MODEL_URL=https://huggingface.co/bartowski/Qwen_Qwen3.5-4B-GGUF/resolve/main/Qwen_Qwen3.5-4B-Q4_K_M.gguf
|
||||||
COMPRESSION_MODEL_SHA256=57a1085840f497d764a7fc5d346922dbde961efb54cc792ea81d694fd846a1d8
|
COMPRESSION_MODEL_SHA256=13c16f426047e2de38cd075bdade4a7bcbc8c774384876f677740cda65f8a983
|
||||||
COMPRESSION_CONTEXT=65536
|
COMPRESSION_CONTEXT=65536
|
||||||
COMPRESSION_GPU_DEVICE=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
|
COMPRESSION_GPU_DEVICE=GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b
|
||||||
# All standard profiles use the MTP tensor embedded in their GGUF. A separate
|
# All standard profiles use the MTP tensor embedded in their GGUF. A separate
|
||||||
|
|||||||
@@ -30,19 +30,26 @@ und einer bewussten Aktualisierung dieser Datei.
|
|||||||
- Das gesonderte Experimentalprofil gehört nicht zur Benutzer-Matrix und wird
|
- Das gesonderte Experimentalprofil gehört nicht zur Benutzer-Matrix und wird
|
||||||
in Open WebUI nicht als reguläres Modell angeboten.
|
in Open WebUI nicht als reguläres Modell angeboten.
|
||||||
|
|
||||||
## Hermes-Kompressionsmodell
|
## Experimentelles Hermes-Kompressionsmodell
|
||||||
|
|
||||||
Das Hilfsmodell `qwen-compression` ist kein Chatprofil. Qwen3.5-2B Q4_K_M
|
Der Alias `qwen-compression` ist kein Chatprofil. Qwen3.5-4B Q4_K_M läuft
|
||||||
läuft unabhängig mit 65.536 Token Kontext auf der RTX 3060 und ist nur über
|
unabhängig mit 65.536 Token Kontext auf der RTX 3060 und ist nur über
|
||||||
WireGuard auf Port 8099 erreichbar. Hermes erzwingt dafür Non-Thinking.
|
WireGuard auf Port 8099 erreichbar. Non-Thinking und ein 12K-Notausgangslimit
|
||||||
|
verhindern bekannte Wiederholungsschleifen. Das Modell ist absichtlich **nicht
|
||||||
|
global in Hermes aktiviert**; produktive Kompression nutzt das gewählte
|
||||||
|
27B-Hauptprofil.
|
||||||
|
|
||||||
- 30.245 Token, kalter Prompt: 4.519,7 Prefill-Tok/s.
|
- 12 technische Fälle: 79,27 % exakte Anker im reinen Modellteil und 98,78 %
|
||||||
- Ausgabe im 30K-Qualitätstest: 117,1 Tok/s; vier von fünf ungefilterten
|
im vollständigen Hermes-Lean-Ergebnis; 10 von 12 formal bestanden.
|
||||||
Ankern bewahrt.
|
- Echter 205K-Token-End-to-End-Lauf: Kompression auf 22,8K in 199 Sekunden,
|
||||||
- Echter Hermes-Hilfsmodellaufruf: 158 Token in 4,6 Sekunden; alle geforderten
|
aber ältere `BUILD_ID`- und kurze Git-HEAD-Zustände gingen verloren.
|
||||||
Pfade, Ports, Verbote und der Commit-Anker wurden bewahrt.
|
- Lange technische Einzeltests benötigten 120 bis 185 Sekunden; Ausgabe bei
|
||||||
- Zusätzlicher Speicherbedarf auf der RTX 3060: rund 1,8 GiB; nach dem Laden
|
großem Kontext etwa 50 Tok/s.
|
||||||
blieben im gemessenen Produktionszustand rund 4,5 GiB frei.
|
- Speicherbelegung der RTX 3060 im gemessenen Gesamtzustand: 9.450 MiB,
|
||||||
|
verbleibend rund 2.460 MiB.
|
||||||
|
- Sicherheitsentscheidung: als reproduzierbarer Testdienst behalten, aber
|
||||||
|
erst nach zuverlässig vollständiger technischer Zustandsbewahrung als
|
||||||
|
globalen Kompressor freigeben.
|
||||||
|
|
||||||
## Nachweise
|
## Nachweise
|
||||||
|
|
||||||
|
|||||||
@@ -86,7 +86,7 @@ compression:
|
|||||||
threshold: 0.82
|
threshold: 0.82
|
||||||
target_ratio: 0.35
|
target_ratio: 0.35
|
||||||
tail_mode: "lean"
|
tail_mode: "lean"
|
||||||
protect_last_n: 12
|
protect_last_n: 20
|
||||||
protect_first_n: 0
|
protect_first_n: 0
|
||||||
proactive_prune_tokens: 50000
|
proactive_prune_tokens: 50000
|
||||||
proactive_prune_min_result_chars: 4000
|
proactive_prune_min_result_chars: 4000
|
||||||
@@ -99,16 +99,6 @@ auxiliary:
|
|||||||
# actual chat, so keep the original timestamp/session id instead.
|
# actual chat, so keep the original timestamp/session id instead.
|
||||||
title_generation:
|
title_generation:
|
||||||
enabled: false
|
enabled: false
|
||||||
compression:
|
|
||||||
provider: "openai-api"
|
|
||||||
model: "qwen-compression"
|
|
||||||
base_url: "http://192.168.1.212:8099/v1"
|
|
||||||
api_key: "local"
|
|
||||||
reasoning_effort: "none"
|
|
||||||
extra_body:
|
|
||||||
chat_template_kwargs:
|
|
||||||
enable_thinking: false
|
|
||||||
timeout: 600
|
|
||||||
|
|
||||||
# The full hermes-cli preset injects several large, overlapping schemas on
|
# The full hermes-cli preset injects several large, overlapping schemas on
|
||||||
# every turn. Athena already exposes browsing, orchestration and host services
|
# every turn. Athena already exposes browsing, orchestration and host services
|
||||||
|
|||||||
@@ -33,16 +33,13 @@ create_profile() {
|
|||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config unset compression.threshold_tokens || true
|
docker exec "$HERMES_CONTAINER" hermes -p "$name" config unset compression.threshold_tokens || true
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.threshold 0.82
|
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.threshold 0.82
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.target_ratio 0.35
|
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.target_ratio 0.35
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.protect_last_n 12
|
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.protect_last_n 20
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.protect_first_n 0
|
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set compression.protect_first_n 0
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.provider openai-api
|
# The small compressor remains an opt-in benchmark endpoint. It is not safe
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.model qwen-compression
|
# as a global technical-session compressor: end-to-end tests showed that it
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.base_url http://192.168.1.212:8099/v1
|
# can omit older exact KEY=value state. Remove stale opt-in settings so the
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.api_key local
|
# selected 27B profile performs production compaction.
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.reasoning_effort none
|
docker exec "$HERMES_CONTAINER" hermes -p "$name" config unset auxiliary.compression || true
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.extra_body \
|
|
||||||
'{"chat_template_kwargs":{"enable_thinking":false}}'
|
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set auxiliary.compression.timeout 600
|
|
||||||
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set platform_toolsets.cli \
|
docker exec "$HERMES_CONTAINER" hermes -p "$name" config set platform_toolsets.cli \
|
||||||
'["web","terminal","file","skills","todo","memory","vision","tts"]'
|
'["web","terminal","file","skills","todo","memory","vision","tts"]'
|
||||||
[[ $(docker exec "$HERMES_CONTAINER" hermes -p "$name" config get model.default) == "$model" ]] || \
|
[[ $(docker exec "$HERMES_CONTAINER" hermes -p "$name" config get model.default) == "$model" ]] || \
|
||||||
|
|||||||
Reference in New Issue
Block a user