Enable cross-chat llama prompt cache reuse

This commit is contained in:
Mikei386
2026-08-25 19:19:49 +02:00
parent 57afdd4f15
commit d921dfaf69
7 changed files with 60 additions and 5 deletions
+33
View File
@@ -90,6 +90,14 @@ services:
- q4_0 - q4_0
- --cache-type-v - --cache-type-v
- q4_0 - q4_0
# Keep the cross-chat prefix cache explicit. cache-reuse tolerates small
# changes after Hermes' stable system-prompt prefix without enabling the
# currently unreliable on-disk slot restore path.
- --cache-prompt
- --cache-reuse
- "${LLAMA_CACHE_REUSE:-256}"
- --cache-ram
- "${LLAMA_CACHE_RAM_MIB:-8192}"
- --threads - --threads
- "${LLAMA_THREADS:-6}" - "${LLAMA_THREADS:-6}"
- --threads-batch - --threads-batch
@@ -163,6 +171,11 @@ services:
- q4_0 - q4_0
- --cache-type-v - --cache-type-v
- q4_0 - q4_0
- --cache-prompt
- --cache-reuse
- "${LLAMA_CACHE_REUSE:-256}"
- --cache-ram
- "${LLAMA_CACHE_RAM_MIB:-8192}"
- --threads - --threads
- "${LLAMA_THREADS:-6}" - "${LLAMA_THREADS:-6}"
- --threads-batch - --threads-batch
@@ -247,6 +260,11 @@ services:
- q4_0 - q4_0
- --cache-type-v - --cache-type-v
- q4_0 - q4_0
- --cache-prompt
- --cache-reuse
- "${LLAMA_CACHE_REUSE:-256}"
- --cache-ram
- "${LLAMA_CACHE_RAM_MIB:-8192}"
- --threads - --threads
- "${LLAMA_THREADS:-6}" - "${LLAMA_THREADS:-6}"
- --threads-batch - --threads-batch
@@ -319,6 +337,11 @@ services:
- q4_0 - q4_0
- --cache-type-v - --cache-type-v
- q4_0 - q4_0
- --cache-prompt
- --cache-reuse
- "${LLAMA_CACHE_REUSE:-256}"
- --cache-ram
- "${LLAMA_CACHE_RAM_MIB:-8192}"
- --threads - --threads
- "${LLAMA_THREADS:-6}" - "${LLAMA_THREADS:-6}"
- --threads-batch - --threads-batch
@@ -397,6 +420,11 @@ services:
- q4_0 - q4_0
- --cache-type-v - --cache-type-v
- q4_0 - q4_0
- --cache-prompt
- --cache-reuse
- "${LLAMA_CACHE_REUSE:-256}"
- --cache-ram
- "${LLAMA_CACHE_RAM_MIB:-8192}"
- --threads - --threads
- "${LLAMA_THREADS:-6}" - "${LLAMA_THREADS:-6}"
- --threads-batch - --threads-batch
@@ -469,6 +497,11 @@ services:
- q4_0 - q4_0
- --cache-type-v - --cache-type-v
- q4_0 - q4_0
- --cache-prompt
- --cache-reuse
- "${LLAMA_CACHE_REUSE:-256}"
- --cache-ram
- "${LLAMA_CACHE_RAM_MIB:-8192}"
- --parallel - --parallel
- "1" - "1"
- --jinja - --jinja
+15
View File
@@ -43,6 +43,9 @@ Zielplattform.
- RTX 5080 + RTX 3060 im Verhältnis 90:10 - RTX 5080 + RTX 3060 im Verhältnis 90:10
- Flash Attention - Flash Attention
- KV-Cache Q4_0 für K und V - KV-Cache Q4_0 für K und V
- explizites Prompt-Caching mit 8.192 MiB profilinternem RAM-Cache
- `--cache-reuse 256` für die Wiederverwendung langer stabiler Präfixe trotz
kleiner späterer Abweichungen
- MTP Draft, maximal drei Tokens - MTP Draft, maximal drei Tokens
- MTP-Akzeptanzschwelle 0,05; im Referenzlauf 77,26 statt 73,88 Tok/s - MTP-Akzeptanzschwelle 0,05; im Referenzlauf 77,26 statt 73,88 Tok/s
- sechs Threads und sechs Batch-Threads - sechs Threads und sechs Batch-Threads
@@ -96,6 +99,18 @@ Der Router übernimmt:
## Hermes Agent ## Hermes Agent
- Hermes 0.20.5 baut den Systemprompt bereits in drei geordneten Bereichen:
einen chatübergreifend stabilen Präfix, sitzungsstabilen Kontext und einen
variablen Nachlauf. Der stabile Präfix bleibt bei gleicher Profil- und
Werkzeugkonfiguration wortgleich und kann dadurch vom llama.cpp-RAM-Cache
wiederverwendet werden.
- Der RAM-Promptcache lebt nur so lange wie der jeweilige llama.cpp-Prozess.
Ein Profilwechsel entlädt das bisherige Modell und damit dessen Cache.
- Persistente Slot-Dateien (`--slot-save-path`) sind vorerst bewusst nicht
aktiviert. Die aktuelle llama.cpp-Linie hat offene Restore-Fehler; ein
gemeldetes erfolgreiches Restore kann trotzdem einen vollständigen Prefill
auslösen. Erst nach einem isolierten Regressionstest aktivieren.
- Kontextkompression läuft spätestens bei 60.000 Token; die relative - Kontextkompression läuft spätestens bei 60.000 Token; die relative
65-Prozent-Grenze greift nur, wenn sie noch früher erreicht wird. Damit gilt 65-Prozent-Grenze greift nur, wenn sie noch früher erreicht wird. Damit gilt
dieselbe Obergrenze auch für Medium, Large und Ultra und ein Profilwechsel dieselbe Obergrenze auch für Medium, Large und Ultra und ein Profilwechsel
+8 -1
View File
@@ -51,6 +51,14 @@ Der Cache `/var/lib/docker/unraid-update-status.json` ist nur ein
Kandidatenhinweis. Er darf nie allein eine Neuerstellung auslösen. Autoritativ Kandidatenhinweis. Er darf nie allein eine Neuerstellung auslösen. Autoritativ
ist der Image-ID-Vergleich nach dem Pull. ist der Image-ID-Vergleich nach dem Pull.
Die Unraid-Weboberfläche und mobile Ansichten lesen weiterhin diesen separaten
Cache. Ein technisch verifizierter Pull/Rebuild aktualisiert dessen Anzeige
nicht zwingend sofort. Deshalb kann dort weiterhin „Apply Update“ stehen,
obwohl der lokale Image-ID-Vergleich bereits `already-current` ergeben hat.
Für eine frische Anzeige muss Unraids eigener Statuslauf
`dynamix.docker.manager/scripts/dockerupdate check` abgeschlossen sein. Das ist
eine Aktualisierung der Anzeige und kein erneuter Container-Rebuild.
Ein wiederholter Lauf muss bei einem aktuellen Image folgendes melden: Ein wiederholter Lauf muss bei einem aktuellen Image folgendes melden:
```text ```text
@@ -86,4 +94,3 @@ einzelne Update-Aufrufe. Dieser Pfad ist ersetzt.
4. Ein ausdrücklich schreibender synthetischer Auftrag muss beide MUA-Zugänge 4. Ein ausdrücklich schreibender synthetischer Auftrag muss beide MUA-Zugänge
bereitstellen und das Batch-Werkzeug wählen. bereitstellen und das Batch-Werkzeug wählen.
5. Ein Wiederholungstest mit aktuellem Image darf keine Neuerstellung auslösen. 5. Ein Wiederholungstest mit aktuellem Image darf keine Neuerstellung auslösen.
+1 -1
View File
@@ -3,4 +3,4 @@ Description=Local AI llama.cpp - Qwen Fast 76.8K MTP2 with CPU Vision
[Service] [Service]
ExecStart= ExecStart=
ExecStart=/opt/mike-ai/llama.cpp-nvfp4/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-mix/Qwen3.8-27B-IQ4-MIX.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen38-27b-iq4mix-76k-mtp2-vision --ctx-size 76800 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0 --split-mode none --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16 ExecStart=/opt/mike-ai/llama.cpp-nvfp4/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-mix/Qwen3.8-27B-IQ4-MIX.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen38-27b-iq4mix-76k-mtp2-vision --ctx-size 76800 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0 --split-mode none --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
+1 -1
View File
@@ -3,4 +3,4 @@ Description=Legacy native Qwen Large 192K profile (Docker is the production path
[Service] [Service]
ExecStart= ExecStart=
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-large --ctx-size 192000 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 86,14 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-type-k f16 --spec-draft-type-v f16 ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-large --ctx-size 192000 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 86,14 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-type-k f16 --spec-draft-type-v f16
+1 -1
View File
@@ -3,4 +3,4 @@ Description=Local AI llama.cpp - Qwen Medium 160K IQ4_XS Pure with CPU Vision
[Service] [Service]
ExecStart= ExecStart=
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-medium --ctx-size 160000 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --reasoning-budget 8192 --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 1.0 --top-p 0.95 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 90,10 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-type-k f16 --spec-draft-type-v f16 ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-medium --ctx-size 160000 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --reasoning-budget 8192 --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 1.0 --top-p 0.95 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 90,10 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-type-k f16 --spec-draft-type-v f16
+1 -1
View File
@@ -3,4 +3,4 @@ Description=Legacy native Qwen Ultra 256K profile (Docker is the production path
[Service] [Service]
ExecStart= ExecStart=
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --alias qwen-ultra --ctx-size 262144 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16 ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --alias qwen-ultra --ctx-size 262144 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16