Enable cross-chat llama prompt cache reuse
This commit is contained in:
@@ -90,6 +90,14 @@ services:
|
|||||||
- q4_0
|
- q4_0
|
||||||
- --cache-type-v
|
- --cache-type-v
|
||||||
- q4_0
|
- q4_0
|
||||||
|
# Keep the cross-chat prefix cache explicit. cache-reuse tolerates small
|
||||||
|
# changes after Hermes' stable system-prompt prefix without enabling the
|
||||||
|
# currently unreliable on-disk slot restore path.
|
||||||
|
- --cache-prompt
|
||||||
|
- --cache-reuse
|
||||||
|
- "${LLAMA_CACHE_REUSE:-256}"
|
||||||
|
- --cache-ram
|
||||||
|
- "${LLAMA_CACHE_RAM_MIB:-8192}"
|
||||||
- --threads
|
- --threads
|
||||||
- "${LLAMA_THREADS:-6}"
|
- "${LLAMA_THREADS:-6}"
|
||||||
- --threads-batch
|
- --threads-batch
|
||||||
@@ -163,6 +171,11 @@ services:
|
|||||||
- q4_0
|
- q4_0
|
||||||
- --cache-type-v
|
- --cache-type-v
|
||||||
- q4_0
|
- q4_0
|
||||||
|
- --cache-prompt
|
||||||
|
- --cache-reuse
|
||||||
|
- "${LLAMA_CACHE_REUSE:-256}"
|
||||||
|
- --cache-ram
|
||||||
|
- "${LLAMA_CACHE_RAM_MIB:-8192}"
|
||||||
- --threads
|
- --threads
|
||||||
- "${LLAMA_THREADS:-6}"
|
- "${LLAMA_THREADS:-6}"
|
||||||
- --threads-batch
|
- --threads-batch
|
||||||
@@ -247,6 +260,11 @@ services:
|
|||||||
- q4_0
|
- q4_0
|
||||||
- --cache-type-v
|
- --cache-type-v
|
||||||
- q4_0
|
- q4_0
|
||||||
|
- --cache-prompt
|
||||||
|
- --cache-reuse
|
||||||
|
- "${LLAMA_CACHE_REUSE:-256}"
|
||||||
|
- --cache-ram
|
||||||
|
- "${LLAMA_CACHE_RAM_MIB:-8192}"
|
||||||
- --threads
|
- --threads
|
||||||
- "${LLAMA_THREADS:-6}"
|
- "${LLAMA_THREADS:-6}"
|
||||||
- --threads-batch
|
- --threads-batch
|
||||||
@@ -319,6 +337,11 @@ services:
|
|||||||
- q4_0
|
- q4_0
|
||||||
- --cache-type-v
|
- --cache-type-v
|
||||||
- q4_0
|
- q4_0
|
||||||
|
- --cache-prompt
|
||||||
|
- --cache-reuse
|
||||||
|
- "${LLAMA_CACHE_REUSE:-256}"
|
||||||
|
- --cache-ram
|
||||||
|
- "${LLAMA_CACHE_RAM_MIB:-8192}"
|
||||||
- --threads
|
- --threads
|
||||||
- "${LLAMA_THREADS:-6}"
|
- "${LLAMA_THREADS:-6}"
|
||||||
- --threads-batch
|
- --threads-batch
|
||||||
@@ -397,6 +420,11 @@ services:
|
|||||||
- q4_0
|
- q4_0
|
||||||
- --cache-type-v
|
- --cache-type-v
|
||||||
- q4_0
|
- q4_0
|
||||||
|
- --cache-prompt
|
||||||
|
- --cache-reuse
|
||||||
|
- "${LLAMA_CACHE_REUSE:-256}"
|
||||||
|
- --cache-ram
|
||||||
|
- "${LLAMA_CACHE_RAM_MIB:-8192}"
|
||||||
- --threads
|
- --threads
|
||||||
- "${LLAMA_THREADS:-6}"
|
- "${LLAMA_THREADS:-6}"
|
||||||
- --threads-batch
|
- --threads-batch
|
||||||
@@ -469,6 +497,11 @@ services:
|
|||||||
- q4_0
|
- q4_0
|
||||||
- --cache-type-v
|
- --cache-type-v
|
||||||
- q4_0
|
- q4_0
|
||||||
|
- --cache-prompt
|
||||||
|
- --cache-reuse
|
||||||
|
- "${LLAMA_CACHE_REUSE:-256}"
|
||||||
|
- --cache-ram
|
||||||
|
- "${LLAMA_CACHE_RAM_MIB:-8192}"
|
||||||
- --parallel
|
- --parallel
|
||||||
- "1"
|
- "1"
|
||||||
- --jinja
|
- --jinja
|
||||||
|
|||||||
@@ -43,6 +43,9 @@ Zielplattform.
|
|||||||
- RTX 5080 + RTX 3060 im Verhältnis 90:10
|
- RTX 5080 + RTX 3060 im Verhältnis 90:10
|
||||||
- Flash Attention
|
- Flash Attention
|
||||||
- KV-Cache Q4_0 für K und V
|
- KV-Cache Q4_0 für K und V
|
||||||
|
- explizites Prompt-Caching mit 8.192 MiB profilinternem RAM-Cache
|
||||||
|
- `--cache-reuse 256` für die Wiederverwendung langer stabiler Präfixe trotz
|
||||||
|
kleiner späterer Abweichungen
|
||||||
- MTP Draft, maximal drei Tokens
|
- MTP Draft, maximal drei Tokens
|
||||||
- MTP-Akzeptanzschwelle 0,05; im Referenzlauf 77,26 statt 73,88 Tok/s
|
- MTP-Akzeptanzschwelle 0,05; im Referenzlauf 77,26 statt 73,88 Tok/s
|
||||||
- sechs Threads und sechs Batch-Threads
|
- sechs Threads und sechs Batch-Threads
|
||||||
@@ -96,6 +99,18 @@ Der Router übernimmt:
|
|||||||
|
|
||||||
## Hermes Agent
|
## Hermes Agent
|
||||||
|
|
||||||
|
- Hermes 0.20.5 baut den Systemprompt bereits in drei geordneten Bereichen:
|
||||||
|
einen chatübergreifend stabilen Präfix, sitzungsstabilen Kontext und einen
|
||||||
|
variablen Nachlauf. Der stabile Präfix bleibt bei gleicher Profil- und
|
||||||
|
Werkzeugkonfiguration wortgleich und kann dadurch vom llama.cpp-RAM-Cache
|
||||||
|
wiederverwendet werden.
|
||||||
|
- Der RAM-Promptcache lebt nur so lange wie der jeweilige llama.cpp-Prozess.
|
||||||
|
Ein Profilwechsel entlädt das bisherige Modell und damit dessen Cache.
|
||||||
|
- Persistente Slot-Dateien (`--slot-save-path`) sind vorerst bewusst nicht
|
||||||
|
aktiviert. Die aktuelle llama.cpp-Linie hat offene Restore-Fehler; ein
|
||||||
|
gemeldetes erfolgreiches Restore kann trotzdem einen vollständigen Prefill
|
||||||
|
auslösen. Erst nach einem isolierten Regressionstest aktivieren.
|
||||||
|
|
||||||
- Kontextkompression läuft spätestens bei 60.000 Token; die relative
|
- Kontextkompression läuft spätestens bei 60.000 Token; die relative
|
||||||
65-Prozent-Grenze greift nur, wenn sie noch früher erreicht wird. Damit gilt
|
65-Prozent-Grenze greift nur, wenn sie noch früher erreicht wird. Damit gilt
|
||||||
dieselbe Obergrenze auch für Medium, Large und Ultra und ein Profilwechsel
|
dieselbe Obergrenze auch für Medium, Large und Ultra und ein Profilwechsel
|
||||||
|
|||||||
@@ -51,6 +51,14 @@ Der Cache `/var/lib/docker/unraid-update-status.json` ist nur ein
|
|||||||
Kandidatenhinweis. Er darf nie allein eine Neuerstellung auslösen. Autoritativ
|
Kandidatenhinweis. Er darf nie allein eine Neuerstellung auslösen. Autoritativ
|
||||||
ist der Image-ID-Vergleich nach dem Pull.
|
ist der Image-ID-Vergleich nach dem Pull.
|
||||||
|
|
||||||
|
Die Unraid-Weboberfläche und mobile Ansichten lesen weiterhin diesen separaten
|
||||||
|
Cache. Ein technisch verifizierter Pull/Rebuild aktualisiert dessen Anzeige
|
||||||
|
nicht zwingend sofort. Deshalb kann dort weiterhin „Apply Update“ stehen,
|
||||||
|
obwohl der lokale Image-ID-Vergleich bereits `already-current` ergeben hat.
|
||||||
|
Für eine frische Anzeige muss Unraids eigener Statuslauf
|
||||||
|
`dynamix.docker.manager/scripts/dockerupdate check` abgeschlossen sein. Das ist
|
||||||
|
eine Aktualisierung der Anzeige und kein erneuter Container-Rebuild.
|
||||||
|
|
||||||
Ein wiederholter Lauf muss bei einem aktuellen Image folgendes melden:
|
Ein wiederholter Lauf muss bei einem aktuellen Image folgendes melden:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
@@ -86,4 +94,3 @@ einzelne Update-Aufrufe. Dieser Pfad ist ersetzt.
|
|||||||
4. Ein ausdrücklich schreibender synthetischer Auftrag muss beide MUA-Zugänge
|
4. Ein ausdrücklich schreibender synthetischer Auftrag muss beide MUA-Zugänge
|
||||||
bereitstellen und das Batch-Werkzeug wählen.
|
bereitstellen und das Batch-Werkzeug wählen.
|
||||||
5. Ein Wiederholungstest mit aktuellem Image darf keine Neuerstellung auslösen.
|
5. Ein Wiederholungstest mit aktuellem Image darf keine Neuerstellung auslösen.
|
||||||
|
|
||||||
|
|||||||
@@ -3,4 +3,4 @@ Description=Local AI llama.cpp - Qwen Fast 76.8K MTP2 with CPU Vision
|
|||||||
|
|
||||||
[Service]
|
[Service]
|
||||||
ExecStart=
|
ExecStart=
|
||||||
ExecStart=/opt/mike-ai/llama.cpp-nvfp4/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-mix/Qwen3.8-27B-IQ4-MIX.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen38-27b-iq4mix-76k-mtp2-vision --ctx-size 76800 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0 --split-mode none --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
|
ExecStart=/opt/mike-ai/llama.cpp-nvfp4/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-mix/Qwen3.8-27B-IQ4-MIX.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen38-27b-iq4mix-76k-mtp2-vision --ctx-size 76800 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0 --split-mode none --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
|
||||||
|
|||||||
@@ -3,4 +3,4 @@ Description=Legacy native Qwen Large 192K profile (Docker is the production path
|
|||||||
|
|
||||||
[Service]
|
[Service]
|
||||||
ExecStart=
|
ExecStart=
|
||||||
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-large --ctx-size 192000 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 86,14 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-type-k f16 --spec-draft-type-v f16
|
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-large --ctx-size 192000 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 86,14 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-type-k f16 --spec-draft-type-v f16
|
||||||
|
|||||||
@@ -3,4 +3,4 @@ Description=Local AI llama.cpp - Qwen Medium 160K IQ4_XS Pure with CPU Vision
|
|||||||
|
|
||||||
[Service]
|
[Service]
|
||||||
ExecStart=
|
ExecStart=
|
||||||
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-medium --ctx-size 160000 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --reasoning-budget 8192 --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 1.0 --top-p 0.95 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 90,10 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-type-k f16 --spec-draft-type-v f16
|
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-medium --ctx-size 160000 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --reasoning-budget 8192 --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 1.0 --top-p 0.95 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 90,10 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-type-k f16 --spec-draft-type-v f16
|
||||||
|
|||||||
@@ -3,4 +3,4 @@ Description=Legacy native Qwen Ultra 256K profile (Docker is the production path
|
|||||||
|
|
||||||
[Service]
|
[Service]
|
||||||
ExecStart=
|
ExecStart=
|
||||||
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --alias qwen-ultra --ctx-size 262144 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
|
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --alias qwen-ultra --ctx-size 262144 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --no-mmap --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
|
||||||
|
|||||||
Reference in New Issue
Block a user