Enable CPU vision projector for Ultra profile
This commit is contained in:
@@ -32,7 +32,7 @@ gateway, UI, CPU-STT, backup and operator containers may remain active.
|
||||
- Fast: Qwen3.8-27B IQ4-MIX, 76,800 tokens.
|
||||
- Medium: Qwen3.8-27B IQ4_XS-pure, 160,000 tokens, vision.
|
||||
- Large: the same Q4 model, 192,000 tokens, vision.
|
||||
- Ultra: the same Q4 model, 262,144 tokens, no vision projector.
|
||||
- Ultra: the same Q4 model, 262,144 tokens, vision projector on CPU.
|
||||
- Uncensored: Abliterated Q4_K_M, 80,000 tokens, vision.
|
||||
|
||||
Medium, Large, Ultra and Beta distribute their runtime across both GPUs. Do not
|
||||
|
||||
@@ -16,10 +16,10 @@ nach Standardbenchmark, Tool-Calling-Test und Kontexttest übernommen.
|
||||
|
||||
- Fast: Qwen3.8-27B IQ4-MIX mit MTP2
|
||||
- Medium und Large: Qwen3.8-27B IQ4_XS Pure mit MTP3
|
||||
- Ultra: Qwen3.8-27B IQ4_XS Pure mit MTP2 und maximalem Textkontext
|
||||
- Ultra: Qwen3.8-27B IQ4_XS Pure mit MTP2, maximalem Kontext und Vision-Projektor auf CPU
|
||||
- Uncensored: Blackfrost Qwen3.8-27B Abliterated Q4_K_M mit MTP2
|
||||
- Fast, Medium, Large und Uncensored: integrierte Vision; der jeweilige
|
||||
Projektor liegt vollständig auf der RTX 3060
|
||||
- Alle fünf Profile: integrierte Vision. Bei Fast, Medium, Large und
|
||||
Uncensored liegt der Projektor auf der RTX 3060; bei Ultra auf der CPU.
|
||||
|
||||
Die produktive Runtime ist seit dem 12. September 2026 auf llama.cpp
|
||||
**Build 10930**, Commit `56381e407c0ccfb3a6f71e668a27a901001d22ce`,
|
||||
|
||||
@@ -3,4 +3,4 @@ Description=Legacy native Qwen Ultra 256K profile (Docker is the production path
|
||||
|
||||
[Service]
|
||||
ExecStart=
|
||||
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --alias qwen-ultra --ctx-size 262144 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --load-mode none --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
|
||||
ExecStart=/opt/mike-ai/llama.cpp/build/bin/llama-server --model /opt/mike-ai/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf --mmproj /opt/mike-ai/models/qwen3.8-27b-nvfp4/mmproj-BF16.gguf --no-mmproj-offload --alias qwen-ultra --ctx-size 262144 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --cache-prompt --cache-reuse 256 --cache-ram 8192 --threads 6 --threads-batch 6 --batch-size 64 --ubatch-size 32 --parallel 1 --jinja --reasoning auto --host 127.0.0.1 --port 8080 --metrics --fit off --n-gpu-layers all --load-mode none --temperature 0.2 --top-p 0.8 --top-k 20 --device CUDA0,CUDA1 --main-gpu 0 --split-mode layer --tensor-split 80,20 --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k f16 --spec-draft-type-v f16
|
||||
|
||||
Reference in New Issue
Block a user