Add tested 256K Ultra profile

This commit is contained in:
Mikei386 committed 2026-08-22 10:44:14 +02:00
1 parent a194a2941d
commit f52cf14069
12 files changed
+111 -13

No files matched your search

+1
View File
@@ -22,6 +22,7 @@ Heimnetz / VPN-Clients
+-- llama-medium > exakt einer aktiv
+-- llama-long --/
+-- llama-experimental
+-- llama-ultra (256K, text-only, dual GPU)
+-- Piper-TTS (CPU, nur intern)
+-- internes MCP-Netz
+-- Web-MCP + TinySearch + SearXNG
+5 -2
View File
@@ -74,9 +74,12 @@ Community Store nachgeladen.
| Fast | `qwen-fast` | 76.800 | Alltag, Agenten, hohe Geschwindigkeit, integrierte Vision |
| Medium | `qwen-medium` | 94.208 | mehr Kontext, reine IQ4_XS-Variante |
| Long | `qwen-long` | 131.072 | lange Hermes-/MCP-Sitzungen |
| Ultra | `qwen-ultra` | 262.144 | maximaler Textkontext, IQ4-MIX auf RTX 5080 + RTX 3060 (80:20), ohne Vision-Projektor |
Manuell wird mit `llama-profile fast|medium|long` gewechselt. Über HTTP stehen
`POST /fast`, `/medium` und `/long` zur Verfügung. Für eine spätere Version ist
Manuell wird mit `llama-profile fast|medium|long|ultra` gewechselt. Über HTTP stehen
`POST /fast`, `/medium`, `/long` und `/ultra` zur Verfügung. Ultra erreichte im
Referenzlauf etwa 59 Token/s; ein Prompt-Fülltest mit rund 220.000 Tokens war
erfolgreich. Für eine spätere Version ist
`large` als Alias für `long` vorgesehen; bestehende Namen bleiben kompatibel.
## Clients