Add Qwen GSQ-RCO Beta 1 profile

This commit is contained in:
Mikei386
2026-09-04 14:36:50 +02:00
parent 744a207e5a
commit f553108912
11 changed files with 172 additions and 13 deletions
+32
View File
@@ -0,0 +1,32 @@
# Qwen Beta 1 – GSQ-RCO
`qwen-beta-1` ist ein zusätzliches, nicht standardmäßig aktives Router-Profil.
Die bestehenden Profile und das Standardprofil `qwen-medium` bleiben unverändert.
## Laufzeitkonfiguration
- Modell: Qwen3.8-27B GSQ-RCO IQ3_XXS MTP
- Kontext: 192.000 Token
- Textmodell und KV-Cache: vollständig RTX 5080
- Vision-Projektor: RTX 3060
- KV-Quantisierung: Q4_0 für K und V
- MTP: 3 Draft-Token
- Batch / Micro-Batch: 2048 / 128
## Gemessene Kontextgrenze
Eine echte Bildanfrage mit einer 2,3-MB-JPEG-Datei wurde zur Bestimmung der
VRAM-Grenze verwendet.
| Kontext | Ergebnis | Rest auf RTX 5080 nach Bildlauf |
|---:|---|---:|
| 192.000 | bestanden | ca. 129 MiB |
| 196.608 | bestanden, harte Kante | ca. 9 MiB |
| 197.120 | CUDA Out of Memory | ca. 1 MiB vor Abbruch |
Der produktive Beta-Modus verwendet deshalb 192.000 Token. 196.608 ist nur
die gemessene technische Obergrenze und besitzt keine ausreichende Reserve
für einen verlässlichen Dauerbetrieb.
Beim erfolgreichen 196.608-Test erreichte die Bildanfrage rund 366 Prompt-
Token/s und 85 Ausgabe-Token/s. Das erkannte Bild wurde korrekt beschrieben.
+2
View File
@@ -8,6 +8,7 @@ Standardprofil: **medium** · globales Ausgabelimit: **8192 Token**
|---|---|---:|---:|---|---|---|---:|
| fast | `qwen-fast` | 76,800 | 1 | Qwen3.8-27B IQ4 Mix | 5080 only | ja | 2 |
| medium | `qwen-medium` | 160,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 85:15 | ja | 3 |
| beta1 | `qwen-beta-1` | 192,000 | 1 | Qwen3.8-27B GSQ-RCO IQ3_XXS MTP | 5080 model / 3060 vision | ja | 3 |
| large | `qwen-large` | 192,000 | 1 | Qwen3.8-27B IQ4 XS Pure | 86:14 | ja | 3 |
| ultra | `qwen-ultra` | 262,144 | 1 | Qwen3.8-27B IQ4 XS Pure | 80:20 | nein | 2 |
| uncensored | `qwen-uncensored` | 80,000 | 1 | Qwen3.8-27B Abliterated Q4_K_M | 90:10 | ja | 2 |
@@ -16,6 +17,7 @@ Standardprofil: **medium** · globales Ausgabelimit: **8192 Token**
- **fast**: Schnelles Profil für kurze Chats und zügige Werkzeugaufgaben.
- **medium**: Ausgewogenes Standardprofil für Alltag und lange agentische Aufgaben.
- **beta1**: Beta 1: schnelles GSQ-RCO-Testprofil mit 192K Kontext und Vision-Projektor auf der RTX 3060.
- **large**: Großes Profil für umfangreiche Dokumente und lange technische Arbeiten.
- **ultra**: Maximaler Textkontext; bewusst ohne Vision-Projektor.
- **uncensored**: Weniger restriktives Spezialprofil; Werkzeugrechte bleiben unverändert.