Tune prefill batching across Qwen profiles

This commit is contained in:
Mikei386
2026-09-15 16:57:53 +02:00
parent 78055996f5
commit b48fcc63d3
5 changed files with 37 additions and 12 deletions
+22
View File
@@ -0,0 +1,22 @@
# Prefill-Batch-Tuning vom 15. September 2026
Alle fünf Qwen-Profile wurden auf Athena mit demselben kalten Textprompt von
rund 54.000 Tokens getestet. Prompt-Cache-Treffer wurden ausgeschlossen. Als
Zielwert gilt der schnellste stabile Lauf, nicht die größte startbare Zahl.
| Profil | vorher | getestet | produktiv | Prefill vorher | Prefill produktiv |
|---|---:|---:|---:|---:|---:|
| Fast | 64 / 32 | 128/64, 2048/64, 2048/96, 2048/128 | **2048 / 64** | 806 tok/s | 1.176 tok/s |
| Medium | 2048 / 128 | 2048/512, 2048/1024 | **2048 / 512** | 1.232 tok/s | 1.597 tok/s |
| Large | 2048 / 128 | 2048/256, 2048/512 | **2048 / 256** | 1.231 tok/s | 1.542 tok/s |
| Ultra | 2048 / 128 | 2048/256, 2048/512 | **2048 / 128** | 1.407 tok/s | 1.407 tok/s |
| Uncensored | 2048 / 128 | 2048/256, 2048/512 | **2048 / 256** | 1.166 tok/s | 1.475 tok/s |
Die Werte sind `Batch / Micro-Batch`. Fast OOMte mit Micro-Batch 96 während
des langen Prompts und mit 128 beim Start. Ultra OOMte bereits beim Start mit
256 und 512. Medium 1024 sowie Large und Uncensored 512 liefen, waren aber
langsamer als der jeweils kleinere produktive Wert und benötigten mehr Reserve.
Die Messung prüft Prefill und VRAM unter einem langen Textprompt. Änderungen an
Modell, Kontextgröße, Vision-Projektor, Tensor-Split oder llama.cpp-Build
erfordern einen neuen Vergleichstest.