From 535bd751b522a8cf680b2d17ed76f08a866701e7 Mon Sep 17 00:00:00 2001 From: Mikei386 <44135113+Mikei386@users.noreply.github.com> Date: Wed, 9 Sep 2026 12:27:36 +0200 Subject: [PATCH] Update llama.cpp runtime to build 10872 --- platform/llama/LLAMA_CPP_COMMIT | 2 +- platform/llama/README.md | 11 ++++++----- 2 files changed, 7 insertions(+), 6 deletions(-) diff --git a/platform/llama/LLAMA_CPP_COMMIT b/platform/llama/LLAMA_CPP_COMMIT index b802600..29b5af9 100644 --- a/platform/llama/LLAMA_CPP_COMMIT +++ b/platform/llama/LLAMA_CPP_COMMIT @@ -1 +1 @@ -c7bda030e7faee594dbe7550185e857351ad405d +b31b71f3a076bfc4278daad442203a9c51c6e676 diff --git a/platform/llama/README.md b/platform/llama/README.md index 414d84a..9f361e6 100644 --- a/platform/llama/README.md +++ b/platform/llama/README.md @@ -21,11 +21,12 @@ nach Standardbenchmark, Tool-Calling-Test und Kontexttest übernommen. - Fast, Medium, Large und Uncensored: integrierte Vision; der jeweilige Projektor liegt vollständig auf der RTX 3060 -Die produktive Runtime ist auf llama.cpp Build 10781, Commit -`c7bda030e7faee594dbe7550185e857351ad405d`, festgeschrieben. Dieser Stand -enthält die ab Build 10751 verfügbare Korrektur für eine zwischenzeitliche -MTP-/KV-Cache-Initialisierungsregression. Der vorherige produktive Stand war -Build 10718, Commit `41ef91f7c8046087cdfbb276b79bff311ecf1c6d`. +Die produktive Runtime ist auf llama.cpp Build 10872, Commit +`b31b71f3a076bfc4278daad442203a9c51c6e676`, festgeschrieben. Dieser Stand +enthält unter anderem die Korrekturen für die Qwen3.8-GDN-Normalisierung, +CUDA-Races, eine divergente Barriere im F16-Flash-Attention-Kernel sowie die +Checkpoint-Verdrängung im Prompt-Cache. Der vorherige produktive Stand war +Commit `c7bda030e7faee594dbe7550185e857351ad405d`. Die Dateien selbst sind nicht Bestandteil des Repositories. Pfade und Hashes werden im lokalen Modellmanifest verwaltet.