diff --git a/docs/INFERENCE_OPTIMIZATION_TODO_20260920.md b/docs/INFERENCE_OPTIMIZATION_TODO_20260920.md index d43db57..ad10914 100644 --- a/docs/INFERENCE_OPTIMIZATION_TODO_20260920.md +++ b/docs/INFERENCE_OPTIMIZATION_TODO_20260920.md @@ -47,3 +47,12 @@ Ein-Slot-Folgeversuche: [Bericht](DFLASH2_ONE_SLOT_20260920.md). Draft auf 5080 [512 gegen 256](MEDIUM_MICROBATCH_20260920.md): 256 lädt ohne vorherige Pufferfehler-Meldung, lässt aber nur 461 MiB auf der 5080 frei. Schutzabbruch vor Inferenz; kein Geschwindigkeitsgewinn belegt. 512 wiederhergestellt. Für einen eventuellen Folgetest zuerst mehr 5080-Reserve durch angepassten Split schaffen. Nach expliziter Freigabe: [Reserve448-Folgetest](MEDIUM_MICROBATCH_RESERVE448_20260920.md) bestanden. 256 liefert +7,1 % Prefill im einmaligen 24K-Kurzvergleich; kurze Decodes nahezu gleich. Alle drei Fakten korrekt, keine volle Langkontext-/Vision-/Parallelitätsfreigabe. Produktiv weiter512. + +## Punkt 2: GPU-Split/MTP-Kurzvergleich + +[Vier Varianten getestet](MEDIUM_SPLIT_MTP_20260920.md): bei Microbatch256 ist +85:15/MTP2 der interessanteste Kandidat (+10,3% Deutsch, Code/24K etwa gleich, +268MiB weniger Peak5080). MTP4 und83:17 ohne allgemeinen Vorteil. Je ein Lauf, +unterschiedliche Texte, Recall überall3/3; breite Qualität und lange/visuelle/ +parallele Last offen. Produktiv weiterhin85:15/MTP3/Microbatch512. Punkt2 teilweise +bearbeitet, keine pauschale Freigabe oder höhere Kontextgrenze. diff --git a/docs/MEDIUM_SPLIT_MTP_20260920.md b/docs/MEDIUM_SPLIT_MTP_20260920.md new file mode 100644 index 0000000..690d102 --- /dev/null +++ b/docs/MEDIUM_SPLIT_MTP_20260920.md @@ -0,0 +1,55 @@ +# Medium: GPU-Split und MTP – Kurzvergleich 20.09.2026 + +Vier isolierte Läufe mit Pure IQ4_XS, 160.000 konfigurierten Kontexttokens, +zwei Slots, Vision geladen, TTS auf der 3060, Microbatch 256. Produktionsprofil +nachher unverändert wiederhergestellt (85:15, MTP3, Microbatch512). + +| Split 5080:3060 / MTP | Deutsch tok/s | Code tok/s | 24K Prefill tok/s | Decode nach 24K tok/s | Peak 5080 MiB | Peak 3060 MiB | +|---|---:|---:|---:|---:|---:|---:| +| 85:15 / 3 Kontrolle | 57,39 | 75,37 | 1907,30 | 59,96 | 15860 | 10646 | +| 85:15 / 2 | 63,32 | 75,35 | 1906,40 | 59,85 | 15592 | 10614 | +| 85:15 / 4 | 51,36 | 78,87 | 1813,80 | 50,33 | 15872 | 10420 | +| 83:17 / 3 | 57,54 | 73,99 | 1930,21 | 56,12 | 15274 | 11230 | + +## Einordnung + +MTP2 ist der interessanteste Folgekandidat: +10,3 % beim deutschen Text, +praktisch gleiche Code-/24K-Raten und 268 MiB weniger gesampelte Spitzenbelegung +auf der 5080. MTP4 ist beim Code +4,6 %, aber bei Deutsch −10,5 % und beim Decode +nach 24K −16,1 %. 83:17 verschafft mehr Reserve auf der 5080, liefert aber keinen +klaren Gesamtgewinn und belegt die langsamere 3060 stärker. + +Dies sind je ein Lauf und unterschiedliche erzeugte Tokenfolgen. Identische +Prompts, Samplingwerte und Seeds erzwingen über veränderte MTP-/GPU-Konfigurationen +keine identischen Antworten. Die Werte sind keine isolierten Messungen derselben +Tokenfolge und kein Nachweis eines allgemeinen 10-%-Gewinns. + +Alle vier Läufe finden die drei eingebetteten Fakten korrekt (3/3). Modellgewichte +und Quantisierung unverändert. Keine allgemeine Qualitätsgleichheit nachgewiesen: +Die beiden Decode-Aufgaben wurden absichtlich bei 768 Ausgabetokens begrenzt +(`finish_reason=length`); kein vollständiger Code-Test oder breiter Qualitätstest. +Die Recall-Ausgaben enthalten korrekte Werte, aber auch überflüssige Erläuterungen. +Keine Vision-, Parallelitäts- oder volle160K-Prüfung; größte tatsächliche Eingabe +24.674 Tokens, gefolgt von512 Ausgabetokens. Kein höheres Kontextmaximum ermittelt. + +## Sicherheit und Reproduzierbarkeit + +Runner kopiert die echten Produktionsargumente. Variiert werden nur Microbatch, +Split, MTP-Tiefe sowie isolierter Port/Host und Log-Verbosity. Der neue Kontrolllauf +ist für diesen direkten Vergleich erforderlich, keine Wiederholung der breiten +archivierten Qwen-Modellreferenz. Kein Prompt-Cache in den Messanfragen. + +Nach laufenden Anfragen erfolgte ein begrenzter Testbetrieb mit automatischer +Wiederherstellung. 448MiB Mindestreserve nach Laden, 85°C Temperatur- und3GiB +Host-RAM-Grenze. Container26GiB RAM ohne zusätzliches Swapbudget; keine Host-, +Treiber-, Netzwerk- oder Kerneländerungen. Maximal75°C5080/58°C3060 gemessen. +PCIe unter Last: 5080Gen4x16, 3060Gen3x4. Keine Kernel-/Xid-/OOM-Kill-Meldungen. +Router/Medium/Controller/Gateway/TTS danach gesund, readiness und OK-Antwort bestanden. + +Rohdaten einschließlich Antworten, tatsächlichen Argumenten, GPU-Samples und +komprimierten Serverlogs: `experiments/medium-split-mtp-20260920/`. +Athena: `/data/benchmarks/medium-split-mtp-20260920/`. + +Empfehlung: vor dauerhafter Umstellung MTP2/Microbatch256 mit vollständigen +Qualitätsantworten und relevanter Langkontext-/Visionlast prüfen. Kein automatischer +Folgetest und keine Produktionsumstellung durch diesen Kurzvergleich. diff --git a/experiments/medium-split-mtp-20260920/83-mtp3/config.json b/experiments/medium-split-mtp-20260920/83-mtp3/config.json new file mode 100644 index 0000000..2e0967b --- /dev/null +++ b/experiments/medium-split-mtp-20260920/83-mtp3/config.json @@ -0,0 +1,10 @@ +{ + "label": "83-mtp3", + "ubatch": 256, + "split": "83,17", + "mtp": 3, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 +} diff --git a/experiments/medium-split-mtp-20260920/83-mtp3/gpu.json b/experiments/medium-split-mtp-20260920/83-mtp3/gpu.json new file mode 100644 index 0000000..a7bfb88 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/83-mtp3/gpu.json @@ -0,0 +1,623 @@ +[ + { + "time": 1789936766.905543, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "7346", + "total": "12288", + "temp": "50", + "util": "100", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "10888", + "total": "16303", + "temp": "56", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936768.9340215, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "9024", + "total": "12288", + "temp": "50", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15228", + "total": "16303", + "temp": "55", + "util": "1", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936770.9613502, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11222", + "total": "12288", + "temp": "49", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15256", + "total": "16303", + "temp": "54", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936772.992449, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11222", + "total": "12288", + "temp": "48", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15256", + "total": "16303", + "temp": "53", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936775.0237627, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "54", + "util": "48", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "59", + "util": "44", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936777.056265, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "54", + "util": "60", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "62", + "util": "34", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936779.0883763, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "55", + "util": "59", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "59", + "util": "46", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936781.11879, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "55", + "util": "49", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "59", + "util": "41", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936783.1498864, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "56", + "util": "51", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "61", + "util": "48", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936785.180021, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "55", + "util": "49", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "62", + "util": "34", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936787.2098007, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "56", + "util": "54", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "63", + "util": "41", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936789.240345, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "56", + "util": "52", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "63", + "util": "39", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936791.2699392, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "56", + "util": "59", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "64", + "util": "36", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936793.3014712, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "57", + "util": "56", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "60", + "util": "46", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936795.3310409, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "57", + "util": "61", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "62", + "util": "50", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936797.3619807, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11228", + "total": "12288", + "temp": "56", + "util": "55", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15272", + "total": "16303", + "temp": "60", + "util": "45", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936799.3926547, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "57", + "util": "56", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "59", + "util": "15", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936801.422532, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "57", + "util": "32", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "72", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936803.4543598, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "57", + "util": "45", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "73", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936805.4877064, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "55", + "util": "70", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "73", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936807.533202, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "57", + "util": "68", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "72", + "util": "12", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936809.5705192, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "57", + "util": "70", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "75", + "util": "46", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936811.6039462, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "56", + "util": "55", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "64", + "util": "40", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936813.633929, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "57", + "util": "51", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "62", + "util": "39", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936815.6666515, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "57", + "util": "54", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "62", + "util": "40", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936817.7020009, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "56", + "util": "46", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "62", + "util": "40", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936819.7318478, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11230", + "total": "12288", + "temp": "58", + "util": "51", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15274", + "total": "16303", + "temp": "61", + "util": "40", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + } +] diff --git a/experiments/medium-split-mtp-20260920/83-mtp3/loaded.json b/experiments/medium-split-mtp-20260920/83-mtp3/loaded.json new file mode 100644 index 0000000..2d62df0 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/83-mtp3/loaded.json @@ -0,0 +1,138 @@ +{ + "case": { + "label": "83-mtp3", + "ubatch": 256, + "split": "83,17", + "mtp": 3, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + "started": 1789936764.8736556, + "idle_gpu": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11222", + "total": "12288", + "temp": "48", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15256", + "total": "16303", + "temp": "53", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ], + "props": { + "default_generation_settings": { + "params": { + "seed": 4294967295, + "temperature": 1.0, + "dynatemp_range": 0.0, + "dynatemp_exponent": 1.0, + "top_k": 20, + "top_p": 0.949999988079071, + "min_p": 0.05000000074505806, + "top_n_sigma": -1.0, + "xtc_probability": 0.0, + "xtc_threshold": 0.10000000149011612, + "typical_p": 1.0, + "repeat_last_n": 64, + "repeat_penalty": 1.0, + "presence_penalty": 0.0, + "frequency_penalty": 0.0, + "dry_multiplier": 0.0, + "dry_base": 1.75, + "dry_allowed_length": 2, + "dry_penalty_last_n": 64, + "mirostat": 0, + "mirostat_tau": 5.0, + "mirostat_eta": 0.10000000149011612, + "adaptive_target": -1.0, + "adaptive_decay": 0.8999999761581421, + "max_tokens": -1, + "n_predict": -1, + "n_keep": 0, + "n_discard": 0, + "ignore_eos": false, + "stream": false, + "n_probs": 0, + "min_keep": 0, + "chat_format": "Content-only", + "reasoning_format": "none", + "reasoning_in_content": false, + "generation_prompt": "", + "samplers": [ + "penalties", + "dry", + "top_n_sigma", + "top_k", + "typ_p", + "top_p", + "min_p", + "xtc", + "temperature" + ], + "speculative.types": "none", + "timings_per_token": false, + "post_sampling_probs": false, + "backend_sampling": false, + "lora": [] + }, + "n_ctx": 160000 + }, + "total_slots": 2, + "model_alias": "qwen-medium", + "model_ftype": "IQ4_XS - 4.25 bpw", + "model_path": "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "modalities": { + "vision": true, + "video": true, + "audio": false + }, + "media_marker": "<__media_6beqV5u0Vr11b3oODbUSmmU1ZGgPZj1J__>", + "endpoint_slots": true, + "endpoint_props": false, + "endpoint_metrics": true, + "ui": false, + "ui_settings": {}, + "chat_template": "{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- macro render_content(content, do_vision_count, is_system_content=false) %}\n {%- if content is string %}\n {{- content }}\n {%- elif content is iterable and content is not mapping %}\n {%- for item in content %}\n {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain images.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set image_count.value = image_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Picture ' ~ image_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|image_pad|><|vision_end|>' }}\n {%- elif 'video' in item or item.type == 'video' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain videos.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set video_count.value = video_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Video ' ~ video_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|video_pad|><|vision_end|>' }}\n {%- elif 'text' in item %}\n {{- item.text }}\n {%- else %}\n {{- raise_exception('Unexpected item type in content.') }}\n {%- endif %}\n {%- endfor %}\n {%- elif content is none or content is undefined %}\n {{- '' }}\n {%- else %}\n {{- raise_exception('Unexpected content type.') }}\n {%- endif %}\n{%- endmacro %}\n{%- if not messages %}\n {{- raise_exception('No messages provided.') }}\n{%- endif %}\n{%- set sysns = namespace(count=0, text='') %}\n{%- for message in messages %}\n {%- if sysns.count == loop.index0 and (message.role == 'system' or message.role == 'developer') %}\n {%- set sys_content = render_content(message.content, false, true)|trim %}\n {%- if sys_content %}\n {%- set sysns.text = sysns.text + ('\\n' if sysns.text else '') + sys_content %}\n {%- endif %}\n {%- set sysns.count = sysns.count + 1 %}\n {%- endif %}\n{%- endfor %}\n{%- set num_sys = sysns.count %}\n{%- set merged_system = sysns.text %}\n{%- set reasoning_instructions = '' %}\n{%- if enable_thinking is undefined or enable_thinking is true %}\n {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}\n {%- if resolved_reasoning_effort == 'high' %}\n {%- set resolved_reasoning_effort = 'xhigh' %}\n {%- endif %}\n {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}\n {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}\n {%- endif %}\n {%- if resolved_reasoning_effort == 'xhigh' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}\n {%- elif resolved_reasoning_effort == 'low' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}\n {%- endif %}\n{%- endif %}\n{%- if tools and tools is iterable and tools is not mapping %}\n {{- '<|im_start|>system\\n' }}\n {%- if reasoning_instructions %}\n {{- reasoning_instructions + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou have access to the following functions:\\n\\n\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n\" }}\n {{- '\\n\\nIf you choose to call a function ONLY reply in the following format with NO suffix:\\n\\n\\n\\n\\nvalue_1\\n\\n\\nThis is the value for the second parameter\\nthat can span\\nmultiple lines\\n\\n\\n\\n\\n\\nReminder:\\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\\n- Required parameters MUST be specified\\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\\n' }}\n {%- if merged_system %}\n {{- '\\n\\n' + merged_system }}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n{%- else %}\n {%- if merged_system %}\n {{- '<|im_start|>system\\n' + (reasoning_instructions + '\\n\\n' if reasoning_instructions else '') + merged_system + '<|im_end|>\\n' }}\n {%- elif reasoning_instructions %}\n {{- '<|im_start|>system\\n' + reasoning_instructions + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" %}\n {%- set content = render_content(message.content, false)|trim %}\n {%- if not(content.startswith('') and content.endswith('')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if loop.index0 >= num_sys %}\n {%- set content = render_content(message.content, true)|trim %}\n {%- if message.role == \"system\" or message.role == \"developer\" %}\n {{- raise_exception('System message must be at the beginning.') }}\n {%- elif message.role == \"user\" %}\n {{- '<|im_start|>' + message.role + '\\n' + content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is string %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- endif %}\n {%- set reasoning_content = reasoning_content|trim %}\n {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}\n {{- '<|im_start|>' + message.role + '\\n\\n' + reasoning_content + '\\n\\n\\n' + content }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {%- if tool_call.name is not defined or tool_call.name is none %}\n {{- raise_exception('Tool call is missing a function name.') }}\n {%- endif %}\n {%- if loop.first %}\n {%- if content|trim %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n\\n' }}\n {%- endif %}\n {%- else %}\n {{- '\\n\\n\\n' }}\n {%- endif %}\n {%- if tool_call.arguments is mapping %}\n {%- for args_name, args_value in tool_call.arguments|items %}\n {{- '\\n' }}\n {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}\n {{- args_value }}\n {{- '\\n\\n' }}\n {%- endfor %}\n {%- elif tool_call.arguments is string %}\n {%- if tool_call.arguments|trim %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" were passed as a JSON string. Parse them into an object before calling apply_chat_template.') }}\n {%- endif %}\n {%- elif tool_call.arguments is defined and tool_call.arguments is not none %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" must be an object/mapping or a JSON string.') }}\n {%- endif %}\n {{- '\\n' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.previtem and loop.previtem.role != \"tool\" %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n\\n' }}\n {{- content }}\n {{- '\\n' }}\n {%- if not loop.last and loop.nextitem.role != \"tool\" %}\n {{- '<|im_end|>\\n' }}\n {%- elif loop.last %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- else %}\n {{- raise_exception('Unexpected message role.') }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n' }}\n {%- endif %}\n{%- endif %}\n{#- Unsloth fixes - developer role, merged system messages, tool calling #}", + "chat_template_caps": { + "supports_object_arguments": true, + "supports_parallel_tool_calls": true, + "supports_preserve_reasoning": true, + "supports_reasoning_effort": true, + "supports_string_content": true, + "supports_system_role": true, + "supports_tool_calls": true, + "supports_tools": true, + "supports_typed_content": true + }, + "bos_token": "<|endoftext|>", + "eos_token": "<|im_end|>", + "build_info": "b10964-b29c606e2", + "is_sleeping": false, + "cors_proxy_enabled": false + }, + "slots": [ + { + "id": 0, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + }, + { + "id": 1, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + } + ] +} diff --git a/experiments/medium-split-mtp-20260920/83-mtp3/result.json b/experiments/medium-split-mtp-20260920/83-mtp3/result.json new file mode 100644 index 0000000..a0c871f --- /dev/null +++ b/experiments/medium-split-mtp-20260920/83-mtp3/result.json @@ -0,0 +1,307 @@ +{ + "case": { + "label": "83-mtp3", + "ubatch": 256, + "split": "83,17", + "mtp": 3, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + "started": 1789936764.8736556, + "idle_gpu": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "11222", + "total": "12288", + "temp": "48", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15256", + "total": "16303", + "temp": "53", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ], + "props": { + "default_generation_settings": { + "params": { + "seed": 4294967295, + "temperature": 1.0, + "dynatemp_range": 0.0, + "dynatemp_exponent": 1.0, + "top_k": 20, + "top_p": 0.949999988079071, + "min_p": 0.05000000074505806, + "top_n_sigma": -1.0, + "xtc_probability": 0.0, + "xtc_threshold": 0.10000000149011612, + "typical_p": 1.0, + "repeat_last_n": 64, + "repeat_penalty": 1.0, + "presence_penalty": 0.0, + "frequency_penalty": 0.0, + "dry_multiplier": 0.0, + "dry_base": 1.75, + "dry_allowed_length": 2, + "dry_penalty_last_n": 64, + "mirostat": 0, + "mirostat_tau": 5.0, + "mirostat_eta": 0.10000000149011612, + "adaptive_target": -1.0, + "adaptive_decay": 0.8999999761581421, + "max_tokens": -1, + "n_predict": -1, + "n_keep": 0, + "n_discard": 0, + "ignore_eos": false, + "stream": false, + "n_probs": 0, + "min_keep": 0, + "chat_format": "Content-only", + "reasoning_format": "none", + "reasoning_in_content": false, + "generation_prompt": "", + "samplers": [ + "penalties", + "dry", + "top_n_sigma", + "top_k", + "typ_p", + "top_p", + "min_p", + "xtc", + "temperature" + ], + "speculative.types": "none", + "timings_per_token": false, + "post_sampling_probs": false, + "backend_sampling": false, + "lora": [] + }, + "n_ctx": 160000 + }, + "total_slots": 2, + "model_alias": "qwen-medium", + "model_ftype": "IQ4_XS - 4.25 bpw", + "model_path": "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "modalities": { + "vision": true, + "video": true, + "audio": false + }, + "media_marker": "<__media_6beqV5u0Vr11b3oODbUSmmU1ZGgPZj1J__>", + "endpoint_slots": true, + "endpoint_props": false, + "endpoint_metrics": true, + "ui": false, + "ui_settings": {}, + "chat_template": "{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- macro render_content(content, do_vision_count, is_system_content=false) %}\n {%- if content is string %}\n {{- content }}\n {%- elif content is iterable and content is not mapping %}\n {%- for item in content %}\n {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain images.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set image_count.value = image_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Picture ' ~ image_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|image_pad|><|vision_end|>' }}\n {%- elif 'video' in item or item.type == 'video' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain videos.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set video_count.value = video_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Video ' ~ video_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|video_pad|><|vision_end|>' }}\n {%- elif 'text' in item %}\n {{- item.text }}\n {%- else %}\n {{- raise_exception('Unexpected item type in content.') }}\n {%- endif %}\n {%- endfor %}\n {%- elif content is none or content is undefined %}\n {{- '' }}\n {%- else %}\n {{- raise_exception('Unexpected content type.') }}\n {%- endif %}\n{%- endmacro %}\n{%- if not messages %}\n {{- raise_exception('No messages provided.') }}\n{%- endif %}\n{%- set sysns = namespace(count=0, text='') %}\n{%- for message in messages %}\n {%- if sysns.count == loop.index0 and (message.role == 'system' or message.role == 'developer') %}\n {%- set sys_content = render_content(message.content, false, true)|trim %}\n {%- if sys_content %}\n {%- set sysns.text = sysns.text + ('\\n' if sysns.text else '') + sys_content %}\n {%- endif %}\n {%- set sysns.count = sysns.count + 1 %}\n {%- endif %}\n{%- endfor %}\n{%- set num_sys = sysns.count %}\n{%- set merged_system = sysns.text %}\n{%- set reasoning_instructions = '' %}\n{%- if enable_thinking is undefined or enable_thinking is true %}\n {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}\n {%- if resolved_reasoning_effort == 'high' %}\n {%- set resolved_reasoning_effort = 'xhigh' %}\n {%- endif %}\n {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}\n {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}\n {%- endif %}\n {%- if resolved_reasoning_effort == 'xhigh' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}\n {%- elif resolved_reasoning_effort == 'low' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}\n {%- endif %}\n{%- endif %}\n{%- if tools and tools is iterable and tools is not mapping %}\n {{- '<|im_start|>system\\n' }}\n {%- if reasoning_instructions %}\n {{- reasoning_instructions + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou have access to the following functions:\\n\\n\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n\" }}\n {{- '\\n\\nIf you choose to call a function ONLY reply in the following format with NO suffix:\\n\\n\\n\\n\\nvalue_1\\n\\n\\nThis is the value for the second parameter\\nthat can span\\nmultiple lines\\n\\n\\n\\n\\n\\nReminder:\\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\\n- Required parameters MUST be specified\\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\\n' }}\n {%- if merged_system %}\n {{- '\\n\\n' + merged_system }}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n{%- else %}\n {%- if merged_system %}\n {{- '<|im_start|>system\\n' + (reasoning_instructions + '\\n\\n' if reasoning_instructions else '') + merged_system + '<|im_end|>\\n' }}\n {%- elif reasoning_instructions %}\n {{- '<|im_start|>system\\n' + reasoning_instructions + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" %}\n {%- set content = render_content(message.content, false)|trim %}\n {%- if not(content.startswith('') and content.endswith('')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if loop.index0 >= num_sys %}\n {%- set content = render_content(message.content, true)|trim %}\n {%- if message.role == \"system\" or message.role == \"developer\" %}\n {{- raise_exception('System message must be at the beginning.') }}\n {%- elif message.role == \"user\" %}\n {{- '<|im_start|>' + message.role + '\\n' + content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is string %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- endif %}\n {%- set reasoning_content = reasoning_content|trim %}\n {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}\n {{- '<|im_start|>' + message.role + '\\n\\n' + reasoning_content + '\\n\\n\\n' + content }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {%- if tool_call.name is not defined or tool_call.name is none %}\n {{- raise_exception('Tool call is missing a function name.') }}\n {%- endif %}\n {%- if loop.first %}\n {%- if content|trim %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n\\n' }}\n {%- endif %}\n {%- else %}\n {{- '\\n\\n\\n' }}\n {%- endif %}\n {%- if tool_call.arguments is mapping %}\n {%- for args_name, args_value in tool_call.arguments|items %}\n {{- '\\n' }}\n {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}\n {{- args_value }}\n {{- '\\n\\n' }}\n {%- endfor %}\n {%- elif tool_call.arguments is string %}\n {%- if tool_call.arguments|trim %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" were passed as a JSON string. Parse them into an object before calling apply_chat_template.') }}\n {%- endif %}\n {%- elif tool_call.arguments is defined and tool_call.arguments is not none %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" must be an object/mapping or a JSON string.') }}\n {%- endif %}\n {{- '\\n' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.previtem and loop.previtem.role != \"tool\" %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n\\n' }}\n {{- content }}\n {{- '\\n' }}\n {%- if not loop.last and loop.nextitem.role != \"tool\" %}\n {{- '<|im_end|>\\n' }}\n {%- elif loop.last %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- else %}\n {{- raise_exception('Unexpected message role.') }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n' }}\n {%- endif %}\n{%- endif %}\n{#- Unsloth fixes - developer role, merged system messages, tool calling #}", + "chat_template_caps": { + "supports_object_arguments": true, + "supports_parallel_tool_calls": true, + "supports_preserve_reasoning": true, + "supports_reasoning_effort": true, + "supports_string_content": true, + "supports_system_role": true, + "supports_tool_calls": true, + "supports_tools": true, + "supports_typed_content": true + }, + "bos_token": "<|endoftext|>", + "eos_token": "<|im_end|>", + "build_info": "b10964-b29c606e2", + "is_sleeping": false, + "cors_proxy_enabled": false + }, + "slots": [ + { + "id": 0, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + }, + { + "id": 1, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + } + ], + "smoke": { + "choices": [ + { + "finish_reason": "stop", + "index": 0, + "message": { + "role": "assistant", + "content": "OK" + } + } + ], + "created": 1789936774, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 2, + "prompt_tokens": 19, + "total_tokens": 21, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-crkNjSblJZQlUtkCBAwAiGOYZ2MavxPK", + "timings": { + "cache_n": 0, + "prompt_n": 19, + "prompt_ms": 782.694, + "prompt_per_token_ms": 41.194421052631576, + "prompt_per_second": 24.275131788412843, + "predicted_n": 2, + "predicted_ms": 293.483, + "predicted_per_token_ms": 293.483, + "predicted_per_second": 3.407352384976302, + "draft_n": 3, + "draft_n_accepted": 3 + }, + "wall_seconds": 1.0789340379997157 + }, + "decode": [ + { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "Ein Reverse Proxy ist ein zentrales Element in modernen IT-Infrastrukturen, das als Vermittler zwischen Clienten und Backend-Servern agiert. Im Gegensatz zu einem Forward Proxy, der Anfragen von internen Clients nach außen vertritt, stellt der Reverse Proxy eine Eindepunkt-Schnittstelle (Single Point of Entry) für externe Anfragen dar. Typische Vertreter sind nginx, HAProxy oder Envoy. Der primäre Vorteil dieser Architektur liegt in der Entkopplung des Clients vom eigentlichen Server: Der Client muss die konkrete IP-Adresse oder den Hostnamen des Backend-Servers nicht kennen. Der Proxy übernimmt die Lastverteilung (Load Balancing), Terminiert TLS-Verbindungen, bündelt Traffic und kann zusätzliche Sicherheitsmechanismen wie Rate-Limiting oder Firewall-Regeln implementieren.\n\nDie Funktionsweise basiert im Kern auf der Weiterleitung von Anfragen. Wenn ein HTTP-Request auf dem Reverse Proxy eingeht, wird dieser inspiziert, bevor er an den internen Server weitergeleitet wird. Der Proxy ersetzt dabei oft den eigenen Hostnamen durch den internen Zielsystem-Hostnamen und passt Header-Informationen an. Eine entscheidende Rolle spielt dabei die Kommunikation über TCP/IP. Der Proxy öffnet eine Verbindung zum Backend-Server, sendet die Anfrage und wartet auf die Antwort, die er anschließend an den Client zurückgibt. In dynamischen Umgebungen, insbesondere bei der Nutzung von Containern (z. B. via Docker oder Kubernetes), ist diese Architektur besonders anfällig für eine spezifische Klasse von Fehlern: die IP-Adressen-Änderung.\n\nContainer sind durch ihre Effizienz und Skalierbarkeit zwar ideal für moderne Anwendungen, sie haben aber die Eigenschaft, dass ihre IP-Adressen nicht statisch sind. Wenn ein Container neu gestartet, migriert oder wegen Ressourcenengpässen evakuiert wird, wird ihm häufig eine neue IP-Adresse zugewiesen. Wenn ein Reverse Proxy die Ziel-IP-Adresse hartkodiert oder nicht dynamisch auf Veränderungen reagiert, kommt es zu schwerwiegenden Kommunikationsausfällen. Die häufigsten Symptome hierfür sind „Connection Refused“-Fehler, wenn der Proxy versucht, auf eine alte, nicht mehr existierende IP zuzugreifen, oder Timeouts, wenn Pakete an eine Adresse gerichtet werden, auf der kein Prozess mehr lauscht. Ein weiteres subtiles Problem ist das „Stale Connection“-Phänomen: Lange bestehende TCP-Verbindungen können hängen bleiben, obwohl das Backend bereits neu gestartet wurde.\n\nUm solche Fehler präzise einzugrenzen, ist eine systematische Analyse der Logs essenziell, da man zwischen Client-, Proxy- und Server-Seite unterscheiden muss. Der erste Schritt ist die Prüfung der Access Logs des Reverse Proxys. Hier sucht man nach HTTP-Statuscodes, die auf Backend-Fehler hindeuten, insbesondere 502 (Bad Gateway), 503 (Service Unavailable) oder 504 (Gateway Timeout). Ein 502-Fehler bedeutet oft, dass der Proxy die Anfrage an das Backend geschickt hat, aber keine gültige HTTP-Antwort erhalten hat – typischerweise, weil die TCP-Verbindung zum Backend abgelehnt wurde (Connection Reset oder Refused).\n\nIn den Error-Logs des Proxys (z. B. bei nginx in der Datei `error.log`) finden sich die spezifischen technischen Ursachen. Man sieht Meldungen wie „connect() failed (111: Connection refused) to [IP]:[PORT]“. Diese IP-Adresse ist der entscheidende Hinweis. Vergleicht man diese IP mit der aktuell aktiven IP des Containers (z." + } + } + ], + "created": 1789936787, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 768, + "prompt_tokens": 59, + "total_tokens": 827, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-MInUDwt7oMjqbagxGKpU2FNOn00er50H", + "timings": { + "cache_n": 0, + "prompt_n": 59, + "prompt_ms": 124.032, + "prompt_per_token_ms": 2.1022372881355933, + "prompt_per_second": 475.68369453044374, + "predicted_n": 768, + "predicted_ms": 13329.771, + "predicted_per_token_ms": 17.379101694915256, + "predicted_per_second": 57.540373349249585, + "draft_n": 1024, + "draft_n_accepted": 425 + }, + "wall_seconds": 13.5130456770421 + }, + { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "```python\nimport asyncio\nfrom typing import Any, Awaitable, Callable, List, Optional, Union\n\n\nasync def first_success(\n awaitables: List[Awaitable],\n return_exceptions: bool = True\n) -> Any:\n \"\"\"\n Start all awaitables concurrently, return the first successful result,\n cancel and await remaining tasks, and collect exceptions if all fail.\n\n Args:\n awaitables: A list of awaitables (coroutines, tasks, or other awaitables).\n return_exceptions: If True, raise an ExceptionGroup containing all exceptions\n when all awaitables fail. If False, raise the first exception.\n\n Returns:\n The result of the first successfully completed awaitable.\n\n Raises:\n ExceptionGroup: If all awaitables fail and return_exceptions is True.\n Exception: If all awaitables fail and return_exceptions is False (first exception).\n \"\"\"\n if not awaitables:\n raise ValueError(\"At least one awaitable is required\")\n\n # Convert all awaitables to tasks\n tasks = [asyncio.ensure_future(av) if not isinstance(av, asyncio.Task) else av\n for av in awaitables]\n\n # Use asyncio.wait with FIRST_COMPLETED to get results as they complete\n pending = set(tasks)\n results = {} # task -> result or exception\n exceptions = []\n\n while pending:\n done, pending = await asyncio.wait(pending, return_when=asyncio.FIRST_COMPLETED)\n\n for task in done:\n try:\n result = task.result()\n # Success: cancel all other pending tasks and return\n for t in pending:\n t.cancel()\n # Await all cancelled tasks to ensure they finish cancelling\n if pending:\n await asyncio.wait(pending)\n return result\n except asyncio.CancelledError:\n # This task was cancelled (unlikely in this context since we\n # only cancel pending ones after a success, but handle for safety)\n pass\n except Exception as e:\n # Record the exception for this task\n exceptions.append(e)\n\n # All tasks have failed\n if not return_exceptions:\n if exceptions:\n raise exceptions[0]\n else:\n raise RuntimeError(\"All awaitables failed with no exception recorded\")\n\n # Return an ExceptionGroup with all collected exceptions\n if len(exceptions) == 1:\n raise exceptions[0]\n else:\n # Create an ExceptionGroup-like behavior. In Python 3.11+, we can use ExceptionGroup.\n # For broader compatibility, we'll create a custom one or use the built-in.\n if hasattr(asyncio, 'ExceptionGroup') or hasattr(Builtins, 'ExceptionGroup'):\n try:\n # Python 3.11+ has ExceptionGroup in builtins\n raise ExceptionGroup(\"All awaitables failed\", exceptions)\n except NameError:\n # Fallback: raise the last exception\n raise exceptions[-1]\n else:\n raise exceptions[-1]\n\n\n# For Python 3.10 and below compatibility, provide a simple ExceptionGroup alternative\nclass CustomExceptionGroup(Exception):\n def __init__(self, message: str, exceptions: List[Exception]):\n self.exceptions = exceptions\n super().__init__(message)\n\n def __str__(self):\n msgs = [str(e) for e in self.exceptions]\n return f\"{self.args[0]}:\\n\" + \"\\n\".join(f\" {m}\" for m in msgs" + } + } + ], + "created": 1789936798, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 768, + "prompt_tokens": 54, + "total_tokens": 822, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-QpTMAiaFXqAAp22cfxOcjJb7A0v6toBp", + "timings": { + "cache_n": 0, + "prompt_n": 54, + "prompt_ms": 122.928, + "prompt_per_token_ms": 2.2764444444444445, + "prompt_per_second": 439.2815306520891, + "predicted_n": 768, + "predicted_ms": 10366.471, + "predicted_per_token_ms": 13.515607561929595, + "predicted_per_second": 73.98853476752117, + "draft_n": 789, + "draft_n_accepted": 503 + }, + "wall_seconds": 10.554032287036534 + } + ], + "prefill": [ + { + "target": 24576, + "response": { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "```json\n{\n \"alpha\": \"RAVEN-417\",\n \"beta\": \"CEDAR-928\",\n \"gamma\": \"ORBIT-563\"\n}\n```\n\nDie Verifizierung der identifizierten Datensätze erfordert einen strengen, methodischen Ansatz, der auf der genauen Analyse der bereitgestellten Textdaten basiert, ohne auf externe Annahmen oder erfundene Beweise zurückzugreifen. Der erste Schritt besteht darin, die Struktur des vorliegenden Datensatzes zu verstehen. Die meisten Einträge folgen einem einheitlichen Muster: \"Record [ID]: cobalt lantern maple orbit quartz river silver tango.\" Diese Wiederholung dient als Baseline, von der jede Abweichung abweicht.\n\nUm die Nadelwerte (NEEDLE values) zu verifizieren, muss man gezielt nach Stellen im Text suchen, an denen dieses Muster durchbrochen wird. Bei \"NEEDLE_ALPHA=RAVEN-417\" sieht man, dass zwischen Record 000146 und Record 000147 eine unterbrochene Zeile steht: \"Record 000146: cob\" gefolgt von \"NEEDLE_ALPHA=RAVEN-417\" und dann \"alt lantern maple orbit quartz river silver tango.\" Hier ist der Wert klar als separater String definiert, der das Wort \"cobalt\" teilt und ersetzt. Die Verifizierung erfolgt, indem man prüft, ob der String \"RAVEN-417\" exakt an dieser Stelle zwischen den Fragmenten \"cob\" und \"alt\" steht, was die Integrität der Einbettung bestätigt.\n\nÄhnlich verhält es sich mit \"NEEDLE_BETA=CEDAR-928\". Hier wird Record 000584 in zwei Teile geteilt: \"Recor\" und \"d 000585...\". Dazwischen ist der Wert \"CEDAR-928\" positioniert. Die Überprüfung erfordert das Abgleichen des Anfangs \"Recor\" mit dem Fortsetzungsanfang \"d 000585\", um sicherzustellen, dass der Datensatz nur durch die Einspeisung des NEEDLE-Werts unterbrochen wurde. Es ist wichtig, nicht zu versuchen, den fehlenden Teil des Wortes \"Record\" zu rekonstruieren oder zu ändern, sondern" + } + } + ], + "created": 1789936820, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 512, + "prompt_tokens": 24674, + "total_tokens": 25186, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-FcH5bKbslbNLAiLQGhilCzDIBppAAnN0", + "timings": { + "cache_n": 0, + "prompt_n": 24674, + "prompt_ms": 12783.075, + "prompt_per_token_ms": 0.518078746859042, + "prompt_per_second": 1930.208498346446, + "predicted_n": 512, + "predicted_ms": 9104.903, + "predicted_per_token_ms": 17.81781409001957, + "predicted_per_second": 56.123607247655464, + "draft_n": 598, + "draft_n_accepted": 311 + }, + "wall_seconds": 21.976672343036626, + "recall": { + "RAVEN-417": true, + "CEDAR-928": true, + "ORBIT-563": true + } + } + } + ], + "finished": 1789936820.3087893 +} diff --git a/experiments/medium-split-mtp-20260920/83-mtp3/server-args.json b/experiments/medium-split-mtp-20260920/83-mtp3/server-args.json new file mode 100644 index 0000000..3cf9a32 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/83-mtp3/server-args.json @@ -0,0 +1,75 @@ +[ + "--model", + "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "--mmproj", + "/models/qwen/mmproj-BF16.gguf", + "--mmproj-offload", + "--mmproj-device", + "CUDA1", + "--alias", + "qwen-medium", + "--ctx-size", + "160000", + "--flash-attn", + "on", + "--cache-type-k", + "q4_0", + "--cache-type-v", + "q4_0", + "--cache-prompt", + "--cache-ram", + "32768", + "--threads", + "6", + "--threads-batch", + "6", + "--batch-size", + "2048", + "--ubatch-size", + "256", + "--parallel", + "2", + "--kv-unified", + "--jinja", + "--reasoning", + "auto", + "--reasoning-preserve", + "--host", + "127.0.0.1", + "--port", + "5005", + "--metrics", + "--fit", + "off", + "--n-gpu-layers", + "all", + "--load-mode", + "none", + "--no-ui", + "--temperature", + "1.0", + "--top-p", + "0.95", + "--top-k", + "20", + "--device", + "CUDA0,CUDA1", + "--main-gpu", + "0", + "--split-mode", + "layer", + "--tensor-split", + "83,17", + "--spec-type", + "draft-mtp", + "--spec-draft-n-max", + "3", + "--spec-draft-type-k", + "f16", + "--spec-draft-type-v", + "f16", + "--spec-draft-p-min", + "0.05", + "--verbosity", + "3" +] diff --git a/experiments/medium-split-mtp-20260920/83-mtp3/server.log.gz b/experiments/medium-split-mtp-20260920/83-mtp3/server.log.gz new file mode 100644 index 0000000..f4fe19e Binary files /dev/null and b/experiments/medium-split-mtp-20260920/83-mtp3/server.log.gz differ diff --git a/experiments/medium-split-mtp-20260920/85-mtp2/config.json b/experiments/medium-split-mtp-20260920/85-mtp2/config.json new file mode 100644 index 0000000..fa53184 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/85-mtp2/config.json @@ -0,0 +1,10 @@ +{ + "label": "85-mtp2", + "ubatch": 256, + "split": "85,15", + "mtp": 2, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 +} diff --git a/experiments/medium-split-mtp-20260920/85-mtp2/gpu.json b/experiments/medium-split-mtp-20260920/85-mtp2/gpu.json new file mode 100644 index 0000000..5398d87 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/85-mtp2/gpu.json @@ -0,0 +1,577 @@ +[ + { + "time": 1789936657.4383829, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "6964", + "total": "12288", + "temp": "48", + "util": "100", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "11272", + "total": "16303", + "temp": "55", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936659.4671788, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "8424", + "total": "12288", + "temp": "48", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15574", + "total": "16303", + "temp": "54", + "util": "17", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936661.4983554, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10608", + "total": "12288", + "temp": "47", + "util": "2", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15574", + "total": "16303", + "temp": "53", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936663.5305703, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "52", + "util": "44", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "60", + "util": "51", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936665.5614212, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "53", + "util": "43", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "63", + "util": "50", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936667.5952804, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "54", + "util": "50", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "59", + "util": "51", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936669.627116, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "54", + "util": "48", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "61", + "util": "51", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936671.6623967, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "53", + "util": "44", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "60", + "util": "50", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936673.6930187, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "54", + "util": "47", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "63", + "util": "51", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936675.7236755, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "54", + "util": "44", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "60", + "util": "50", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936677.7546844, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "53", + "util": "45", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "64", + "util": "51", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936679.785267, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "54", + "util": "43", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "61", + "util": "50", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936681.8165836, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "54", + "util": "44", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "62", + "util": "50", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936683.8488667, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10612", + "total": "12288", + "temp": "54", + "util": "45", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15590", + "total": "16303", + "temp": "64", + "util": "50", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936685.8805652, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "52", + "util": "53", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "71", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936687.911105, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "54", + "util": "52", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "73", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936689.9424007, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "54", + "util": "55", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "73", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936691.973218, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "55", + "util": "48", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "72", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936694.0040557, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "53", + "util": "65", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "74", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936696.0368607, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "54", + "util": "64", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "74", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936698.0835536, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "57", + "util": "88", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "63", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936700.1248791, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "55", + "util": "41", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "65", + "util": "45", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936702.155853, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "54", + "util": "47", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "62", + "util": "60", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936704.1904633, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "55", + "util": "51", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "62", + "util": "56", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936706.2270036, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10614", + "total": "12288", + "temp": "55", + "util": "42", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15592", + "total": "16303", + "temp": "64", + "util": "44", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + } +] diff --git a/experiments/medium-split-mtp-20260920/85-mtp2/loaded.json b/experiments/medium-split-mtp-20260920/85-mtp2/loaded.json new file mode 100644 index 0000000..cd39ac4 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/85-mtp2/loaded.json @@ -0,0 +1,138 @@ +{ + "case": { + "label": "85-mtp2", + "ubatch": 256, + "split": "85,15", + "mtp": 2, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + "started": 1789936655.4056249, + "idle_gpu": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10608", + "total": "12288", + "temp": "47", + "util": "2", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15574", + "total": "16303", + "temp": "53", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ], + "props": { + "default_generation_settings": { + "params": { + "seed": 4294967295, + "temperature": 1.0, + "dynatemp_range": 0.0, + "dynatemp_exponent": 1.0, + "top_k": 20, + "top_p": 0.949999988079071, + "min_p": 0.05000000074505806, + "top_n_sigma": -1.0, + "xtc_probability": 0.0, + "xtc_threshold": 0.10000000149011612, + "typical_p": 1.0, + "repeat_last_n": 64, + "repeat_penalty": 1.0, + "presence_penalty": 0.0, + "frequency_penalty": 0.0, + "dry_multiplier": 0.0, + "dry_base": 1.75, + "dry_allowed_length": 2, + "dry_penalty_last_n": 64, + "mirostat": 0, + "mirostat_tau": 5.0, + "mirostat_eta": 0.10000000149011612, + "adaptive_target": -1.0, + "adaptive_decay": 0.8999999761581421, + "max_tokens": -1, + "n_predict": -1, + "n_keep": 0, + "n_discard": 0, + "ignore_eos": false, + "stream": false, + "n_probs": 0, + "min_keep": 0, + "chat_format": "Content-only", + "reasoning_format": "none", + "reasoning_in_content": false, + "generation_prompt": "", + "samplers": [ + "penalties", + "dry", + "top_n_sigma", + "top_k", + "typ_p", + "top_p", + "min_p", + "xtc", + "temperature" + ], + "speculative.types": "none", + "timings_per_token": false, + "post_sampling_probs": false, + "backend_sampling": false, + "lora": [] + }, + "n_ctx": 160000 + }, + "total_slots": 2, + "model_alias": "qwen-medium", + "model_ftype": "IQ4_XS - 4.25 bpw", + "model_path": "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "modalities": { + "vision": true, + "video": true, + "audio": false + }, + "media_marker": "<__media_uTU13sY3gUv1cIKm43dbCGG9ONDyqVyZ__>", + "endpoint_slots": true, + "endpoint_props": false, + "endpoint_metrics": true, + "ui": false, + "ui_settings": {}, + "chat_template": "{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- macro render_content(content, do_vision_count, is_system_content=false) %}\n {%- if content is string %}\n {{- content }}\n {%- elif content is iterable and content is not mapping %}\n {%- for item in content %}\n {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain images.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set image_count.value = image_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Picture ' ~ image_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|image_pad|><|vision_end|>' }}\n {%- elif 'video' in item or item.type == 'video' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain videos.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set video_count.value = video_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Video ' ~ video_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|video_pad|><|vision_end|>' }}\n {%- elif 'text' in item %}\n {{- item.text }}\n {%- else %}\n {{- raise_exception('Unexpected item type in content.') }}\n {%- endif %}\n {%- endfor %}\n {%- elif content is none or content is undefined %}\n {{- '' }}\n {%- else %}\n {{- raise_exception('Unexpected content type.') }}\n {%- endif %}\n{%- endmacro %}\n{%- if not messages %}\n {{- raise_exception('No messages provided.') }}\n{%- endif %}\n{%- set sysns = namespace(count=0, text='') %}\n{%- for message in messages %}\n {%- if sysns.count == loop.index0 and (message.role == 'system' or message.role == 'developer') %}\n {%- set sys_content = render_content(message.content, false, true)|trim %}\n {%- if sys_content %}\n {%- set sysns.text = sysns.text + ('\\n' if sysns.text else '') + sys_content %}\n {%- endif %}\n {%- set sysns.count = sysns.count + 1 %}\n {%- endif %}\n{%- endfor %}\n{%- set num_sys = sysns.count %}\n{%- set merged_system = sysns.text %}\n{%- set reasoning_instructions = '' %}\n{%- if enable_thinking is undefined or enable_thinking is true %}\n {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}\n {%- if resolved_reasoning_effort == 'high' %}\n {%- set resolved_reasoning_effort = 'xhigh' %}\n {%- endif %}\n {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}\n {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}\n {%- endif %}\n {%- if resolved_reasoning_effort == 'xhigh' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}\n {%- elif resolved_reasoning_effort == 'low' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}\n {%- endif %}\n{%- endif %}\n{%- if tools and tools is iterable and tools is not mapping %}\n {{- '<|im_start|>system\\n' }}\n {%- if reasoning_instructions %}\n {{- reasoning_instructions + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou have access to the following functions:\\n\\n\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n\" }}\n {{- '\\n\\nIf you choose to call a function ONLY reply in the following format with NO suffix:\\n\\n\\n\\n\\nvalue_1\\n\\n\\nThis is the value for the second parameter\\nthat can span\\nmultiple lines\\n\\n\\n\\n\\n\\nReminder:\\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\\n- Required parameters MUST be specified\\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\\n' }}\n {%- if merged_system %}\n {{- '\\n\\n' + merged_system }}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n{%- else %}\n {%- if merged_system %}\n {{- '<|im_start|>system\\n' + (reasoning_instructions + '\\n\\n' if reasoning_instructions else '') + merged_system + '<|im_end|>\\n' }}\n {%- elif reasoning_instructions %}\n {{- '<|im_start|>system\\n' + reasoning_instructions + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" %}\n {%- set content = render_content(message.content, false)|trim %}\n {%- if not(content.startswith('') and content.endswith('')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if loop.index0 >= num_sys %}\n {%- set content = render_content(message.content, true)|trim %}\n {%- if message.role == \"system\" or message.role == \"developer\" %}\n {{- raise_exception('System message must be at the beginning.') }}\n {%- elif message.role == \"user\" %}\n {{- '<|im_start|>' + message.role + '\\n' + content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is string %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- endif %}\n {%- set reasoning_content = reasoning_content|trim %}\n {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}\n {{- '<|im_start|>' + message.role + '\\n\\n' + reasoning_content + '\\n\\n\\n' + content }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {%- if tool_call.name is not defined or tool_call.name is none %}\n {{- raise_exception('Tool call is missing a function name.') }}\n {%- endif %}\n {%- if loop.first %}\n {%- if content|trim %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n\\n' }}\n {%- endif %}\n {%- else %}\n {{- '\\n\\n\\n' }}\n {%- endif %}\n {%- if tool_call.arguments is mapping %}\n {%- for args_name, args_value in tool_call.arguments|items %}\n {{- '\\n' }}\n {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}\n {{- args_value }}\n {{- '\\n\\n' }}\n {%- endfor %}\n {%- elif tool_call.arguments is string %}\n {%- if tool_call.arguments|trim %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" were passed as a JSON string. Parse them into an object before calling apply_chat_template.') }}\n {%- endif %}\n {%- elif tool_call.arguments is defined and tool_call.arguments is not none %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" must be an object/mapping or a JSON string.') }}\n {%- endif %}\n {{- '\\n' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.previtem and loop.previtem.role != \"tool\" %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n\\n' }}\n {{- content }}\n {{- '\\n' }}\n {%- if not loop.last and loop.nextitem.role != \"tool\" %}\n {{- '<|im_end|>\\n' }}\n {%- elif loop.last %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- else %}\n {{- raise_exception('Unexpected message role.') }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n' }}\n {%- endif %}\n{%- endif %}\n{#- Unsloth fixes - developer role, merged system messages, tool calling #}", + "chat_template_caps": { + "supports_object_arguments": true, + "supports_parallel_tool_calls": true, + "supports_preserve_reasoning": true, + "supports_reasoning_effort": true, + "supports_string_content": true, + "supports_system_role": true, + "supports_tool_calls": true, + "supports_tools": true, + "supports_typed_content": true + }, + "bos_token": "<|endoftext|>", + "eos_token": "<|im_end|>", + "build_info": "b10964-b29c606e2", + "is_sleeping": false, + "cors_proxy_enabled": false + }, + "slots": [ + { + "id": 0, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + }, + { + "id": 1, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + } + ] +} diff --git a/experiments/medium-split-mtp-20260920/85-mtp2/result.json b/experiments/medium-split-mtp-20260920/85-mtp2/result.json new file mode 100644 index 0000000..72720aa --- /dev/null +++ b/experiments/medium-split-mtp-20260920/85-mtp2/result.json @@ -0,0 +1,307 @@ +{ + "case": { + "label": "85-mtp2", + "ubatch": 256, + "split": "85,15", + "mtp": 2, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + "started": 1789936655.4056249, + "idle_gpu": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10608", + "total": "12288", + "temp": "47", + "util": "2", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15574", + "total": "16303", + "temp": "53", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ], + "props": { + "default_generation_settings": { + "params": { + "seed": 4294967295, + "temperature": 1.0, + "dynatemp_range": 0.0, + "dynatemp_exponent": 1.0, + "top_k": 20, + "top_p": 0.949999988079071, + "min_p": 0.05000000074505806, + "top_n_sigma": -1.0, + "xtc_probability": 0.0, + "xtc_threshold": 0.10000000149011612, + "typical_p": 1.0, + "repeat_last_n": 64, + "repeat_penalty": 1.0, + "presence_penalty": 0.0, + "frequency_penalty": 0.0, + "dry_multiplier": 0.0, + "dry_base": 1.75, + "dry_allowed_length": 2, + "dry_penalty_last_n": 64, + "mirostat": 0, + "mirostat_tau": 5.0, + "mirostat_eta": 0.10000000149011612, + "adaptive_target": -1.0, + "adaptive_decay": 0.8999999761581421, + "max_tokens": -1, + "n_predict": -1, + "n_keep": 0, + "n_discard": 0, + "ignore_eos": false, + "stream": false, + "n_probs": 0, + "min_keep": 0, + "chat_format": "Content-only", + "reasoning_format": "none", + "reasoning_in_content": false, + "generation_prompt": "", + "samplers": [ + "penalties", + "dry", + "top_n_sigma", + "top_k", + "typ_p", + "top_p", + "min_p", + "xtc", + "temperature" + ], + "speculative.types": "none", + "timings_per_token": false, + "post_sampling_probs": false, + "backend_sampling": false, + "lora": [] + }, + "n_ctx": 160000 + }, + "total_slots": 2, + "model_alias": "qwen-medium", + "model_ftype": "IQ4_XS - 4.25 bpw", + "model_path": "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "modalities": { + "vision": true, + "video": true, + "audio": false + }, + "media_marker": "<__media_uTU13sY3gUv1cIKm43dbCGG9ONDyqVyZ__>", + "endpoint_slots": true, + "endpoint_props": false, + "endpoint_metrics": true, + "ui": false, + "ui_settings": {}, + "chat_template": "{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- macro render_content(content, do_vision_count, is_system_content=false) %}\n {%- if content is string %}\n {{- content }}\n {%- elif content is iterable and content is not mapping %}\n {%- for item in content %}\n {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain images.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set image_count.value = image_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Picture ' ~ image_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|image_pad|><|vision_end|>' }}\n {%- elif 'video' in item or item.type == 'video' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain videos.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set video_count.value = video_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Video ' ~ video_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|video_pad|><|vision_end|>' }}\n {%- elif 'text' in item %}\n {{- item.text }}\n {%- else %}\n {{- raise_exception('Unexpected item type in content.') }}\n {%- endif %}\n {%- endfor %}\n {%- elif content is none or content is undefined %}\n {{- '' }}\n {%- else %}\n {{- raise_exception('Unexpected content type.') }}\n {%- endif %}\n{%- endmacro %}\n{%- if not messages %}\n {{- raise_exception('No messages provided.') }}\n{%- endif %}\n{%- set sysns = namespace(count=0, text='') %}\n{%- for message in messages %}\n {%- if sysns.count == loop.index0 and (message.role == 'system' or message.role == 'developer') %}\n {%- set sys_content = render_content(message.content, false, true)|trim %}\n {%- if sys_content %}\n {%- set sysns.text = sysns.text + ('\\n' if sysns.text else '') + sys_content %}\n {%- endif %}\n {%- set sysns.count = sysns.count + 1 %}\n {%- endif %}\n{%- endfor %}\n{%- set num_sys = sysns.count %}\n{%- set merged_system = sysns.text %}\n{%- set reasoning_instructions = '' %}\n{%- if enable_thinking is undefined or enable_thinking is true %}\n {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}\n {%- if resolved_reasoning_effort == 'high' %}\n {%- set resolved_reasoning_effort = 'xhigh' %}\n {%- endif %}\n {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}\n {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}\n {%- endif %}\n {%- if resolved_reasoning_effort == 'xhigh' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}\n {%- elif resolved_reasoning_effort == 'low' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}\n {%- endif %}\n{%- endif %}\n{%- if tools and tools is iterable and tools is not mapping %}\n {{- '<|im_start|>system\\n' }}\n {%- if reasoning_instructions %}\n {{- reasoning_instructions + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou have access to the following functions:\\n\\n\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n\" }}\n {{- '\\n\\nIf you choose to call a function ONLY reply in the following format with NO suffix:\\n\\n\\n\\n\\nvalue_1\\n\\n\\nThis is the value for the second parameter\\nthat can span\\nmultiple lines\\n\\n\\n\\n\\n\\nReminder:\\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\\n- Required parameters MUST be specified\\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\\n' }}\n {%- if merged_system %}\n {{- '\\n\\n' + merged_system }}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n{%- else %}\n {%- if merged_system %}\n {{- '<|im_start|>system\\n' + (reasoning_instructions + '\\n\\n' if reasoning_instructions else '') + merged_system + '<|im_end|>\\n' }}\n {%- elif reasoning_instructions %}\n {{- '<|im_start|>system\\n' + reasoning_instructions + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" %}\n {%- set content = render_content(message.content, false)|trim %}\n {%- if not(content.startswith('') and content.endswith('')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if loop.index0 >= num_sys %}\n {%- set content = render_content(message.content, true)|trim %}\n {%- if message.role == \"system\" or message.role == \"developer\" %}\n {{- raise_exception('System message must be at the beginning.') }}\n {%- elif message.role == \"user\" %}\n {{- '<|im_start|>' + message.role + '\\n' + content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is string %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- endif %}\n {%- set reasoning_content = reasoning_content|trim %}\n {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}\n {{- '<|im_start|>' + message.role + '\\n\\n' + reasoning_content + '\\n\\n\\n' + content }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {%- if tool_call.name is not defined or tool_call.name is none %}\n {{- raise_exception('Tool call is missing a function name.') }}\n {%- endif %}\n {%- if loop.first %}\n {%- if content|trim %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n\\n' }}\n {%- endif %}\n {%- else %}\n {{- '\\n\\n\\n' }}\n {%- endif %}\n {%- if tool_call.arguments is mapping %}\n {%- for args_name, args_value in tool_call.arguments|items %}\n {{- '\\n' }}\n {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}\n {{- args_value }}\n {{- '\\n\\n' }}\n {%- endfor %}\n {%- elif tool_call.arguments is string %}\n {%- if tool_call.arguments|trim %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" were passed as a JSON string. Parse them into an object before calling apply_chat_template.') }}\n {%- endif %}\n {%- elif tool_call.arguments is defined and tool_call.arguments is not none %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" must be an object/mapping or a JSON string.') }}\n {%- endif %}\n {{- '\\n' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.previtem and loop.previtem.role != \"tool\" %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n\\n' }}\n {{- content }}\n {{- '\\n' }}\n {%- if not loop.last and loop.nextitem.role != \"tool\" %}\n {{- '<|im_end|>\\n' }}\n {%- elif loop.last %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- else %}\n {{- raise_exception('Unexpected message role.') }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n' }}\n {%- endif %}\n{%- endif %}\n{#- Unsloth fixes - developer role, merged system messages, tool calling #}", + "chat_template_caps": { + "supports_object_arguments": true, + "supports_parallel_tool_calls": true, + "supports_preserve_reasoning": true, + "supports_reasoning_effort": true, + "supports_string_content": true, + "supports_system_role": true, + "supports_tool_calls": true, + "supports_tools": true, + "supports_typed_content": true + }, + "bos_token": "<|endoftext|>", + "eos_token": "<|im_end|>", + "build_info": "b10964-b29c606e2", + "is_sleeping": false, + "cors_proxy_enabled": false + }, + "slots": [ + { + "id": 0, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + }, + { + "id": 1, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + } + ], + "smoke": { + "choices": [ + { + "finish_reason": "stop", + "index": 0, + "message": { + "role": "assistant", + "content": "OK" + } + } + ], + "created": 1789936662, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 2, + "prompt_tokens": 19, + "total_tokens": 21, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-H5kwhDH81uL8rJBQx1R7dHv1AmwUitq5", + "timings": { + "cache_n": 0, + "prompt_n": 19, + "prompt_ms": 781.219, + "prompt_per_token_ms": 41.116789473684214, + "prompt_per_second": 24.320965055893417, + "predicted_n": 2, + "predicted_ms": 286.485, + "predicted_per_token_ms": 286.485, + "predicted_per_second": 3.490584149257378, + "draft_n": 2, + "draft_n_accepted": 2 + }, + "wall_seconds": 1.0702720619738102 + }, + "decode": [ + { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "Ein Reverse Proxy ist ein zentrales Element in modernen IT-Infrastrukturen, das als Vermittler zwischen Clienten und Backend-Servern agiert. Im Gegensatz zu einem Forward Proxy, der Anfragen von internen Clients nach außen vertritt, stellt der Reverse Proxy eine Eindepunkt-Schnittstelle (Single Point of Entry) für externe Anfragen dar. Typische Vertreter sind nginx, HAProxy oder Apache HTTP Server. Die Grundprinzipie des Reverse Proxys besteht darin, eingehende HTTP- oder HTTPS-Anfragen entgegenzunehmen und diese basierend auf definierten Routing-Regeln an die geeigneten Backend-Server weiterzuleiten. Diese Architektur bietet erhebliche Vorteile wie Lastverteilung, SSL/TLS-Terminierung, Caching und eine vereinfachte Verwaltung mehrerer hinterlegter Dienste. Durch die Abstraktion der Backend-Infrastruktur wird sichergestellt, dass Clients die Details der internen Topologie, wie IP-Adressen oder Server-Namen, nicht kennen müssen.\n\nIn dynamischen Umgebungen, insbesondere bei der Nutzung von Container-Orchestrierungsplattformen wie Kubernetes oder Docker Compose, ist die Stabilität der Netzwerkanbindung eine große Herausforderung. Container sind per Definition flüchtig; sie können bei Updates, Skalierungsoperationen oder Fehlern neu gestartet werden. Bei jedem Neustart oder einer Neuanlage eines Containers kann die zugewiesene IP-Adresse im internen Docker- oder Kubernetes-Netzwerk wechseln. Wenn ein Reverse Proxy statische Konfigurationen verwendet, die auf bestimmte IP-Adressen oder Ports verweisen, führt dieser Wechsel direkt zu Verbindungsproblemen.\n\nDer häufigste Fehler in solchen Szenarien ist die \"Connection Refused\"- oder \"Connection Timeout\"-Situation. Wenn der Reverse Proxy versucht, eine Anfrage an eine IP-Adresse weiterzugeben, die einem Container zugeordnet war, dieser aber mittlerweile unter einer neuen Adresse erreichbar ist oder gar nicht mehr existiert, schlägt das TCP-Handshake-Protokoll fehl. Der Client sieht daraufhin typischerweise HTTP-Fehler 502 (Bad Gateway) oder 504 (Gateway Timeout). Ein 502-Fehler deutet oft darauf hin, dass das Backend-Server-Protokoll eine ungültige Antwort geliefert hat oder die Verbindung komplett abgelehnt wurde. Ein 504-Fehler tritt auf, wenn das Backend nicht rechtzeitig auf die Anfrage des Proxys antwortet, was bei einem hängenden oder falsch konfigurierten Container häufig vorkommt.\n\nUm solche Probleme anhand von Logs effektiv einzugrenzen, muss man systematisch vorgehen. Der erste Schritt ist die Analyse der Logs des Reverse Proxys selbst. Bei nginx befindet sich die Fehlerinformation im `error.log`. Hier sucht man nach Einträgen wie `connect() failed (111: Connection refused)` oder `upstream timed out`. Diese Meldungen bestätigen, dass der Proxy die Verbindung zum definierten Backend nicht herstellen konnte. Wichtig ist dabei, die Ziel-IP und den Port im Log-Eintrag zu notieren.\n\nIm zweiten Schritt muss man die Logs des Orchestrierers oder des Container-Runtimes (z. B. `docker logs` oder `kubectl describe pod`) prüfen. Man vergleicht die im Proxy-Log festgestellte IP-Adresse mit der aktuellen IP des laufenden Containers. Stimmt die IP nicht überein, liegt ein Problem mit der dynamischen DNS- oder Service-Auflösung vor. In Kubernetes-Umgebungen ist es daher ratsam, statt fester IPs die internen Service-Namen (z. B. `http://api-service`) in der Proxy-Konfiguration" + } + } + ], + "created": 1789936674, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 768, + "prompt_tokens": 59, + "total_tokens": 827, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-n5iknizKxzo961SEpq50X0OcfWWIPgyV", + "timings": { + "cache_n": 0, + "prompt_n": 59, + "prompt_ms": 118.721, + "prompt_per_token_ms": 2.012220338983051, + "prompt_per_second": 496.96346897347564, + "predicted_n": 768, + "predicted_ms": 12112.526, + "predicted_per_token_ms": 15.792080834419817, + "predicted_per_second": 63.322877490624165, + "draft_n": 732, + "draft_n_accepted": 401 + }, + "wall_seconds": 12.29133506200742 + }, + { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "```python\nimport asyncio\nfrom typing import Any, Awaitable, Callable, List, Optional, Union\n\n\nasync def first_success(\n awaitables: List[Awaitable],\n return_exceptions: bool = True\n) -> Any:\n \"\"\"\n Start all awaitables concurrently, return the first successful result,\n cancel and await remaining tasks, and collect exceptions if all fail.\n\n Args:\n awaitables: A list of awaitables to run concurrently.\n return_exceptions: If True, when all awaitables fail, raise an\n ExceptionGroup (Python 3.11+) or re-raise the first exception\n with others attached. If False, simply re-raise the first exception.\n\n Returns:\n The result of the first successfully completed awaitable.\n\n Raises:\n ExceptionGroup (or the first exception if return_exceptions=False):\n If all awaitables fail.\n \"\"\"\n if not awaitables:\n raise ValueError(\"awaitables list must not be empty\")\n\n # Create tasks from all awaitables\n tasks = [asyncio.ensure_future(av) for av in awaitables]\n\n try:\n # Use asyncio.wait to get the first completed task\n # We need to keep polling until we find a successful one\n pending = set(tasks)\n\n while pending:\n done, pending = await asyncio.wait(pending, return_when=asyncio.FIRST_COMPLETED)\n\n for task in done:\n try:\n result = task.result()\n # Success! Cancel all remaining tasks and return result\n for t in pending:\n t.cancel()\n # Await the cancellation of remaining tasks to ensure clean shutdown\n if pending:\n await asyncio.gather(*pending, return_exceptions=True)\n return result\n except asyncio.CancelledError:\n # This task was cancelled, treat as failure\n continue\n except Exception:\n # Task raised an exception, continue checking others\n continue\n\n # If we get here, all done tasks failed; continue waiting for remaining\n\n # If we exhaust all tasks and none succeeded\n # Collect all exceptions\n exceptions = []\n for task in tasks:\n try:\n task.result()\n except Exception as e:\n exceptions.append(e)\n except asyncio.CancelledError:\n exceptions.append(asyncio.CancelledError())\n\n if not exceptions:\n # All tasks were cancelled somehow, but no exceptions captured\n raise RuntimeError(\"All awaitables were cancelled without explicit exceptions\")\n\n if return_exceptions:\n if hasattr(asyncio, 'ExceptionGroup'): # Python 3.11+\n raise asyncio.ExceptionGroup(\"All awaitables failed\", exceptions)\n else:\n # Fallback: raise the first exception, attach others\n first_exc = exceptions[0]\n # Add other exceptions as context\n for exc in exceptions[1:]:\n first_exc.__context__ = exc\n raise first_exc\n else:\n raise exceptions[0]\n\n except asyncio.CancelledError:\n # If the entire first_success coroutine is cancelled, cancel all tasks\n for t in tasks:\n t.cancel()\n await asyncio.gather(*tasks, return_exceptions=True)\n raise\n\n\n# Usage example\nasync def example():\n \"\"\"Demonstrate first_success with a practical example.\"\"\"\n\n async def slow_success(delay: float) -> str:\n \"\"\"Simulate a slow operation that eventually succeeds.\"\"\"\n await asyncio.sleep(delay)\n return f\"success after {delay}s\"\n\n" + } + } + ], + "created": 1789936685, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 768, + "prompt_tokens": 54, + "total_tokens": 822, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-JW9n3P53QA4c7sgHVdEmONCUEeTj2Dpv", + "timings": { + "cache_n": 0, + "prompt_n": 54, + "prompt_ms": 117.561, + "prompt_per_token_ms": 2.1770555555555555, + "prompt_per_second": 459.336004287136, + "predicted_n": 768, + "predicted_ms": 10178.693, + "predicted_per_token_ms": 13.270786179921773, + "predicted_per_second": 75.353485953452, + "draft_n": 614, + "draft_n_accepted": 460 + }, + "wall_seconds": 10.361086986027658 + } + ], + "prefill": [ + { + "target": 24576, + "response": { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "```json\n{\n \"alpha\": \"RAVEN-417\",\n \"beta\": \"CEDAR-928\",\n \"gamma\": \"ORBIT-563\"\n}\n```\n\nDie Verifizierung der identifizierten Nadelwerte (`RAVEN-417`, `CEDAR-928`, `ORBIT-563`) innerhalb des vorliegenden Datensatzes erfordert einen streng systematischen und evidenzbasierten Ansatz, der auf der präzisen Analyse der Textstruktur und der Positionierung der Markierungen beruht. Zunächst ist es entscheidend, die Integrität der Quelldaten zu wahren, indem man keine extrapolierten oder hypothetischen Datenpunkte einführt. Der vorliegende Datenset besteht überwiegend aus identischen, wiederholten Zeilen, die den String „cobalt lantern maple orbit quartz river silver tango.“ enthalten, wobei die Einträge von `Record 000000` bis `Record 001169` nummeriert sind. Die Nadeln wurden jedoch nicht als Teil des eigentlichen Datensatzinhalts (wie etwa in der Spalte der Werte) platziert, sondern als isolierte Metadaten-Zeilen zwischen den regulären Aufzeichnungen.\n\nUm `NEEDLE_ALPHA=RAVEN-417` zu verifizieren, muss der Fokus auf die Region um die Aufzeichnungen 000146 bis 000147 gelegt werden. Hier findet sich die Zeile, die mit „Record 000146: cob...“ beginnt und abgebrochen ist, gefolgt direkt von der isolierten Zeile „NEEDLE_ALPHA=RAVEN-417“. Die Verifikation erfolgt durch die Bestätigung, dass diese spezifische String-Kombination exakt an dieser Position im Textstrom erscheint und nicht in anderen Teilen des Dokuments vorkommt. Ein ähnlicher Prozess gilt für `NEEDLE_BETA=CEDAR-928`, der sich zwischen `Record 000584` und `Record 000585` befindet. Hier ist besondere Sorgfalt geboten, da die Zeile für Record 000585 aufgrund der Platzierung der Nadel ungewöhnlich formatiert oder fragmentiert erscheint („Recor... d 000585“). Die Verifikation muss daher sicherstellen," + } + } + ], + "created": 1789936706, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 512, + "prompt_tokens": 24674, + "total_tokens": 25186, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-ri1lA3zBowOkArxgp0aNkq6PGla5VmBI", + "timings": { + "cache_n": 0, + "prompt_n": 24674, + "prompt_ms": 12942.719, + "prompt_per_token_ms": 0.5245488773607846, + "prompt_per_second": 1906.4000385081374, + "predicted_n": 512, + "predicted_ms": 8537.934, + "predicted_per_token_ms": 16.70828571428571, + "predicted_per_second": 59.850544639956226, + "draft_n": 441, + "draft_n_accepted": 290 + }, + "wall_seconds": 21.566215507977176, + "recall": { + "RAVEN-417": true, + "CEDAR-928": true, + "ORBIT-563": true + } + } + } + ], + "finished": 1789936707.0146825 +} diff --git a/experiments/medium-split-mtp-20260920/85-mtp2/server-args.json b/experiments/medium-split-mtp-20260920/85-mtp2/server-args.json new file mode 100644 index 0000000..d9cc2d9 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/85-mtp2/server-args.json @@ -0,0 +1,75 @@ +[ + "--model", + "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "--mmproj", + "/models/qwen/mmproj-BF16.gguf", + "--mmproj-offload", + "--mmproj-device", + "CUDA1", + "--alias", + "qwen-medium", + "--ctx-size", + "160000", + "--flash-attn", + "on", + "--cache-type-k", + "q4_0", + "--cache-type-v", + "q4_0", + "--cache-prompt", + "--cache-ram", + "32768", + "--threads", + "6", + "--threads-batch", + "6", + "--batch-size", + "2048", + "--ubatch-size", + "256", + "--parallel", + "2", + "--kv-unified", + "--jinja", + "--reasoning", + "auto", + "--reasoning-preserve", + "--host", + "127.0.0.1", + "--port", + "5005", + "--metrics", + "--fit", + "off", + "--n-gpu-layers", + "all", + "--load-mode", + "none", + "--no-ui", + "--temperature", + "1.0", + "--top-p", + "0.95", + "--top-k", + "20", + "--device", + "CUDA0,CUDA1", + "--main-gpu", + "0", + "--split-mode", + "layer", + "--tensor-split", + "85,15", + "--spec-type", + "draft-mtp", + "--spec-draft-n-max", + "2", + "--spec-draft-type-k", + "f16", + "--spec-draft-type-v", + "f16", + "--spec-draft-p-min", + "0.05", + "--verbosity", + "3" +] diff --git a/experiments/medium-split-mtp-20260920/85-mtp2/server.log.gz b/experiments/medium-split-mtp-20260920/85-mtp2/server.log.gz new file mode 100644 index 0000000..792ab9b Binary files /dev/null and b/experiments/medium-split-mtp-20260920/85-mtp2/server.log.gz differ diff --git a/experiments/medium-split-mtp-20260920/85-mtp4/config.json b/experiments/medium-split-mtp-20260920/85-mtp4/config.json new file mode 100644 index 0000000..58c4927 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/85-mtp4/config.json @@ -0,0 +1,10 @@ +{ + "label": "85-mtp4", + "ubatch": 256, + "split": "85,15", + "mtp": 4, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 +} diff --git a/experiments/medium-split-mtp-20260920/85-mtp4/gpu.json b/experiments/medium-split-mtp-20260920/85-mtp4/gpu.json new file mode 100644 index 0000000..78f17bd --- /dev/null +++ b/experiments/medium-split-mtp-20260920/85-mtp4/gpu.json @@ -0,0 +1,623 @@ +[ + { + "time": 1789936709.8619335, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "6964", + "total": "12288", + "temp": "49", + "util": "100", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "11272", + "total": "16303", + "temp": "56", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936711.8901343, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "8230", + "total": "12288", + "temp": "48", + "util": "7", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15852", + "total": "16303", + "temp": "55", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936713.9222171, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10414", + "total": "12288", + "temp": "48", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15854", + "total": "16303", + "temp": "54", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936715.9548185, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "54", + "util": "50", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "63", + "util": "35", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936717.9897373, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "55", + "util": "63", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "60", + "util": "43", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936720.0238986, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "55", + "util": "49", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "63", + "util": "47", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936722.0577638, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "55", + "util": "51", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "65", + "util": "35", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936724.089602, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "55", + "util": "57", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "61", + "util": "36", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936726.1215844, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "55", + "util": "63", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "60", + "util": "35", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936728.155224, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "56", + "util": "61", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "60", + "util": "49", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936730.1939929, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "55", + "util": "9", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "60", + "util": "19", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936732.233054, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "57", + "util": "58", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "60", + "util": "36", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936734.2694385, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "55", + "util": "51", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "62", + "util": "35", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936736.3035913, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "56", + "util": "54", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "60", + "util": "36", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936738.3328967, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "56", + "util": "49", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15870", + "total": "16303", + "temp": "63", + "util": "41", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936740.3630881, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10418", + "total": "12288", + "temp": "53", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "57", + "util": "33", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936742.3925698, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "57", + "util": "1", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "59", + "util": "51", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936744.422326, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "55", + "util": "69", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "70", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936746.4554908, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "56", + "util": "37", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "64", + "util": "12", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936748.4865391, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "56", + "util": "36", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "74", + "util": "90", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936750.516579, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "57", + "util": "44", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "75", + "util": "91", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936752.547624, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "53", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "74", + "util": "81", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936754.5767446, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "57", + "util": "56", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "62", + "util": "44", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936756.6081572, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "56", + "util": "50", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "64", + "util": "39", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936758.6397696, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "56", + "util": "56", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "67", + "util": "43", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936760.67214, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "56", + "util": "51", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "62", + "util": "43", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936762.7054853, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10420", + "total": "12288", + "temp": "57", + "util": "56", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15872", + "total": "16303", + "temp": "61", + "util": "44", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + } +] diff --git a/experiments/medium-split-mtp-20260920/85-mtp4/loaded.json b/experiments/medium-split-mtp-20260920/85-mtp4/loaded.json new file mode 100644 index 0000000..1dae725 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/85-mtp4/loaded.json @@ -0,0 +1,138 @@ +{ + "case": { + "label": "85-mtp4", + "ubatch": 256, + "split": "85,15", + "mtp": 4, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + "started": 1789936707.8295877, + "idle_gpu": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10414", + "total": "12288", + "temp": "48", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15854", + "total": "16303", + "temp": "54", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ], + "props": { + "default_generation_settings": { + "params": { + "seed": 4294967295, + "temperature": 1.0, + "dynatemp_range": 0.0, + "dynatemp_exponent": 1.0, + "top_k": 20, + "top_p": 0.949999988079071, + "min_p": 0.05000000074505806, + "top_n_sigma": -1.0, + "xtc_probability": 0.0, + "xtc_threshold": 0.10000000149011612, + "typical_p": 1.0, + "repeat_last_n": 64, + "repeat_penalty": 1.0, + "presence_penalty": 0.0, + "frequency_penalty": 0.0, + "dry_multiplier": 0.0, + "dry_base": 1.75, + "dry_allowed_length": 2, + "dry_penalty_last_n": 64, + "mirostat": 0, + "mirostat_tau": 5.0, + "mirostat_eta": 0.10000000149011612, + "adaptive_target": -1.0, + "adaptive_decay": 0.8999999761581421, + "max_tokens": -1, + "n_predict": -1, + "n_keep": 0, + "n_discard": 0, + "ignore_eos": false, + "stream": false, + "n_probs": 0, + "min_keep": 0, + "chat_format": "Content-only", + "reasoning_format": "none", + "reasoning_in_content": false, + "generation_prompt": "", + "samplers": [ + "penalties", + "dry", + "top_n_sigma", + "top_k", + "typ_p", + "top_p", + "min_p", + "xtc", + "temperature" + ], + "speculative.types": "none", + "timings_per_token": false, + "post_sampling_probs": false, + "backend_sampling": false, + "lora": [] + }, + "n_ctx": 160000 + }, + "total_slots": 2, + "model_alias": "qwen-medium", + "model_ftype": "IQ4_XS - 4.25 bpw", + "model_path": "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "modalities": { + "vision": true, + "video": true, + "audio": false + }, + "media_marker": "<__media_vvNyoughnjBZXLAsLjpPKvrcxf6qc9qc__>", + "endpoint_slots": true, + "endpoint_props": false, + "endpoint_metrics": true, + "ui": false, + "ui_settings": {}, + "chat_template": "{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- macro render_content(content, do_vision_count, is_system_content=false) %}\n {%- if content is string %}\n {{- content }}\n {%- elif content is iterable and content is not mapping %}\n {%- for item in content %}\n {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain images.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set image_count.value = image_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Picture ' ~ image_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|image_pad|><|vision_end|>' }}\n {%- elif 'video' in item or item.type == 'video' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain videos.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set video_count.value = video_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Video ' ~ video_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|video_pad|><|vision_end|>' }}\n {%- elif 'text' in item %}\n {{- item.text }}\n {%- else %}\n {{- raise_exception('Unexpected item type in content.') }}\n {%- endif %}\n {%- endfor %}\n {%- elif content is none or content is undefined %}\n {{- '' }}\n {%- else %}\n {{- raise_exception('Unexpected content type.') }}\n {%- endif %}\n{%- endmacro %}\n{%- if not messages %}\n {{- raise_exception('No messages provided.') }}\n{%- endif %}\n{%- set sysns = namespace(count=0, text='') %}\n{%- for message in messages %}\n {%- if sysns.count == loop.index0 and (message.role == 'system' or message.role == 'developer') %}\n {%- set sys_content = render_content(message.content, false, true)|trim %}\n {%- if sys_content %}\n {%- set sysns.text = sysns.text + ('\\n' if sysns.text else '') + sys_content %}\n {%- endif %}\n {%- set sysns.count = sysns.count + 1 %}\n {%- endif %}\n{%- endfor %}\n{%- set num_sys = sysns.count %}\n{%- set merged_system = sysns.text %}\n{%- set reasoning_instructions = '' %}\n{%- if enable_thinking is undefined or enable_thinking is true %}\n {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}\n {%- if resolved_reasoning_effort == 'high' %}\n {%- set resolved_reasoning_effort = 'xhigh' %}\n {%- endif %}\n {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}\n {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}\n {%- endif %}\n {%- if resolved_reasoning_effort == 'xhigh' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}\n {%- elif resolved_reasoning_effort == 'low' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}\n {%- endif %}\n{%- endif %}\n{%- if tools and tools is iterable and tools is not mapping %}\n {{- '<|im_start|>system\\n' }}\n {%- if reasoning_instructions %}\n {{- reasoning_instructions + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou have access to the following functions:\\n\\n\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n\" }}\n {{- '\\n\\nIf you choose to call a function ONLY reply in the following format with NO suffix:\\n\\n\\n\\n\\nvalue_1\\n\\n\\nThis is the value for the second parameter\\nthat can span\\nmultiple lines\\n\\n\\n\\n\\n\\nReminder:\\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\\n- Required parameters MUST be specified\\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\\n' }}\n {%- if merged_system %}\n {{- '\\n\\n' + merged_system }}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n{%- else %}\n {%- if merged_system %}\n {{- '<|im_start|>system\\n' + (reasoning_instructions + '\\n\\n' if reasoning_instructions else '') + merged_system + '<|im_end|>\\n' }}\n {%- elif reasoning_instructions %}\n {{- '<|im_start|>system\\n' + reasoning_instructions + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" %}\n {%- set content = render_content(message.content, false)|trim %}\n {%- if not(content.startswith('') and content.endswith('')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if loop.index0 >= num_sys %}\n {%- set content = render_content(message.content, true)|trim %}\n {%- if message.role == \"system\" or message.role == \"developer\" %}\n {{- raise_exception('System message must be at the beginning.') }}\n {%- elif message.role == \"user\" %}\n {{- '<|im_start|>' + message.role + '\\n' + content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is string %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- endif %}\n {%- set reasoning_content = reasoning_content|trim %}\n {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}\n {{- '<|im_start|>' + message.role + '\\n\\n' + reasoning_content + '\\n\\n\\n' + content }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {%- if tool_call.name is not defined or tool_call.name is none %}\n {{- raise_exception('Tool call is missing a function name.') }}\n {%- endif %}\n {%- if loop.first %}\n {%- if content|trim %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n\\n' }}\n {%- endif %}\n {%- else %}\n {{- '\\n\\n\\n' }}\n {%- endif %}\n {%- if tool_call.arguments is mapping %}\n {%- for args_name, args_value in tool_call.arguments|items %}\n {{- '\\n' }}\n {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}\n {{- args_value }}\n {{- '\\n\\n' }}\n {%- endfor %}\n {%- elif tool_call.arguments is string %}\n {%- if tool_call.arguments|trim %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" were passed as a JSON string. Parse them into an object before calling apply_chat_template.') }}\n {%- endif %}\n {%- elif tool_call.arguments is defined and tool_call.arguments is not none %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" must be an object/mapping or a JSON string.') }}\n {%- endif %}\n {{- '\\n' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.previtem and loop.previtem.role != \"tool\" %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n\\n' }}\n {{- content }}\n {{- '\\n' }}\n {%- if not loop.last and loop.nextitem.role != \"tool\" %}\n {{- '<|im_end|>\\n' }}\n {%- elif loop.last %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- else %}\n {{- raise_exception('Unexpected message role.') }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n' }}\n {%- endif %}\n{%- endif %}\n{#- Unsloth fixes - developer role, merged system messages, tool calling #}", + "chat_template_caps": { + "supports_object_arguments": true, + "supports_parallel_tool_calls": true, + "supports_preserve_reasoning": true, + "supports_reasoning_effort": true, + "supports_string_content": true, + "supports_system_role": true, + "supports_tool_calls": true, + "supports_tools": true, + "supports_typed_content": true + }, + "bos_token": "<|endoftext|>", + "eos_token": "<|im_end|>", + "build_info": "b10964-b29c606e2", + "is_sleeping": false, + "cors_proxy_enabled": false + }, + "slots": [ + { + "id": 0, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + }, + { + "id": 1, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + } + ] +} diff --git a/experiments/medium-split-mtp-20260920/85-mtp4/result.json b/experiments/medium-split-mtp-20260920/85-mtp4/result.json new file mode 100644 index 0000000..be20d82 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/85-mtp4/result.json @@ -0,0 +1,307 @@ +{ + "case": { + "label": "85-mtp4", + "ubatch": 256, + "split": "85,15", + "mtp": 4, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + "started": 1789936707.8295877, + "idle_gpu": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10414", + "total": "12288", + "temp": "48", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15854", + "total": "16303", + "temp": "54", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ], + "props": { + "default_generation_settings": { + "params": { + "seed": 4294967295, + "temperature": 1.0, + "dynatemp_range": 0.0, + "dynatemp_exponent": 1.0, + "top_k": 20, + "top_p": 0.949999988079071, + "min_p": 0.05000000074505806, + "top_n_sigma": -1.0, + "xtc_probability": 0.0, + "xtc_threshold": 0.10000000149011612, + "typical_p": 1.0, + "repeat_last_n": 64, + "repeat_penalty": 1.0, + "presence_penalty": 0.0, + "frequency_penalty": 0.0, + "dry_multiplier": 0.0, + "dry_base": 1.75, + "dry_allowed_length": 2, + "dry_penalty_last_n": 64, + "mirostat": 0, + "mirostat_tau": 5.0, + "mirostat_eta": 0.10000000149011612, + "adaptive_target": -1.0, + "adaptive_decay": 0.8999999761581421, + "max_tokens": -1, + "n_predict": -1, + "n_keep": 0, + "n_discard": 0, + "ignore_eos": false, + "stream": false, + "n_probs": 0, + "min_keep": 0, + "chat_format": "Content-only", + "reasoning_format": "none", + "reasoning_in_content": false, + "generation_prompt": "", + "samplers": [ + "penalties", + "dry", + "top_n_sigma", + "top_k", + "typ_p", + "top_p", + "min_p", + "xtc", + "temperature" + ], + "speculative.types": "none", + "timings_per_token": false, + "post_sampling_probs": false, + "backend_sampling": false, + "lora": [] + }, + "n_ctx": 160000 + }, + "total_slots": 2, + "model_alias": "qwen-medium", + "model_ftype": "IQ4_XS - 4.25 bpw", + "model_path": "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "modalities": { + "vision": true, + "video": true, + "audio": false + }, + "media_marker": "<__media_vvNyoughnjBZXLAsLjpPKvrcxf6qc9qc__>", + "endpoint_slots": true, + "endpoint_props": false, + "endpoint_metrics": true, + "ui": false, + "ui_settings": {}, + "chat_template": "{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- macro render_content(content, do_vision_count, is_system_content=false) %}\n {%- if content is string %}\n {{- content }}\n {%- elif content is iterable and content is not mapping %}\n {%- for item in content %}\n {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain images.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set image_count.value = image_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Picture ' ~ image_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|image_pad|><|vision_end|>' }}\n {%- elif 'video' in item or item.type == 'video' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain videos.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set video_count.value = video_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Video ' ~ video_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|video_pad|><|vision_end|>' }}\n {%- elif 'text' in item %}\n {{- item.text }}\n {%- else %}\n {{- raise_exception('Unexpected item type in content.') }}\n {%- endif %}\n {%- endfor %}\n {%- elif content is none or content is undefined %}\n {{- '' }}\n {%- else %}\n {{- raise_exception('Unexpected content type.') }}\n {%- endif %}\n{%- endmacro %}\n{%- if not messages %}\n {{- raise_exception('No messages provided.') }}\n{%- endif %}\n{%- set sysns = namespace(count=0, text='') %}\n{%- for message in messages %}\n {%- if sysns.count == loop.index0 and (message.role == 'system' or message.role == 'developer') %}\n {%- set sys_content = render_content(message.content, false, true)|trim %}\n {%- if sys_content %}\n {%- set sysns.text = sysns.text + ('\\n' if sysns.text else '') + sys_content %}\n {%- endif %}\n {%- set sysns.count = sysns.count + 1 %}\n {%- endif %}\n{%- endfor %}\n{%- set num_sys = sysns.count %}\n{%- set merged_system = sysns.text %}\n{%- set reasoning_instructions = '' %}\n{%- if enable_thinking is undefined or enable_thinking is true %}\n {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}\n {%- if resolved_reasoning_effort == 'high' %}\n {%- set resolved_reasoning_effort = 'xhigh' %}\n {%- endif %}\n {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}\n {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}\n {%- endif %}\n {%- if resolved_reasoning_effort == 'xhigh' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}\n {%- elif resolved_reasoning_effort == 'low' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}\n {%- endif %}\n{%- endif %}\n{%- if tools and tools is iterable and tools is not mapping %}\n {{- '<|im_start|>system\\n' }}\n {%- if reasoning_instructions %}\n {{- reasoning_instructions + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou have access to the following functions:\\n\\n\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n\" }}\n {{- '\\n\\nIf you choose to call a function ONLY reply in the following format with NO suffix:\\n\\n\\n\\n\\nvalue_1\\n\\n\\nThis is the value for the second parameter\\nthat can span\\nmultiple lines\\n\\n\\n\\n\\n\\nReminder:\\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\\n- Required parameters MUST be specified\\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\\n' }}\n {%- if merged_system %}\n {{- '\\n\\n' + merged_system }}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n{%- else %}\n {%- if merged_system %}\n {{- '<|im_start|>system\\n' + (reasoning_instructions + '\\n\\n' if reasoning_instructions else '') + merged_system + '<|im_end|>\\n' }}\n {%- elif reasoning_instructions %}\n {{- '<|im_start|>system\\n' + reasoning_instructions + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" %}\n {%- set content = render_content(message.content, false)|trim %}\n {%- if not(content.startswith('') and content.endswith('')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if loop.index0 >= num_sys %}\n {%- set content = render_content(message.content, true)|trim %}\n {%- if message.role == \"system\" or message.role == \"developer\" %}\n {{- raise_exception('System message must be at the beginning.') }}\n {%- elif message.role == \"user\" %}\n {{- '<|im_start|>' + message.role + '\\n' + content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is string %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- endif %}\n {%- set reasoning_content = reasoning_content|trim %}\n {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}\n {{- '<|im_start|>' + message.role + '\\n\\n' + reasoning_content + '\\n\\n\\n' + content }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {%- if tool_call.name is not defined or tool_call.name is none %}\n {{- raise_exception('Tool call is missing a function name.') }}\n {%- endif %}\n {%- if loop.first %}\n {%- if content|trim %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n\\n' }}\n {%- endif %}\n {%- else %}\n {{- '\\n\\n\\n' }}\n {%- endif %}\n {%- if tool_call.arguments is mapping %}\n {%- for args_name, args_value in tool_call.arguments|items %}\n {{- '\\n' }}\n {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}\n {{- args_value }}\n {{- '\\n\\n' }}\n {%- endfor %}\n {%- elif tool_call.arguments is string %}\n {%- if tool_call.arguments|trim %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" were passed as a JSON string. Parse them into an object before calling apply_chat_template.') }}\n {%- endif %}\n {%- elif tool_call.arguments is defined and tool_call.arguments is not none %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" must be an object/mapping or a JSON string.') }}\n {%- endif %}\n {{- '\\n' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.previtem and loop.previtem.role != \"tool\" %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n\\n' }}\n {{- content }}\n {{- '\\n' }}\n {%- if not loop.last and loop.nextitem.role != \"tool\" %}\n {{- '<|im_end|>\\n' }}\n {%- elif loop.last %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- else %}\n {{- raise_exception('Unexpected message role.') }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n' }}\n {%- endif %}\n{%- endif %}\n{#- Unsloth fixes - developer role, merged system messages, tool calling #}", + "chat_template_caps": { + "supports_object_arguments": true, + "supports_parallel_tool_calls": true, + "supports_preserve_reasoning": true, + "supports_reasoning_effort": true, + "supports_string_content": true, + "supports_system_role": true, + "supports_tool_calls": true, + "supports_tools": true, + "supports_typed_content": true + }, + "bos_token": "<|endoftext|>", + "eos_token": "<|im_end|>", + "build_info": "b10964-b29c606e2", + "is_sleeping": false, + "cors_proxy_enabled": false + }, + "slots": [ + { + "id": 0, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + }, + { + "id": 1, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + } + ], + "smoke": { + "choices": [ + { + "finish_reason": "stop", + "index": 0, + "message": { + "role": "assistant", + "content": "OK" + } + } + ], + "created": 1789936715, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 2, + "prompt_tokens": 19, + "total_tokens": 21, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-PtJ8eyPFdnPWhQS8gKDk3kJLqmWcCHZ8", + "timings": { + "cache_n": 0, + "prompt_n": 19, + "prompt_ms": 776.965, + "prompt_per_token_ms": 40.89289473684211, + "prompt_per_second": 24.454125990231223, + "predicted_n": 2, + "predicted_ms": 297.386, + "predicted_per_token_ms": 297.386, + "predicted_per_second": 3.3626330762039904, + "draft_n": 4, + "draft_n_accepted": 4 + }, + "wall_seconds": 1.0767567759612575 + }, + "decode": [ + { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "Ein Reverse Proxy ist ein zentrales Element in modernen IT-Infrastrukturen, das als Vermittler zwischen Clienten und Backend-Servern agiert. Im Gegensatz zu einem Forward Proxy, der Anfragen von internen Clients nach außen vertritt, positioniert sich der Reverse Proxy hinter der Firewall oder vor dem eigentlichen Webserver. Wenn ein Client eine Anfrage an eine Domain wie `api.example.com` sendet, trifft diese zunächst beim Reverse Proxy ein. Dieser leitet die Anfrage anschließend an den tatsächlichen Backend-Server (z. B. eine Anwendung in einem Docker-Container oder Kubernetes-Cluster) weiter. Die Antwort des Backends wird vom Proxy zurück zum Client gesendet, wobei die Client-Daten unverändert bleiben, der Backend-Server jedoch für den Client unsichtbar bleibt.\n\nDie Funktionsweise eines Reverse Proxys beruht stark auf dem Lastverteilungskonzept (Load Balancing) und dem Session-Management. Der Proxy pflegt eine Konfigurationsdatei oder ein API-schnittstelle, in der definiert ist, welche IP-Adresse oder welcher Port für bestimmte URI-Pfade zuständig ist. In statischen Umgebungen ist diese Zuordnung fix. Das Problem entsteht jedoch in dynamischen Umgebungen, insbesondere wenn man Container-Technologien wie Docker Compose oder Kubernetes einsetzt. Container sind effizient, aber volatil: Sie können neu gestartet, skalierend (Scale-up/Scale-down) oder verschiebbar sein. Jede dieser Operationen kann dazu führen, dass die interne IP-Adresse eines Containers wechselt. Wenn der Reverse Proxy noch auf die alte IP-Adresse zeigt, während der Container bereits auf der neuen IP erreichbar ist, kommt es zu einer Verbindungsunterbrechung.\n\nDie typischsten Fehler, die bei einem IP-Wechsel auftreten, sind `502 Bad Gateway` oder `504 Gateway Timeout`. Ein `502` bedeutet, dass der Reverse Proxy den Backend-Server zwar gefunden hat (DNS-Auflösung erfolgreich), aber die Verbindung zu dieser spezifischen IP/Port-Kombination fehlgeschlagen ist – oft weil dort einfach kein Prozess mehr lauscht oder die IP bereits freigegeben wurde. Ein `504` tritt auf, wenn der Proxy den Server zwar erreicht, aber das Backend nicht rechtzeitig auf die Anfrage reagiert, was bei kurzzeitigen Netzwerkebenen-Problemen oder beim Neustart des Containers passieren kann.\n\nUm solche Probleme anhand von Logs einzugrenzen, muss man systematisch vorgehen. Zuerst sollte man die Logs des Reverse Proxys selbst untersuchen. Bei Nginx oder Traefik findet man in den Access-Logs häufig den HTTP-Statuscode und in den Error-Logs detaillierte Informationen. Ein typischer Eintrag im Error-Log eines Nginx-Reverse-Proxys könnte so aussehen: `connect() failed (111: Connection refused) while connecting to upstream`. Dies ist ein klares Indiz dafür, dass der Proxy auf eine IP-Adresse zugreifen wollte, auf der aber kein Dienst lauscht. Die in der Logzeile angegebene IP-Adresse ist der entscheidende Hinweis.\n\nIm zweiten Schritt muss man diese IP-Adresse mit dem aktuellen Zustand des Container-Orchestriers abgleichen. Bei Docker kann man `docker inspect | grep -A 6 \"NetworkSettings\"` ausführen, um die aktuelle IP zu ermitteln. Wenn die IP im Nginx-Error-Log eine andere ist als die aktuelle IP des Containers, ist die Ursache bestätigt: Der Proxy referenziert eine veraltete IP. Bei Kubernetes ist die Sache komplexer, da dort oft Services mit Stable IPs verwendet werden. In diesem" + } + } + ], + "created": 1789936730, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 768, + "prompt_tokens": 59, + "total_tokens": 827, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-sxqpPIFaJLoewzQQrdBhC2w8tKKJRyxj", + "timings": { + "cache_n": 0, + "prompt_n": 59, + "prompt_ms": 121.677, + "prompt_per_token_ms": 2.062322033898305, + "prompt_per_second": 484.8903243834085, + "predicted_n": 768, + "predicted_ms": 14934.509, + "predicted_per_token_ms": 19.47132855280313, + "predicted_per_second": 51.357563881075706, + "draft_n": 1392, + "draft_n_accepted": 417 + }, + "wall_seconds": 15.119467873009853 + }, + { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "```python\nimport asyncio\nfrom typing import Any, Awaitable, Callable, List, Optional, Union\n\n\nasync def first_success(\n awaitables: List[Awaitable],\n return_exceptions: bool = True\n) -> Any:\n \"\"\"\n Start all awaitables concurrently, return the first successful result,\n cancel and await remaining tasks, and collect exceptions if all fail.\n\n Args:\n awaitables: A list of awaitables (coroutines, tasks, or other awaitables).\n return_exceptions: If True, raise an ExceptionGroup containing all exceptions\n when all awaitables fail. If False, raise the first exception.\n\n Returns:\n The result of the first successfully completed awaitable.\n\n Raises:\n ExceptionGroup: If all awaitables fail and return_exceptions is True.\n Exception: If all awaitables fail and return_exceptions is False (first exception).\n \"\"\"\n if not awaitables:\n raise ValueError(\"At least one awaitable is required\")\n\n # Convert awaitables to tasks so we can manage them\n tasks = []\n for i, awaitable in enumerate(awaitables):\n task = asyncio.ensure_future(awaitable)\n tasks.append(task)\n\n # Use asyncio.as_completed to get results in completion order\n # We need to handle both successful completions and exceptions\n results = []\n exceptions = []\n\n try:\n for coro in asyncio.as_completed(tasks, return_exceptions=True):\n try:\n result = await coro\n # If we got here without an exception being raised, it's a success\n # Cancel all other pending tasks\n for task in tasks:\n if not task.done():\n task.cancel()\n # Await cancellation of remaining tasks\n await asyncio.gather(*tasks, return_exceptions=True)\n return result\n except BaseException as e:\n # Check if this is an actual exception from an awaitable\n # asyncio.as_completed with return_exceptions=True wraps exceptions\n # but they are still raised as-is in some versions\n exceptions.append(e)\n continue\n except BaseException as e:\n # In case there's some other unexpected issue\n # Cancel remaining tasks\n for task in tasks:\n if not task.done():\n task.cancel()\n await asyncio.gather(*tasks, return_exceptions=True)\n raise e\n\n # All awaitables failed\n # Cancel any remaining tasks (shouldn't be any, but just in case)\n for task in tasks:\n if not task.done():\n task.cancel()\n await asyncio.gather(*tasks, return_exceptions=True)\n\n # Clean up exception list - only keep actual exceptions\n actual_exceptions = []\n for exc in exceptions:\n if isinstance(exc, BaseException):\n actual_exceptions.append(exc)\n elif isinstance(exc, Exception):\n actual_exceptions.append(exc)\n\n if not actual_exceptions:\n # This shouldn't happen, but handle edge case\n raise RuntimeError(\"No results and no exceptions collected\")\n\n if return_exceptions:\n # Raise an ExceptionGroup with all the exceptions\n if len(actual_exceptions) == 1:\n # For a single exception, raise it directly or wrap in a group\n # Using ExceptionGroup for consistency\n raise ExceptionGroup(\"All awaitables failed\", actual_exceptions)\n else:\n raise ExceptionGroup(\"All awaitables failed\", actual_exceptions)\n else:\n # Raise the first exception\n raise actual_exceptions[0]\n\n\n# More robust implementation using a different approach" + } + } + ], + "created": 1789936740, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 768, + "prompt_tokens": 54, + "total_tokens": 822, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-2NRWiSfXbyELu6VOfTgwtbnCFhfZBEdx", + "timings": { + "cache_n": 0, + "prompt_n": 54, + "prompt_ms": 117.801, + "prompt_per_token_ms": 2.1815, + "prompt_per_second": 458.4001833600733, + "predicted_n": 768, + "predicted_ms": 9724.536, + "predicted_per_token_ms": 12.678664928292047, + "predicted_per_second": 78.87265778027867, + "draft_n": 903, + "draft_n_accepted": 539 + }, + "wall_seconds": 9.906276727968361 + } + ], + "prefill": [ + { + "target": 24576, + "response": { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "```json\n{\n \"alpha\": \"RAVEN-417\",\n \"beta\": \"CEDAR-928\",\n \"gamma\": \"ORBIT-563\"\n}\n```\n\nDie Verifizierung der oben genannten Datensätze erfordert einen streng systematischen und faktenbasierten Ansatz, der darauf abzielt, die Integrität der Daten zu wahren, ohne auf unbegründete Annahmen oder erfundene Beweise zurückzugreifen. Der erste Schritt besteht darin, die spezifischen „Needle“-Werte – in diesem Fall `RAVEN-417`, `CEDAR-928` und `ORBIT-563` – als Suchkriterien zu definieren. Es ist crucial, diese exakten Zeichenketten zu verwenden, da selbst geringfügige Abweichungen in der Schreibweise oder Interpunktion zu Fehlerschlägen oder dem Übersehen relevanter Einträge führen können. Da die bereitgestellten Datensätze von 000000 bis 001169 reichen und eine hohe Redundanz aufweisen, wo jeder Eintrag den Text „cobalt lantern maple orbit quartz river silver tango“ enthält, ist eine manuelle Überprüfung ineffizient und fehleranfällig.\n\nStattdessen sollte ein automatisierter Prozess eingesetzt werden, der jede einzelne Zeile des Datensatzes durchsucht, um auf das Vorhandensein dieser spezifischen Marker zu prüfen. In der Praxis bedeutet dies die Verwendung von Skripten oder Suchwerkzeugen, die in der Lage sind, große Datenvolumen schnell zu durchforsten. Wichtig ist, dass bei der Suche zwischen Groß- und Kleinschreibung unterschieden wird, da die Needles in Großbuchstaben mit Bindestrichen formatiert sind (`NEEDLE_ALPHA=RAVEN-417`). Die Verifizierung muss bestätigen, ob diese Marker tatsächlich in den jeweiligen Datensatzeinträgen vorhanden sind oder ob sie nur als Referenzen innerhalb des Kontextes genannt werden.\n\nDarüber hinaus ist es entscheidend, die Integrität der umgebenden Daten zu prüfen. Da die meisten Datensätze identisch sind, könnte die Anwesenheit der Needles darauf hindeuten, dass bestimmte Einträge als „gespickt“ oder gekennzeichnete Datenpunkte innerhalb eines Testdatensatzes dienen. Um sicherzustellen, dass keine Beweise erfunden werden" + } + } + ], + "created": 1789936764, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 512, + "prompt_tokens": 24674, + "total_tokens": 25186, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-d9poeTCrw0F8VkoJCX4xtjHliBsUyutc", + "timings": { + "cache_n": 0, + "prompt_n": 24674, + "prompt_ms": 13603.499, + "prompt_per_token_ms": 0.5513292939936776, + "prompt_per_second": 1813.798052986221, + "predicted_n": 512, + "predicted_ms": 10153.0, + "predicted_per_token_ms": 19.868884540117417, + "predicted_per_second": 50.32995173840244, + "draft_n": 813, + "draft_n_accepted": 307 + }, + "wall_seconds": 23.843644770036917, + "recall": { + "RAVEN-417": true, + "CEDAR-928": true, + "ORBIT-563": true + } + } + } + ], + "finished": 1789936764.098406 +} diff --git a/experiments/medium-split-mtp-20260920/85-mtp4/server-args.json b/experiments/medium-split-mtp-20260920/85-mtp4/server-args.json new file mode 100644 index 0000000..34daee1 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/85-mtp4/server-args.json @@ -0,0 +1,75 @@ +[ + "--model", + "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "--mmproj", + "/models/qwen/mmproj-BF16.gguf", + "--mmproj-offload", + "--mmproj-device", + "CUDA1", + "--alias", + "qwen-medium", + "--ctx-size", + "160000", + "--flash-attn", + "on", + "--cache-type-k", + "q4_0", + "--cache-type-v", + "q4_0", + "--cache-prompt", + "--cache-ram", + "32768", + "--threads", + "6", + "--threads-batch", + "6", + "--batch-size", + "2048", + "--ubatch-size", + "256", + "--parallel", + "2", + "--kv-unified", + "--jinja", + "--reasoning", + "auto", + "--reasoning-preserve", + "--host", + "127.0.0.1", + "--port", + "5005", + "--metrics", + "--fit", + "off", + "--n-gpu-layers", + "all", + "--load-mode", + "none", + "--no-ui", + "--temperature", + "1.0", + "--top-p", + "0.95", + "--top-k", + "20", + "--device", + "CUDA0,CUDA1", + "--main-gpu", + "0", + "--split-mode", + "layer", + "--tensor-split", + "85,15", + "--spec-type", + "draft-mtp", + "--spec-draft-n-max", + "4", + "--spec-draft-type-k", + "f16", + "--spec-draft-type-v", + "f16", + "--spec-draft-p-min", + "0.05", + "--verbosity", + "3" +] diff --git a/experiments/medium-split-mtp-20260920/85-mtp4/server.log.gz b/experiments/medium-split-mtp-20260920/85-mtp4/server.log.gz new file mode 100644 index 0000000..95316e1 Binary files /dev/null and b/experiments/medium-split-mtp-20260920/85-mtp4/server.log.gz differ diff --git a/experiments/medium-split-mtp-20260920/control-85-mtp3/config.json b/experiments/medium-split-mtp-20260920/control-85-mtp3/config.json new file mode 100644 index 0000000..86b9418 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/control-85-mtp3/config.json @@ -0,0 +1,10 @@ +{ + "label": "control-85-mtp3", + "ubatch": 256, + "split": "85,15", + "mtp": 3, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 +} diff --git a/experiments/medium-split-mtp-20260920/control-85-mtp3/gpu.json b/experiments/medium-split-mtp-20260920/control-85-mtp3/gpu.json new file mode 100644 index 0000000..c38ab7a --- /dev/null +++ b/experiments/medium-split-mtp-20260920/control-85-mtp3/gpu.json @@ -0,0 +1,669 @@ +[ + { + "time": 1789936597.676435, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "6964", + "total": "12288", + "temp": "43", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "11272", + "total": "16303", + "temp": "50", + "util": "14", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936599.706274, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "6964", + "total": "12288", + "temp": "43", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "11272", + "total": "16303", + "temp": "49", + "util": "14", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936601.736262, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "6964", + "total": "12288", + "temp": "43", + "util": "94", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "11272", + "total": "16303", + "temp": "49", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936603.763663, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "8442", + "total": "12288", + "temp": "43", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15814", + "total": "16303", + "temp": "49", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936605.7941704, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10384", + "total": "12288", + "temp": "43", + "util": "43", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15842", + "total": "16303", + "temp": "49", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936607.8243515, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10640", + "total": "12288", + "temp": "43", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15842", + "total": "16303", + "temp": "49", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936609.854398, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "49", + "util": "52", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "55", + "util": "48", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936611.8838413, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "50", + "util": "57", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "55", + "util": "52", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936613.917109, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "50", + "util": "50", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "60", + "util": "46", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936615.9517796, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "52", + "util": "47", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "60", + "util": "40", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936617.9818773, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "51", + "util": "55", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "59", + "util": "47", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936620.0144792, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "51", + "util": "56", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "58", + "util": "53", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936622.0455701, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "52", + "util": "49", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "61", + "util": "47", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936624.0782108, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "52", + "util": "49", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "58", + "util": "50", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936626.1088305, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "52", + "util": "45", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "60", + "util": "40", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936628.1410763, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "53", + "util": "54", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "58", + "util": "41", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936630.1718125, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "54", + "util": "47", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "60", + "util": "44", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936632.203748, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10644", + "total": "12288", + "temp": "53", + "util": "51", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15858", + "total": "16303", + "temp": "58", + "util": "38", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936634.2333362, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "51", + "util": "54", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "70", + "util": "95", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936636.2640924, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "54", + "util": "47", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "71", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936638.294401, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "52", + "util": "44", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "72", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936640.3246953, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "53", + "util": "48", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "72", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936642.3555238, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "53", + "util": "58", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "66", + "util": "73", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936644.3847382, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "54", + "util": "57", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "74", + "util": "98", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936646.419878, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "54", + "util": "57", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "64", + "util": "45", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936648.4510555, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "54", + "util": "54", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "65", + "util": "45", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936650.4859831, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "54", + "util": "46", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "61", + "util": "50", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936652.518545, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "53", + "util": "53", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "61", + "util": "41", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + }, + { + "time": 1789936654.5496242, + "gpus": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10646", + "total": "12288", + "temp": "54", + "util": "56", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15860", + "total": "16303", + "temp": "63", + "util": "51", + "pcie_gen": "4", + "pcie_width": "16" + } + ] + } +] diff --git a/experiments/medium-split-mtp-20260920/control-85-mtp3/loaded.json b/experiments/medium-split-mtp-20260920/control-85-mtp3/loaded.json new file mode 100644 index 0000000..02fb9c4 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/control-85-mtp3/loaded.json @@ -0,0 +1,138 @@ +{ + "case": { + "label": "control-85-mtp3", + "ubatch": 256, + "split": "85,15", + "mtp": 3, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + "started": 1789936595.6476705, + "idle_gpu": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10640", + "total": "12288", + "temp": "43", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15842", + "total": "16303", + "temp": "49", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ], + "props": { + "default_generation_settings": { + "params": { + "seed": 4294967295, + "temperature": 1.0, + "dynatemp_range": 0.0, + "dynatemp_exponent": 1.0, + "top_k": 20, + "top_p": 0.949999988079071, + "min_p": 0.05000000074505806, + "top_n_sigma": -1.0, + "xtc_probability": 0.0, + "xtc_threshold": 0.10000000149011612, + "typical_p": 1.0, + "repeat_last_n": 64, + "repeat_penalty": 1.0, + "presence_penalty": 0.0, + "frequency_penalty": 0.0, + "dry_multiplier": 0.0, + "dry_base": 1.75, + "dry_allowed_length": 2, + "dry_penalty_last_n": 64, + "mirostat": 0, + "mirostat_tau": 5.0, + "mirostat_eta": 0.10000000149011612, + "adaptive_target": -1.0, + "adaptive_decay": 0.8999999761581421, + "max_tokens": -1, + "n_predict": -1, + "n_keep": 0, + "n_discard": 0, + "ignore_eos": false, + "stream": false, + "n_probs": 0, + "min_keep": 0, + "chat_format": "Content-only", + "reasoning_format": "none", + "reasoning_in_content": false, + "generation_prompt": "", + "samplers": [ + "penalties", + "dry", + "top_n_sigma", + "top_k", + "typ_p", + "top_p", + "min_p", + "xtc", + "temperature" + ], + "speculative.types": "none", + "timings_per_token": false, + "post_sampling_probs": false, + "backend_sampling": false, + "lora": [] + }, + "n_ctx": 160000 + }, + "total_slots": 2, + "model_alias": "qwen-medium", + "model_ftype": "IQ4_XS - 4.25 bpw", + "model_path": "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "modalities": { + "vision": true, + "video": true, + "audio": false + }, + "media_marker": "<__media_CclCoD7CO3rlBd673CRwhz9GjHIXhKmN__>", + "endpoint_slots": true, + "endpoint_props": false, + "endpoint_metrics": true, + "ui": false, + "ui_settings": {}, + "chat_template": "{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- macro render_content(content, do_vision_count, is_system_content=false) %}\n {%- if content is string %}\n {{- content }}\n {%- elif content is iterable and content is not mapping %}\n {%- for item in content %}\n {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain images.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set image_count.value = image_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Picture ' ~ image_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|image_pad|><|vision_end|>' }}\n {%- elif 'video' in item or item.type == 'video' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain videos.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set video_count.value = video_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Video ' ~ video_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|video_pad|><|vision_end|>' }}\n {%- elif 'text' in item %}\n {{- item.text }}\n {%- else %}\n {{- raise_exception('Unexpected item type in content.') }}\n {%- endif %}\n {%- endfor %}\n {%- elif content is none or content is undefined %}\n {{- '' }}\n {%- else %}\n {{- raise_exception('Unexpected content type.') }}\n {%- endif %}\n{%- endmacro %}\n{%- if not messages %}\n {{- raise_exception('No messages provided.') }}\n{%- endif %}\n{%- set sysns = namespace(count=0, text='') %}\n{%- for message in messages %}\n {%- if sysns.count == loop.index0 and (message.role == 'system' or message.role == 'developer') %}\n {%- set sys_content = render_content(message.content, false, true)|trim %}\n {%- if sys_content %}\n {%- set sysns.text = sysns.text + ('\\n' if sysns.text else '') + sys_content %}\n {%- endif %}\n {%- set sysns.count = sysns.count + 1 %}\n {%- endif %}\n{%- endfor %}\n{%- set num_sys = sysns.count %}\n{%- set merged_system = sysns.text %}\n{%- set reasoning_instructions = '' %}\n{%- if enable_thinking is undefined or enable_thinking is true %}\n {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}\n {%- if resolved_reasoning_effort == 'high' %}\n {%- set resolved_reasoning_effort = 'xhigh' %}\n {%- endif %}\n {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}\n {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}\n {%- endif %}\n {%- if resolved_reasoning_effort == 'xhigh' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}\n {%- elif resolved_reasoning_effort == 'low' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}\n {%- endif %}\n{%- endif %}\n{%- if tools and tools is iterable and tools is not mapping %}\n {{- '<|im_start|>system\\n' }}\n {%- if reasoning_instructions %}\n {{- reasoning_instructions + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou have access to the following functions:\\n\\n\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n\" }}\n {{- '\\n\\nIf you choose to call a function ONLY reply in the following format with NO suffix:\\n\\n\\n\\n\\nvalue_1\\n\\n\\nThis is the value for the second parameter\\nthat can span\\nmultiple lines\\n\\n\\n\\n\\n\\nReminder:\\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\\n- Required parameters MUST be specified\\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\\n' }}\n {%- if merged_system %}\n {{- '\\n\\n' + merged_system }}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n{%- else %}\n {%- if merged_system %}\n {{- '<|im_start|>system\\n' + (reasoning_instructions + '\\n\\n' if reasoning_instructions else '') + merged_system + '<|im_end|>\\n' }}\n {%- elif reasoning_instructions %}\n {{- '<|im_start|>system\\n' + reasoning_instructions + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" %}\n {%- set content = render_content(message.content, false)|trim %}\n {%- if not(content.startswith('') and content.endswith('')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if loop.index0 >= num_sys %}\n {%- set content = render_content(message.content, true)|trim %}\n {%- if message.role == \"system\" or message.role == \"developer\" %}\n {{- raise_exception('System message must be at the beginning.') }}\n {%- elif message.role == \"user\" %}\n {{- '<|im_start|>' + message.role + '\\n' + content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is string %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- endif %}\n {%- set reasoning_content = reasoning_content|trim %}\n {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}\n {{- '<|im_start|>' + message.role + '\\n\\n' + reasoning_content + '\\n\\n\\n' + content }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {%- if tool_call.name is not defined or tool_call.name is none %}\n {{- raise_exception('Tool call is missing a function name.') }}\n {%- endif %}\n {%- if loop.first %}\n {%- if content|trim %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n\\n' }}\n {%- endif %}\n {%- else %}\n {{- '\\n\\n\\n' }}\n {%- endif %}\n {%- if tool_call.arguments is mapping %}\n {%- for args_name, args_value in tool_call.arguments|items %}\n {{- '\\n' }}\n {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}\n {{- args_value }}\n {{- '\\n\\n' }}\n {%- endfor %}\n {%- elif tool_call.arguments is string %}\n {%- if tool_call.arguments|trim %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" were passed as a JSON string. Parse them into an object before calling apply_chat_template.') }}\n {%- endif %}\n {%- elif tool_call.arguments is defined and tool_call.arguments is not none %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" must be an object/mapping or a JSON string.') }}\n {%- endif %}\n {{- '\\n' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.previtem and loop.previtem.role != \"tool\" %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n\\n' }}\n {{- content }}\n {{- '\\n' }}\n {%- if not loop.last and loop.nextitem.role != \"tool\" %}\n {{- '<|im_end|>\\n' }}\n {%- elif loop.last %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- else %}\n {{- raise_exception('Unexpected message role.') }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n' }}\n {%- endif %}\n{%- endif %}\n{#- Unsloth fixes - developer role, merged system messages, tool calling #}", + "chat_template_caps": { + "supports_object_arguments": true, + "supports_parallel_tool_calls": true, + "supports_preserve_reasoning": true, + "supports_reasoning_effort": true, + "supports_string_content": true, + "supports_system_role": true, + "supports_tool_calls": true, + "supports_tools": true, + "supports_typed_content": true + }, + "bos_token": "<|endoftext|>", + "eos_token": "<|im_end|>", + "build_info": "b10964-b29c606e2", + "is_sleeping": false, + "cors_proxy_enabled": false + }, + "slots": [ + { + "id": 0, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + }, + { + "id": 1, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + } + ] +} diff --git a/experiments/medium-split-mtp-20260920/control-85-mtp3/result.json b/experiments/medium-split-mtp-20260920/control-85-mtp3/result.json new file mode 100644 index 0000000..7af9675 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/control-85-mtp3/result.json @@ -0,0 +1,307 @@ +{ + "case": { + "label": "control-85-mtp3", + "ubatch": 256, + "split": "85,15", + "mtp": 3, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + "started": 1789936595.6476705, + "idle_gpu": [ + { + "name": "NVIDIA GeForce RTX 3060", + "used": "10640", + "total": "12288", + "temp": "43", + "util": "0", + "pcie_gen": "3", + "pcie_width": "4" + }, + { + "name": "NVIDIA GeForce RTX 5080", + "used": "15842", + "total": "16303", + "temp": "49", + "util": "0", + "pcie_gen": "4", + "pcie_width": "16" + } + ], + "props": { + "default_generation_settings": { + "params": { + "seed": 4294967295, + "temperature": 1.0, + "dynatemp_range": 0.0, + "dynatemp_exponent": 1.0, + "top_k": 20, + "top_p": 0.949999988079071, + "min_p": 0.05000000074505806, + "top_n_sigma": -1.0, + "xtc_probability": 0.0, + "xtc_threshold": 0.10000000149011612, + "typical_p": 1.0, + "repeat_last_n": 64, + "repeat_penalty": 1.0, + "presence_penalty": 0.0, + "frequency_penalty": 0.0, + "dry_multiplier": 0.0, + "dry_base": 1.75, + "dry_allowed_length": 2, + "dry_penalty_last_n": 64, + "mirostat": 0, + "mirostat_tau": 5.0, + "mirostat_eta": 0.10000000149011612, + "adaptive_target": -1.0, + "adaptive_decay": 0.8999999761581421, + "max_tokens": -1, + "n_predict": -1, + "n_keep": 0, + "n_discard": 0, + "ignore_eos": false, + "stream": false, + "n_probs": 0, + "min_keep": 0, + "chat_format": "Content-only", + "reasoning_format": "none", + "reasoning_in_content": false, + "generation_prompt": "", + "samplers": [ + "penalties", + "dry", + "top_n_sigma", + "top_k", + "typ_p", + "top_p", + "min_p", + "xtc", + "temperature" + ], + "speculative.types": "none", + "timings_per_token": false, + "post_sampling_probs": false, + "backend_sampling": false, + "lora": [] + }, + "n_ctx": 160000 + }, + "total_slots": 2, + "model_alias": "qwen-medium", + "model_ftype": "IQ4_XS - 4.25 bpw", + "model_path": "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "modalities": { + "vision": true, + "video": true, + "audio": false + }, + "media_marker": "<__media_CclCoD7CO3rlBd673CRwhz9GjHIXhKmN__>", + "endpoint_slots": true, + "endpoint_props": false, + "endpoint_metrics": true, + "ui": false, + "ui_settings": {}, + "chat_template": "{%- set image_count = namespace(value=0) %}\n{%- set video_count = namespace(value=0) %}\n{%- macro render_content(content, do_vision_count, is_system_content=false) %}\n {%- if content is string %}\n {{- content }}\n {%- elif content is iterable and content is not mapping %}\n {%- for item in content %}\n {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain images.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set image_count.value = image_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Picture ' ~ image_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|image_pad|><|vision_end|>' }}\n {%- elif 'video' in item or item.type == 'video' %}\n {%- if is_system_content %}\n {{- raise_exception('System message cannot contain videos.') }}\n {%- endif %}\n {%- if do_vision_count %}\n {%- set video_count.value = video_count.value + 1 %}\n {%- endif %}\n {%- if add_vision_id %}\n {{- 'Video ' ~ video_count.value ~ ': ' }}\n {%- endif %}\n {{- '<|vision_start|><|video_pad|><|vision_end|>' }}\n {%- elif 'text' in item %}\n {{- item.text }}\n {%- else %}\n {{- raise_exception('Unexpected item type in content.') }}\n {%- endif %}\n {%- endfor %}\n {%- elif content is none or content is undefined %}\n {{- '' }}\n {%- else %}\n {{- raise_exception('Unexpected content type.') }}\n {%- endif %}\n{%- endmacro %}\n{%- if not messages %}\n {{- raise_exception('No messages provided.') }}\n{%- endif %}\n{%- set sysns = namespace(count=0, text='') %}\n{%- for message in messages %}\n {%- if sysns.count == loop.index0 and (message.role == 'system' or message.role == 'developer') %}\n {%- set sys_content = render_content(message.content, false, true)|trim %}\n {%- if sys_content %}\n {%- set sysns.text = sysns.text + ('\\n' if sysns.text else '') + sys_content %}\n {%- endif %}\n {%- set sysns.count = sysns.count + 1 %}\n {%- endif %}\n{%- endfor %}\n{%- set num_sys = sysns.count %}\n{%- set merged_system = sysns.text %}\n{%- set reasoning_instructions = '' %}\n{%- if enable_thinking is undefined or enable_thinking is true %}\n {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}\n {%- if resolved_reasoning_effort == 'high' %}\n {%- set resolved_reasoning_effort = 'xhigh' %}\n {%- endif %}\n {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}\n {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}\n {%- endif %}\n {%- if resolved_reasoning_effort == 'xhigh' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}\n {%- elif resolved_reasoning_effort == 'low' %}\n {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}\n {%- endif %}\n{%- endif %}\n{%- if tools and tools is iterable and tools is not mapping %}\n {{- '<|im_start|>system\\n' }}\n {%- if reasoning_instructions %}\n {{- reasoning_instructions + '\\n\\n' }}\n {%- endif %}\n {{- \"# Tools\\n\\nYou have access to the following functions:\\n\\n\" }}\n {%- for tool in tools %}\n {{- \"\\n\" }}\n {{- tool | tojson }}\n {%- endfor %}\n {{- \"\\n\" }}\n {{- '\\n\\nIf you choose to call a function ONLY reply in the following format with NO suffix:\\n\\n\\n\\n\\nvalue_1\\n\\n\\nThis is the value for the second parameter\\nthat can span\\nmultiple lines\\n\\n\\n\\n\\n\\nReminder:\\n- Function calls MUST follow the specified format: an inner block must be nested within XML tags\\n- Required parameters MUST be specified\\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\\n' }}\n {%- if merged_system %}\n {{- '\\n\\n' + merged_system }}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n{%- else %}\n {%- if merged_system %}\n {{- '<|im_start|>system\\n' + (reasoning_instructions + '\\n\\n' if reasoning_instructions else '') + merged_system + '<|im_end|>\\n' }}\n {%- elif reasoning_instructions %}\n {{- '<|im_start|>system\\n' + reasoning_instructions + '<|im_end|>\\n' }}\n {%- endif %}\n{%- endif %}\n{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}\n{%- for message in messages[::-1] %}\n {%- set index = (messages|length - 1) - loop.index0 %}\n {%- if ns.multi_step_tool and message.role == \"user\" %}\n {%- set content = render_content(message.content, false)|trim %}\n {%- if not(content.startswith('') and content.endswith('')) %}\n {%- set ns.multi_step_tool = false %}\n {%- set ns.last_query_index = index %}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- for message in messages %}\n {%- if loop.index0 >= num_sys %}\n {%- set content = render_content(message.content, true)|trim %}\n {%- if message.role == \"system\" or message.role == \"developer\" %}\n {{- raise_exception('System message must be at the beginning.') }}\n {%- elif message.role == \"user\" %}\n {{- '<|im_start|>' + message.role + '\\n' + content + '<|im_end|>' + '\\n' }}\n {%- elif message.role == \"assistant\" %}\n {%- set reasoning_content = '' %}\n {%- if message.reasoning_content is string %}\n {%- set reasoning_content = message.reasoning_content %}\n {%- endif %}\n {%- set reasoning_content = reasoning_content|trim %}\n {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}\n {{- '<|im_start|>' + message.role + '\\n\\n' + reasoning_content + '\\n\\n\\n' + content }}\n {%- else %}\n {{- '<|im_start|>' + message.role + '\\n' + content }}\n {%- endif %}\n {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}\n {%- for tool_call in message.tool_calls %}\n {%- if tool_call.function is defined %}\n {%- set tool_call = tool_call.function %}\n {%- endif %}\n {%- if tool_call.name is not defined or tool_call.name is none %}\n {{- raise_exception('Tool call is missing a function name.') }}\n {%- endif %}\n {%- if loop.first %}\n {%- if content|trim %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n\\n' }}\n {%- endif %}\n {%- else %}\n {{- '\\n\\n\\n' }}\n {%- endif %}\n {%- if tool_call.arguments is mapping %}\n {%- for args_name, args_value in tool_call.arguments|items %}\n {{- '\\n' }}\n {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}\n {{- args_value }}\n {{- '\\n\\n' }}\n {%- endfor %}\n {%- elif tool_call.arguments is string %}\n {%- if tool_call.arguments|trim %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" were passed as a JSON string. Parse them into an object before calling apply_chat_template.') }}\n {%- endif %}\n {%- elif tool_call.arguments is defined and tool_call.arguments is not none %}\n {{- raise_exception('Tool call arguments for function \"' + (tool_call.name | string) + '\" must be an object/mapping or a JSON string.') }}\n {%- endif %}\n {{- '\\n' }}\n {%- endfor %}\n {%- endif %}\n {{- '<|im_end|>\\n' }}\n {%- elif message.role == \"tool\" %}\n {%- if loop.previtem and loop.previtem.role != \"tool\" %}\n {{- '<|im_start|>user' }}\n {%- endif %}\n {{- '\\n\\n' }}\n {{- content }}\n {{- '\\n' }}\n {%- if not loop.last and loop.nextitem.role != \"tool\" %}\n {{- '<|im_end|>\\n' }}\n {%- elif loop.last %}\n {{- '<|im_end|>\\n' }}\n {%- endif %}\n {%- else %}\n {{- raise_exception('Unexpected message role.') }}\n {%- endif %}\n {%- endif %}\n{%- endfor %}\n{%- if add_generation_prompt %}\n {{- '<|im_start|>assistant\\n' }}\n {%- if enable_thinking is defined and enable_thinking is false %}\n {{- '\\n\\n\\n\\n' }}\n {%- else %}\n {{- '\\n' }}\n {%- endif %}\n{%- endif %}\n{#- Unsloth fixes - developer role, merged system messages, tool calling #}", + "chat_template_caps": { + "supports_object_arguments": true, + "supports_parallel_tool_calls": true, + "supports_preserve_reasoning": true, + "supports_reasoning_effort": true, + "supports_string_content": true, + "supports_system_role": true, + "supports_tool_calls": true, + "supports_tools": true, + "supports_typed_content": true + }, + "bos_token": "<|endoftext|>", + "eos_token": "<|im_end|>", + "build_info": "b10964-b29c606e2", + "is_sleeping": false, + "cors_proxy_enabled": false + }, + "slots": [ + { + "id": 0, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + }, + { + "id": 1, + "n_ctx": 160000, + "speculative": true, + "is_processing": false + } + ], + "smoke": { + "choices": [ + { + "finish_reason": "stop", + "index": 0, + "message": { + "role": "assistant", + "content": "OK" + } + } + ], + "created": 1789936608, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 2, + "prompt_tokens": 19, + "total_tokens": 21, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-M3uwqFa6U3Cy6PocFxJn9saK8NrAxAPr", + "timings": { + "cache_n": 0, + "prompt_n": 19, + "prompt_ms": 811.12, + "prompt_per_token_ms": 42.69052631578948, + "prompt_per_second": 23.42440082848407, + "predicted_n": 2, + "predicted_ms": 306.127, + "predicted_per_token_ms": 306.127, + "predicted_per_second": 3.2666181029442027, + "draft_n": 3, + "draft_n_accepted": 3 + }, + "wall_seconds": 1.1197090409696102 + }, + "decode": [ + { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "Ein Reverse Proxy ist ein zentrales Element in modernen IT-Infrastrukturen, das als Vermittler zwischen Clienten und Backend-Servern agiert. Im Gegensatz zu einem Forward Proxy, der Anfragen von internen Clients nach außen vertritt, stellt der Reverse Proxy eine Eindepunkt-Schnittstelle (Single Point of Entry) für externe Anfragen dar. Typische Vertreter sind nginx, HAProxy oder Apache HTTP Server. Die Grundprinzipie des Reverse Proxys besteht darin, eingehende HTTP- oder HTTPS-Anfragen entgegenzunehmen und diese basierend auf definierten Routing-Regeln an die geeigneten Backend-Server weiterzuleiten. Diese Architektur bietet erhebliche Vorteile wie Lastverteilung, SSL/TLS-Terminierung, Caching und eine vereinfachte Verwaltung mehrerer hinterlegter Dienste. Durch die Abstraktion der Backend-Infrastruktur wird sichergestellt, dass Clients die Details der internen Topologie, wie IP-Adressen oder Server-Namen, nicht kennen müssen.\n\nIn dynamischen Umgebungen, insbesondere bei der Nutzung von Container-Orchestrierungsplattformen wie Kubernetes oder Docker Compose, ist die Stabilität der IP-Adressen ein wesentlicher Faktor. Container sind per Design oft ephemer; sie können bei Neustarts, Skalierungsoperationen oder Fehlerbehebungen neu erstellt werden. Dies führt dazu, dass ihre zugewiesenen IP-Adressen ändern können. Wenn ein Reverse Proxy statisch konfiguriert ist und auf feste IP-Adressen oder Ports zeigt, die durch den Container-Wechsel ungültig werden, treten häufig Verbindungsfehler auf. Ein typisches Szenario ist, dass der Proxy weiterhin versucht, den Datenverkehr an eine alte, nicht mehr existierende IP-Adresse zu senden. In diesem Fall scheitert die Verbindung, da es keinen Host mit dieser Adresse im Netzwerksegment gibt.\n\nEin weiterer häufiger Fehlermechanismus entsteht durch das Timing von DNS- oder Service-Auflösungen. In Container-Netzwerken werden oft Dienstnamen anstelle fester IPs verwendet. Wenn sich ein Container neu registriert, kann es eine Zeitspanne geben, in der die DNS-Einträge noch nicht aktualisiert sind. Der Reverse Proxy könnte dann eine veraltete Adresse auflösen, was zu einem \"Connection Refused\"-Fehler oder einer Timeout-Situation führt. Zudem können Probleme auftreten, wenn der Backend-Container zwar erreichbar ist, aber der Dienst innerhalb des Containers noch nicht bereit ist, Anfragen anzunehmen (beispielsweise während des Boot-Prozesses). Der Proxy meldet dann oft einen \"Bad Gateway\" (HTTP 502), obwohl die Netzwerkverbindung technisch möglich wäre, aber die Applikation nicht antwortet.\n\nUm solche Fehler effizient zu eingrenzen, ist die Analyse der Logs sowohl des Proxys als auch der umgebenden Infrastruktur unerlässlich. Bei nginx beispielsweise beginnt die Fehlersuche meist im `error.log`. Hier sucht man nach Meldungen wie `connect() failed (111: Connection refused)`, was darauf hindeutet, dass der Proxy eine IP erreicht hat, aber kein Prozess auf dem angegebenen Port lauscht, oder dass die IP falsch ist. Ein weiterer Hinweis ist `recv() failed (104: Connection reset by peer)`, was oft auf Netzwerkprobleme oder einen Abbruch durch den Backend-Server hindeutet.\n\nEs ist jedoch wichtig zu beachten, dass ein Reverse Proxy nicht immer die IP-Adresse loggt, an die er sich wendet, sondern eher den Dienstnamen. Daher muss man die Logs mit den Events der Orchestrierungsplattform korrelieren. In Kubernetes etwa kann man" + } + } + ], + "created": 1789936622, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 768, + "prompt_tokens": 59, + "total_tokens": 827, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-vmyiY20xbf8VvEpwoqU8ZaRalfjB5agL", + "timings": { + "cache_n": 0, + "prompt_n": 59, + "prompt_ms": 119.305, + "prompt_per_token_ms": 2.0221186440677967, + "prompt_per_second": 494.53082435773854, + "predicted_n": 768, + "predicted_ms": 13364.148, + "predicted_per_token_ms": 17.423921773142112, + "predicted_per_second": 57.39236051561237, + "draft_n": 1057, + "draft_n_accepted": 414 + }, + "wall_seconds": 13.542396485980134 + }, + { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "```python\nimport asyncio\nfrom typing import Any, Awaitable, Callable, List, Optional, Union\n\n\nasync def first_success(\n awaitables: List[Awaitable],\n return_exceptions: bool = True\n) -> Any:\n \"\"\"\n Start all awaitables concurrently, return the first successful result,\n cancel and await remaining tasks, and collect exceptions if all fail.\n\n Args:\n awaitables: A list of awaitables to run concurrently.\n return_exceptions: If True, when all awaitables fail, raise an\n ExceptionGroup (Python 3.11+) or re-raise the first exception\n with others attached. If False, simply re-raise the first exception.\n\n Returns:\n The result of the first successfully completed awaitable.\n\n Raises:\n ExceptionGroup (or the first exception if return_exceptions=False):\n If all awaitables fail.\n \"\"\"\n if not awaitables:\n raise ValueError(\"awaitables list must not be empty\")\n\n # Create tasks from all awaitables\n tasks = [asyncio.ensure_future(av) for av in awaitables]\n\n try:\n # Use asyncio.wait to get the first completed task\n # We need to keep polling until we find a successful one\n pending = set(tasks)\n\n while pending:\n done, pending = await asyncio.wait(\n pending,\n return_when=asyncio.FIRST_COMPLETED\n )\n\n if not done:\n break\n\n for task in done:\n try:\n result = task.result()\n # Success: cancel all remaining pending tasks and return result\n for remaining in pending:\n remaining.cancel()\n # Await cancellation of remaining tasks\n if pending:\n await asyncio.wait(pending)\n return result\n except Exception:\n # This task failed; continue checking other done tasks\n # (if multiple completed simultaneously)\n pass\n # If we get here, all completed tasks in this batch failed\n # Continue looping to check if any other pending tasks complete\n # All tasks are done and none succeeded\n # Collect all exceptions\n exceptions = []\n for task in tasks:\n try:\n task.result()\n except Exception as e:\n exceptions.append(e)\n\n if exceptions:\n if return_exceptions:\n if hasattr(asyncio, 'ExceptionGroup'):\n # Python 3.11+\n raise ExceptionGroup(\"All awaitables failed\", exceptions)\n else:\n # Fallback for older Python: raise first exception\n # Attach other exceptions as context if possible\n raise exceptions[0] from exceptions[1:] if len(exceptions) > 1 else None\n else:\n raise exceptions[0]\n else:\n # No exceptions but no success either (shouldn't happen normally)\n raise RuntimeError(\"No successful result and no exceptions captured\")\n\n except (asyncio.CancelledError, KeyboardInterrupt):\n # Cancel all tasks on interruption\n for task in tasks:\n if not task.done():\n task.cancel()\n if tasks:\n await asyncio.gather(*tasks, return_exceptions=True)\n raise\n\n\n# Usage example\nif __name__ == \"__main__\":\n async def slow_success(delay: float, value: Any) -> Any:\n \"\"\"Simulate an async operation that succeeds after a delay.\"\"\"\n await asyncio.sleep(delay)\n return value\n\n async def failing_op(delay: float, msg: str) -> Any:\n " + } + } + ], + "created": 1789936632, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 768, + "prompt_tokens": 54, + "total_tokens": 822, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-KoYAO3KmvB8ZbjUzmpfY0gidKhQtmqXD", + "timings": { + "cache_n": 0, + "prompt_n": 54, + "prompt_ms": 117.614, + "prompt_per_token_ms": 2.178037037037037, + "prompt_per_second": 459.1290152532862, + "predicted_n": 768, + "predicted_ms": 10175.854, + "predicted_per_token_ms": 13.26708474576271, + "predicted_per_second": 75.37450910754026, + "draft_n": 803, + "draft_n_accepted": 499 + }, + "wall_seconds": 10.357252335990779 + } + ], + "prefill": [ + { + "target": 24576, + "response": { + "choices": [ + { + "finish_reason": "length", + "index": 0, + "message": { + "role": "assistant", + "content": "```json\n{\n \"alpha\": \"RAVEN-417\",\n \"beta\": \"CEDAR-928\",\n \"gamma\": \"ORBIT-563\"\n}\n```\n\nDie Verifizierung der identifizierten Nadelwerte (`RAVEN-417`, `CEDAR-928`, `ORBIT-563`) innerhalb des vorliegenden Datensatzes erfordert einen streng systematischen und evidenzbasierten Ansatz, der auf der präzisen Analyse der Textstruktur und der Positionierung der Markierungen beruht. Zunächst ist es entscheidend, die Integrität der Quelldaten zu wahren, indem man keine extrapolierten oder hypothetischen Datenpunkte einführt. Der vorliegende Datenset besteht überwiegend aus identischen, wiederholten Zeilen, die den String „cobalt lantern maple orbit quartz river silver tango.“ enthalten, wobei die Einträge von `Record 000000` bis `Record 001169` nummeriert sind. Die Nadeln wurden jedoch nicht als Teil des eigentlichen Datensatzinhalts (wie etwa in der Spalte der Werte) platziert, sondern als isolierte Metadaten-Zeilen zwischen den regulären Aufzeichnungen.\n\nUm `NEEDLE_ALPHA=RAVEN-417` zu verifizieren, muss der Fokus auf die Region um die Aufzeichnungen 000146 bis 000147 gelegt werden. Hier findet sich die Zeile, die mit „Record 000146: cob...“ beginnt und abgebrochen ist, gefolgt direkt von der isolierten Zeile „NEEDLE_ALPHA=RAVEN-417“. Die Verifikation erfolgt durch die Bestätigung, dass diese spezifische String-Kombination exakt an dieser Position im Textstrom erscheint und nicht in anderen Teilen des Dokuments vorkommt. Ein ähnlicher Prozess gilt für `NEEDLE_BETA=CEDAR-928`, der sich zwischen `Record 000584` und `Record 000585` befindet. Hier ist besondere Sorgfalt geboten, da die Zeile für Record 000585 aufgrund der Platzierung der Nadel ungewöhnlich formatiert oder fragmentiert erscheint („Recor... d 000585“). Die Verifikation verlangt die sorgfält" + } + } + ], + "created": 1789936654, + "model": "qwen-medium", + "system_fingerprint": "b10964-b29c606e2", + "object": "chat.completion", + "usage": { + "completion_tokens": 512, + "prompt_tokens": 24674, + "total_tokens": 25186, + "prompt_tokens_details": { + "cached_tokens": 0 + } + }, + "id": "chatcmpl-SdiBR3b0wLujVUJv7evtigcnXq0O6zVk", + "timings": { + "cache_n": 0, + "prompt_n": 24674, + "prompt_ms": 12936.629, + "prompt_per_token_ms": 0.5243020588473697, + "prompt_per_second": 1907.2974883951606, + "predicted_n": 512, + "predicted_ms": 8522.374, + "predicted_per_token_ms": 16.677835616438355, + "predicted_per_second": 59.95981870779199, + "draft_n": 579, + "draft_n_accepted": 317 + }, + "wall_seconds": 21.544975613011047, + "recall": { + "RAVEN-417": true, + "CEDAR-928": true, + "ORBIT-563": true + } + } + } + ], + "finished": 1789936654.5839925 +} diff --git a/experiments/medium-split-mtp-20260920/control-85-mtp3/server-args.json b/experiments/medium-split-mtp-20260920/control-85-mtp3/server-args.json new file mode 100644 index 0000000..ba504b3 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/control-85-mtp3/server-args.json @@ -0,0 +1,75 @@ +[ + "--model", + "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "--mmproj", + "/models/qwen/mmproj-BF16.gguf", + "--mmproj-offload", + "--mmproj-device", + "CUDA1", + "--alias", + "qwen-medium", + "--ctx-size", + "160000", + "--flash-attn", + "on", + "--cache-type-k", + "q4_0", + "--cache-type-v", + "q4_0", + "--cache-prompt", + "--cache-ram", + "32768", + "--threads", + "6", + "--threads-batch", + "6", + "--batch-size", + "2048", + "--ubatch-size", + "256", + "--parallel", + "2", + "--kv-unified", + "--jinja", + "--reasoning", + "auto", + "--reasoning-preserve", + "--host", + "127.0.0.1", + "--port", + "5005", + "--metrics", + "--fit", + "off", + "--n-gpu-layers", + "all", + "--load-mode", + "none", + "--no-ui", + "--temperature", + "1.0", + "--top-p", + "0.95", + "--top-k", + "20", + "--device", + "CUDA0,CUDA1", + "--main-gpu", + "0", + "--split-mode", + "layer", + "--tensor-split", + "85,15", + "--spec-type", + "draft-mtp", + "--spec-draft-n-max", + "3", + "--spec-draft-type-k", + "f16", + "--spec-draft-type-v", + "f16", + "--spec-draft-p-min", + "0.05", + "--verbosity", + "3" +] diff --git a/experiments/medium-split-mtp-20260920/control-85-mtp3/server.log.gz b/experiments/medium-split-mtp-20260920/control-85-mtp3/server.log.gz new file mode 100644 index 0000000..0ce6d8d Binary files /dev/null and b/experiments/medium-split-mtp-20260920/control-85-mtp3/server.log.gz differ diff --git a/experiments/medium-split-mtp-20260920/production.json b/experiments/medium-split-mtp-20260920/production.json new file mode 100644 index 0000000..2f2b1fb --- /dev/null +++ b/experiments/medium-split-mtp-20260920/production.json @@ -0,0 +1,76 @@ +{ + "image": "sha256:5e3c12c145b8045e5731b44b6b97033f24b327ae3d4a3fa85ecdd159cc844907", + "args": [ + "--model", + "/models/qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf", + "--mmproj", + "/models/qwen/mmproj-BF16.gguf", + "--mmproj-offload", + "--mmproj-device", + "CUDA1", + "--alias", + "qwen-medium", + "--ctx-size", + "160000", + "--flash-attn", + "on", + "--cache-type-k", + "q4_0", + "--cache-type-v", + "q4_0", + "--cache-prompt", + "--cache-ram", + "32768", + "--threads", + "6", + "--threads-batch", + "6", + "--batch-size", + "2048", + "--ubatch-size", + "512", + "--parallel", + "2", + "--kv-unified", + "--jinja", + "--reasoning", + "auto", + "--reasoning-preserve", + "--host", + "0.0.0.0", + "--port", + "8080", + "--metrics", + "--fit", + "off", + "--n-gpu-layers", + "all", + "--load-mode", + "none", + "--no-ui", + "--temperature", + "1.0", + "--top-p", + "0.95", + "--top-k", + "20", + "--device", + "CUDA0,CUDA1", + "--main-gpu", + "0", + "--split-mode", + "layer", + "--tensor-split", + "85,15", + "--spec-type", + "draft-mtp", + "--spec-draft-n-max", + "3", + "--spec-draft-type-k", + "f16", + "--spec-draft-type-v", + "f16", + "--spec-draft-p-min", + "0.05" + ] +} diff --git a/experiments/medium-split-mtp-20260920/quick.json b/experiments/medium-split-mtp-20260920/quick.json new file mode 100644 index 0000000..77964ee --- /dev/null +++ b/experiments/medium-split-mtp-20260920/quick.json @@ -0,0 +1,42 @@ +[ + { + "label": "control-85-mtp3", + "ubatch": 256, + "split": "85,15", + "mtp": 3, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + { + "label": "85-mtp2", + "ubatch": 256, + "split": "85,15", + "mtp": 2, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + { + "label": "85-mtp4", + "ubatch": 256, + "split": "85,15", + "mtp": 4, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + }, + { + "label": "83-mtp3", + "ubatch": 256, + "split": "83,17", + "mtp": 3, + "prompts": [ + 24576 + ], + "minimum_headroom_mib": 448 + } +] diff --git a/experiments/medium-split-mtp-20260920/restore-verification.json b/experiments/medium-split-mtp-20260920/restore-verification.json new file mode 100644 index 0000000..f4a2a22 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/restore-verification.json @@ -0,0 +1,61 @@ +{ + "checked_at": 1789936855.966608, + "uptime": "22:40:55 up 3 days, 11:36, 1 user, load average: 1.59, 1.19, 1.06", + "containers": { + "mike-ai-llama-medium": { + "running": true, + "health": "healthy", + "image": "sha256:5e3c12c145b8045e5731b44b6b97033f24b327ae3d4a3fa85ecdd159cc844907", + "started": "2026-09-20T20:40:20.923606083Z" + }, + "mike-ai-router": { + "running": true, + "health": "healthy", + "image": "sha256:0448758bec6968b29263bcac0f8b4682c6d3029d626f7c12334698c11ef096bb", + "started": "2026-09-20T20:40:31.403318094Z" + }, + "mike-ai-profile-controller": { + "running": true, + "health": "healthy", + "image": "sha256:a5f156d94c4921e671fafa524cf9c0fe91e1cbec0113c1f19c136204d63f243f", + "started": "2026-09-20T20:40:31.237404223Z" + }, + "mike-ai-wireguard-gateway": { + "running": true, + "health": "healthy", + "image": "sha256:0d24e93c85a1c420b52b17666ede5fbd4672d92ab8dc28ea4ac5ba144ebc41d0", + "started": "2026-09-17T09:05:07.164533904Z" + }, + "mike-ai-qwen3-tts": { + "running": true, + "health": "healthy", + "image": "sha256:b363a01d08b1bbecbfc3ca6f585368fae2cfdc591f9ecca6643738369f9a9d98", + "started": "2026-09-20T20:16:34.144444642Z" + } + }, + "router": { + "/health": { + "status": "ok", + "router": "alive" + }, + "/ready": { + "status": "ok", + "router": "alive", + "upstream": "ready" + } + }, + "smoke": { + "content": "OK", + "usage": { + "completion_tokens": 2, + "prompt_tokens": 19, + "total_tokens": 21, + "prompt_tokens_details": { + "cached_tokens": 0 + } + } + }, + "kernel_errors": [], + "gpu": "name, memory.used [MiB], memory.total [MiB], temperature.gpu\nNVIDIA GeForce RTX 3060, 10918 MiB, 12288 MiB, 46\nNVIDIA GeForce RTX 5080, 15714 MiB, 16303 MiB, 49", + "disk": "Filesystem Size Used Avail Use% Mounted on\n/dev/nvme0n1p2 868G 187G 637G 23% /\n/dev/nvme1n1p1 916G 822G 49G 95% /data" +} diff --git a/experiments/medium-split-mtp-20260920/run.py b/experiments/medium-split-mtp-20260920/run.py new file mode 100644 index 0000000..be138b8 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/run.py @@ -0,0 +1,171 @@ +#!/usr/bin/env python3 +"""Bounded, isolated Qwen quantization benchmark. Supervisor restores production.""" +import json, pathlib, subprocess, sys, time, urllib.request, threading, signal +ROOT = pathlib.Path('/data/benchmarks/medium-split-mtp-20260920') +NAME = 'mike-ai-split-mtp-test' +BASE = 'http://127.0.0.1:5005' +GPU0 = 'GPU-8ad38c6c-5a01-9d8e-1dfa-ed662ad78fbe' +GPU1 = 'GPU-4834d9d7-5b61-3004-1fb3-4ae49d482d4b' +IMAGE = 'sha256:5e3c12c145b8045e5731b44b6b97033f24b327ae3d4a3fa85ecdd159cc844907' +MODELS = {'mix':'qwen3.8-27b-iq4-mix/Qwen3.8-27B-IQ4-MIX.gguf', 'pure':'qwen3.8-27b-iq4-xs-pure/qwen3.8-27b-IQ4_XS-pure.gguf', 'byteshape':'byteshape-qwen38-gpu5/model.gguf'} + +def cmd(*args, check=True, timeout=90): + r = subprocess.run(args, capture_output=True, text=True, timeout=timeout) + if check and r.returncode: raise RuntimeError(str(args[:3])+': '+r.stderr[-2000:]) + return r.stdout + +def api(path, data=None, timeout=900): + req = urllib.request.Request(BASE+path, data=None if data is None else json.dumps(data).encode(), headers={'Content-Type':'application/json'}) + with urllib.request.urlopen(req, timeout=timeout) as r: return json.load(r) + +def save(path, data): + path.write_text(json.dumps(data, indent=2, ensure_ascii=False)+'\n') + +def gpu(): + rows = cmd('nvidia-smi','--query-gpu=name,memory.used,memory.total,temperature.gpu,utilization.gpu,pcie.link.gen.current,pcie.link.width.current','--format=csv,noheader,nounits',timeout=15) + return [dict(zip(['name','used','total','temp','util','pcie_gen','pcie_width'], [v.strip() for v in row.split(',')])) for row in rows.splitlines()] + +def health_check(): + rows=gpu() + if any(int(x['temp']) >= 85 for x in rows): raise RuntimeError('GPU temperature limit') + mem = dict((a.split(':')[0],int(a.split()[1])) for a in pathlib.Path('/proc/meminfo').read_text().splitlines()) + if mem['MemAvailable'] < 3*1024*1024: raise RuntimeError('Host RAM reserve below 3 GiB') + return rows + +def chat(prompt, max_tokens=512, effort='none', seed=42, tools=None): + health_check() + p={'model':'qwen-medium','messages':[{'role':'user','content':prompt}], 'max_tokens':max_tokens,'temperature':1.0,'top_p':0.95,'top_k':20,'min_p':0.0,'seed':seed,'reasoning_effort':effort,'cache_prompt':False} + if tools: p.update(tools=tools,tool_choice='auto') + start=time.monotonic(); r=api('/v1/chat/completions',p); r['wall_seconds']=time.monotonic()-start + health_check() + return r + +def prefill(n, seed): + import gzip + payload=json.loads(gzip.decompress(pathlib.Path('/opt/mike-ai/stack/benchmarks/athena-qwen38-reference-20260920/frozen-requests.json.gz').read_bytes()))['prefill-'+str(n)] + r=chat(payload['messages'][0]['content'],payload['max_tokens'],seed=payload['seed']) + content=r['choices'][0]['message'].get('content','') + r['recall']={x:x in content for x in ['RAVEN-417','CEDAR-928','ORBIT-563']} + return r + +def run_case(case): + label=case['label']; out=ROOT/label; out.mkdir(exist_ok=True) + if (out/'result.json').exists(): raise RuntimeError('Refusing to overwrite completed case '+label) + save(out/'config.json',case) + print('START',label,flush=True) + single=case.get('single',False) + production=json.loads((ROOT/'production.json').read_text()) + assert production['image']==IMAGE + args=list(production['args']) + for flag,value in [('--ubatch-size',str(case['ubatch'])),('--tensor-split',case['split']),('--spec-draft-n-max',str(case['mtp'])),('--host','127.0.0.1'),('--port','5005')]: + args[args.index(flag)+1]=value + args += ['--verbosity','3'] + save(out/'server-args.json',args) + cmd('docker','run','-d','--name',NAME,'--gpus','all','--network','host','--read-only','--tmpfs','/tmp:rw,nosuid,nodev,size=256m','--security-opt','no-new-privileges:true','--cap-drop','ALL','--pids-limit','512','--ulimit','core=0','--memory','26g','--memory-swap','26g','--shm-size','1g','--log-opt','max-size=32m','--log-opt','max-file=1','-e','NVIDIA_VISIBLE_DEVICES='+GPU0+','+GPU1,'-e','NVIDIA_DRIVER_CAPABILITIES=compute,utility','-v','/data/models:/models:ro',IMAGE,*args) + stop=threading.Event(); samples=[] + def monitor(): + while not stop.wait(2): + try: + rows=health_check() + samples.append({'time':time.time(),'gpus':rows}) + except RuntimeError as e: + samples.append({'error':str(e),'aborted':True}) + cmd('docker','stop','-t','10',NAME,check=False) + return + except Exception as e: samples.append({'error':str(e)}) + thread=threading.Thread(target=monitor,daemon=True); thread.start() + result={'case':case,'started':time.time()} + try: + for _ in range(150): + try: + if api('/health',timeout=3).get('status')=='ok': break + except Exception: pass + if cmd('docker','inspect',NAME,'--format','{{.State.Running}}').strip()!='true': raise RuntimeError('Test container exited during load') + time.sleep(2) + else: raise RuntimeError('Startup exceeded 300s') + result['idle_gpu']=health_check(); result['props']=api('/props'); result['slots']=api('/slots') + save(out/'loaded.json',result) + # Added after the 88:12 trial: model loading alone can succeed while + # the first real attention graph still needs more CUDA workspace. + minimum=case.get('minimum_headroom_mib',512) + used_devices=['5080'] if single else ['5080','3060'] + for g in result['idle_gpu']: + if any(device in g['name'] for device in used_devices): + free=int(g['total'])-int(g['used']) + if free=262144: break + previous=context + context=min(262144,context+growth) + final={**case,'ctx':context,'label':case['label']+'-validated-'+str(context),'load_only':False,'prompts':[49152,context-1024]} + return run_case(final) + +if __name__=='__main__': + for case in json.loads(pathlib.Path(sys.argv[1]).read_text()): + if case.get('capacity_search'): capacity_case(case) + else: run_case(case) diff --git a/experiments/medium-split-mtp-20260920/supervise.py b/experiments/medium-split-mtp-20260920/supervise.py new file mode 100644 index 0000000..8b60e8e --- /dev/null +++ b/experiments/medium-split-mtp-20260920/supervise.py @@ -0,0 +1,52 @@ +#!/usr/bin/env python3 +"""Stop only existing router/controller/model; always restore the same containers.""" +import json, pathlib, subprocess, sys, time, signal +ROOT=pathlib.Path('/data/benchmarks/medium-split-mtp-20260920') +NAMES=['mike-ai-router','mike-ai-profile-controller','mike-ai-llama-medium'] +def run(*args,check=True,timeout=90): + return subprocess.run(args,capture_output=True,text=True,check=check,timeout=timeout) +def stop_signal(*_): raise RuntimeError('Supervisor interrupted') +signal.signal(signal.SIGTERM,stop_signal); signal.signal(signal.SIGINT,stop_signal) +# Refuse if the known production state has changed, or if requests are active. +for name in NAMES: + assert run('docker','inspect',name,'--format','{{.State.Running}}').stdout.strip()=='true',name +for attempt in range(60): + slots=json.loads(run('docker','exec',NAMES[-1],'curl','-fsS','http://127.0.0.1:8080/slots').stdout) + if not any(s['is_processing'] for s in slots): break + if attempt==0: print('WAIT production request active; no interruption',flush=True) + time.sleep(3) +else: raise RuntimeError('Production remained busy for 180s; no services stopped') +assert not run('docker','ps','-q','--filter','name=^mike-ai-split-mtp-test$').stdout.strip(),'Existing experiment' +production=json.loads(run('docker','inspect',NAMES[-1]).stdout)[0] +(ROOT/'production.json').write_text(json.dumps({'image':production['Image'],'args':production['Args']},indent=2)+'\n') +child=None +try: + run('docker','stop','-t','30',*NAMES[:2]) + # Drain requests already handed to the model, before unloading it. + for _ in range(120): + slots=json.loads(run('docker','exec',NAMES[-1],'curl','-fsS','http://127.0.0.1:8080/slots').stdout) + if not any(s['is_processing'] for s in slots): break + time.sleep(2) + else: raise RuntimeError('Model did not drain') + run('docker','stop','-t','30',NAMES[-1]) + child=subprocess.Popen(['python3',str(ROOT/'run.py'),sys.argv[1]]) + code=child.wait(timeout=600) + if code: raise RuntimeError('Benchmark failed: '+str(code)) +finally: + if child is not None and child.poll() is None: + child.terminate() + try: child.wait(timeout=20) + except subprocess.TimeoutExpired: child.kill(); child.wait(timeout=10) + run('docker','rm','-f','mike-ai-split-mtp-test',check=False) + run('docker','start',NAMES[-1]) + for _ in range(150): + if run('docker','inspect',NAMES[-1],'--format','{{.State.Health.Status}}').stdout.strip()=='healthy': break + time.sleep(2) + else: raise RuntimeError('Restored Medium did not become healthy') + run('docker','start',NAMES[1],NAMES[0]) + for _ in range(60): + statuses=[run('docker','inspect',name,'--format','{{.State.Health.Status}}').stdout.strip() for name in NAMES[:2]] + if all(status=='healthy' for status in statuses): break + time.sleep(2) + else: raise RuntimeError('Restored router/controller did not become healthy') + print('RESTORED existing medium/controller/router; all healthy',flush=True) diff --git a/experiments/medium-split-mtp-20260920/verify_restore.py b/experiments/medium-split-mtp-20260920/verify_restore.py new file mode 100644 index 0000000..7cc2ae4 --- /dev/null +++ b/experiments/medium-split-mtp-20260920/verify_restore.py @@ -0,0 +1,29 @@ +#!/usr/bin/env python3 +"""Read-only restoration verification plus a two-token model smoke request.""" +import json,pathlib,re,subprocess,time +ROOT=pathlib.Path('/data/benchmarks/medium-split-mtp-20260920') +def run(*args):return subprocess.check_output(args,text=True,timeout=30) +report={'checked_at':time.time(),'uptime':run('uptime').strip(),'containers':{}} +for name in ['mike-ai-llama-medium','mike-ai-router','mike-ai-profile-controller','mike-ai-wireguard-gateway','mike-ai-qwen3-tts']: + d=json.loads(run('docker','inspect',name))[0] + report['containers'][name]={'running':d['State']['Running'],'health':d['State'].get('Health',{}).get('Status'),'image':d['Image'],'started':d['State']['StartedAt']} + assert d['State']['Running'],name + if name in ['mike-ai-llama-medium','mike-ai-router','mike-ai-profile-controller','mike-ai-wireguard-gateway']: + assert d['State'].get('Health',{}).get('Status')=='healthy',name +assert report['containers']['mike-ai-llama-medium']['image']=='sha256:5e3c12c145b8045e5731b44b6b97033f24b327ae3d4a3fa85ecdd159cc844907' +assert not run('docker','ps','-q','--filter','name=^mike-ai-split-mtp-test$').strip() +probe='import urllib.request,json; print(json.dumps({p:json.load(urllib.request.urlopen("http://127.0.0.1:8081"+p,timeout=20)) for p in ["/health","/ready"]}))' +report['router']=json.loads(run('docker','exec','mike-ai-router','python','-c',probe)) +payload={'model':'qwen-medium','messages':[{'role':'user','content':'Antworte ausschließlich mit OK.'}],'reasoning_effort':'none','max_tokens':8,'temperature':0} +r=json.loads(run('docker','exec','mike-ai-llama-medium','curl','-fsS','--max-time','20','-H','Content-Type: application/json','--data',json.dumps(payload),'http://127.0.0.1:8080/v1/chat/completions')) +report['smoke']={'content':r['choices'][0]['message'].get('content',''),'usage':r.get('usage')} +assert report['smoke']['content'].strip()=='OK',report['smoke'] +started=min(json.loads(p.read_text())['started'] for p in ROOT.glob('*/result.json')) +journal=run('journalctl','-k','--since','@'+str(int(started)-60),'--no-pager') +pattern=re.compile(r'NVRM.*Xid|oom-kill|Out of memory: Killed process|Kernel panic|GPU has fallen off',re.I) +report['kernel_errors']=[line for line in journal.splitlines() if pattern.search(line)] +report['gpu']=run('nvidia-smi','--query-gpu=name,memory.used,memory.total,temperature.gpu','--format=csv').strip() +report['disk']=run('df','-h','/','/data').strip() +(ROOT/'restore-verification.json').write_text(json.dumps(report,indent=2)+'\n') +print(json.dumps(report,indent=2)) +assert not report['kernel_errors'],'Kernel/GPU errors require review'