Mikei386
|
b3e86cc7ae
|
Make reasoning levels enforce real token budgets
|
2026-09-01 13:55:11 +02:00 |
|
Mikei386
|
c25e57af57
|
Limit prompt cache reuse to text-only profiles
|
2026-08-25 19:26:48 +02:00 |
|
Mikei386
|
d921dfaf69
|
Enable cross-chat llama prompt cache reuse
|
2026-08-25 19:19:49 +02:00 |
|
Mikei386
|
335c9a8501
|
fix: prevent Qwen reasoning and tool loops
|
2026-08-24 20:06:26 +02:00 |
|
Mikei386
|
41b9c17f0f
|
Define four-profile production matrix with Medium default
|
2026-08-22 11:30:38 +02:00 |
|
Mikei386
|
f0d552ef58
|
Containerize MCP tool services
|
2026-08-20 23:54:02 +02:00 |
|
Mikei386
|
cb07779f5a
|
Replace vision hotswap with native multimodal profiles
|
2026-08-20 14:24:56 +02:00 |
|
Mikei386
|
8a45a0d851
|
Optimize Qwen fast profile for vision and 76K context
|
2026-08-20 14:11:25 +02:00 |
|
Mikei386
|
0e4a9de5ba
|
Expand router into reproducible local AI platform
|
2026-08-20 12:56:53 +02:00 |
|