-
Stefy Lanza (nextime / spora ) authored
The tooltip matched loaded_info entries (raw engine keys) against canonicalized display names client-side and mostly missed — footprints now canonicalize server-side with the same mapping as the model list, keyed by canonical id, so every loaded model shows its VRAM (+RAM when offloading). llama-vl: bound mtmd image_max_tokens (config `image_max_tokens`, default 1024) — large photos otherwise expand to thousands of vision tokens whose transient compute buffer evicts co-resident models on a small card. Co-Authored-By:
Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
1920e6aa
| Name |
Last commit
|
Last update |
|---|---|---|
| .. | ||
| admin | ||
| api | ||
| backends | ||
| broker | ||
| frontproxy | ||
| models | ||
| openai | ||
| pydantic | ||
| queue | ||
| tasks | ||
| __init__.py | ||
| cli.py | ||
| config.py | ||
| main.py | ||
| platform_paths.py | ||
| system_app.py |