-
Stefy Lanza (nextime / spora ) authored
Teardown leak: _free_pipeline_vram now records the pipeline's CUDA storage pointers up front and, after the normal teardown+gc, walks gc.get_objects() and nulls any surviving tensors out of their list/dict referrers (the proven pass from the CUDA text backend's cleanup) — the "~22 GB untracked (teardown leak; referenced elsewhere)" VRAM is now reclaimed instead of poisoning every later load until a restart. Names the holders when it fires so future leak sources identify themselves. Tracking: a video pipeline only registers in manager.models AFTER its multi-minute load, so concurrent loads raced the free-VRAM check into mutual OOM. note_loading()/clear_loading() reserve the model's estimate for the whole load window; _get_free_vram_gb subtracts reservations and the orphan-VRAM check counts them (no more false "teardown leak" for an in-progress load). The embeddings OOM-retry waits out active reservations (bounded 300s) before evicting and retrying. Co-Authored-By:
Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
2ea24b0b
| Name |
Last commit
|
Last update |
|---|---|---|
| .. | ||
| admin | ||
| api | ||
| backends | ||
| broker | ||
| frontproxy | ||
| models | ||
| openai | ||
| pydantic | ||
| queue | ||
| tasks | ||
| __init__.py | ||
| cli.py | ||
| config.py | ||
| main.py | ||
| platform_paths.py | ||
| system_app.py |