-
Stefy Lanza (nextime / spora ) authored
H3 and Krea 2 need diffusers >= 0.40 while the server runs 0.38. Upgrading in place is not free: measured on this box, SDXL comes out bit-identical but Z-Image drifts (mean |diff| 0.0146, max 0.32, against a same-version control that is bit-identical) — so the upgrade would silently change image output for a model that is working fine. Since diffusers is imported once per process, the split only has to be per-process. The H3 venv is now created with --system-site-packages and given nothing but diffusers>=0.40 and huggingface_hub>=1.23; torch, transformers and the rest are inherited from the server's venv. That is 75 MB instead of the several GB a standalone venv costs, and keeps one torch on the machine. CODERAI_H3_STANDALONE_VENV=1 still builds a fully independent venv from requirements-h3.txt, for the day the parent's torch is too old. Training follows the same boundary: tools/lora_train_worker.py runs one job in the overlay — references written out as PNGs, request flattened to JSON, progress returned as JSON lines and mirrored onto the usual /v1/loras/progress so polling is unchanged. _train_via_overlay() drives it. Routing is by capability, not hardcoded: _diffusers_has(arch) tries to load the classes, so wan and ltx2 train in-process (0.38 has them) while h3 and krea go out-of-process — and if the main venv is ever upgraded they move back in-process with no config change. Verified: the overlay resolves diffusers 0.40 with torch/transformers inherited and leaves the parent untouched; a probe job ran under it, reached _train_flow_dit and reported back through the job file. The trainers themselves are still unrun against real weights. Co-Authored-By:
Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0182DhR1eFNPQyGmDrpHbedY
3d53da50