-
Stefy Lanza (nextime / spora ) authored
Plan item H.2. Also fixes a real bug shipped in 0.2.35. THE BUG: load_quantized_dit was imported from longcat_video.modules.longcat_video_dit, where it does not exist — it is in longcat_video.modules.quantization (beside QuantizedLinear and quantize_model), and it returns the AVATAR transformer, consistent with INT8 being an avatar-1.5-only build. The wrong import would have raised at the exact moment someone asked for INT8. TRAINING. coderai does not implement this loop, and the reasons are the design: 1. LongCat's venv is standalone (Python 3.10 / torch 2.6 / transformers 4.41), so a trainer there cannot `from codai.api import loras` the way tools/lora_train_worker.py does — that overlay venv inherits the parent's site-packages, this one does not. 2. The upstream repo has NO way to create trainable LoRA layers. Its DiT exposes load_lora() / enable_loras() / disable_all_loras() and nothing that initialises one, and the adapters it loads (cfg_step_lora, refinement_lora, dmd_lora) use its own key layout. Reimplementing that layout unverified would produce adapters the pipeline cannot load: support in appearance only. SimpleTuner implements LongCat-Video properly — model_family "longcat_video", LoRA and quantised LoRA via int8-quanto / int4-quanto / fp8-torchao with no extra installs — so coderai generates its config and drives it: - lora_archs gains the arch with an `external` marker; load_classes refuses with that reason instead of an AttributeError that reads like a diffusers version problem. `target: video` still resolves Wan weights to Wan, so existing configs are untouched. - tools/longcat_train.py writes the config, runs SimpleTuner, parses its steps into the same JSON-lines protocol the overlay trainer uses, and reports the adapter it actually wrote rather than assuming one appeared. - Two VAE constraints are checked BEFORE a long run starts: (num_frames - 1) divisible by 4 — the same 4n+1 rule coderai already applies to Wan — and both sides divisible by 16. - A venv of its own, because SimpleTuner pins its own torch and would fight the inference venv's. train_auto_build is off by default, and the refusal carries the manual command. - QLoRA defaults ON (int8-quanto): a 13.6B transformer plus optimiser state does not fit a consumer card at bf16, so the quantised base is the normal path. `dataset_config` is REQUIRED and refused when absent or missing: SimpleTuner wants captioned video clips, and synthesising a dataset from the `images` field other targets use would train something other than a video LoRA, silently. 34 tests. No training run has been executed — the driver and its constraints are tested, the run against a real dataset is not. Co-Authored-By:
Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0182DhR1eFNPQyGmDrpHbedY
15d208a7