• Stefy Lanza (nextime / spora )'s avatar
    longcat: LoRA / QLoRA training, driven through SimpleTuner (v0.2.37) · 15d208a7
    Stefy Lanza (nextime / spora ) authored
    Plan item H.2. Also fixes a real bug shipped in 0.2.35.
    
    THE BUG: load_quantized_dit was imported from longcat_video.modules.longcat_video_dit,
    where it does not exist — it is in longcat_video.modules.quantization (beside
    QuantizedLinear and quantize_model), and it returns the AVATAR transformer, consistent
    with INT8 being an avatar-1.5-only build. The wrong import would have raised at the exact
    moment someone asked for INT8.
    
    TRAINING. coderai does not implement this loop, and the reasons are the design:
    
    1. LongCat's venv is standalone (Python 3.10 / torch 2.6 / transformers 4.41), so a
       trainer there cannot `from codai.api import loras` the way tools/lora_train_worker.py
       does — that overlay venv inherits the parent's site-packages, this one does not.
    2. The upstream repo has NO way to create trainable LoRA layers. Its DiT exposes
       load_lora() / enable_loras() / disable_all_loras() and nothing that initialises one,
       and the adapters it loads (cfg_step_lora, refinement_lora, dmd_lora) use its own key
       layout. Reimplementing that layout unverified would produce adapters the pipeline
       cannot load: support in appearance only.
    
    SimpleTuner implements LongCat-Video properly — model_family "longcat_video", LoRA and
    quantised LoRA via int8-quanto / int4-quanto / fp8-torchao with no extra installs — so
    coderai generates its config and drives it:
    
    - lora_archs gains the arch with an `external` marker; load_classes refuses with that
      reason instead of an AttributeError that reads like a diffusers version problem.
      `target: video` still resolves Wan weights to Wan, so existing configs are untouched.
    - tools/longcat_train.py writes the config, runs SimpleTuner, parses its steps into the
      same JSON-lines protocol the overlay trainer uses, and reports the adapter it actually
      wrote rather than assuming one appeared.
    - Two VAE constraints are checked BEFORE a long run starts: (num_frames - 1) divisible by
      4 — the same 4n+1 rule coderai already applies to Wan — and both sides divisible by 16.
    - A venv of its own, because SimpleTuner pins its own torch and would fight the inference
      venv's. train_auto_build is off by default, and the refusal carries the manual command.
    - QLoRA defaults ON (int8-quanto): a 13.6B transformer plus optimiser state does not fit
      a consumer card at bf16, so the quantised base is the normal path.
    
    `dataset_config` is REQUIRED and refused when absent or missing: SimpleTuner wants
    captioned video clips, and synthesising a dataset from the `images` field other targets
    use would train something other than a video LoRA, silently.
    
    34 tests. No training run has been executed — the driver and its constraints are tested,
    the run against a real dataset is not.
    Co-Authored-By: 's avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    Claude-Session: https://claude.ai/code/session_0182DhR1eFNPQyGmDrpHbedY
    15d208a7