• Stefy Lanza (nextime / spora )'s avatar
    longcat: quantise the text encoder, so nothing has to leave the card (v0.2.77) · 3b384924
    Stefy Lanza (nextime / spora ) authored
    The DiT has upstream's INT8 path. The UMT5-XXL text encoder does not, and at
    bf16 it is ~11 GB — 45% of why a 13.6 GB INT8 DiT still overflows a 24 GB card
    (13.6 + 11.2 + VAE = 25.1 GB before a single activation). Offload worked around
    it by making the DiT and the encoder take turns; quantising the encoder removes
    the reason to.
    
    Measured against the real checkpoint, what the card must hold:
      bf16 encoder, offloaded   13.9 GB
      bf16 encoder, resident    25.2 GB   <- the configuration that OOMs
      int8 encoder, resident    19.5 GB
      nf4  encoder, resident    16.7 GB
    
    So `text_encoder_quant` is a setting: none | int8 | nf4 | fp4, per model or
    server-wide, validated at startup rather than after a seven-minute load.
    bitsandbytes is pinned in requirements-longcat.txt and asserted by both the pod
    image build and the local image's smoke test.
    
    A quantised encoder is RESIDENT and never offloaded: bitsandbytes quantises as
    the weights land on CUDA and is not built to shuttle them back — and at 2.8 GB
    it does not need to. That is detected from the MODEL, not the config, so a
    pipeline loaded quantised cannot be offloaded by a config that changed since,
    and defensively, so a component that cannot be walked is simply not quantised.
    
    Inserting these helpers by anchor dropped one of them between
    @contextlib.contextmanager and the function it was meant to decorate, so
    _encoder_is_quantised returned a context manager — truthy, making every pipeline
    look quantised — and _bsa_for quietly stopped being one. Both now have a test
    asserting their shape, because the symptom was five unrelated failures.
    
    Unverified: bitsandbytes has not executed on this GPU. A throwaway container
    cannot reach the card here (torch reports no CUDA devices under both --gpus all
    and --runtime=nvidia, though the production container's own CUDA is fine), so
    the NF4 path is proven only by its unit tests until a service run exercises it.
    Co-Authored-By: 's avatarClaude Opus 5 (1M context) <noreply@anthropic.com>
    Claude-Session: https://claude.ai/code/session_0185zEBiy3KWB37xNGSrxL7w
    3b384924