• Stefy Lanza (nextime / spora )'s avatar
    images: stop video's flash-attn backend leaking to image models (Z-Image attn_mask crash) · 0043eb2a
    Stefy Lanza (nextime / spora ) authored
    diffusers' Model.set_attention_backend() doesn't just set per-processor backends
    — it ALSO flips a process-wide active backend (attention_dispatch's
    _active_backend). The video path sets that to flash-attn for the Wan transformer;
    image and video share the nvidia-engine process, so the global stayed flash and
    leaked to the next image model. Z-Image's transformer sets no backend of its own
    (passes backend=None → uses the global) and its attention is masked, so it
    crashed with "`attn_mask` is not supported for flash-attn 2" → image/environment
    generation 400. reset_attention_backend() clears per-processor backends but NOT
    the global, so it didn't help.
    
    Fix: restore the diffusers global backend to the env default (native/SDPA)
    (a) before every image generation — bulletproof against a leaked flash backend —
    and (b) in the video pipeline teardown (_free_pipeline_vram), so it can't persist
    after a video pipe is freed. Masked image attention (SDPA) now always works; the
    video transformer keeps its own per-processor backend.
    Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
    Claude-Session: https://claude.ai/code/session_01RdMufYvtTbtGDWsiZVoXce
    0043eb2a
Name
Last commit
Last update
..
admin Loading commit data...
api Loading commit data...
backends Loading commit data...
broker Loading commit data...
frontproxy Loading commit data...
models Loading commit data...
openai Loading commit data...
pydantic Loading commit data...
queue Loading commit data...
tasks Loading commit data...
__init__.py Loading commit data...
cli.py Loading commit data...
config.py Loading commit data...
main.py Loading commit data...
platform_paths.py Loading commit data...
system_app.py Loading commit data...