1. 28 Jul, 2026 2 commits
    • Stefy Lanza (nextime / spora )'s avatar
      colibri: GLM-5.2 tool-call parsing + treat colibri as a normal VRAM-eviction citizen · e2081488
      Stefy Lanza (nextime / spora ) authored
      - parser: add GLMParser + parse_glm_tool_calls/strip_glm_tool_calls for GLM-5.2's
        <tool_call>name<arg_key>k</arg_key><arg_value>v</arg_value>...</tool_call> format
        (byte-compatible with colibri's own parser). Gated on the <arg_key> marker so a
        generic <tool_call>{json} from other families is never hijacked. Declared-type
        coercion keeps string args verbatim (no "12345"->int). Unclosed-box recovery for
        budget-truncated calls. Wired into family selection ('glm'/'colibri'), the
        model-agnostic ToolCallParser path, and both strip_tool_calls_from_content paths.
      - manager: colibri no longer seizes the whole GPU like ds4. It pins only a
        configurable expert tier (CUDA_EXPERT_GB) and streams the rest, so it coexists and
        is evicted like any other model. VRAM footprint estimated from cuda_expert_gb
        (+overhead) until measured.
      
      Verified: typed/untyped parse, strip, unclosed recovery, non-GLM gating, and the
      by-name dispatcher all correct.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01LoSpEthysqmseCc6Geizty
      e2081488
    • Stefy Lanza (nextime / spora )'s avatar
      colibri: integrate GLM-5.2 native C engine (driven directly, no colibri Python) · 50f54eb3
      Stefy Lanza (nextime / spora ) authored
      Add JustVugg/colibri as a managed engine, mirroring the ds4 integration but
      driving the pure-C `colibri` binary DIRECTLY over its stdin/stdout mux protocol
      (docs/serve_protocol.md) instead of proxying to a Python gateway — coderai
      reproduces openai_server.py's engine client and GLM-5.2 chat template itself.
      
      - config: ColibriConfig (disabled by default; model is a directory container,
        not a GGUF — routes by model_id/alias/name, no arch sniff)
      - codai/api/colibri_worker.py: clone+build (make colibri CUDA=1), MuxEngine
        protocol client (READY handshake, SUBMIT/DATA/DONE, KV-slot pool, CANCEL,
        stderr log pump), per-container engine registry
      - codai/backends/colibri.py: ColibriBackend + render_chat (byte-exact port of
        colibri's GLM-5.2 template; the engine tokenizes what we send)
      - front routing: `colibri` capability on nvidia/cuda/auto; router/assignment/
        app/engine_supervisor thread config.colibri
      - manager: colibri_should_handle, backend selection, /v1/models surfacing,
        exclusive-VRAM eviction (wants the whole GPU like ds4)
      - admin: config get/set + per-model overrides; settings.html card + models.html row
      - packaging: build.sh --colibri, OCI bundle (repo+binary, no 372GB model),
        entrypoint seed, CODERAI_COLIBRI_DIR, smoke test
      
      Ships OFF: no routing changes until colibri.enabled + colibri.model_path are set.
      Verified offline: config round-trip, routing predicate, render_chat byte-match
      vs colibri's own, MuxEngine end-to-end against a fake engine, CPU engine builds.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01LoSpEthysqmseCc6Geizty
      50f54eb3
  2. 26 Jul, 2026 3 commits
    • Stefy Lanza (nextime / spora )'s avatar
      feat: add DINOv2-SALAD VPR (8448-d) as an isolated-subprocess embedding model · 319fb44a
      Stefy Lanza (nextime / spora ) authored
      Second VPR option for REQUEST-vpr-embeddings.md, alongside the in-process
      EigenPlaces. SALAD is near-SOTA for place recognition but its deps
      (pytorch_lightning, pytorch_metric_learning) are heavy and version-sensitive, so
      it runs in a SEPARATE process with those deps quarantined in a
      --system-site-packages venv (shares the main torch, adds only the extras). The
      main engine venv never imports them — verified clean.
      
      - codai/api/vpr_salad_server.py: stdin image-path -> stdout JSON embedding server
        (same protocol as dinov2-embed), GPU-preferred with CPU fallback, announces its
        true width (8448) on ready.
      - embeddings.py: 'salad'/'dinov2-salad' -> 'vpr-server' backend that spawns the
        server via /cache/salad-venv/bin/python; reuses the dinov2cpp reader/cleanup.
      - Self-healing: _EmbeddingModel gains a respawn hook. If the manager kills the
        child under VRAM pressure but keeps the cached model, the embed path now
        respawns the server in place instead of failing forever with a stale handle
        (this was causing persistent "process died" 500s once salad was evicted).
      - packaging/setup-salad-venv.sh: idempotent creator for the isolated venv.
      
      Both VPR models pass the decisive ordering test on same-place/different-place
      image sets (zero overlap between same-place min and different-place max); SALAD's
      margin (0.516 vs 0.150) is wider than EigenPlaces' (0.479 vs 0.349). Verified:
      8448-d L2=1.0, deterministic, ~0.2s warm / ~8s cold, GPU, subprocess isolation
      intact. Config: models.json 'salad' + alias 'dinov2-salad', engine nvidia.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01LoSpEthysqmseCc6Geizty
      319fb44a
    • Stefy Lanza (nextime / spora )'s avatar
      feat: serve VPR (EigenPlaces) via /v1/embeddings as 'vpr'/'eigenplaces' · 4e8e4f04
      Stefy Lanza (nextime / spora ) authored
      Implements REQUEST-vpr-embeddings.md: a visual place recognition model that turns
      a photo into ONE L2-normalised descriptor trained so two images of the SAME place
      land close — the building-identity discrimination dinov2/gme/geoclip lack (they
      rate similar-looking houses as matches). HomeHunter uses it to match listing
      exteriors against Street View panoramas.
      
      Backed by EigenPlaces (gmberton, ResNet50, 2048-d) loaded via torch.hub. Chosen
      over SALAD/MixVPR for a clean dependency footprint: torch + torchvision only (no
      pytorch_lightning). ImageNet-normalised, 512x512 eval transform (matches training
      crops); output L2-normalised (idempotent — the net already ends in an L2 layer).
      On GPU (nvidia engine device); CPU only as fallback.
      
      The image data URI arrives in `input` (per the contract) or the `image` field.
      torch.hub cache pinned to a persistent TORCH_HOME (/cache/torchhub) with the repo
      + weights + trusted_list pre-seeded, so the engine loads offline and
      non-interactively (a cold torch.hub.load would otherwise hit an interactive trust
      prompt that EOF-crashes a server).
      
      Verified: 2048-d L2=1.0, deterministic, one vector/image; and the decisive
      ordering test — same-place pairs (min cos 0.479) rank strictly above every
      different-place pair (max 0.349), zero overlap. Config: models.json 'eigenplaces'
      + alias 'vpr', engine nvidia. SALAD (pytorch_lightning, subprocess-isolated) to
      follow as a second, higher-accuracy option.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01LoSpEthysqmseCc6Geizty
      4e8e4f04
    • Stefy Lanza (nextime / spora )'s avatar
      feat: serve GeoCLIP visual geolocation via /v1/embeddings (geoclip + geoclip-location) · fd3221c6
      Stefy Lanza (nextime / spora ) authored
      Implements the HomeHunter request (REQUEST-geoclip-embeddings.md): two model ids
      on the existing embeddings endpoint, both returning 512-d L2-normalised vectors in
      ONE shared space so a property photo can be scored directly against candidate GPS
      coordinates.
      
        geoclip           — image encoder (CLIP ViT-L/14 + projection MLP), on GPU.
                            Accepts the image data URI in EITHER `input` (per the request)
                            or the `image` field (like dinov2).
        geoclip-location  — GPS encoder (equal-earth + RFF MLP), pinned to CPU: it is a
                            tiny MLP, so CPU is fast, frees VRAM, and is bit-exact across
                            batch sizes (GPU reduction order otherwise perturbs a cached
                            coordinate by ~1e-7). Batches an array of "lat,lon" strings to
                            one vector per element, in input order.
      
      Both sides load from the SAME bundled GeoCLIP checkpoint, so the spaces coincide by
      construction — the mismatched-checkpoint failure the request warns about cannot
      happen. Loaded as separate _EmbeddingModel('geoclip', ...) instances so each is
      independently evictable.
      
      transformers>=5 compat: CLIPModel.get_image_features now returns a
      BaseModelOutputWithPooling (768-d projected features in pooler_output) instead of a
      bare tensor, which breaks GeoCLIP's internal mlp() call. The image path extracts the
      768-d tensor itself (tolerant of both the old tensor and the new object) before the
      projection MLP.
      
      geoclip (+geopy, geographiclib) added to requirements.txt; its heavy deps (torch,
      transformers, Pillow, pandas, numpy) are already pinned, so no ML-stack churn.
      
      Verified end-to-end: 512-d L2-normalised both sides; batching + index order; same
      request bit-exact; both ids in GET /v1/models; image·location scoring varies by image.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01LoSpEthysqmseCc6Geizty
      fd3221c6
  3. 24 Jul, 2026 17 commits
    • Stefy Lanza (nextime / spora )'s avatar
      oci: expose the NVIDIA GPU to Vulkan (NVIDIA_DRIVER_CAPABILITIES) for docker --nvidia · dbb0676d
      Stefy Lanza (nextime / spora ) authored
      The nvidia container runtime only injects the NVIDIA Vulkan ICD when the
      'graphics' driver capability is requested. --nvidia ran with the runtime default
      (compute,utility), so inside the container Vulkan could see only the AMD card
      (RADV) and llvmpipe — never the NVIDIA GPU. A Vulkan-only GGUF embedder
      (dinov2-embed, built with GGML_VULKAN) therefore could not run on the NVIDIA card
      at all; it was pinned to the AMD RX 580 or forced to CPU.
      
      Add `-e NVIDIA_DRIVER_CAPABILITIES=all` (honouring any caller-exported value) to
      the docker --nvidia args so the NVIDIA Vulkan ICD is injected and Vulkan
      enumerates the NVIDIA GPU. Enables running the GGUF dinov2 embedder on the 3090
      via Vulkan under the nvidia-gguf engine (device pinned via GGML_VK_VISIBLE_DEVICES
      / VK_ICD_FILENAMES; the build hardcodes ggml_backend_vk_init(0)). Takes effect on
      the next container (re)launch.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01LoSpEthysqmseCc6Geizty
      dbb0676d
    • Stefy Lanza (nextime / spora )'s avatar
      fix: Z-Image image-gen 500 — flash-attn-2 rejects attn_mask (shared-backend race) · 643fa155
      Stefy Lanza (nextime / spora ) authored
      POST /v1/images/generations for a masked image transformer (Z-Image) crashed
      with "`attn_mask` is not supported for flash-attn 2." → HTTP 500.
      
      The diffusers "active attention backend" is PROCESS-WIDE global mutable state
      (_AttentionBackendRegistry class attribute) shared by the image and video paths.
      The video path sets it to flash-attn-2. A transformer whose per-module backend
      is None reads that global at dispatch time; flash-attn-2 rejects attn_mask, so
      Z-Image (which uses caption masks) crashes. The previous fix reset the global to
      native before generating, but that reset is NOT atomic with the pipeline() call
      (it runs in a to_thread worker), so a concurrent/subsequent video generation
      races and flips the shared global back to flash before image attention dispatches.
      
      Fix: pin the image denoiser's per-module attention backend to native (SDPA,
      mask-supporting) via set_attention_backend("native"). The dispatcher passes
      processor._attention_backend explicitly, bypassing the shared global — so image
      attention is immune to the video path flipping it. Video is undisturbed (it sets
      its own explicit per-module 'flash'). Image pipelines were never intentionally on
      flash, so this is non-regressive. Global reset kept as a fallback.
      
      Verified: unsloth/Z-Image-Turbo-unsloth-bnb-4bit now returns 200 with an image;
      no "attn_mask is not supported" in the log.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01LoSpEthysqmseCc6Geizty
      643fa155
    • Stefy Lanza (nextime / spora )'s avatar
      fix: GME vision embeddings 500 on large images (KV context overflow) · e362e56c
      Stefy Lanza (nextime / spora ) authored
      POST /v1/embeddings to the GME-Qwen2-VL GGUF embedder returned HTTP 500 for
      large photos. The qwen2-vl vision tower turns a big image into >n_ctx image
      tokens (observed batches of 1012–2048), and mtmd's decode then fails —
      "decode: failed to find a memory slot for batch of size N" / "failed to eval
      chunk 1" — because the image tokens plus the chat-prompt frame exceed the
      model's KV context. Small images (few hundred tokens) were unaffected.
      
      mtmd's own image_max_tokens budget is NOT honoured by this llama.cpp build
      (2048-token batches slipped through), so add a hard, Python-side backstop:
      _resize_image_to_token_budget() downscales any oversize PIL image (aspect
      preserved, 28px patches) to at most n_ctx-128 tokens BEFORE it reaches mtmd,
      so a request can never overflow the context regardless of input size. The
      budget derives from the live n_ctx, so it self-adjusts to any model config.
      
      Pairs with raising the GME model's n_ctx (instance models.json) so a
      full-size image (~2048 tokens) fits comfortably; n_batch/n_ubatch follow
      n_ctx in the loader.
      
      Verified on the radeon Vulkan embedder: 336² / 896² / 1400² / 1680² images
      all return 200 (the 1400²/1680² are logged resizing to ~2400 tokens to fit
      the 2432 budget); previously the ≥1024-token cases 500'd.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01LoSpEthysqmseCc6Geizty
      e362e56c
    • Stefy Lanza (nextime / spora )'s avatar
      fix: front-driven thermal pause indicator no longer erased by engine poll · a7a09b13
      Stefy Lanza (nextime / spora ) authored
      On a global CPU cooldown the thermal supervisor already pauses ALL engines
      (want_pause = gpu_hot or cpu_hot, applied per-engine), and the logs confirm
      every engine gets a pause. But the Tasks page often showed only ONE engine
      as cooling while the others looked like they were still running.
      
      Cause: two threads write engine.cooling. The thermal loop sets it when it
      pauses an engine; the health poll loop overwrites it every tick with the
      engine's OWN self-reported cooldown (cooling=d.get("cooling")). An engine
      only self-reports cooling while sitting in its own wait_until_safe loop (it
      has an in-flight request). An engine the front paused while IDLE reports
      cooling=None, so the poll loop cleared the front's pause indicator — the UI
      then showed that engine as running mid-cooldown.
      
      Fix: while the front holds an engine paused (engine.therm_paused), pass the
      update_state sentinel (cooling=False, "don't touch") instead of the engine's
      self-report, so the front's indicator survives. Once the front resumes the
      engine, the engine's self-report flows through again as before.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01LoSpEthysqmseCc6Geizty
      a7a09b13
    • Stefy Lanza (nextime / spora )'s avatar
      fix: config-key shadow dropped mmproj on GGUF vision model reload · 1126bdbc
      Stefy Lanza (nextime / spora ) authored
      A GGUF vision model (e.g. Gemma-4-14B) served correct image descriptions
      on its FIRST load after a restart but hallucinated an identical answer for
      every image on every subsequent load — the image was silently flattened to
      a "[image_url content]" text placeholder and the model never saw pixels.
      
      Root cause was a config-key mismatch that self-polluted the in-memory
      config. On-demand loads arrive as a basename (Gemma-...gguf) while the real
      models.json entry is keyed by full path. record_vram_delta() resolved the
      write target via _config_for_model_key(), which — unlike _config_for_model()
      — did NOT fall back to basename/alias matching, so it returned {} and then
      persisted a NEW basename-keyed entry holding ONLY the measured_* fields (no
      mmproj, no n_ctx). On the next load _config_for_model()'s exact-match hit
      that stripped basename entry FIRST, before the basename loop that would have
      found the real full-path config, so mmproj was dropped, supports_vision went
      False, and the vision projector never loaded.
      
      Fix:
      - Add _resolve_config_key(): returns the actual self.config key for a model
        (exact -> alias -> basename), the single canonical key readers and writers
        must agree on.
      - Route _config_for_model() through it; give _config_for_model_key() the same
        basename/alias fallback so it can no longer return {} for a basename.
      - record_vram_delta()/_persist() now read and persist measured fields under
        the canonical key, merging into the real config instead of spawning a
        stripped shadow entry.
      
      Verified: after restart the 14B loads with "mmproj ... (vision enabled)" on
      every reload and three distinct test images produce three distinct, accurate
      descriptions; the measured-VRAM writeback now logs "(force_vram_update)"
      (real config resolved) instead of the old "(no used_vram_gb)" (empty config).
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01LoSpEthysqmseCc6Geizty
      1126bdbc
    • Stefy Lanza (nextime / spora )'s avatar
      front: rate limit now guarantees an idle gap AFTER each request + live-applies · 086c6543
      Stefy Lanza (nextime / spora ) authored
      Reworked engine_request_min_interval_ms from start-spacing to a proper
      post-completion gap: _rate_acquire holds a per-engine lock for the whole
      request, _rate_release frees it only `interval` ms AFTER completion (via
      loop.call_later, non-blocking) — so consecutive requests to the engine
      are ALWAYS separated by at least that idle GPU time regardless of request
      duration. Wired acquire/release into all 3 inference dispatch paths with
      release in every finally/early-return. _rate_acquire refreshes config on
      mtime change so a value saved in the web UI applies to the next request.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      086c6543
    • Stefy Lanza (nextime / spora )'s avatar
      admin: per-engine request rate limit configurable in the web UI · e1bd2277
      Stefy Lanza (nextime / spora ) authored
      Adds a "Rate limit (ms)" column to the Settings per-engine overrides
      table for engine_request_min_interval_ms (0 = no limit). Backend GET
      returns the map, POST saves it via the int-override sanitizer.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      e1bd2277
    • Stefy Lanza (nextime / spora )'s avatar
      front: configurable per-engine request rate throttle (min interval ms) · 19241dc1
      Stefy Lanza (nextime / spora ) authored
      New server.engine_request_min_interval_ms (engine name → ms, 0/unset =
      off). _rate_gate spaces inference dispatch STARTS to an engine by at
      least the interval, capping request rate and inserting idle time between
      GPU submissions — a stability lever for a marginal card (e.g. RX 580)
      that wedges under sustained back-to-back Vulkan compute. Wired into all
      three inference dispatch paths after the swap-gate; the request itself
      runs unthrottled, only the start cadence is gated.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      19241dc1
    • Stefy Lanza (nextime / spora )'s avatar
      config: per-engine env overrides (RADV driver tuning for Polaris hangs) · b6170681
      Stefy Lanza (nextime / spora ) authored
      New server.engine_env_overrides (engine name → {VAR: val}), merged into
      the engine's process env at spawn. Lets low-level driver knobs be set
      without hardcoding — set radeon → RADV_DEBUG=syncshaders to serialize
      RADV shader dispatch and test whether the Polaris compute-ring async
      race (ring comp_x.y.z timeout) is what wedges the RX 580 under sustained
      Vulkan compute.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      b6170681
    • Stefy Lanza (nextime / spora )'s avatar
      frontproxy: a wedged GPU can no longer hang the whole front · 09558635
      Stefy Lanza (nextime / spora ) authored
      At front startup _build_engines runs vulkaninfo to enumerate GPUs; on a
      wedged Polaris card that process blocks UNINTERRUPTIBLY (D-state), and
      subprocess.run(timeout=) can't kill a D-state child — so the front's
      uvicorn startup hung forever ("Waiting for application startup"),
      taking down the healthy NVIDIA engine too (502 everywhere).
      
      Add _call_with_timeout: runs the probe in a daemon thread and STOPS
      WAITING after a wall-clock timeout, leaving the un-killable call
      orphaned instead of blocking. Wrap vulkan_devices() (12s) so startup
      always completes, and _amd_stats() (6s) so a hung Radeon can't freeze
      the health/thermal poll threads that also monitor the healthy card.
      A wedged engine now degrades to "that engine down", not "everything down".
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      09558635
    • Stefy Lanza (nextime / spora )'s avatar
      admin: make per-engine GPU power lock configurable in the web UI · 49b66f0e
      Stefy Lanza (nextime / spora ) authored
      The per-engine concurrency-overrides table on the Settings page already
      exposes max_parallel_requests per engine; add a "GPU power lock" column
      (dpm level select) so dpm_force_performance_level_overrides is editable
      too. Backend: GET returns the map, POST saves it via a level-validating
      sanitizer. Takes effect on engine restart.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      49b66f0e
    • Stefy Lanza (nextime / spora )'s avatar
      packaging: rc.local helper to lock AMD GPU power state (eudev bind unsupported) · 91ee14d0
      Stefy Lanza (nextime / spora ) authored
      Devuan's udev rejects the 'bind' trigger action, so the udev-rule path
      can't reliably apply the DPM lock at boot. sysvinit's /etc/rc.local runs
      race-free after amdgpu binds — this root-owned helper, called from
      rc.local, sets power_dpm_force_performance_level=high on every AMD card
      by PCI vendor id.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      91ee14d0
    • Stefy Lanza (nextime / spora )'s avatar
      packaging: fix amdgpu DPM udev rule — fire on bind via RUN, not add ATTR · f1f33a10
      Stefy Lanza (nextime / spora ) authored
      An ATTR{}= assignment on the add event races the amdgpu probe (the
      power_dpm attribute doesn't exist yet), so it silently no-ops at boot
      (observed: level stayed 'auto'). Fire on the bind action — driver bound
      and sysfs attrs created — and write via RUN with %p devpath.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      f1f33a10
    • Stefy Lanza (nextime / spora )'s avatar
      packaging: udev rule to persist AMD GPU high power state across boots · 2b46b02d
      Stefy Lanza (nextime / spora ) authored
      Devuan (sysvinit/OpenRC, no systemd) — a udev rule is the init-agnostic
      way to make power_dpm_force_performance_level=high stick. Matches by
      DRIVER==amdgpu so it survives DRM card renumbering across GPU resets.
      Complements the in-container best-effort apply (unprivileged, can't
      write root sysfs) shipped in 0.1.53.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      2b46b02d
    • Stefy Lanza (nextime / spora )'s avatar
      config: per-engine AMD GPU power-state lock (Polaris compute-hang mitigation) · 20b55838
      Stefy Lanza (nextime / spora ) authored
      New server.dpm_force_performance_level_overrides (engine name → level,
      e.g. {"radeon":"high"}). At engine startup the engine writes the level
      to every amdgpu card's power_dpm_force_performance_level — locking a
      Polaris/GCN card to fixed top clocks avoids the DPM power-state
      transitions that hang these cards under sustained Vulkan compute
      (current default is 'auto', the hang-prone mode). Card is matched by PCI
      vendor id (0x1002), robust to DRM card renumbering across resets.
      Best-effort: an unprivileged container logs the exact host command on
      PermissionError.
      
      Pairs with two config-only stability levers applied for the radeon:
      n_ctx reduction (qwen3 1024→512, gme 1536→1024) to keep all three
      embedders in real VRAM with headroom (no GTT-over-PCIe spill during
      compute — another Polaris hang trigger), and max_parallel_requests
      override {"radeon":1} to serialize Vulkan submissions.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      20b55838
    • Stefy Lanza (nextime / spora )'s avatar
      vision: transcode AVIF/WebP images to PNG for the llama.cpp chat handler · 0c69e4bc
      Stefy Lanza (nextime / spora ) authored
      A vision-capable GGUF model (gemma-4-26B + mmproj) replied "none" to
      real-estate photos because the client sends AVIF, and llama.cpp's mtmd
      chat handler decodes data-URIs with stb_image, which has no AVIF/WebP
      support — the model silently received no image. _normalize_vision_content
      now transcodes any non-stb format to PNG via PIL (which has libavif) before
      handing it to the handler. The radeon GME embedder was unaffected (it
      already decodes via PIL). Text-only models (gemma-2-9b, no mmproj) are
      untouched — the image is still flattened, since they have no vision tower.
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      0c69e4bc
    • Stefy Lanza (nextime / spora )'s avatar
      thermal: filter impossible GPU temps at the source (gpu_detect + read chokepoint) · 9ff0f709
      Stefy Lanza (nextime / spora ) authored
      The 0.1.50 guard only covered _read_gpu_temp_uncached's FIRST probe path;
      its rocm/psutil fallbacks and gpu_eval() (which reads engine_gpu_stats
      directly) still saw the raw 511°C and cooling-waited a healthy card
      forever. Filter at the real source — the amdgpu sysfs temp1_input read
      in gpu_detect — so temp is None for any ≥150°C reading and NO downstream
      reader can act on it; plus a final chokepoint in read_gpu_temp().
      Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
      9ff0f709
  4. 23 Jul, 2026 9 commits
  5. 22 Jul, 2026 9 commits