1. 11 Oct, 2026 20 commits
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a pod is ready when the model is loaded, not when the port opens (v0.3.20) · 74aca6e1
      Stefy Lanza (nextime / spora ) authored
      /v1/models answers the moment a vLLM process starts — minutes before it can
      serve, because the checkpoint still has to be downloaded (every pod
      re-downloads it without a network volume) and the KV cache built on first use.
      The pool called that ready, every client read it as usable, and the whole model
      load landed inside somebody's first request, where it is indistinguishable from
      a hang.
      
      That is what cost today. One OCR page took 807s end to end — ~400s image pull,
      ~400s weights — while the client timed out at 180s and then 300s and concluded
      the serving was broken. It was not: nobody had ever waited long enough, and the
      only reason we know is a probe run with a 1400s timeout, which came back 200
      with text and conf 0.9019.
      
      So the pool now sends one token against the served model before calling the pod
      ready. The pod bills through the load either way; this only decides whether the
      wait is visible as a pod that is not ready yet, or hidden inside a request that
      looks stuck. `pods_ready` becomes a signal a client can act on, which is what
      every client already assumed it was.
      
      A failed warm-up is reported and the pod used anyway: it may still serve (an
      engine that takes no chat completions, a model whose warm-up shape we guessed
      wrong), and rejecting a pod we have already paid to boot over a diagnostic
      request would be worse than the hidden latency this removes. Off with
      warmup_on_boot=false or warmup_timeout_s=0.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      74aca6e1
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a pod's log has to outlive the pod (v0.3.19) · 9572ef56
      Stefy Lanza (nextime / spora ) authored
      A rented pod is the one machine in the system whose logs cannot be read. RunPod
      publishes no pod-log API — its REST spec lists 23 routes and none serve logs,
      which is why every attempt returns HTTP 400 — and boot.sh writes its phases to
      a file and then execs uvicorn, whose output goes to the container's stdout and
      nowhere else. /boot therefore covered the way up and nothing after it, and
      _dump_pod_logs only ever ran on a failure, printing into the orchestrator's log.
      
      That cost an evening. Three pods were investigated in one session and two were
      already destroyed by the time there was a question to ask — one reclaimed by
      RunPod mid-request, taking the only evidence with it — so "the serving hangs"
      stayed a hypothesis for an hour. It was not hanging: a page took 807s, most of
      it the weights download every pod repeats because global volumes are
      console-only, and every client had given up at 180-300s.
      
      So the pod keeps a bounded ring of its own stdout and stderr (codai/api/podlog)
      and serves it at /logs with a cursor, and the pool drains that on every scaler
      tick into <config_dir>/pod-logs/<pod_id>.log — kept AFTER the pod dies, which is
      the whole point. It also drains once more before tearing a pod down, because the
      interesting lines are the last ones and they are only readable while the pod
      exists.
      
      A pull, not a push: the orchestrator already reaches pods to serve requests, so
      the bearer token and pinned TLS apply unchanged and a pod with no outbound route
      still gets its log kept. Bounded at both ends, since neither disk is free, and a
      gap in the pod's ring is written into the file instead of passing silently. The
      stdout tee writes through to the real stream on every failure path — a logging
      aid that could break printing would be worse than none.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      9572ef56
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: a pod that vanished must not take the queue with it (v0.3.18) · b2f91079
      Stefy Lanza (nextime / spora ) authored
      RunPod reclaimed an A40 mid-document today. The pool went on reporting it
      healthy, so every page chose the same corpse: the in-flight count climbed, the
      pod was never dropped, and nothing drained. From the client that reads as
      "requests neither complete nor time out", and it is what wedged the test.
      
      RunpodBackend has always handled this — _unreachable() drops the pod and the
      request goes again. The fan-out gateway never did: it returned 502 and left the
      dead pod in the pool as healthy. So a lease can now report its upstream dead,
      which drops the pod, clears the cached candidate set and releases the slot; the
      page is retried ONCE against whatever replaces it, and a third attempt would
      just be a retry loop against a real outage.
      
      Only connect-level failures count as dead. A read timeout is a model taking its
      time over a dense page, and discarding a healthy pod for that would be an
      outage of its own.
      
      Also: the first version of that handler could not work. It used a metaclass
      with __instancecheck__ so `except _DEAD_UPSTREAM` would match httpx's errors —
      but `except` matches on the class hierarchy at C level and never consults
      __instancecheck__, so it caught nothing. A test caught it; the check is now an
      explicit isinstance() where it cannot lie.
      
      And the registry path again: 0.3.15 keyed the fallback off XDG_CONFIG_HOME,
      which the container launcher exports only into the processes IT starts. A
      diagnostic script entering the container another way printed a THIRD registry
      path and created a second pod TLS CA. The fallback now reads
      CODERAI_CONFIG_DIR, the variable the launcher actually sets, so every process
      agrees however it was started — and since pod_tls.tls_dir() derives from the
      registry path, the CA follows.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      b2f91079
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a release that cannot be honoured yet is remembered, and a pod that · a1e95f99
      Stefy Lanza (nextime / spora ) authored
      refuses our token is not an asset (v0.3.17)
      
      Both from watching a real OCR run.
      
      THE LOST INTENT. A client asked for the fleet back while a page was still being
      read. The pods correctly stayed — work is never dropped to save money — but the
      intent was then forgotten: the pool fell back to idle_timeout_s (900s there) and
      billed for a fleet nobody had asked to keep. It took a second teardown call,
      four minutes later, to actually release it. The pool now remembers: pods kept
      only because they are working are retired the moment they fall idle, the same
      "do not wait out the timeout" treatment a closed schedule window already gets.
      Asking for warm pods again cancels it.
      
      Also fixes the reply and the log, which said "floor 1" when the 1 came from work
      in flight. That reads as though the operator had asked for the pod — the opposite
      of what it meant. Demand and configuration are now reported separately
      ("needed" and "floor"), because they are two different reasons a pod survived.
      
      THE ADOPTED POD THAT 401s. After a restart the pool adopts the pod it left
      running instead of renting another, and carries that pod's bearer token across
      through the registry — when one was recorded. With none recorded it fell back to
      our own token and never checked the result, so a pod launched with a different
      one answered 401 to every request for the rest of its paid life. Seen today,
      right after a restart: "reusing the pod already running" followed by 401 on
      every OCR page. The health probe that found it proves nothing here — a coderai
      pod answers /healthz to anybody.
      
      Adoption now probes an AUTHENTICATED path with the token it would use, and a pod
      that refuses it is terminated rather than adopted: with one registry per
      deployment (0.3.15) an unrecorded key is a lost key, so that pod can never serve
      anyone, and paying for it is the only outcome worse than killing it. An
      unreachable pod is NOT treated as unauthorised — that stays with the health
      machinery, which has a budget for it.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      a1e95f99
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: the gateway has to trust a pod's certificate like everything else (v0.3.16) · 0dd0f9a6
      Stefy Lanza (nextime / spora ) authored
      First real document, first real pod: it came up at https://194.68.245.79:22198
      and every page came back
      
        502 — OCR upstream pod failed: [SSL: CERTIFICATE_VERIFY_FAILED]
              unable to get local issuer certificate
      
      A direct-TCP pod serves the certificate coderai's OWN CA signed, on a bare IP
      that certificate can never name. codai.api.pod_http has always known this: it
      pins to pod_tls.ensure_ca() and switches hostname checking off. The fan-out
      gateway is async and could not reuse that requests-based module, and I gave it
      a bare httpx client — so the one hop that now carries every OCR page was the
      one hop with no pod trust.
      
      It keeps TWO clients per event loop, chosen by pod_http.is_pinned_url: the
      pinned context trusts only our CA, so using it for everything would reject a
      cluster node whose certificate a public CA signed, and sharing one client would
      let whichever upstream connected first decide the policy for both.
      
      Worth recording: this was never fan-out-specific. With serve_fanout off, surya
      attaches straight to that same https:// IP with its own httpx and fails
      identically — `model` mode has never worked against a TLS pod. The gateway is
      what makes it fixable in one place.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      0dd0f9a6
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: one registry for every process, and a client may give pods back now (v0.3.15) · 7548c70e
      Stefy Lanza (nextime / spora ) authored
      Two things, both from today's live incident.
      
      THE SPLIT REGISTRY. _pod_registry_path() fell back to ~/.coderai whenever
      config_manager was not initialised in the calling process. The container's
      launcher puts the real config dir one level deeper — it exports
      XDG_CONFIG_HOME="$CODERAI_CONFIG_DIR" and uses "$CODERAI_CONFIG_DIR/coderai",
      which is also where platform_paths.user_config_dir() lands. So a process with
      the config manager wrote one registry and a process without it read another,
      and the pods each owned were orphans to the other. That is what made the
      documented mutual-kill loop recur: pod created 19:24:56, terminated as
      untracked 19:24:57. The fallback is now the same XDG-aware path the launcher
      computes, and the resolved path is logged once — a registry that splits
      silently is invisible until it costs a GPU.
      
      CLIENT-FORCED TEARDOWN. `pods: 0` released a client's lease but gave nothing
      back: the pods then waited out scale_down_after_s, or idle_timeout_s before
      that. The automatic path is slow on purpose — it infers from a momentary
      reading, and a cold pod costs minutes to replace — but a client that has just
      finished a batch is not inferring. `teardown: true` sheds the surplus at once,
      skipping both the window and the one-per-pass limit, and answers with what
      happened. It still cannot scale through the operator's min_pods, and a pod with
      requests on it drains rather than dying under them.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      7548c70e
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: the reaper's grace period has never protected a single pod (v0.3.14) · 5fa59441
      Stefy Lanza (nextime / spora ) authored
      Observed live on the Digesta orchestrator while testing OCR: pod created
      19:24:56, "terminated STALE pod … not tracked by any pool" 19:24:57. The pool
      then waited out its 900s boot timeout on a pod that no longer existed, every
      OCR request 503'd, and the next attempt rented another.
      
      This is the race the source already documents from an earlier incident — "the
      nvidia engine created a pod and the radeon engine terminated it four seconds
      later, then the reverse — a mutual kill loop that rented and destroyed pods
      until stopped" — and the defence against it was supposed to be the age guard,
      "never reap a pod younger than this, whatever the registry says".
      
      It was written as `age is not None and age < REAP_GRACE_SECONDS`, and
      list_pods() returned {id, name, status, cost_per_hr, image} with no timestamp
      of any kind. So age was None for every pod, the condition was never true, and
      the guard abstained on exactly the pods it existed to protect. A guard that
      fails open is worse than none: it reads as covered.
      
      Two changes. The guard now fails CLOSED — an undatable pod is left alone and
      says so — and the listing carries runtime.uptimeInSeconds so it can measure
      instead of abstaining. Not reaping costs one more cycle; reaping a live pod
      costs the work on it and bills in a loop that never serves a request.
      
      What made it recur here is worth recording: this install has TWO config dirs,
      so the cross-process registry that shields a new pod exists twice and each
      process reads the one the other does not write. That is the deeper fault and is
      not fixed by this commit.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      5fa59441
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: surya's local mode means a GPU here, and has to be asked for (v0.3.13) · a754f45d
      Stefy Lanza (nextime / spora ) authored
      Two faults, one cause: `local` was both the default and the only engine mode
      that would quietly settle for the CPU.
      
      Surya's own settings pick the device and fall back cuda -> mps -> cpu
      (surya/settings.py TORCH_DEVICE_MODEL), so `local` on a machine with no CUDA
      ran a VLM pipeline on the processor: minutes a page, a result that looks
      perfectly normal, and nothing anywhere saying why. paddle and docTR already
      refuse in exactly that situation — surya was the exception. It now refuses too,
      with a 503 that names both ways out, and the worker is TOLD its device rather
      than left to guess. Detection fails closed: torch is asked first (it is what
      surya will ask), nvidia-smi second, and if neither can answer the answer is no.
      
      And `surya_serve` no longer defaults to `local`. An install that said nothing
      got the one mode that runs the model on this box — on a GPU-less orchestrator
      whose whole job is renting one, that is the worst possible default. It now
      defaults to `model`, which adapts to wherever the checkpoint is served and asks
      for surya_model_id when it has not been told. Safe to change: both production
      installs set the mode explicitly (`model` on Aruba, `vllm` on zeiss), so
      nothing was relying on the old default.
      
      Running on the CPU is still fully supported — surya_use_gpu = false — it just
      has to be a decision now.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      a754f45d
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: optionally give the dearest pod back first (v0.3.12) · d321ad48
      Stefy Lanza (nextime / spora ) authored
      A pool rents at whatever the market offered when each pod was started, so one
      fleet can hold a $0.35 card and a $0.79 card doing identical work. Shedding the
      cheap one saves less than half as much, and which to prefer is the operator's
      call, not ours — hence a checkbox rather than a new default.
      
      With scale_down_costliest_first the dearest surplus pod goes first, tie-broken
      by the idlest and then the coldest cache. Cost decides WHICH pod goes, never
      whether one goes, and a dear pod with requests on it is drained rather than
      dropped — so choosing it costs time, not work. Worth weighing before switching
      on: an on-demand pod costs more than a spot pod, so "dearest first" prefers to
      keep the interruptible ones.
      
      Also fixes the default tie-break, which was inverted when scale-down was
      written: `-last_used` picked the pod used MOST recently, discarding the warmest
      prefix cache of the set. It now retires the coldest.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      d321ad48
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a fleet that grew for a burst has to give the pods back (v0.3.11) · de706ddb
      Stefy Lanza (nextime / spora ) authored
      Only half the scaler existed. The pool adds a pod once every pod it has carries
      scale_up_inflight_per_pod requests, and pods went away again only through the
      idle reaper — which fires on a pod with NO work and no request for
      idle_timeout_s.
      
      That misses the expensive case. A pool that grew for a burst and is now serving
      a trickle keeps every pod: each is touched often enough that last_used never
      ages out, each is a fraction busy, and the bill stays at the peak for as long as
      the trickle lasts. With max_pods raised to 64 that is the difference between a
      few dollars an hour and a few hundred.
      
      So the pool also asks how many pods the CURRENT concurrency justifies and
      retires the excess. The rules that matter are not "fewer pods":
      
        * work is never dropped to save money. The pod retired is the least loaded,
          and if it still has requests it is marked draining — no new work, and
          terminated only once it has finished what it had;
        * one pod per pass, after the surplus has persisted for scale_down_after_s,
          and the window restarts after each retirement. A dip of a few seconds must
          not cost a pod wanted again seconds later: a cold one takes minutes;
        * a quiet pool is NOT this routine's business. With nothing in flight,
          teardown stays entirely with idle_timeout_s — a number the operator chose,
          weighing the bill against a cold start — so this never silently shortens it.
      
      A draining pod is also excluded from the growth decision, which reads the LEAST
      loaded pod: an emptying drainer would otherwise answer "no need to grow" with
      the one pod that cannot take the work.
      
      Two existing tests had to change. A _Pod stub stood in for PodHandle without
      the new field, and a source-text assertion used a 600-character window that a
      new field plus its comment pushed data_center out of — now asserted against the
      whole class, which is what it was always about.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      de706ddb
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: one handshake per pod, not one per page · e1156bbd
      Stefy Lanza (nextime / spora ) authored
      The gateway built an httpx.AsyncClient for every request, so each page paid a
      fresh TCP and TLS handshake to the upstream. With surya_instances: 24 that is
      24 streams reconnecting per page, against a round trip that already dominates
      the cost of reading one.
      
      One pooled client with keep-alive instead, so a document's pages reuse the
      connection. Correctness-neutral; it only ever cost time.
      
      Keyed by the event loop, because an AsyncClient binds to the loop that created
      it and coderai runs more than one (the OCR pool keeps its own). The key is the
      loop OBJECT held weakly, not id(): keying by id would leak an entry per dead
      loop and, since CPython reuses addresses, could hand a new loop a client bound
      to the closed one that used to live at that address — a fault that would only
      appear under the load this change exists to handle.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      e1156bbd
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: the attach path must not be one keystroke from a public one · 23b49d60
      Stefy Lanza (nextime / spora ) authored
      The fan-out gateway lived at /v1/ocr/engine/<lease>/… and the new catalogue at
      /v1/ocr/engines. They do not collide, but both auth gates exempt the gateway by
      PREFIX — so any later route under /v1/ocr/engine* would have inherited that
      exemption silently, and the two names are indistinguishable at a glance in a
      middleware condition.
      
      The gateway is now /v1/ocr/attach/<lease>/…: internal by name, and nothing a
      public path family will grow into. Tests assert /v1/ocr/engines still requires
      ordinary authentication at both gates, and that a plausible future sibling
      (/v1/ocr/engine-stats) does too.
      
      Free to rename because nothing consumes the lease URL yet — it is minted per
      attachment at run time and never written down.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      23b49d60
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: let a client ask what this install can read pages with (v0.3.10) · 6970e4e1
      Stefy Lanza (nextime / spora ) authored
      Naming an engine an install had not enabled was a 400, and there was no way to
      find out beforehand — the client had to know, or guess and handle the refusal.
      
      GET /v1/ocr/engines answers it from this install's own config: every engine
      with `enabled` and, when it is not, a `reason` that names the setting to
      change (an unaccepted licence is reported as such rather than as merely
      disabled, because that is a different checkbox). Each row also carries `kind`
      — an engine owns its weights, a pipeline drives a model — the `aliases` that
      also resolve to it, and `returns`, so a caller that needs coordinates can see
      that olmOCR has none rather than discovering it in an empty field. The top
      level reports both what the operator wrote as the default and what it resolves
      to.
      
      Where the model runs is deliberately not in the answer. It is configuration,
      it can differ between two requests, and a client that could see it would start
      depending on it. A test asserts the payload mentions no pod, node or address.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      6970e4e1
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: no engine moves GPU work to the CPU on its own · 52dfaca3
      Stefy Lanza (nextime / spora ) authored
      docTR called .cuda() when torch happened to see a card and otherwise simply
      stayed on the CPU, so a GPU-less install read every page many times slower
      with nothing in the response to say why — the result looked normal, it just
      arrived minutes late. Nothing else in coderai quietly relocates GPU work, and
      paddle refuses outright in exactly this situation; docTR was the odd one out.
      
      `doctr_use_gpu` is now a requirement: no CUDA device means a 503 that names
      the setting. Running docTR on the CPU is still fully supported and unchanged —
      it just has to be asked for, with doctr_use_gpu = false.
      
      The test drives the real load() with docTR and torch stubbed, rather than
      re-implementing the placement rules beside it; reintroducing the fallback
      makes it fail, which is the only reason to trust it.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      52dfaca3
    • Stefy Lanza (nextime / spora )'s avatar
      placement: where a model runs is configured, not discovered · 8120e267
      Stefy Lanza (nextime / spora ) authored
      Three corrections to the fan-out, all the same mistake in different places:
      the code was deciding things the configuration had already decided.
      
      Nodes are no longer discovered. Asking every cluster node what it holds and
      using whoever answered is configuration by probe: it puts work on a machine
      nobody chose and moves it the moment a different node replies first. A node
      serves a model because the model is PINNED to it — `engine` on the entry, the
      same hard constraint the front router already honours, comma-separated when
      several nodes should share the work. A pinned node that does not hold the
      model is reported, not silently replaced.
      
      The decision is not an OCR decision, so it does not live in codai/ocr: any
      model, any caller that needs an endpoint rather than a proxy hop, gets the
      same answer from codai/models/placement.py. For requests that pass THROUGH
      coderai this was already true — the router honours the pin, treats a node as
      an engine by name, reaches RunPod through the backend. This is that decision
      for a caller that must hold an address itself, which is why OCR needed it.
      
      And "how busy is too busy" is the model's own scale_up_inflight_per_pod, on
      the model page, instead of the parallel ocr.serve_spill_inflight knob I had
      added — one number per model, already there, already meaning exactly this. A
      model with no RunPod block has nothing to rent, so load alone never moves its
      pages.
      
      Finally, the caller: resolve_engine accepted the four pipeline names only, so
      a client naming the model it can see in /v1/models got "unknown OCR engine".
      Both now work — engine=surya and engine=datalab-to/surya-ocr-2 reach the same
      place — because which pipeline reads a page, and where its model runs, are
      ours to know and not the client's.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      8120e267
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: one address for the engine, every server behind it (v0.3.9) · 39145dfb
      Stefy Lanza (nextime / spora ) authored
      An OCR pipeline is handed one server address and drives it itself, which is
      what keeps its page images, its strict json_schema and its logprobs intact.
      One address cannot spread load, and three things followed from that:
      
        * a pool grows on in-flight requests counted as they pass through it. An
          attached engine never passed through, so the pool saw no load however hard
          OCR ran, never grew, and the engine had one pod's URL anyway;
        * several cluster nodes serving the model resolved to whichever answered
          first, for ever, however busy it became;
        * a local model configured to burst to RunPod never did: that trigger lives
          on the proxy path an attached engine bypasses.
      
      So the engine gets a lease on an address of ours, and codai/ocr/dispatch.py
      chooses per request: a pool that is the model's configured home takes every
      page (picking the least-loaded pod, renting when they are all at the
      threshold); otherwise this machine and the nodes holding the model compete on
      load — including the load a node reports for its whole install, so a node busy
      with someone else's work is left alone — and the overflow past
      ocr.serve_spill_inflight rents instead of queueing. An explicit `backend` pin
      still wins over discovery: that was a decision, and load-aware picking must
      not quietly repeal it.
      
      The hop is allowed to exist only because it decides nothing: body bytes in,
      body bytes out. The losses it is there to avoid came from PARSING a request
      and re-emitting it through the backend abstraction.
      
      Two gates stood in the way and neither could be satisfied the usual way: the
      engine has no API token, and surya's /health and /v1/models probes send no
      headers at all (bare httpx GETs), while refusing to attach if /health is not
      200. The lease is therefore read from the URL, and the engine-only internal
      gate exempts that path alone — tested to be no wider than that.
      
      olmOCR's model mode asks the same dispatcher, and still takes the in-process
      route when the answer is local, where the model manager's loading, eviction
      and queueing belong.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      39145dfb
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: a confidence nobody measured is worse than none (v0.3.8) · 903797e0
      Stefy Lanza (nextime / spora ) authored
      Surya asks every request for logprobs and averages exp(logprob) over the
      generated tokens to score each line. coderai accepted the parameter, never
      sent it, and hardcoded `logprobs: None` in every response builder.
      
      Surya's fallback is what makes that expensive: with no logprobs it does not
      report "unknown", it reports a literal 1.0 (recognition/__init__.py:248, :302,
      layout/__init__.py:97). Every line of every page comes back perfectly
      confident, the guesses included, and nothing downstream can tell which
      judicial document needs a human to look at it.
      
      So the round trip is wired where it can exist: the proxy backends (vLLM,
      RunPod) forward `logprobs`/`top_logprobs` and hand back what the server
      measured, buffered and streamed, accumulating the streamed pieces into one
      response-shaped object. The manager asks the backend's signature rather than
      negotiating by TypeError, which would have called it twice for a kwarg it
      never had. A request that did not ask is sent exactly as before, and a
      backend that cannot answer still returns null.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      903797e0
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: a pipeline asks where the engine is, and a node is one of the answers · 2f20f8e9
      Stefy Lanza (nextime / spora ) authored
      Everywhere else in coderai an engine is what RUNS a model. `ocr.default_engine`
      borrows the word for two different kinds of thing: paddle and docTR are runners
      with their own weights, while surya-2 and olmOCR are pipelines over a model that
      vLLM runs. Only the second kind has to be told where a server is — and that is
      the same question the model layer already answers for everything else.
      
      So answer it in all the places a model can actually be served. A cluster node
      was the one that was missing: a node is a whole coderai registered with the head
      as a remote ENGINE, so nothing in the backend layer knows it exists, and a VLM
      hosted there came back as "not served over HTTP" from an install that was
      serving it perfectly well. The node's own catalogue is the authority, so it is
      asked. An explicit backend pin still wins — discovery is the fallback, not an
      override — and the nodes are not probed at all when clustering is off, which
      would have added a timeout to the first page of every document.
      
      Also: one test here was passing for the wrong reason. `from codai.models import
      manager` reads the package attribute and never consults sys.modules, so a fake
      installed there was simply not seen. Patched as attributes of the real module,
      which is what made the precedence test fail honestly.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      2f20f8e9
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a proxy that edits the request decides for the server · a484119e
      Stefy Lanza (nextime / spora ) authored
      Three silent drops on the way to a rented pod, all found while wiring Surya to
      the model layer, all of the same shape: coderai accepted a field, forwarded a
      reduced version of it, and the pod answered something plausible.
      
      * Images. The front flattens multipart content to the literal text
        "[image_url content]" unless the backend claims vision, and only the llama.cpp
        backend ever did. So every page sent to a VLM served on RunPod — olmOCR in
        "model" mode, or Surya through a coderai-vllm pod — arrived as that
        placeholder, and the model answered fluently about nothing. A proxy is the
        wrong place to judge: forwarded, a text-only server returns a plain 400 and a
        VLM server reads the page.
      
      * Structured output. Only {"type":"json_object"} survived, so a json_schema —
        the strict form current clients send, Surya's guided decoding among them — was
        dropped and the reply came back as free text the caller could not parse. The
        streamed path ignored the field entirely, json_object included.
      
      * extra_args. Whitelisted per engine on a model's runpod block since pods
        shipped, parsed, stored, and never put on a command line. Setting it looked
        like tuning and did nothing, which is worse than refusing it: the pod boots,
        serves, and ignores the decision. A page-reading model wants prefix caching
        and the Qwen-VL pixel bounds, and until now the only way in was docker_args,
        which replaces the whole generated line.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      a484119e
    • Stefy Lanza (nextime / spora )'s avatar
      ocr: surya should ask who serves the model, not where the GPU is (v0.3.7) · 51aa0d2f
      Stefy Lanza (nextime / spora ) authored
      OCR on the Digesta install answered 503: "vLLM isolated venv not found at
      /cache/vllm_venv". That host has no GPU at all — renting A40 pods is its whole
      job — so the venv it was told to build could never have run, while the Surya-2
      checkpoint it needed was already a configured model (backend: runpod,
      engine: coderai-vllm) nothing in the OCR path could ask for.
      
      The model layer was never the problem; `ocr.surya_serve` was asking the wrong
      question. It enumerated server PROCESSES — torch in a local venv, our own vLLM,
      a URL typed in by hand — when the answer the operator had already written down
      was a MODEL. olmOCR has had that bridge since it shipped ("model" mode, through
      our own chat path); Surya never got one, because Surya2 drives its server
      directly with guided-JSON decoding and logprobs that a proxy hop would drop.
      
      So it gets the same bridge, in the only shape it can take: resolve the model
      through whichever backend owns it and hand the worker the address. Three values,
      because Surya refuses a server whose /v1/models disagrees with the checkpoint it
      expects, and a pod URL is useless without its token — which only Surya's vLLM
      client reads, so a bearer-locked endpoint is always attached to as "vllm".
      
      The address is a lease, not a location: a min_pods=0 pod is reaped when OCR goes
      quiet, and the page that finds it gone re-resolves once — renting a fresh one —
      instead of every later page failing against a dead URL. A failure against a
      server that still answers /health is a real error and is reported as one.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
      51aa0d2f
  2. 10 Oct, 2026 5 commits
    • Stefy Lanza (nextime / spora )'s avatar
      oom: the server should be the last thing the kernel kills, not the first · 860b3dd1
      Stefy Lanza (nextime / spora ) authored
      When zeiss ran out of memory this morning the OOM killer picked coderai-nvidia.
      Not because it was at fault — a runaway test process was — but because oom_score
      is essentially "how much memory does this task hold", and an engine with a model
      loaded is always near the top of that list. The leak carried on; the thing that
      was serving died; the box needed a hard reset.
      
      Linux offers one lever, /proc/<pid>/oom_score_adj, and two rules that decide
      where it can be pulled from. Both were verified on the two production hosts
      rather than assumed:
      
        raising is free, lowering needs CAP_SYS_RESOURCE — and a docker container does
        not have it (tested: in-container `echo -500` is denied, `echo 200` succeeds).
        So the protected base has to come from the engine daemon, via
        --oom-score-adj on the run, and everything inside can only step back UP from it.
      
        rootless podman cannot protect at all: a negative value is clamped to 0. 0 is
        still worth having on the Digesta host, where coderai currently sits at +200 —
        inherited from systemd's user manager, which makes it a PREFERRED victim with a
        kernel score of 800 against zeiss's 670.
      
      So: the container gets -500 (CODERAI_OOM_SCORE_ADJ, 0 disables), and inside it
      codai/util/oom.py states the order as three tiers relative to that base —
      engines +100, workers +300, training +500. Relative, because the base differs by
      host and a tier should keep its distance whatever it turns out to be. Engines are
      marked in the existing _engine_preexec hook, which covers the isolated model
      workers too: they are spawned by the engine, not the front, so they inherit it.
      Training is marked from the parent after Popen — a preexec_fn would run Python
      between fork and exec in a threaded server, which is the documented way to
      deadlock.
      
      The ordering is the point. Protecting everything equally would just move the
      bullet to an innocent bystander; this way the kernel sheds a training run (hours,
      restartable) before an engine (loses its in-flight work, and the front notices
      and respawns it) before the front (loses everything, including the ability to
      respawn anything).
      
      Every helper can only ever make a process MORE killable: nudge() refuses a
      negative delta rather than attempting a privileged write, because a process that
      believes it is protected and is not is worse than one that knows it is exposed.
      All of it is best-effort — a process that cannot adjust its score still serves.
      
      16 tests, including that the number reaches the kernel and moves the score, that
      a child can be marked from the parent, and that the engine mark sits inside a try
      (it runs between fork and exec, so anything that raises there fails the spawn).
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_019sRvvg5WSdLgcFUM96pmE4
      860b3dd1
    • Stefy Lanza (nextime / spora )'s avatar
      autoupgrade: a server that is down is not a server that is busy · 99cff39e
      Stefy Lanza (nextime / spora ) authored
      zeiss came back from this morning's hard reset with no container at all, and the
      hourly upgrade then skipped every tick — so the host sat on old code and, worse,
      stayed DOWN, with nothing in the loop that would ever have recovered it. Both
      idle probes are HTTP, so an absent server is indistinguishable from a wrong
      token: inflight() said "unknown" and the fail-closed guard refused to act. The
      Digesta host, whose container was up, upgraded itself fourteen seconds after the
      push; zeiss needed hands.
      
      container_state() answers the one question the probes cannot — is there a server
      at all — by name, then by any running container of the image being upgraded.
      Only a POSITIVE "not there" counts as idle: an engine that cannot be asked
      (daemon down, wrong socket) stays unknown and still fails closed, because the
      point is to recognise an absent server, not to assume one. Checked before the
      token, since a host with no server is in the same position as one with no token.
      Both installs run their engine's ps as the service user and name the container
      `coderai`, so the default matches; CONTAINER_NAME covers an install that does not.
      
      The second half: every knob reads `VAR="${VAR:-default}"`, which looks like the
      environment may override it, but the conf is sourced first and assigns
      unconditionally — so `MAX_DEFERRALS=1 coderai-autoupgrade` did nothing at all,
      with no line in the log to say why. I hit that landing this morning's restart and
      had to source the conf from a wrapper to get round it. The environment is now
      snapshotted before the conf and restored after, for that run only, and a test
      checks the snapshot list against every knob the script resolves so the next one
      added cannot quietly fall out. require-idle and max-deferrals are also logged
      now: "it skipped again" and "it was told to skip" used to look identical.
      
      The conf example claimed that without a token "the restart happens whenever the
      timer says so". It is the opposite — the check fails closed, so a running server
      is then never upgraded. Corrected, and REQUIRE_IDLE_CHECK documented next to it.
      
      tests/test_autoupgrade_recovers_a_down_server.py RUNS the script against a stubbed
      engine, runner, restart command and health file: the env restore is eval-based and
      the container check branches on an exit code, neither of which reading the source
      can confirm. 12 new tests, 56 across the three autoupgrade files.
      
      No version bump: nothing in the image changed, and bumping would send both
      production installs through an image upgrade and a restart for a host-side script.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_019sRvvg5WSdLgcFUM96pmE4
      99cff39e
    • Stefy Lanza (nextime / spora )'s avatar
      broker: a swallowed cancel spun until the host died (v0.3.5) · 63efd6cf
      Stefy Lanza (nextime / spora ) authored
      A full `pytest tests/` run reached 46 GB RSS plus 20 GB of swap on a 54 GB
      machine. Swap filled, the OOM killer took coderai-nvidia — which was serving —
      and the box needed a hard reset, which lost the last forty minutes of
      kern.log with it. atop's history is what identified the process: a pytest
      worker, carrying the `coderai-front` process title because importing the front
      sets it, growing 7.8 GB → 34.8 GB → 46.6 GB across three ten-minute samples.
      
      The defect is `except asyncio.CancelledError: pass` wrapped around an await
      whose job is to collect a CHILD task. asyncio delivers a cancellation as a
      CancelledError at the next await point, so that handler catches two unrelated
      things: the child reporting it stopped (fine to swallow) and the caller asking
      us to stop (never fine). run_forever's reconnect handler sat on exactly such an
      await, so `task.cancel()` was absorbed and the loop carried on. With the
      reconnect delay mocked to zero by the test driving it, that became a busy loop
      allocating ~78 MB/s — AsyncMock records every call it receives — and 400,000
      iterations in the five seconds I measured.
      
      being_cancelled() (Task.cancelling(), 3.11+) tells the two apart, and all four
      sites now re-raise ours: _stop_heartbeat_task, _cancel_inflight_tasks, the two
      keepalive collects in handle_message, and BrokerService.stop. The inflight ones
      matter in production independently of the leak — absorbing a cancel there means
      a shutdown waits out a whole video generation instead of stopping it, and
      run_forever simply reconnected forever instead of exiting. run_forever's own
      cancellation path clears the cancel for the duration of its cleanup and
      re-raises at the end, so honouring the cancel does not cost the websocket close.
      
      The test that drove it now brakes itself: the zero-delay sleep stand-in counts,
      and raises a BaseException (anything else is caught by the reconnect handler)
      well past the two reconnects it needs. With the product bug reintroduced the
      file fails in 1.5 s at 441 MB instead of eating the machine.
      
      tests/conftest.py adds what was missing in general: an RSS ceiling (6 GB,
      CODERAI_TEST_RSS_LIMIT_MB) and a per-test stall ceiling (180 s,
      CODERAI_TEST_STALL_SECONDS), both dumping every thread's stack to a file before
      exiting — pytest captures stderr per test and drops the buffer when the process
      does not unwind, so a watchdog that only printed would have died silently. No
      new dependency. "A test can consume the host" is a property of the run, not of
      one test.
      
      tests/test_broker_protocol.py could not finish before this — it hung at the
      14th test. The whole suite now completes for the first time: 78 s, 1.57 GB peak,
      1635 passed. The 25 failures are the 17 that already failed plus the 8 stale
      register-ack expectations in that file, which were unreachable behind the hang
      and are a separate question (the fixtures send `accepted: true`; the client has
      required `status: "ok"` for some time).
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      Claude-Session: https://claude.ai/code/session_019sRvvg5WSdLgcFUM96pmE4
      63efd6cf
    • Stefy Lanza (nextime / spora )'s avatar
      video: a model says how long its clips can be (v0.3.4) · 0dc6084e
      Stefy Lanza (nextime / spora ) authored
      Township's clip lengths were tuned when the only video model could hold ~81
      frames in one call, so a 70-second match became a dozen short cuts. LongCat was
      pretrained on continuation and the server renders a long shot in segments, so
      the same match wants a handful of long takes — and the frame numbers that suit
      one are wrong for the other by a factor of four. Only the two starter templates
      knew that, and only because someone typed it in.
      
      Worse, nothing snapped a picked count onto what the model can emit:
      random.randint handed the server lengths like 250, which is not 4n+1 for any
      VAE and not on LongCat's 93+80k segment grid, so the server rounded and the
      plan's own numbers became fiction. On LongCat that also silently costs
      block-sparse attention, because every point of 93+80k is a 16n+13 count BSA
      can tile and nothing between them is.
      
      Server — codai/models/video_geometry.py is the geometry once: native rate, the
      legal frame grid, how much ONE render holds, whether continuation is native,
      and the side multiple below which the fast attention path is lost. Published
      per model on /v1/models (ModelInfo.video) and used by the render path itself,
      so what the server advertises and what it rounds to cannot drift. The numbers
      were spread across three modules that could not tell a client; the literals
      here are pinned to each source by tests, since the LongCat and H3 workers run
      in 3.10 venvs this cannot import. models.json overrides every field.
      
      Township — the five frame fields now default to 0 = AUTO: derive the budget
      from the model's geometry and this tool's seconds band, then snap both ends and
      every pick onto the model's grid. Seconds, not frames, because clip count falls
      out of long_target / (frames/fps). An explicit number still wins, so a saved
      config keeps exactly the budget it had. Entrances and the face-off get their own
      shorter band instead of borrowing the fight budget (on LongCat that made every
      entrance an 11-22s take). Outcome shares are snapped per shot and the total
      rewritten, the per-render chunk comes from the model (Wan keeps its
      conservative 50), and the chaining decision reads the published continuation
      rather than matching the model's name.
      
      At the defaults a 70s match goes from 9-11 Wan cuts to 4-6 LongCat takes, each
      one request with native continuation, with BSA still on.
      
      1607 tests pass; the 17 failures and the test_broker_protocol hang are
      pre-existing and unrelated (verified with the diff stashed).
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      0dc6084e
    • Stefy Lanza (nextime / spora )'s avatar
      longcat: a conditioning pass tiles twice (v0.3.3) · 69240b20
      Stefy Lanza (nextime / spora ) authored
      An i2v generation died in the first denoising step with a bare AssertionError out
      of flash_attn_bsa_3d, on a geometry the guard had just declared valid: 93 frames,
      sides divisible by 64.
      
      bsa_problems only ever checked the noise branch. A conditioning pass does not tile
      the latent once — Attention.forward splits the sequence into a conditioning block
      and a noise block and calls flash_attn_bsa_3d on each (modules/attention.py:123-134),
      so the cond depth and the remaining noise depth must divide by the chunk as well.
      generate_i2v hardcodes one conditioning latent (pipeline:796) and 1 % 4 != 0, which
      means block-sparse attention has never once been able to run an i2v pass, at any
      frame count or resolution. The avatar pipeline hardcodes the same 1 in four places,
      and generate_vc computes a count it never rounds. Only generate_refine gets it
      right, by padding both counts to the granularity itself (pipeline:1245-1250).
      
      (a) The guard now knows about the split. _cond_latent_passes enumerates every
      conditioning count the request can produce — segment 0 uses the task's own entry
      point and every later segment uses generate_vc, so even t2v runs a conditioning
      pass once num_segments > 1, and BSA is one flag for the whole request — and a
      count that cannot tile puts the pass on dense attention with a line saying which
      condition failed.
      
      (b) bsa_pad_cond, OFF by default, rounds the count up to the granularity on the way
      into the DiT, which is what generate_refine does for itself, so an i2v or
      continuation pass can keep BSA. Patched onto the DiT instance rather than the
      vendored source, the same way the denoising bars are rebound: the count is born
      inside generate_i2v's loop, so there is no outer argument to correct. Off by
      default because the pipeline marks only the real conditioning frames clean
      (timestep[:, :1] = 0), so rounding 1 up to 4 hands the conditioning block three
      latent frames that are still noise — well-defined attention, but not what the model
      was trained on, and on a video model that kind of error shows up as temporal drift.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      69240b20
  3. 09 Oct, 2026 15 commits
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a configured global volume must not fail every pod (v0.3.2) · 608127c8
      Stefy Lanza (nextime / spora ) authored
      Probed every API surface coderai can use for attaching a global volume, each
      with a body valid enough that the extra-key check actually runs:
      
        POST /v1/pods            key … not in input schema: 'globalNetworkVolumeId'
        podFindAndDeployOnDemand field not defined by the input type
        POST /v2/pods            additional properties 'globalVolumeId' not allowed
        templates (mounts)       every candidate key refused
      
      There is no public API for it: global volumes are console-only, as the docs
      say. Which makes `global_volume_id` a loaded gun — sending the field fails the
      create, so configuring a global volume would have stopped the install renting
      pods at all.
      
      Both create paths now recognise that specific rejection, drop the field and
      retry. The pod boots without the volume and downloads its weights: slower and
      it costs bandwidth, and far better than serving nothing. Logged once per
      process, naming the consequence and the override to set when RunPod publishes
      the field name. An unrelated failure is NOT treated this way — a capacity error
      retried with the volume stripped would bring a pod up silently wrong.
      
      Earlier probes that reported the field as "accepted" were wrong: with an
      otherwise-invalid body the validators stop before the extra-key check. Worth
      remembering next time I probe a schema.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      608127c8
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a preferred GPU, not the only GPU (v0.3.1) · 50caed7b
      Stefy Lanza (nextime / spora ) authored
      `gpu_type` was a filter: "NVIDIA A40" meant an A40 or nothing, so when RunPod
      had no A40 free the request failed rather than taking a card that would have
      served it. Now `gpu_types` is an ordered PREFERENCE and `allow_other_gpus` lets
      the search fall through to anything else that still satisfies min_vram_gb and
      max_hourly_usd.
      
      A preferred card ranks ahead of every fallback whatever selection_criteria
      says — otherwise "cheaper" would put the fallback first and naming a card would
      mean nothing. Each ranked option carries `preferred`, and provisioning a card
      nobody asked for prints "[FALLBACK — no preferred card available]": correct
      behaviour, but never silent.
      
      Off by default. A pinned card is sometimes pinned for a reason, and an install
      that quietly started renting something else would be a surprise on the invoice.
      
      Name matching is exact on the id or display name, case-insensitive, tolerating
      a missing "NVIDIA " prefix, and deliberately NOT a substring: "A40" would
      otherwise match "NVIDIA RTX A4000" — 16 GB for a job that asked for 48. GPU
      names split on commas only, because "NVIDIA RTX A6000" has spaces in it.
      
      Also: the admin handler rebuilds the runpod block from a whitelist, so all six
      keys added today (global_volume_id, data_centers, gpu_types, allow_other_gpus,
      allow_client_scale, client_scale_ttl_s) were being dropped the next time anyone
      saved a model from the GUI — taking the warm schedule and the volume with them.
      They are in the whitelist now, with the two ordered lists kept in order.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      50caed7b
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: global volumes and a region allow-list (v0.3.0) · 88edb84f
      Stefy Lanza (nextime / spora ) authored
      Weights in one place, pods wherever a card is free.
      
      `global_volume_id` attaches a RunPod GLOBAL volume (beta since Sept 2026):
      region-independent storage, so unlike `network_volume_id` it does NOT pin the
      pod's data centre — which is the only reason to use one. Resolved by its own
      function precisely so the two cannot be confused: the region lookup stays under
      the network-volume branch, and a test holds them apart.
      
      `data_centers` is the other half. A single `data_center` pins one region and a
      blank one lets RunPod pick any region on earth, so there was no way to say
      "anywhere in the EU" — which is what data that may not leave the EU requires.
      The list is tried in order, and each (card, region) pair becomes its own
      candidate so the capacity fallback that already walks GPU options walks regions
      too, with no new control flow. The pod handle records the region it actually
      got, not the pool's first choice.
      
      Two honest limits, both documented in the code and the guide. A global volume
      must be created in the RunPod CONSOLE: POST /v2/network-volumes allows only
      STANDARD|HIGH_PERFORMANCE and demands a dataCenter. And attaching is
      console-documented only — no field exists in the Pod API reference, and the v1
      REST validator accepts unknown body keys silently, so a wrong name cannot be
      caught by probing; it yields a pod with no volume that re-downloads its
      weights. We send globalNetworkVolumeId, overridable by
      CODERAI_RUNPOD_GLOBAL_VOLUME_FIELD, and it must be verified on a real pod.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      88edb84f
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a pid is not a process identity in a container (v0.2.99) · 7b87efd7
      Stefy Lanza (nextime / spora ) authored
      The root cause under the last three fixes. The pod registry recorded
      os.getpid() and both readers skipped an entry whose pid matched ours as "our
      own, the local pool handles it". Inside a container the engine comes up as the
      SAME pid on every boot — 74, on the production host — so after a restart the
      previous container's pod read as ours and was skipped by both
      find_shared_pod (nothing to adopt) and sibling_is_provisioning (nothing
      booting). The warm path then rented another A40 beside the one already running.
      Three restarts this afternoon, three extra A40s, each billing until it was
      terminated by hand.
      
      Entries now carry a per-process nonce and identity is that. The pid is still
      recorded, and an entry written before the nonce falls back to comparing it,
      which is the old behaviour and no worse.
      
      test_a_surviving_pods_token_is_remembered_across_a_restart simulated the new
      process by monkeypatching os.getpid — the very thing that does not change. It
      patches the nonce now, which is why it failed on this commit and passes on it.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      7b87efd7
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: do not rent beside a pod that is still booting (v0.2.98) · c9fc3741
      Stefy Lanza (nextime / spora ) authored
      The last hole in the warm path. sibling_is_provisioning already recognises a
      pod registered at creation with no URL yet — and after a restart the recorded
      pid no longer matches, so our own booting pod qualifies — but ensure_ready
      never asked. Seen live: a pod created at 18:57 was six minutes into its image
      pull when the 19:03 upgrade restarted the service, and the warm path rented a
      second A40 beside it. The upgrade runs hourly and an image pull is ~6 minutes,
      so that is not a rare alignment.
      
      Order in ensure_ready is now: a healthy pod of our own, else adopt a ready one,
      else wait for one that is booting, else rent.
      
      Also prune the pod registry in reap_orphans: an id RunPod no longer has stayed
      "known", which both shielded it from reaping and let the file grow for the life
      of the install — including pods terminated out of band, of which this afternoon
      produced five.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      c9fc3741
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a restart rented a second pod beside its own (v0.2.97) · 3770472d
      Stefy Lanza (nextime / spora ) authored
      A warm floor is re-established at every boot, and ensure_ready() went straight
      to _provision_one(). The adoption logic — "a pod for this pool is already
      running, take it over" — existed only inside acquire(), so the warm path never
      saw it. Three restarts on the production orchestrator this afternoon rented
      three A40s and left each previous one billing until the reaper's 20-minute
      grace expired. An hourly unattended upgrade would have done that every hour.
      
      find_shared_pod already covers the case: after a restart the recorded pid no
      longer matches, so our own previous pod is adoptable. It is now one method,
      _adopt_shared_pod, called from acquire(), ensure_ready() and maintain()'s warm
      top-up.
      
      It also stops lying about the cost. An adopted pod was recorded at $0/hr as
      "billed by its owner", which is true of a sibling engine's pod and false of one
      we are inheriting from our own restart — a real A40 hidden from the $/hr cap and
      from the spend page. register_pod now stores the rate and adoption restores it;
      an older registry entry without one still adopts, at 0.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      3770472d
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a rented pod that was never recorded, and billed on (v0.2.96) · afdc1d28
      Stefy Lanza (nextime / spora ) authored
      NameError in _provision_one: it read `dc`, a local computed inside
      _create_with_fallback, from a function that has no such name. The sequence that
      costs money is this — create_pod succeeds, the pod boots on RunPod, the
      PodHandle raises, the pool never records the pod, the caller provisions again.
      Four A40s existed for a warm floor of one before anyone looked.
      
      Both paths now call pool._data_center(), which keeps the same precedence (the
      volume's region, then the model's, then the account's) in one place.
      
      The two tests that covered this asserted the broken expression as a SOURCE
      STRING — 'data_center=(getattr(self, "_volume_dc", "") or dc or "")' — so they
      passed on a NameError and would have gone on passing. They now call the
      resolver, including with nothing configured anywhere, which is the shape that
      raised.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      afdc1d28
    • Stefy Lanza (nextime / spora )'s avatar
      docs: the production container has two config.json and two readers · e7b4d576
      Stefy Lanza (nextime / spora ) authored
      Found while configuring the Digesta catalogue: the launcher reads
      /root/.coderai/config.json (engine_specs, published port) while the front and
      the engine read /root/.coderai/coderai/config.json. The RunPod api_key and the
      spend caps were in the first one, so the engine refused to provision anything
      and the admin page reported no caps at all. Records what was moved where, and
      that CODERAI_CONFIG_DIR=/config remains the fix that makes it unnecessary.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      e7b4d576
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: two pods for a warm floor of one (v0.2.95) · 5cec0702
      Stefy Lanza (nextime / spora ) authored
      Watched it happen on the production orchestrator: configuring qwen38-awq warm
      during office hours provisioned TWO A40s a second apart.
      
      A pod does not enter self.pods until it answers its health probe, which is
      minutes away. ensure_ready() never set the _provisioning flag, and the warm
      top-up in maintain() never read it — so the 15 s scaler saw "0 healthy, floor 1"
      while a pod was already booting and rented another. It stopped at two only
      because _provision_one blocks the scaler thread for the whole boot; a
      non-blocking provision would have added one every 15 seconds.
      
      The demand path (_should_grow) has always checked the flag. Now all three agree:
      one on the way is one pod, not none, including one a sibling engine sharing the
      pool is starting.
      
      Also: /v1/runpod/scale answered with the internal pool key ("pool:digesta-qwen38")
      rather than the model the client asked about. Several models share a pool, and
      that string is not something the caller sent or can look up — so the answer
      names the model and reports the pool beside it.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      5cec0702
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: a keep_warm pod that is disabled, unnamed or out of hours (v0.2.94) · a198ea93
      Stefy Lanza (nextime / spora ) authored
      Configuring the Digesta catalogue on the production orchestrator turned up three
      faults in warm_configured_pools, all of which spend money:
      
        * `enabled: false` was not consulted. A catalogue entry parked for later with
          keep_warm still set rented a GPU at every boot.
        * the entry was named `alias or path`. A RunPod-only entry has neither — no
          weights live on that machine, and an alias only exists when a second name is
          wanted — so it warmed under the key "" and the pool was never the one that
          went on to serve requests. `id` is now the fallback, and an entry with no
          name at all says so instead of warming nothing quietly.
        * the schedule was ignored. ensure_ready() is the "a request is waiting" path
          and provisions unconditionally; at boot nothing is waiting, so a restart at
          02:00 on a Sunday rented exactly the pod the window exists to avoid. It now
          checks effective_min_pods first, which also means a client lease that
          outlives a restart is honoured rather than dropped.
      
      Found on the host, not by reading: the warm pod the configuration asked for
      never appeared.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      a198ea93
    • Stefy Lanza (nextime / spora )'s avatar
      runpod: let a client lease warm pods up to max_pods (v0.2.93) · 81507a8b
      Stefy Lanza (nextime / spora ) authored
      The autoscaler reacts to load — a pod is added once the least-loaded one
      already has scale_up_inflight_per_pod requests in flight. That is right for
      traffic nobody predicted and wrong for a client that KNOWS what is coming: a
      batch of a thousand documents otherwise discovers the pool one cold start at a
      time.
      
      So a client may now say so: POST /v1/runpod/scale {model, pods}. It is a lease,
      bounded three ways, because a client that can raise the floor can raise the
      bill:
      
        * allow_client_scale, off by default, so no model is client-scalable unless
          someone said so;
        * clamped to max_pods — asking for more returns the ceiling with
          "clamped": true rather than an error, since the client asked for "as many
          as you can" and refusing would leave it with none;
        * client_scale_ttl_s (default 30 min), because the failure that matters is
          not a wrong number, it is a client that asks for four A40s and then dies.
          A lease that never expired would be a bill that never stopped, so a 0 in
          the config falls back to the default instead of meaning "forever".
      
      The lease and a warm-pod schedule are MAXED, never summed: both mean "hold this
      many ready", so a client cannot lower a configured floor and a closed window
      cannot cancel a lease. /v1/runpod/status reports the live lease per model, so an
      unexpected bill has a visible cause.
      
      Not admin-scoped, unlike /v1/runpod/spend: the gate is the per-model flag, not
      the key, and the account's $/hr and spend caps still refuse to provision.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      81507a8b
    • Stefy Lanza (nextime / spora )'s avatar
      autoupgrade: a detached job is not "serving a request" · 05802ce0
      Stefy Lanza (nextime / spora ) authored
      Proven on zeiss with a real LoRA training run: the skip fired correctly but
      read "SKIP: serving 1 request(s) [engines=0 jobs=1]". There was no request --
      that is the entire case this check exists for, since /v1/loras/train with
      wait=false returns at once and the work continues for hours. The log now says
      "work in progress", which is what it means.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      05802ce0
    • Stefy Lanza (nextime / spora )'s avatar
      front: expose active tasks on /metrics so training blocks an upgrade (v0.2.92) · dadb0196
      Stefy Lanza (nextime / spora ) authored
      The unattended upgrade measured idleness as requests in flight, and a LoRA
      training run has none: the request that starts it returns immediately and the
      work continues for hours, so coderai_engine_inflight is 0 throughout. An hourly
      upgrade would have restarted straight through one -- the standing "never restart
      during training" rule, reintroduced by automation.
      
      Nothing was missing in the training path. loras.py already registers a
      `training` task, the engine already ships active tasks in
      /internal/engine-state, and the front already stores them on
      EngineEntry.tasks. The gap was that /metrics never exposed them and the probe
      never looked, so this adds coderai_engine_jobs_active per engine and
      coderai_jobs_active across them, counting running/queued/paused -- paused on
      purpose, because a thermally paused generation has not finished and restarting
      loses it.
      
      The probe now reads both numbers from one scrape and takes the MAXIMUM, never
      the sum: a generation is both an in-flight request and a running task, so adding
      them would double-count and the numbers in the log would match nothing an
      operator can see. A front that predates the gauge reports "-" and the detail
      line says "front too old", rather than claiming zero and silently restoring the
      blind spot.
      
      tests/test_autoupgrade_sees_training.py renders the exposition from a fake
      engine carrying a training task, runs the probe's awk program as the script
      actually has it, and pins both halves of the premise -- the registration in
      loras.py and the engine's reporting of active states -- so the gauge cannot go
      quiet without a test failing.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      dadb0196
    • Stefy Lanza (nextime / spora )'s avatar
      autoupgrade: tell the runner which container engine to use (v0.2.91) · 50f94ca5
      Stefy Lanza (nextime / spora ) authored
      run_oci.sh takes its engine from CONTAINER_ENGINE and defaults to docker, so on
      the rootless-podman host this script was written for it would have looked for a
      docker socket that is not there -- failing on every hourly tick with nothing to
      show but a non-zero exit. Found while installing it on the Digesta production
      server, before it ran.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      50f94ca5
    • Stefy Lanza (nextime / spora )'s avatar
      front: compare forward-auth peers as addresses, not strings (v0.2.90) · 7b570bc1
      Stefy Lanza (nextime / spora ) authored
      The Digesta host added both 10.89.0.178 and ::ffff:10.89.0.178 to
      trusted_proxies and host->8777 still refused. The ::ffff: hunch was right: the
      internal nginx listens on 8776 AND [::]:8776, so a connection landing on the v6
      socket makes $remote_addr -- and therefore X-Forwarded-For, and therefore the
      peer -- the IPv4-mapped form, which a string membership test rejects even when
      the plain address is listed. Nothing in the configuration explains that refusal.
      
      Peers are now compared as addresses: a mapped v6 peer matches a plain v4 entry
      and the reverse, and CIDR entries work, so a container network can be trusted as
      10.89.0.0/24 instead of pinning one address in the quadlet. A non-address entry
      (a socket path) is still compared literally, and a v6 peer never matches a v4
      network.
      
      Their broadened list could not have taken effect anyway -- it went into the
      config.json the app is not reading, which is why `enabled` there did nothing and
      the env var was needed. So that test did not rule the peer check out; the list in
      force was still the default pair.
      
      The rest of the report is answered in the doc rather than in code, because there
      is no bug behind it: `Server: nginx` appears on PROXIED responses too unless
      proxy_pass_header Server is set, and the missing `GET /admin` access-log line is
      _PollNoiseFilter, which drops reads of /admin, /static, /login and /logout from
      the front's access log while leaving /healthz alone -- exactly the asymmetry that
      looked like nginx short-circuiting by source. A 302 to /coderai/login with no
      Set-Cookie is what the front returns when forward-auth declines.
      Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
      7b570bc1