- 11 Oct, 2026 20 commits
-
-
Stefy Lanza (nextime / spora ) authored
/v1/models answers the moment a vLLM process starts — minutes before it can serve, because the checkpoint still has to be downloaded (every pod re-downloads it without a network volume) and the KV cache built on first use. The pool called that ready, every client read it as usable, and the whole model load landed inside somebody's first request, where it is indistinguishable from a hang. That is what cost today. One OCR page took 807s end to end — ~400s image pull, ~400s weights — while the client timed out at 180s and then 300s and concluded the serving was broken. It was not: nobody had ever waited long enough, and the only reason we know is a probe run with a 1400s timeout, which came back 200 with text and conf 0.9019. So the pool now sends one token against the served model before calling the pod ready. The pod bills through the load either way; this only decides whether the wait is visible as a pod that is not ready yet, or hidden inside a request that looks stuck. `pods_ready` becomes a signal a client can act on, which is what every client already assumed it was. A failed warm-up is reported and the pod used anyway: it may still serve (an engine that takes no chat completions, a model whose warm-up shape we guessed wrong), and rejecting a pod we have already paid to boot over a diagnostic request would be worse than the hidden latency this removes. Off with warmup_on_boot=false or warmup_timeout_s=0. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
A rented pod is the one machine in the system whose logs cannot be read. RunPod publishes no pod-log API — its REST spec lists 23 routes and none serve logs, which is why every attempt returns HTTP 400 — and boot.sh writes its phases to a file and then execs uvicorn, whose output goes to the container's stdout and nowhere else. /boot therefore covered the way up and nothing after it, and _dump_pod_logs only ever ran on a failure, printing into the orchestrator's log. That cost an evening. Three pods were investigated in one session and two were already destroyed by the time there was a question to ask — one reclaimed by RunPod mid-request, taking the only evidence with it — so "the serving hangs" stayed a hypothesis for an hour. It was not hanging: a page took 807s, most of it the weights download every pod repeats because global volumes are console-only, and every client had given up at 180-300s. So the pod keeps a bounded ring of its own stdout and stderr (codai/api/podlog) and serves it at /logs with a cursor, and the pool drains that on every scaler tick into <config_dir>/pod-logs/<pod_id>.log — kept AFTER the pod dies, which is the whole point. It also drains once more before tearing a pod down, because the interesting lines are the last ones and they are only readable while the pod exists. A pull, not a push: the orchestrator already reaches pods to serve requests, so the bearer token and pinned TLS apply unchanged and a pod with no outbound route still gets its log kept. Bounded at both ends, since neither disk is free, and a gap in the pod's ring is written into the file instead of passing silently. The stdout tee writes through to the real stream on every failure path — a logging aid that could break printing would be worse than none. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
RunPod reclaimed an A40 mid-document today. The pool went on reporting it healthy, so every page chose the same corpse: the in-flight count climbed, the pod was never dropped, and nothing drained. From the client that reads as "requests neither complete nor time out", and it is what wedged the test. RunpodBackend has always handled this — _unreachable() drops the pod and the request goes again. The fan-out gateway never did: it returned 502 and left the dead pod in the pool as healthy. So a lease can now report its upstream dead, which drops the pod, clears the cached candidate set and releases the slot; the page is retried ONCE against whatever replaces it, and a third attempt would just be a retry loop against a real outage. Only connect-level failures count as dead. A read timeout is a model taking its time over a dense page, and discarding a healthy pod for that would be an outage of its own. Also: the first version of that handler could not work. It used a metaclass with __instancecheck__ so `except _DEAD_UPSTREAM` would match httpx's errors — but `except` matches on the class hierarchy at C level and never consults __instancecheck__, so it caught nothing. A test caught it; the check is now an explicit isinstance() where it cannot lie. And the registry path again: 0.3.15 keyed the fallback off XDG_CONFIG_HOME, which the container launcher exports only into the processes IT starts. A diagnostic script entering the container another way printed a THIRD registry path and created a second pod TLS CA. The fallback now reads CODERAI_CONFIG_DIR, the variable the launcher actually sets, so every process agrees however it was started — and since pod_tls.tls_dir() derives from the registry path, the CA follows. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
refuses our token is not an asset (v0.3.17) Both from watching a real OCR run. THE LOST INTENT. A client asked for the fleet back while a page was still being read. The pods correctly stayed — work is never dropped to save money — but the intent was then forgotten: the pool fell back to idle_timeout_s (900s there) and billed for a fleet nobody had asked to keep. It took a second teardown call, four minutes later, to actually release it. The pool now remembers: pods kept only because they are working are retired the moment they fall idle, the same "do not wait out the timeout" treatment a closed schedule window already gets. Asking for warm pods again cancels it. Also fixes the reply and the log, which said "floor 1" when the 1 came from work in flight. That reads as though the operator had asked for the pod — the opposite of what it meant. Demand and configuration are now reported separately ("needed" and "floor"), because they are two different reasons a pod survived. THE ADOPTED POD THAT 401s. After a restart the pool adopts the pod it left running instead of renting another, and carries that pod's bearer token across through the registry — when one was recorded. With none recorded it fell back to our own token and never checked the result, so a pod launched with a different one answered 401 to every request for the rest of its paid life. Seen today, right after a restart: "reusing the pod already running" followed by 401 on every OCR page. The health probe that found it proves nothing here — a coderai pod answers /healthz to anybody. Adoption now probes an AUTHENTICATED path with the token it would use, and a pod that refuses it is terminated rather than adopted: with one registry per deployment (0.3.15) an unrecorded key is a lost key, so that pod can never serve anyone, and paying for it is the only outcome worse than killing it. An unreachable pod is NOT treated as unauthorised — that stays with the health machinery, which has a budget for it. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
First real document, first real pod: it came up at https://194.68.245.79:22198 and every page came back 502 — OCR upstream pod failed: [SSL: CERTIFICATE_VERIFY_FAILED] unable to get local issuer certificate A direct-TCP pod serves the certificate coderai's OWN CA signed, on a bare IP that certificate can never name. codai.api.pod_http has always known this: it pins to pod_tls.ensure_ca() and switches hostname checking off. The fan-out gateway is async and could not reuse that requests-based module, and I gave it a bare httpx client — so the one hop that now carries every OCR page was the one hop with no pod trust. It keeps TWO clients per event loop, chosen by pod_http.is_pinned_url: the pinned context trusts only our CA, so using it for everything would reject a cluster node whose certificate a public CA signed, and sharing one client would let whichever upstream connected first decide the policy for both. Worth recording: this was never fan-out-specific. With serve_fanout off, surya attaches straight to that same https:// IP with its own httpx and fails identically — `model` mode has never worked against a TLS pod. The gateway is what makes it fixable in one place. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
Two things, both from today's live incident. THE SPLIT REGISTRY. _pod_registry_path() fell back to ~/.coderai whenever config_manager was not initialised in the calling process. The container's launcher puts the real config dir one level deeper — it exports XDG_CONFIG_HOME="$CODERAI_CONFIG_DIR" and uses "$CODERAI_CONFIG_DIR/coderai", which is also where platform_paths.user_config_dir() lands. So a process with the config manager wrote one registry and a process without it read another, and the pods each owned were orphans to the other. That is what made the documented mutual-kill loop recur: pod created 19:24:56, terminated as untracked 19:24:57. The fallback is now the same XDG-aware path the launcher computes, and the resolved path is logged once — a registry that splits silently is invisible until it costs a GPU. CLIENT-FORCED TEARDOWN. `pods: 0` released a client's lease but gave nothing back: the pods then waited out scale_down_after_s, or idle_timeout_s before that. The automatic path is slow on purpose — it infers from a momentary reading, and a cold pod costs minutes to replace — but a client that has just finished a batch is not inferring. `teardown: true` sheds the surplus at once, skipping both the window and the one-per-pass limit, and answers with what happened. It still cannot scale through the operator's min_pods, and a pod with requests on it drains rather than dying under them. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
Observed live on the Digesta orchestrator while testing OCR: pod created 19:24:56, "terminated STALE pod … not tracked by any pool" 19:24:57. The pool then waited out its 900s boot timeout on a pod that no longer existed, every OCR request 503'd, and the next attempt rented another. This is the race the source already documents from an earlier incident — "the nvidia engine created a pod and the radeon engine terminated it four seconds later, then the reverse — a mutual kill loop that rented and destroyed pods until stopped" — and the defence against it was supposed to be the age guard, "never reap a pod younger than this, whatever the registry says". It was written as `age is not None and age < REAP_GRACE_SECONDS`, and list_pods() returned {id, name, status, cost_per_hr, image} with no timestamp of any kind. So age was None for every pod, the condition was never true, and the guard abstained on exactly the pods it existed to protect. A guard that fails open is worse than none: it reads as covered. Two changes. The guard now fails CLOSED — an undatable pod is left alone and says so — and the listing carries runtime.uptimeInSeconds so it can measure instead of abstaining. Not reaping costs one more cycle; reaping a live pod costs the work on it and bills in a loop that never serves a request. What made it recur here is worth recording: this install has TWO config dirs, so the cross-process registry that shields a new pod exists twice and each process reads the one the other does not write. That is the deeper fault and is not fixed by this commit. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
Two faults, one cause: `local` was both the default and the only engine mode that would quietly settle for the CPU. Surya's own settings pick the device and fall back cuda -> mps -> cpu (surya/settings.py TORCH_DEVICE_MODEL), so `local` on a machine with no CUDA ran a VLM pipeline on the processor: minutes a page, a result that looks perfectly normal, and nothing anywhere saying why. paddle and docTR already refuse in exactly that situation — surya was the exception. It now refuses too, with a 503 that names both ways out, and the worker is TOLD its device rather than left to guess. Detection fails closed: torch is asked first (it is what surya will ask), nvidia-smi second, and if neither can answer the answer is no. And `surya_serve` no longer defaults to `local`. An install that said nothing got the one mode that runs the model on this box — on a GPU-less orchestrator whose whole job is renting one, that is the worst possible default. It now defaults to `model`, which adapts to wherever the checkpoint is served and asks for surya_model_id when it has not been told. Safe to change: both production installs set the mode explicitly (`model` on Aruba, `vllm` on zeiss), so nothing was relying on the old default. Running on the CPU is still fully supported — surya_use_gpu = false — it just has to be a decision now. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
A pool rents at whatever the market offered when each pod was started, so one fleet can hold a $0.35 card and a $0.79 card doing identical work. Shedding the cheap one saves less than half as much, and which to prefer is the operator's call, not ours — hence a checkbox rather than a new default. With scale_down_costliest_first the dearest surplus pod goes first, tie-broken by the idlest and then the coldest cache. Cost decides WHICH pod goes, never whether one goes, and a dear pod with requests on it is drained rather than dropped — so choosing it costs time, not work. Worth weighing before switching on: an on-demand pod costs more than a spot pod, so "dearest first" prefers to keep the interruptible ones. Also fixes the default tie-break, which was inverted when scale-down was written: `-last_used` picked the pod used MOST recently, discarding the warmest prefix cache of the set. It now retires the coldest. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
Only half the scaler existed. The pool adds a pod once every pod it has carries scale_up_inflight_per_pod requests, and pods went away again only through the idle reaper — which fires on a pod with NO work and no request for idle_timeout_s. That misses the expensive case. A pool that grew for a burst and is now serving a trickle keeps every pod: each is touched often enough that last_used never ages out, each is a fraction busy, and the bill stays at the peak for as long as the trickle lasts. With max_pods raised to 64 that is the difference between a few dollars an hour and a few hundred. So the pool also asks how many pods the CURRENT concurrency justifies and retires the excess. The rules that matter are not "fewer pods": * work is never dropped to save money. The pod retired is the least loaded, and if it still has requests it is marked draining — no new work, and terminated only once it has finished what it had; * one pod per pass, after the surplus has persisted for scale_down_after_s, and the window restarts after each retirement. A dip of a few seconds must not cost a pod wanted again seconds later: a cold one takes minutes; * a quiet pool is NOT this routine's business. With nothing in flight, teardown stays entirely with idle_timeout_s — a number the operator chose, weighing the bill against a cold start — so this never silently shortens it. A draining pod is also excluded from the growth decision, which reads the LEAST loaded pod: an emptying drainer would otherwise answer "no need to grow" with the one pod that cannot take the work. Two existing tests had to change. A _Pod stub stood in for PodHandle without the new field, and a source-text assertion used a 600-character window that a new field plus its comment pushed data_center out of — now asserted against the whole class, which is what it was always about. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
The gateway built an httpx.AsyncClient for every request, so each page paid a fresh TCP and TLS handshake to the upstream. With surya_instances: 24 that is 24 streams reconnecting per page, against a round trip that already dominates the cost of reading one. One pooled client with keep-alive instead, so a document's pages reuse the connection. Correctness-neutral; it only ever cost time. Keyed by the event loop, because an AsyncClient binds to the loop that created it and coderai runs more than one (the OCR pool keeps its own). The key is the loop OBJECT held weakly, not id(): keying by id would leak an entry per dead loop and, since CPython reuses addresses, could hand a new loop a client bound to the closed one that used to live at that address — a fault that would only appear under the load this change exists to handle. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
The fan-out gateway lived at /v1/ocr/engine/<lease>/… and the new catalogue at /v1/ocr/engines. They do not collide, but both auth gates exempt the gateway by PREFIX — so any later route under /v1/ocr/engine* would have inherited that exemption silently, and the two names are indistinguishable at a glance in a middleware condition. The gateway is now /v1/ocr/attach/<lease>/…: internal by name, and nothing a public path family will grow into. Tests assert /v1/ocr/engines still requires ordinary authentication at both gates, and that a plausible future sibling (/v1/ocr/engine-stats) does too. Free to rename because nothing consumes the lease URL yet — it is minted per attachment at run time and never written down. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
Naming an engine an install had not enabled was a 400, and there was no way to find out beforehand — the client had to know, or guess and handle the refusal. GET /v1/ocr/engines answers it from this install's own config: every engine with `enabled` and, when it is not, a `reason` that names the setting to change (an unaccepted licence is reported as such rather than as merely disabled, because that is a different checkbox). Each row also carries `kind` — an engine owns its weights, a pipeline drives a model — the `aliases` that also resolve to it, and `returns`, so a caller that needs coordinates can see that olmOCR has none rather than discovering it in an empty field. The top level reports both what the operator wrote as the default and what it resolves to. Where the model runs is deliberately not in the answer. It is configuration, it can differ between two requests, and a client that could see it would start depending on it. A test asserts the payload mentions no pod, node or address. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
docTR called .cuda() when torch happened to see a card and otherwise simply stayed on the CPU, so a GPU-less install read every page many times slower with nothing in the response to say why — the result looked normal, it just arrived minutes late. Nothing else in coderai quietly relocates GPU work, and paddle refuses outright in exactly this situation; docTR was the odd one out. `doctr_use_gpu` is now a requirement: no CUDA device means a 503 that names the setting. Running docTR on the CPU is still fully supported and unchanged — it just has to be asked for, with doctr_use_gpu = false. The test drives the real load() with docTR and torch stubbed, rather than re-implementing the placement rules beside it; reintroducing the fallback makes it fail, which is the only reason to trust it. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
Three corrections to the fan-out, all the same mistake in different places: the code was deciding things the configuration had already decided. Nodes are no longer discovered. Asking every cluster node what it holds and using whoever answered is configuration by probe: it puts work on a machine nobody chose and moves it the moment a different node replies first. A node serves a model because the model is PINNED to it — `engine` on the entry, the same hard constraint the front router already honours, comma-separated when several nodes should share the work. A pinned node that does not hold the model is reported, not silently replaced. The decision is not an OCR decision, so it does not live in codai/ocr: any model, any caller that needs an endpoint rather than a proxy hop, gets the same answer from codai/models/placement.py. For requests that pass THROUGH coderai this was already true — the router honours the pin, treats a node as an engine by name, reaches RunPod through the backend. This is that decision for a caller that must hold an address itself, which is why OCR needed it. And "how busy is too busy" is the model's own scale_up_inflight_per_pod, on the model page, instead of the parallel ocr.serve_spill_inflight knob I had added — one number per model, already there, already meaning exactly this. A model with no RunPod block has nothing to rent, so load alone never moves its pages. Finally, the caller: resolve_engine accepted the four pipeline names only, so a client naming the model it can see in /v1/models got "unknown OCR engine". Both now work — engine=surya and engine=datalab-to/surya-ocr-2 reach the same place — because which pipeline reads a page, and where its model runs, are ours to know and not the client's. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
An OCR pipeline is handed one server address and drives it itself, which is what keeps its page images, its strict json_schema and its logprobs intact. One address cannot spread load, and three things followed from that: * a pool grows on in-flight requests counted as they pass through it. An attached engine never passed through, so the pool saw no load however hard OCR ran, never grew, and the engine had one pod's URL anyway; * several cluster nodes serving the model resolved to whichever answered first, for ever, however busy it became; * a local model configured to burst to RunPod never did: that trigger lives on the proxy path an attached engine bypasses. So the engine gets a lease on an address of ours, and codai/ocr/dispatch.py chooses per request: a pool that is the model's configured home takes every page (picking the least-loaded pod, renting when they are all at the threshold); otherwise this machine and the nodes holding the model compete on load — including the load a node reports for its whole install, so a node busy with someone else's work is left alone — and the overflow past ocr.serve_spill_inflight rents instead of queueing. An explicit `backend` pin still wins over discovery: that was a decision, and load-aware picking must not quietly repeal it. The hop is allowed to exist only because it decides nothing: body bytes in, body bytes out. The losses it is there to avoid came from PARSING a request and re-emitting it through the backend abstraction. Two gates stood in the way and neither could be satisfied the usual way: the engine has no API token, and surya's /health and /v1/models probes send no headers at all (bare httpx GETs), while refusing to attach if /health is not 200. The lease is therefore read from the URL, and the engine-only internal gate exempts that path alone — tested to be no wider than that. olmOCR's model mode asks the same dispatcher, and still takes the in-process route when the answer is local, where the model manager's loading, eviction and queueing belong. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
Surya asks every request for logprobs and averages exp(logprob) over the generated tokens to score each line. coderai accepted the parameter, never sent it, and hardcoded `logprobs: None` in every response builder. Surya's fallback is what makes that expensive: with no logprobs it does not report "unknown", it reports a literal 1.0 (recognition/__init__.py:248, :302, layout/__init__.py:97). Every line of every page comes back perfectly confident, the guesses included, and nothing downstream can tell which judicial document needs a human to look at it. So the round trip is wired where it can exist: the proxy backends (vLLM, RunPod) forward `logprobs`/`top_logprobs` and hand back what the server measured, buffered and streamed, accumulating the streamed pieces into one response-shaped object. The manager asks the backend's signature rather than negotiating by TypeError, which would have called it twice for a kwarg it never had. A request that did not ask is sent exactly as before, and a backend that cannot answer still returns null. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
Everywhere else in coderai an engine is what RUNS a model. `ocr.default_engine` borrows the word for two different kinds of thing: paddle and docTR are runners with their own weights, while surya-2 and olmOCR are pipelines over a model that vLLM runs. Only the second kind has to be told where a server is — and that is the same question the model layer already answers for everything else. So answer it in all the places a model can actually be served. A cluster node was the one that was missing: a node is a whole coderai registered with the head as a remote ENGINE, so nothing in the backend layer knows it exists, and a VLM hosted there came back as "not served over HTTP" from an install that was serving it perfectly well. The node's own catalogue is the authority, so it is asked. An explicit backend pin still wins — discovery is the fallback, not an override — and the nodes are not probed at all when clustering is off, which would have added a timeout to the first page of every document. Also: one test here was passing for the wrong reason. `from codai.models import manager` reads the package attribute and never consults sys.modules, so a fake installed there was simply not seen. Patched as attributes of the real module, which is what made the precedence test fail honestly. Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
Three silent drops on the way to a rented pod, all found while wiring Surya to the model layer, all of the same shape: coderai accepted a field, forwarded a reduced version of it, and the pod answered something plausible. * Images. The front flattens multipart content to the literal text "[image_url content]" unless the backend claims vision, and only the llama.cpp backend ever did. So every page sent to a VLM served on RunPod — olmOCR in "model" mode, or Surya through a coderai-vllm pod — arrived as that placeholder, and the model answered fluently about nothing. A proxy is the wrong place to judge: forwarded, a text-only server returns a plain 400 and a VLM server reads the page. * Structured output. Only {"type":"json_object"} survived, so a json_schema — the strict form current clients send, Surya's guided decoding among them — was dropped and the reply came back as free text the caller could not parse. The streamed path ignored the field entirely, json_object included. * extra_args. Whitelisted per engine on a model's runpod block since pods shipped, parsed, stored, and never put on a command line. Setting it looked like tuning and did nothing, which is worse than refusing it: the pod boots, serves, and ignores the decision. A page-reading model wants prefix caching and the Qwen-VL pixel bounds, and until now the only way in was docker_args, which replaces the whole generated line. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
Stefy Lanza (nextime / spora ) authored
OCR on the Digesta install answered 503: "vLLM isolated venv not found at /cache/vllm_venv". That host has no GPU at all — renting A40 pods is its whole job — so the venv it was told to build could never have run, while the Surya-2 checkpoint it needed was already a configured model (backend: runpod, engine: coderai-vllm) nothing in the OCR path could ask for. The model layer was never the problem; `ocr.surya_serve` was asking the wrong question. It enumerated server PROCESSES — torch in a local venv, our own vLLM, a URL typed in by hand — when the answer the operator had already written down was a MODEL. olmOCR has had that bridge since it shipped ("model" mode, through our own chat path); Surya never got one, because Surya2 drives its server directly with guided-JSON decoding and logprobs that a proxy hop would drop. So it gets the same bridge, in the only shape it can take: resolve the model through whichever backend owns it and hand the worker the address. Three values, because Surya refuses a server whose /v1/models disagrees with the checkpoint it expects, and a pod URL is useless without its token — which only Surya's vLLM client reads, so a bearer-locked endpoint is always attached to as "vllm". The address is a lease, not a location: a min_pods=0 pod is reaped when OCR goes quiet, and the page that finds it gone re-resolves once — renting a fresh one — instead of every later page failing against a dead URL. A failure against a server that still answers /health is a real error and is reported as one. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
-
- 10 Oct, 2026 5 commits
-
-
Stefy Lanza (nextime / spora ) authored
When zeiss ran out of memory this morning the OOM killer picked coderai-nvidia. Not because it was at fault — a runaway test process was — but because oom_score is essentially "how much memory does this task hold", and an engine with a model loaded is always near the top of that list. The leak carried on; the thing that was serving died; the box needed a hard reset. Linux offers one lever, /proc/<pid>/oom_score_adj, and two rules that decide where it can be pulled from. Both were verified on the two production hosts rather than assumed: raising is free, lowering needs CAP_SYS_RESOURCE — and a docker container does not have it (tested: in-container `echo -500` is denied, `echo 200` succeeds). So the protected base has to come from the engine daemon, via --oom-score-adj on the run, and everything inside can only step back UP from it. rootless podman cannot protect at all: a negative value is clamped to 0. 0 is still worth having on the Digesta host, where coderai currently sits at +200 — inherited from systemd's user manager, which makes it a PREFERRED victim with a kernel score of 800 against zeiss's 670. So: the container gets -500 (CODERAI_OOM_SCORE_ADJ, 0 disables), and inside it codai/util/oom.py states the order as three tiers relative to that base — engines +100, workers +300, training +500. Relative, because the base differs by host and a tier should keep its distance whatever it turns out to be. Engines are marked in the existing _engine_preexec hook, which covers the isolated model workers too: they are spawned by the engine, not the front, so they inherit it. Training is marked from the parent after Popen — a preexec_fn would run Python between fork and exec in a threaded server, which is the documented way to deadlock. The ordering is the point. Protecting everything equally would just move the bullet to an innocent bystander; this way the kernel sheds a training run (hours, restartable) before an engine (loses its in-flight work, and the front notices and respawns it) before the front (loses everything, including the ability to respawn anything). Every helper can only ever make a process MORE killable: nudge() refuses a negative delta rather than attempting a privileged write, because a process that believes it is protected and is not is worse than one that knows it is exposed. All of it is best-effort — a process that cannot adjust its score still serves. 16 tests, including that the number reaches the kernel and moves the score, that a child can be marked from the parent, and that the engine mark sits inside a try (it runs between fork and exec, so anything that raises there fails the spawn). Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019sRvvg5WSdLgcFUM96pmE4
-
Stefy Lanza (nextime / spora ) authored
zeiss came back from this morning's hard reset with no container at all, and the hourly upgrade then skipped every tick — so the host sat on old code and, worse, stayed DOWN, with nothing in the loop that would ever have recovered it. Both idle probes are HTTP, so an absent server is indistinguishable from a wrong token: inflight() said "unknown" and the fail-closed guard refused to act. The Digesta host, whose container was up, upgraded itself fourteen seconds after the push; zeiss needed hands. container_state() answers the one question the probes cannot — is there a server at all — by name, then by any running container of the image being upgraded. Only a POSITIVE "not there" counts as idle: an engine that cannot be asked (daemon down, wrong socket) stays unknown and still fails closed, because the point is to recognise an absent server, not to assume one. Checked before the token, since a host with no server is in the same position as one with no token. Both installs run their engine's ps as the service user and name the container `coderai`, so the default matches; CONTAINER_NAME covers an install that does not. The second half: every knob reads `VAR="${VAR:-default}"`, which looks like the environment may override it, but the conf is sourced first and assigns unconditionally — so `MAX_DEFERRALS=1 coderai-autoupgrade` did nothing at all, with no line in the log to say why. I hit that landing this morning's restart and had to source the conf from a wrapper to get round it. The environment is now snapshotted before the conf and restored after, for that run only, and a test checks the snapshot list against every knob the script resolves so the next one added cannot quietly fall out. require-idle and max-deferrals are also logged now: "it skipped again" and "it was told to skip" used to look identical. The conf example claimed that without a token "the restart happens whenever the timer says so". It is the opposite — the check fails closed, so a running server is then never upgraded. Corrected, and REQUIRE_IDLE_CHECK documented next to it. tests/test_autoupgrade_recovers_a_down_server.py RUNS the script against a stubbed engine, runner, restart command and health file: the env restore is eval-based and the container check branches on an exit code, neither of which reading the source can confirm. 12 new tests, 56 across the three autoupgrade files. No version bump: nothing in the image changed, and bumping would send both production installs through an image upgrade and a restart for a host-side script. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019sRvvg5WSdLgcFUM96pmE4
-
Stefy Lanza (nextime / spora ) authored
A full `pytest tests/` run reached 46 GB RSS plus 20 GB of swap on a 54 GB machine. Swap filled, the OOM killer took coderai-nvidia — which was serving — and the box needed a hard reset, which lost the last forty minutes of kern.log with it. atop's history is what identified the process: a pytest worker, carrying the `coderai-front` process title because importing the front sets it, growing 7.8 GB → 34.8 GB → 46.6 GB across three ten-minute samples. The defect is `except asyncio.CancelledError: pass` wrapped around an await whose job is to collect a CHILD task. asyncio delivers a cancellation as a CancelledError at the next await point, so that handler catches two unrelated things: the child reporting it stopped (fine to swallow) and the caller asking us to stop (never fine). run_forever's reconnect handler sat on exactly such an await, so `task.cancel()` was absorbed and the loop carried on. With the reconnect delay mocked to zero by the test driving it, that became a busy loop allocating ~78 MB/s — AsyncMock records every call it receives — and 400,000 iterations in the five seconds I measured. being_cancelled() (Task.cancelling(), 3.11+) tells the two apart, and all four sites now re-raise ours: _stop_heartbeat_task, _cancel_inflight_tasks, the two keepalive collects in handle_message, and BrokerService.stop. The inflight ones matter in production independently of the leak — absorbing a cancel there means a shutdown waits out a whole video generation instead of stopping it, and run_forever simply reconnected forever instead of exiting. run_forever's own cancellation path clears the cancel for the duration of its cleanup and re-raises at the end, so honouring the cancel does not cost the websocket close. The test that drove it now brakes itself: the zero-delay sleep stand-in counts, and raises a BaseException (anything else is caught by the reconnect handler) well past the two reconnects it needs. With the product bug reintroduced the file fails in 1.5 s at 441 MB instead of eating the machine. tests/conftest.py adds what was missing in general: an RSS ceiling (6 GB, CODERAI_TEST_RSS_LIMIT_MB) and a per-test stall ceiling (180 s, CODERAI_TEST_STALL_SECONDS), both dumping every thread's stack to a file before exiting — pytest captures stderr per test and drops the buffer when the process does not unwind, so a watchdog that only printed would have died silently. No new dependency. "A test can consume the host" is a property of the run, not of one test. tests/test_broker_protocol.py could not finish before this — it hung at the 14th test. The whole suite now completes for the first time: 78 s, 1.57 GB peak, 1635 passed. The 25 failures are the 17 that already failed plus the 8 stale register-ack expectations in that file, which were unreachable behind the hang and are a separate question (the fixtures send `accepted: true`; the client has required `status: "ok"` for some time). Co-Authored-By:
Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019sRvvg5WSdLgcFUM96pmE4
-
Stefy Lanza (nextime / spora ) authored
Township's clip lengths were tuned when the only video model could hold ~81 frames in one call, so a 70-second match became a dozen short cuts. LongCat was pretrained on continuation and the server renders a long shot in segments, so the same match wants a handful of long takes — and the frame numbers that suit one are wrong for the other by a factor of four. Only the two starter templates knew that, and only because someone typed it in. Worse, nothing snapped a picked count onto what the model can emit: random.randint handed the server lengths like 250, which is not 4n+1 for any VAE and not on LongCat's 93+80k segment grid, so the server rounded and the plan's own numbers became fiction. On LongCat that also silently costs block-sparse attention, because every point of 93+80k is a 16n+13 count BSA can tile and nothing between them is. Server — codai/models/video_geometry.py is the geometry once: native rate, the legal frame grid, how much ONE render holds, whether continuation is native, and the side multiple below which the fast attention path is lost. Published per model on /v1/models (ModelInfo.video) and used by the render path itself, so what the server advertises and what it rounds to cannot drift. The numbers were spread across three modules that could not tell a client; the literals here are pinned to each source by tests, since the LongCat and H3 workers run in 3.10 venvs this cannot import. models.json overrides every field. Township — the five frame fields now default to 0 = AUTO: derive the budget from the model's geometry and this tool's seconds band, then snap both ends and every pick onto the model's grid. Seconds, not frames, because clip count falls out of long_target / (frames/fps). An explicit number still wins, so a saved config keeps exactly the budget it had. Entrances and the face-off get their own shorter band instead of borrowing the fight budget (on LongCat that made every entrance an 11-22s take). Outcome shares are snapped per shot and the total rewritten, the per-render chunk comes from the model (Wan keeps its conservative 50), and the chaining decision reads the published continuation rather than matching the model's name. At the defaults a 70s match goes from 9-11 Wan cuts to 4-6 LongCat takes, each one request with native continuation, with BSA still on. 1607 tests pass; the 17 failures and the test_broker_protocol hang are pre-existing and unrelated (verified with the diff stashed). Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
An i2v generation died in the first denoising step with a bare AssertionError out of flash_attn_bsa_3d, on a geometry the guard had just declared valid: 93 frames, sides divisible by 64. bsa_problems only ever checked the noise branch. A conditioning pass does not tile the latent once — Attention.forward splits the sequence into a conditioning block and a noise block and calls flash_attn_bsa_3d on each (modules/attention.py:123-134), so the cond depth and the remaining noise depth must divide by the chunk as well. generate_i2v hardcodes one conditioning latent (pipeline:796) and 1 % 4 != 0, which means block-sparse attention has never once been able to run an i2v pass, at any frame count or resolution. The avatar pipeline hardcodes the same 1 in four places, and generate_vc computes a count it never rounds. Only generate_refine gets it right, by padding both counts to the granularity itself (pipeline:1245-1250). (a) The guard now knows about the split. _cond_latent_passes enumerates every conditioning count the request can produce — segment 0 uses the task's own entry point and every later segment uses generate_vc, so even t2v runs a conditioning pass once num_segments > 1, and BSA is one flag for the whole request — and a count that cannot tile puts the pass on dense attention with a line saying which condition failed. (b) bsa_pad_cond, OFF by default, rounds the count up to the granularity on the way into the DiT, which is what generate_refine does for itself, so an i2v or continuation pass can keep BSA. Patched onto the DiT instance rather than the vendored source, the same way the denoising bars are rebound: the count is born inside generate_i2v's loop, so there is no outer argument to correct. Off by default because the pipeline marks only the real conditioning frames clean (timestep[:, :1] = 0), so rounding 1 up to 4 hands the conditioning block three latent frames that are still noise — well-defined attention, but not what the model was trained on, and on a video model that kind of error shows up as temporal drift. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
- 09 Oct, 2026 15 commits
-
-
Stefy Lanza (nextime / spora ) authored
Probed every API surface coderai can use for attaching a global volume, each with a body valid enough that the extra-key check actually runs: POST /v1/pods key … not in input schema: 'globalNetworkVolumeId' podFindAndDeployOnDemand field not defined by the input type POST /v2/pods additional properties 'globalVolumeId' not allowed templates (mounts) every candidate key refused There is no public API for it: global volumes are console-only, as the docs say. Which makes `global_volume_id` a loaded gun — sending the field fails the create, so configuring a global volume would have stopped the install renting pods at all. Both create paths now recognise that specific rejection, drop the field and retry. The pod boots without the volume and downloads its weights: slower and it costs bandwidth, and far better than serving nothing. Logged once per process, naming the consequence and the override to set when RunPod publishes the field name. An unrelated failure is NOT treated this way — a capacity error retried with the volume stripped would bring a pod up silently wrong. Earlier probes that reported the field as "accepted" were wrong: with an otherwise-invalid body the validators stop before the extra-key check. Worth remembering next time I probe a schema. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
`gpu_type` was a filter: "NVIDIA A40" meant an A40 or nothing, so when RunPod had no A40 free the request failed rather than taking a card that would have served it. Now `gpu_types` is an ordered PREFERENCE and `allow_other_gpus` lets the search fall through to anything else that still satisfies min_vram_gb and max_hourly_usd. A preferred card ranks ahead of every fallback whatever selection_criteria says — otherwise "cheaper" would put the fallback first and naming a card would mean nothing. Each ranked option carries `preferred`, and provisioning a card nobody asked for prints "[FALLBACK — no preferred card available]": correct behaviour, but never silent. Off by default. A pinned card is sometimes pinned for a reason, and an install that quietly started renting something else would be a surprise on the invoice. Name matching is exact on the id or display name, case-insensitive, tolerating a missing "NVIDIA " prefix, and deliberately NOT a substring: "A40" would otherwise match "NVIDIA RTX A4000" — 16 GB for a job that asked for 48. GPU names split on commas only, because "NVIDIA RTX A6000" has spaces in it. Also: the admin handler rebuilds the runpod block from a whitelist, so all six keys added today (global_volume_id, data_centers, gpu_types, allow_other_gpus, allow_client_scale, client_scale_ttl_s) were being dropped the next time anyone saved a model from the GUI — taking the warm schedule and the volume with them. They are in the whitelist now, with the two ordered lists kept in order. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
Weights in one place, pods wherever a card is free. `global_volume_id` attaches a RunPod GLOBAL volume (beta since Sept 2026): region-independent storage, so unlike `network_volume_id` it does NOT pin the pod's data centre — which is the only reason to use one. Resolved by its own function precisely so the two cannot be confused: the region lookup stays under the network-volume branch, and a test holds them apart. `data_centers` is the other half. A single `data_center` pins one region and a blank one lets RunPod pick any region on earth, so there was no way to say "anywhere in the EU" — which is what data that may not leave the EU requires. The list is tried in order, and each (card, region) pair becomes its own candidate so the capacity fallback that already walks GPU options walks regions too, with no new control flow. The pod handle records the region it actually got, not the pool's first choice. Two honest limits, both documented in the code and the guide. A global volume must be created in the RunPod CONSOLE: POST /v2/network-volumes allows only STANDARD|HIGH_PERFORMANCE and demands a dataCenter. And attaching is console-documented only — no field exists in the Pod API reference, and the v1 REST validator accepts unknown body keys silently, so a wrong name cannot be caught by probing; it yields a pod with no volume that re-downloads its weights. We send globalNetworkVolumeId, overridable by CODERAI_RUNPOD_GLOBAL_VOLUME_FIELD, and it must be verified on a real pod. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
The root cause under the last three fixes. The pod registry recorded os.getpid() and both readers skipped an entry whose pid matched ours as "our own, the local pool handles it". Inside a container the engine comes up as the SAME pid on every boot — 74, on the production host — so after a restart the previous container's pod read as ours and was skipped by both find_shared_pod (nothing to adopt) and sibling_is_provisioning (nothing booting). The warm path then rented another A40 beside the one already running. Three restarts this afternoon, three extra A40s, each billing until it was terminated by hand. Entries now carry a per-process nonce and identity is that. The pid is still recorded, and an entry written before the nonce falls back to comparing it, which is the old behaviour and no worse. test_a_surviving_pods_token_is_remembered_across_a_restart simulated the new process by monkeypatching os.getpid — the very thing that does not change. It patches the nonce now, which is why it failed on this commit and passes on it. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
The last hole in the warm path. sibling_is_provisioning already recognises a pod registered at creation with no URL yet — and after a restart the recorded pid no longer matches, so our own booting pod qualifies — but ensure_ready never asked. Seen live: a pod created at 18:57 was six minutes into its image pull when the 19:03 upgrade restarted the service, and the warm path rented a second A40 beside it. The upgrade runs hourly and an image pull is ~6 minutes, so that is not a rare alignment. Order in ensure_ready is now: a healthy pod of our own, else adopt a ready one, else wait for one that is booting, else rent. Also prune the pod registry in reap_orphans: an id RunPod no longer has stayed "known", which both shielded it from reaping and let the file grow for the life of the install — including pods terminated out of band, of which this afternoon produced five. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
A warm floor is re-established at every boot, and ensure_ready() went straight to _provision_one(). The adoption logic — "a pod for this pool is already running, take it over" — existed only inside acquire(), so the warm path never saw it. Three restarts on the production orchestrator this afternoon rented three A40s and left each previous one billing until the reaper's 20-minute grace expired. An hourly unattended upgrade would have done that every hour. find_shared_pod already covers the case: after a restart the recorded pid no longer matches, so our own previous pod is adoptable. It is now one method, _adopt_shared_pod, called from acquire(), ensure_ready() and maintain()'s warm top-up. It also stops lying about the cost. An adopted pod was recorded at $0/hr as "billed by its owner", which is true of a sibling engine's pod and false of one we are inheriting from our own restart — a real A40 hidden from the $/hr cap and from the spend page. register_pod now stores the rate and adoption restores it; an older registry entry without one still adopts, at 0. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
NameError in _provision_one: it read `dc`, a local computed inside _create_with_fallback, from a function that has no such name. The sequence that costs money is this — create_pod succeeds, the pod boots on RunPod, the PodHandle raises, the pool never records the pod, the caller provisions again. Four A40s existed for a warm floor of one before anyone looked. Both paths now call pool._data_center(), which keeps the same precedence (the volume's region, then the model's, then the account's) in one place. The two tests that covered this asserted the broken expression as a SOURCE STRING — 'data_center=(getattr(self, "_volume_dc", "") or dc or "")' — so they passed on a NameError and would have gone on passing. They now call the resolver, including with nothing configured anywhere, which is the shape that raised. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
Found while configuring the Digesta catalogue: the launcher reads /root/.coderai/config.json (engine_specs, published port) while the front and the engine read /root/.coderai/coderai/config.json. The RunPod api_key and the spend caps were in the first one, so the engine refused to provision anything and the admin page reported no caps at all. Records what was moved where, and that CODERAI_CONFIG_DIR=/config remains the fix that makes it unnecessary. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
Watched it happen on the production orchestrator: configuring qwen38-awq warm during office hours provisioned TWO A40s a second apart. A pod does not enter self.pods until it answers its health probe, which is minutes away. ensure_ready() never set the _provisioning flag, and the warm top-up in maintain() never read it — so the 15 s scaler saw "0 healthy, floor 1" while a pod was already booting and rented another. It stopped at two only because _provision_one blocks the scaler thread for the whole boot; a non-blocking provision would have added one every 15 seconds. The demand path (_should_grow) has always checked the flag. Now all three agree: one on the way is one pod, not none, including one a sibling engine sharing the pool is starting. Also: /v1/runpod/scale answered with the internal pool key ("pool:digesta-qwen38") rather than the model the client asked about. Several models share a pool, and that string is not something the caller sent or can look up — so the answer names the model and reports the pool beside it. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
Configuring the Digesta catalogue on the production orchestrator turned up three faults in warm_configured_pools, all of which spend money: * `enabled: false` was not consulted. A catalogue entry parked for later with keep_warm still set rented a GPU at every boot. * the entry was named `alias or path`. A RunPod-only entry has neither — no weights live on that machine, and an alias only exists when a second name is wanted — so it warmed under the key "" and the pool was never the one that went on to serve requests. `id` is now the fallback, and an entry with no name at all says so instead of warming nothing quietly. * the schedule was ignored. ensure_ready() is the "a request is waiting" path and provisions unconditionally; at boot nothing is waiting, so a restart at 02:00 on a Sunday rented exactly the pod the window exists to avoid. It now checks effective_min_pods first, which also means a client lease that outlives a restart is honoured rather than dropped. Found on the host, not by reading: the warm pod the configuration asked for never appeared. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
The autoscaler reacts to load — a pod is added once the least-loaded one already has scale_up_inflight_per_pod requests in flight. That is right for traffic nobody predicted and wrong for a client that KNOWS what is coming: a batch of a thousand documents otherwise discovers the pool one cold start at a time. So a client may now say so: POST /v1/runpod/scale {model, pods}. It is a lease, bounded three ways, because a client that can raise the floor can raise the bill: * allow_client_scale, off by default, so no model is client-scalable unless someone said so; * clamped to max_pods — asking for more returns the ceiling with "clamped": true rather than an error, since the client asked for "as many as you can" and refusing would leave it with none; * client_scale_ttl_s (default 30 min), because the failure that matters is not a wrong number, it is a client that asks for four A40s and then dies. A lease that never expired would be a bill that never stopped, so a 0 in the config falls back to the default instead of meaning "forever". The lease and a warm-pod schedule are MAXED, never summed: both mean "hold this many ready", so a client cannot lower a configured floor and a closed window cannot cancel a lease. /v1/runpod/status reports the live lease per model, so an unexpected bill has a visible cause. Not admin-scoped, unlike /v1/runpod/spend: the gate is the per-model flag, not the key, and the account's $/hr and spend caps still refuse to provision. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
Proven on zeiss with a real LoRA training run: the skip fired correctly but read "SKIP: serving 1 request(s) [engines=0 jobs=1]". There was no request -- that is the entire case this check exists for, since /v1/loras/train with wait=false returns at once and the work continues for hours. The log now says "work in progress", which is what it means. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
The unattended upgrade measured idleness as requests in flight, and a LoRA training run has none: the request that starts it returns immediately and the work continues for hours, so coderai_engine_inflight is 0 throughout. An hourly upgrade would have restarted straight through one -- the standing "never restart during training" rule, reintroduced by automation. Nothing was missing in the training path. loras.py already registers a `training` task, the engine already ships active tasks in /internal/engine-state, and the front already stores them on EngineEntry.tasks. The gap was that /metrics never exposed them and the probe never looked, so this adds coderai_engine_jobs_active per engine and coderai_jobs_active across them, counting running/queued/paused -- paused on purpose, because a thermally paused generation has not finished and restarting loses it. The probe now reads both numbers from one scrape and takes the MAXIMUM, never the sum: a generation is both an in-flight request and a running task, so adding them would double-count and the numbers in the log would match nothing an operator can see. A front that predates the gauge reports "-" and the detail line says "front too old", rather than claiming zero and silently restoring the blind spot. tests/test_autoupgrade_sees_training.py renders the exposition from a fake engine carrying a training task, runs the probe's awk program as the script actually has it, and pins both halves of the premise -- the registration in loras.py and the engine's reporting of active states -- so the gauge cannot go quiet without a test failing. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
run_oci.sh takes its engine from CONTAINER_ENGINE and defaults to docker, so on the rootless-podman host this script was written for it would have looked for a docker socket that is not there -- failing on every hourly tick with nothing to show but a non-zero exit. Found while installing it on the Digesta production server, before it ran. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-
Stefy Lanza (nextime / spora ) authored
The Digesta host added both 10.89.0.178 and ::ffff:10.89.0.178 to trusted_proxies and host->8777 still refused. The ::ffff: hunch was right: the internal nginx listens on 8776 AND [::]:8776, so a connection landing on the v6 socket makes $remote_addr -- and therefore X-Forwarded-For, and therefore the peer -- the IPv4-mapped form, which a string membership test rejects even when the plain address is listed. Nothing in the configuration explains that refusal. Peers are now compared as addresses: a mapped v6 peer matches a plain v4 entry and the reverse, and CIDR entries work, so a container network can be trusted as 10.89.0.0/24 instead of pinning one address in the quadlet. A non-address entry (a socket path) is still compared literally, and a v6 peer never matches a v4 network. Their broadened list could not have taken effect anyway -- it went into the config.json the app is not reading, which is why `enabled` there did nothing and the env var was needed. So that test did not rule the peer check out; the list in force was still the default pair. The rest of the report is answered in the doc rather than in code, because there is no bug behind it: `Server: nginx` appears on PROXIED responses too unless proxy_pass_header Server is set, and the missing `GET /admin` access-log line is _PollNoiseFilter, which drops reads of /admin, /static, /login and /logout from the front's access log while leaving /healthz alone -- exactly the asymmetry that looked like nginx short-circuiting by source. A 302 to /coderai/login with no Set-Cookie is what the front returns when forward-auth declines. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com>
-