runpod: a pod is ready when the model is loaded, not when the port opens (v0.3.20)
/v1/models answers the moment a vLLM process starts — minutes before it can serve, because the checkpoint still has to be downloaded (every pod re-downloads it without a network volume) and the KV cache built on first use. The pool called that ready, every client read it as usable, and the whole model load landed inside somebody's first request, where it is indistinguishable from a hang. That is what cost today. One OCR page took 807s end to end — ~400s image pull, ~400s weights — while the client timed out at 180s and then 300s and concluded the serving was broken. It was not: nobody had ever waited long enough, and the only reason we know is a probe run with a 1400s timeout, which came back 200 with text and conf 0.9019. So the pool now sends one token against the served model before calling the pod ready. The pod bills through the load either way; this only decides whether the wait is visible as a pod that is not ready yet, or hidden inside a request that looks stuck. `pods_ready` becomes a signal a client can act on, which is what every client already assumed it was. A failed warm-up is reported and the pod used anyway: it may still serve (an engine that takes no chat completions, a model whose warm-up shape we guessed wrong), and rejecting a pod we have already paid to boot over a diagnostic request would be worse than the hidden latency this removes. Off with warmup_on_boot=false or warmup_timeout_s=0. Co-Authored-By:Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
Showing
tests/test_runpod_warmup.py
0 → 100644
Please
register
or
sign in
to comment