• Stefy Lanza (nextime / spora )'s avatar
    runpod: a pod is ready when the model is loaded, not when the port opens (v0.3.20) · 74aca6e1
    Stefy Lanza (nextime / spora ) authored
    /v1/models answers the moment a vLLM process starts — minutes before it can
    serve, because the checkpoint still has to be downloaded (every pod
    re-downloads it without a network volume) and the KV cache built on first use.
    The pool called that ready, every client read it as usable, and the whole model
    load landed inside somebody's first request, where it is indistinguishable from
    a hang.
    
    That is what cost today. One OCR page took 807s end to end — ~400s image pull,
    ~400s weights — while the client timed out at 180s and then 300s and concluded
    the serving was broken. It was not: nobody had ever waited long enough, and the
    only reason we know is a probe run with a 1400s timeout, which came back 200
    with text and conf 0.9019.
    
    So the pool now sends one token against the served model before calling the pod
    ready. The pod bills through the load either way; this only decides whether the
    wait is visible as a pod that is not ready yet, or hidden inside a request that
    looks stuck. `pods_ready` becomes a signal a client can act on, which is what
    every client already assumed it was.
    
    A failed warm-up is reported and the pod used anyway: it may still serve (an
    engine that takes no chat completions, a model whose warm-up shape we guessed
    wrong), and rejecting a pod we have already paid to boot over a diagnostic
    request would be worse than the hidden latency this removes. Off with
    warmup_on_boot=false or warmup_timeout_s=0.
    Co-Authored-By: 's avatarClaude Opus 5 <noreply@anthropic.com>
    Claude-Session: https://claude.ai/code/session_01PxYHWbCAjjjD7A1rhFKtaW
    74aca6e1
Name
Last commit
Last update
codai Loading commit data...
docs Loading commit data...
packaging Loading commit data...
samples Loading commit data...
tests Loading commit data...
third_party Loading commit data...
tools Loading commit data...
.dockerignore Loading commit data...
.gitignore Loading commit data...
AI.PROMPT Loading commit data...
CODERAI_API_DOCUMENTATION.md Loading commit data...
CoderAI.gif Loading commit data...
DISTRIBUTION.md Loading commit data...
LICENSE.md Loading commit data...
MULTIMODAL_CAPABILITIES.md Loading commit data...
MULTIMODAL_UI_EXAMPLES.md Loading commit data...
README.md Loading commit data...
REQUEST-geoclip-embeddings.md Loading commit data...
REQUEST-vpr-embeddings.md Loading commit data...
amdgpu-coredump-05000.bin Loading commit data...
build-oci.sh Loading commit data...
build.ps1 Loading commit data...
build.sh Loading commit data...
coderai Loading commit data...
coderai-broker-implementation-reference.md Loading commit data...
coderai-integration.md Loading commit data...
commands Loading commit data...
osxbuild.sh Loading commit data...
package-oci.sh Loading commit data...
package-tarball.sh Loading commit data...
requirements-crisperwhisper.txt Loading commit data...
requirements-h3.txt Loading commit data...
requirements-longcat-train.txt Loading commit data...
requirements-longcat.txt Loading commit data...
requirements-melotts.txt Loading commit data...
requirements-nemo.txt Loading commit data...
requirements-nvidia.txt Loading commit data...
requirements-ocr-paddle.txt Loading commit data...
requirements-ocr.txt Loading commit data...
requirements-pyannote.txt Loading commit data...
requirements-surya.txt Loading commit data...
requirements-vllm.txt Loading commit data...
requirements-vulkan.txt Loading commit data...
requirements.txt Loading commit data...
run-oci.sh Loading commit data...
smoke-test-oci.sh Loading commit data...
todo.md Loading commit data...