• Stefy Lanza (nextime / spora )'s avatar
    Advertise a provider's configured context window in /v1/models · b1320e45
    Stefy Lanza (nextime / spora ) authored
    Two bugs made codex_think advertise a context window it was not
    configured with, so clients that size their context from the model
    listing refused to run against it.
    
    1. The endpoint-level model cache was keyed on type+endpoint alone.
       codex_think, openai_think and bigscreen are all
       codex:https://api.openai.com/v1 but authenticate as different
       accounts, so whichever prefetched first populated the shared entry
       and the others served its model list -- an OAuth ChatGPT provider
       inherited an API-key provider's generic OpenAI models, contexts and
       all. Same collision across the three kilo-* providers. Key the entry
       on a digest of the provider's credentials too. Also fix
       invalidate_provider_cache(), which tried to find a provider's
       endpoint entry with a substring match on a key that never contains
       the provider id.
    
    2. Provider models were published exactly as the upstream API returned
       them, so default_context_size / default_max_request_tokens never
       reached the listing. Stamp them on, letting explicit configuration
       override the fetched value.
    
    _configured_context_size() deliberately does not call
    get_context_config_for_model(): that helper ends in
    _infer_context_size_from_model(), whose generic 8192 fallback is right
    for sizing a request but would overwrite a real fetched window (272000)
    with a guess when published.
    
    Bump version to 0.99.89.
    Co-Authored-By: 's avatarClaude Opus 4.8 (1M context) <noreply@anthropic.com>
    b1320e45
pyproject.toml 2.12 KB