-
Stefy Lanza (nextime / spora ) authored
Two bugs made codex_think advertise a context window it was not configured with, so clients that size their context from the model listing refused to run against it. 1. The endpoint-level model cache was keyed on type+endpoint alone. codex_think, openai_think and bigscreen are all codex:https://api.openai.com/v1 but authenticate as different accounts, so whichever prefetched first populated the shared entry and the others served its model list -- an OAuth ChatGPT provider inherited an API-key provider's generic OpenAI models, contexts and all. Same collision across the three kilo-* providers. Key the entry on a digest of the provider's credentials too. Also fix invalidate_provider_cache(), which tried to find a provider's endpoint entry with a substring match on a key that never contains the provider id. 2. Provider models were published exactly as the upstream API returned them, so default_context_size / default_max_request_tokens never reached the listing. Stamp them on, letting explicit configuration override the fetched value. _configured_context_size() deliberately does not call get_context_config_for_model(): that helper ends in _infer_context_size_from_model(), whose generic 8192 fallback is right for sizing a request but would overwrite a real fetched window (272000) with a guess when published. Bump version to 0.99.89. Co-Authored-By:
Claude Opus 4.8 (1M context) <noreply@anthropic.com>
b1320e45