• Stefy Lanza (nextime / spora )'s avatar
    Resolve max_tokens per-model + defaults across rotations and autoselect; bump to 0.99.82 · 5dff11f4
    Stefy Lanza (nextime / spora ) authored
    Extend max output token resolution so per-model and default values are
    honored at every layer when the client omits max_tokens.
    
    Effective priority (client value always wins if present):
      autoselect per-model -> rotation per-model -> provider per-model
      -> provider default -> rotation default -> autoselect default
    
    - rotation path: consult the selected provider's config (new
      RotationHandler._get_provider_config) so a provider default_max_tokens
      applies even though rotations can't configure max_tokens themselves
    - AutoselectModelInfo.max_tokens: new per-model override field
    - AutoselectHandler._apply_autoselect_max_tokens applied in both the
      streaming and non-streaming dispatch: per-model override set directly
      (highest), autoselect default threaded via _autoselect_default_max_tokens
      as a lowest-priority fallback
    - rotation and both RequestHandler injections consume the threaded
      autoselect default, covering autoselect->rotation and
      autoselect->provider/model dispatch routes
    Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
    5dff11f4
pyproject.toml 2.12 KB