-
Stefy Lanza (nextime / spora ) authored
Extend max output token resolution so per-model and default values are honored at every layer when the client omits max_tokens. Effective priority (client value always wins if present): autoselect per-model -> rotation per-model -> provider per-model -> provider default -> rotation default -> autoselect default - rotation path: consult the selected provider's config (new RotationHandler._get_provider_config) so a provider default_max_tokens applies even though rotations can't configure max_tokens themselves - AutoselectModelInfo.max_tokens: new per-model override field - AutoselectHandler._apply_autoselect_max_tokens applied in both the streaming and non-streaming dispatch: per-model override set directly (highest), autoselect default threaded via _autoselect_default_max_tokens as a lowest-priority fallback - rotation and both RequestHandler injections consume the threaded autoselect default, covering autoselect->rotation and autoselect->provider/model dispatch routes Co-Authored-By:Claude Opus 4.8 <noreply@anthropic.com>
5dff11f4