• Stefy Lanza (nextime / spora )'s avatar
    ui: document ds4 performance setup in global + per-model settings · 2b2bdc72
    Stefy Lanza (nextime / spora ) authored
    Add a "tuning ds4 for performance" note to the global ds4 Settings section and a
    condensed inline note to the per-model ds4 streaming section on the Models page:
    NVMe/SSD placement (≈10× prefill), capping the expert cache by count, sizing the
    VRAM reserve, DS4_CUDA_WEIGHT_ARENA_CHUNK_MB=512, avoiding DS4_CUDA_WEIGHT_CACHE,
    and that decode of a model larger than VRAM is streaming-bound (smaller quant is
    the real fix). Distilled from measured tuning on the 154GB DeepSeek-V4 MoE.
    Co-Authored-By: 's avatarClaude Opus 4.8 <noreply@anthropic.com>
    2b2bdc72
Name
Last commit
Last update
..
admin Loading commit data...
api Loading commit data...
backends Loading commit data...
broker Loading commit data...
frontproxy Loading commit data...
models Loading commit data...
openai Loading commit data...
pydantic Loading commit data...
queue Loading commit data...
tasks Loading commit data...
__init__.py Loading commit data...
cli.py Loading commit data...
config.py Loading commit data...
main.py Loading commit data...
platform_paths.py Loading commit data...