-
Stefy Lanza (nextime / spora ) authored
Add a "tuning ds4 for performance" note to the global ds4 Settings section and a condensed inline note to the per-model ds4 streaming section on the Models page: NVMe/SSD placement (≈10× prefill), capping the expert cache by count, sizing the VRAM reserve, DS4_CUDA_WEIGHT_ARENA_CHUNK_MB=512, avoiding DS4_CUDA_WEIGHT_CACHE, and that decode of a model larger than VRAM is streaming-bound (smaller quant is the real fix). Distilled from measured tuning on the 154GB DeepSeek-V4 MoE. Co-Authored-By:Claude Opus 4.8 <noreply@anthropic.com>
2b2bdc72
| Name |
Last commit
|
Last update |
|---|---|---|
| .. | ||
| archive.html | ||
| base.html | ||
| change_password.html | ||
| chat.html | ||
| dashboard.html | ||
| login.html | ||
| models.html | ||
| settings.html | ||
| tasks.html | ||
| tokens.html | ||
| users.html |