-
Stefy Lanza (nextime / spora ) authored
After a GPU reset the amdgpu SMU can return a constant garbage reading (observed: 511°C = 0x1FF invalid-ADC, with nonsense voltage/fan values). The thermal supervisor believed it, paused the radeon engine and escalated to SIGSTOP — freezing a healthy card forever. Both readers (engine-side thermal.py and the front's supervisor loop) now discard readings ≥150°C with a one-time warning naming the broken sensor; thermal protection for that card resumes when it reads sane values. Co-Authored-By:
Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014S8VtAvG499SsCbeESRK7V
3c9536e1
| Name |
Last commit
|
Last update |
|---|---|---|
| .. | ||
| admin | ||
| api | ||
| backends | ||
| broker | ||
| frontproxy | ||
| models | ||
| openai | ||
| pydantic | ||
| queue | ||
| tasks | ||
| __init__.py | ||
| cli.py | ||
| config.py | ||
| main.py | ||
| platform_paths.py | ||
| system_app.py |