Rolling reboots across all OSC clusters beginning 09/16/2026

Resolution: 
Unresolved
  • ​​​​​​Date: Wednesday, September 16, 2026

  • Affected Systems: All Clusters (Pitzer, Ascend, Cardinal)

  • Description: OSC systems staff will execute a rolling reboot across all active compute clusters (Pitzer, Ascend, and Cardinal) beginning Wednesday, September 16. Login nodes will be rebooted first, followed by compute nodes. Nodes will be systematically set to a drained state prior to rebooting to prevent a full cluster outage and minimize workflow disruption. As part of this reboot window, staff will upgrade NFS mounts to NFSv4.2 to resolve recent login node stability issues.

  • Impact: The Slurm batch scheduler will automatically hold pending jobs that cannot complete before a target node's scheduled reboot. Held jobs will retain their queue positions and resume scheduling as nodes return to service. Users should expect temporary increases in queue wait times while capacity is reduced. Active jobs on nodes not yet scheduled for reboot will continue running normally.