Known Issues

Title Cat. Res.sort ascending Description Post Upd.
Unscheduled GPFS Outage filesystem Resolved

As of 11:30PM on June 16th, we have removed the GPFS filesystem from service due to a number of hardware failures. At this point, further hardware failures would put a large portion of the entire... (Read more)

2 weeks 1 day ago 1 week 3 days
Oakley login node instability Operations Resolved

Oakley login nodes are seeing some instability related to Lustre. We will reboot the nodes on Thursday, October 2nd 2014 to resolve the issue. If a login node crashes before then and we have the... (Read more)

9 months 1 week ago 8 months 5 days
Statewide Intel compiler license checkout failures Licensing Resolved

This morning (9/10/14) we updated our Intel compiler licenses. We are seeing some unexpected license checkout failures in the logs (please click through to see details):

10:44:... (Read more)          
9 months 3 weeks ago 9 months 1 week
Lustre Updates filesystem Resolved

9/10/14 - We have not seen any additional crashes of the Lustre servers since making this change.

8/26/14 
- Lustre jobs are being accepted as of 10AM this... (Read more)

10 months 1 week ago 9 months 3 weeks
Armstrong offline until Noon Armstrong Resolved

Armstrong will need to be taken down today until Noon.  In the meantime, contact OSCHelp (OSCHelp@osc.edu) for account assistance.

10 months 2 weeks ago 10 months 2 weeks
Lustre jobs suspended filesystem Resolved

The Lustre filesystem ($PFSDIR and /fs/lustre) has crashed several times Friday evening (8/15). We have degraded this service temporarily, while we work to isolate the actions that are triggering... (Read more)

10 months 2 weeks ago 10 months 1 week
issue with OnDemand 6:09 - 8:39 pm Resolved

OnDemand, epi accounting queries, the Viper DB, the Medline DB, the Eweld DB,... (Read more)

10 months 3 weeks ago 10 months 3 weeks
my.osc.edu logins failing Account Management Resolved

Logins to my.osc.edu are failing. This is unrelated to our InfiniBand issue; a router change at OARnet is the believed cause. They are working on re-establishing the necessary routing.

11 months 5 days ago 11 months 5 days
Lustre, Infiniband Operational and Being Monitored Closely filesystem Resolved

UPDATE: Most users should no longer see any issues with Lustre.


Again, please continue to notify OSC Help of any errors you see in job output. For example, you might see "... (Read more)

11 months 5 days ago 10 months 3 weeks
Emergency InfiniBand Shutdown (All systems) Network Resolved

We have returned to service. It appears that we have resolved the networking issues enough to allow jobs to run safely. We will continue working with our vendors to fix any remaining hardware... (Read more)

11 months 1 week ago 11 months 5 days

Pages