Known issues

Unresolved known issues

Known issue with an Unresolved Resolution state is an active problem under investigation; a temporary workaround may be available.

Resolved known issues

A known issue with a Resolved (workaround) Resolution state is an ongoing problem; a permanent workaround is available which may include using different software or hardware.

A known issue with Resolved Resolution state has been corrected.

Known Issues

Title Category Resolution Description Posted Updatedsort ascending
Statewide Intel compiler license checkout failures Licensing Resolved

This morning (9/10/14) we updated our Intel compiler licenses. We are seeing some unexpected license checkout failures in the logs (please click through to see details):

10:44:... Read more          
10 years 1 month ago 10 years 2 weeks ago
Lustre Updates filesystem Resolved

9/10/14 - We have not seen any additional crashes of the Lustre servers since making this change.

8/26/14 
- Lustre jobs are being accepted as of 10AM this... Read more

10 years 1 month ago 10 years 1 month ago
Lustre jobs suspended filesystem Resolved

The Lustre filesystem ($PFSDIR and /fs/lustre) has crashed several times Friday evening (8/15). We have degraded this service temporarily, while we work to isolate the actions that are triggering... Read more

10 years 1 month ago 10 years 1 month ago
Armstrong offline until Noon Resolved

Armstrong will need to be taken down today until Noon.  In the meantime, contact OSCHelp (OSCHelp@osc.edu) for account assistance.

10 years 1 month ago 10 years 1 month ago
Lustre, Infiniband Operational and Being Monitored Closely filesystem Resolved

UPDATE: Most users should no longer see any issues with Lustre.


Again, please continue to notify OSC Help of any errors you see in job output. For example, you might see "... Read more

10 years 2 months ago 10 years 2 months ago
issue with OnDemand 6:09 - 8:39 pm Resolved

OnDemand, epi accounting queries, the Viper DB, the Medline DB, the Eweld DB,... Read more

10 years 2 months ago 10 years 2 months ago
my.osc.edu logins failing Account Management Resolved

Logins to my.osc.edu are failing. This is unrelated to our InfiniBand issue; a router change at OARnet is the believed cause. They are working on re-establishing the necessary routing.

10 years 2 months ago 10 years 2 months ago
Emergency InfiniBand Shutdown (All systems) Network Resolved

We have returned to service. It appears that we have resolved the networking issues enough to allow jobs to run safely. We will continue working with our vendors to fix any remaining hardware... Read more

10 years 2 months ago 10 years 2 months ago
Lustre is still offline. HPC systems back up Maintenance Resolved

Day One of the scheduled downtime has been completed, and HPC operations have resumed. As planned, Lustre work will extend into Day Two. Jobs using /fs/lustre or $PFSDIR cannot run until this work... Read more

10 years 3 months ago 10 years 3 months ago
Certain modules not accessible Software Resolved

Certain modules are not working for all clusters since the downtime.  We have reports specifically that Amber, Gaussian, and Turbomole are not working.  We are working to resolve the issue, but... Read more

10 years 3 months ago 10 years 3 months ago

Pages