A rolling reboot is scheduled to remediate a ptrace vulnerability and safely restore debugger functionalit

Known issues

Unresolved known issues

Known issue with an Unresolved Resolution state is an active problem under investigation; a temporary workaround may be available.

Resolved known issues

A known issue with a Resolved (workaround) Resolution state is an ongoing problem; a permanent workaround is available which may include using different software or hardware.

A known issue with Resolved Resolution state has been corrected.

Known Issues

Title Category Resolutionsort ascending Description Posted Updated
warning: libhwloc.so.1 may conflict with libhwloc.so.5 Resolved

Sometimes when building MPI programs the following warning appears.  It is harmless and can be safely ignored.

ld: warning: libhwloc.so.1, needed by /usr/local/mvapich2/1.7-intel/lib/...
Read more
11 years 2 months ago 10 years 9 months ago
Problems with GPFS filesystem, and OnDemand is not working filesystem Resolved

Updates on 12:00pm June 10:

The issue is fixed. All impacted services return to production.

We apologize for the inconvience. If you have any questions, please...

Read more
5 years 1 month ago 5 years 1 month ago
A bug in the trigger that sends automated emails from client portal client portal Resolved

We deployed a new version to OSC Client Portal (my.osc.edu) at 3 pm Tuesday, July 9th, which involves a bug in the trigger that sends automated emails to some OSC clients with the subject 'Your...

Read more
7 years 1 week ago 7 years 1 week ago
Instability on Clusters after May 13 Downtime Resolved

We've been experiencing some instability on the clusters (particularly Cardinal and Ascend) following the recent May 13 downtime, especially with parallel job processing. If you notice any unusual...

Read more
1 year 2 months ago 1 year 1 month ago
Rolling reboot of compute and login nodes of all clusters, starting from Wednesday morning, March 22, 2017 login, Owens, Ruby Resolved

4:56PM 3/28/2017 Update: The rolling reboots of all systems are completed. 

All compute nodes and login nodes of Owens, Oakley, and Ruby clusters will need to be rebooted...

Read more
9 years 4 months ago 9 years 3 months ago
GPFS problems with /fs/project and possibly /fs/scratch filesystem Resolved

There was an issue with GPFS clients that affected /fs/project and possibly /fs/scratch between around 3:30AM and 8:30AM on Sunday September 4th. Some jobs from clients were also impacted. 

... Read more
3 years 10 months ago 3 years 10 months ago
AlphaFold 3 GPU Out-of-Memory Error During Inference Software Resolved
(workaround)

When you run AlphaFold 3, you may encounter a GPU out-of-memory (OOM) failures during model execution. The job terminated with errors similar to:

Can't...
Read more
7 months 3 weeks ago 7 months 3 weeks ago
Account changes temporarily suspended Account Management Resolved

We are still experiencing some account problems related to Thursday's issue. As a result, we have taken my.osc.edu offline and cannot process email changes or password resets, either via self-...

Read more
12 years 1 month ago 12 years 1 month ago
Storage Problems - GPFS services (Project / Scratch) filesystem Resolved

Updates at 15:52 March 11, 2020:

The issue with the Project file system that causes deletes of file system snapshots to fail has now been resolved. OSC Project file system...

Read more
6 years 4 months ago 6 years 4 months ago
OnDemand unresponsive login Resolved

Some of the login nodes on Owens and Pitzer are in bad states. User can't log into OnDemand. And scratch is unresponsive sometimes. We are working on this issue. We will update when we have more...

Read more
6 years 1 month ago 6 years 1 month ago

Pages