A rolling reboot is scheduled to remediate a ptrace vulnerability and safely restore debugger functionalit

Known issues

Unresolved known issues

Known issue with an Unresolved Resolution state is an active problem under investigation; a temporary workaround may be available.

Resolved known issues

A known issue with a Resolved (workaround) Resolution state is an ongoing problem; a permanent workaround is available which may include using different software or hardware.

A known issue with Resolved Resolution state has been corrected.

Known Issues

Title Category Resolutionsort ascending Description Posted Updated
Inviting user to a project based on project and user types (academic/commercial/classroom) client portal Resolved

A recent update has corrected this error.

6 years 8 months ago 6 years 4 months ago
Slurm to be Upgraded to Version 23.11.4 Owens, Pitzer Resolved

Updates on 04/08/2024:

The rolling reboots are completed. 

Updates:

We will perform rolling reboots on this...

Read more
2 years 4 months ago 2 years 3 months ago
vLLM versions prior to 0.14.1 are deprecated Software Resolved

vLLM versions prior to 0.14.1 are deprecated due to security issue CVE-2026-22778 which can allow remote code execution.  Clients are advised to use versions 0....

Read more
5 months 2 days ago 5 months 2 days ago
pdsh -j broken on Oakley Batch, system software Resolved

pdsh -j is broken on Oakley.  It was broken by updates during the September downtime.  We are currently working on resolving the issue.

Users who require...

Read more
10 years 7 months ago 8 years 1 month ago
All HPC systems are available Login Problems, Operations, Outage Resolved

8/24/16 3:57PM: All HPC systems are availalbe including:

  • Oakley cluster for general access
  • Ruby cluster for restricted access
  • Owens cluster for...
Read more
9 years 11 months ago 9 years 10 months ago
Reached your pull rate limit while pull from Docker hub Software Resolved
(workaround)

You might encounter an error when pulling from Docker hub:

ERROR: toomanyrequests: Too Many Requests.

or

You have reached your...
Read more
5 years 4 weeks ago 7 months 3 weeks ago
Brief disruption of GPFS on 8/28/2013 filesystem Resolved

On the morning August 28th, 2013 we will briefly disrupt the GPFS filesystem to reboot servers. This is necessary to upgrade the GPFS system. The in-place upgrade should only briefly interrupt...

Read more
12 years 10 months ago 12 years 10 months ago
Python version mismatch in Jupyter + Spark instance Software Resolved
(workaround)

You may encounter the following error message when running a Spark instance using a custom kernel in the Jupyter + Spark app:

25/04/25 10:49:01 WARN TaskSetManager:...
Read more
1 year 2 months ago 7 months 6 days ago
Rolling reboot of login nodes of clusters at 7:00AM Dec 19, 2017 login Resolved

We will have rolling reboot of login nodes of clusters at 7:00AM Dec 19, 2017 for GPFS version upgrade. It is supposed to be completed in a short period of time. f you encounter any login issues,...

Read more
8 years 7 months ago 8 years 6 months ago
Symlinks to /fs/project directories were missed filesystem Resolved

Symlinks to /fs/project directories were missed for a short period of time both on Tuesday afternoon October 11(right after downtime) and Wedenday morning October 12. It might result in job...

Read more
3 years 9 months ago 3 years 9 months ago

Pages