A rolling reboot is scheduled to remediate a ptrace vulnerability and safely restore debugger functionalit

Known issues

Unresolved known issues

Known issue with an Unresolved Resolution state is an active problem under investigation; a temporary workaround may be available.

Resolved known issues

A known issue with a Resolved (workaround) Resolution state is an ongoing problem; a permanent workaround is available which may include using different software or hardware.

A known issue with Resolved Resolution state has been corrected.

Known Issues

Titlesort ascending Category Resolution Description Posted Updated
Reboot of Public-Facing Systems for Security Fix Resolved

Updates on May 12:

All  public-facing systems have been rebooted. 

Original Post:

Due to a critical...

Read more
2 months 1 week ago 2 months 14 hours ago
Reached your pull rate limit while pull from Docker hub Software Resolved
(workaround)

You might encounter an error when pulling from Docker hub:

ERROR: toomanyrequests: Too Many Requests.

or

You have reached your...
Read more
5 years 3 weeks ago 7 months 3 weeks ago
quota exceeded error when using chgrp in /fs/ess directories filesystem Resolved

Users may receive an error when using the chgrp command on data in /fs/ess/ locations.

$ chgrp -v PEX1234 my-file.txt
chgrp: changing group of 'my-file.txt': Disk quota exceeded
failed...
Read more
3 years 5 months ago 3 years 4 months ago
qsub filter rejects valid jobs Resolved

Job scripts submitted on Glenn, Oakley, or Ruby all go a submit filter before reaching the resource manager, Torque.  A bug has been discovered in our submit filter which prevents jobs with the...

Read more
11 years 3 months ago 10 years 9 months ago
PyTorch jobs timeout and hanging GPU Resolved

We have observed that many PyTorch users frequently encounter random timeouts, which result in the termination of their jobs but leave the process running on the node....

Read more
3 years 22 hours ago 2 years 6 months ago
PyTorch hangs on dual-gpu node on Ascend Ascend, GPU Resolved
(workaround)

PyTorch can hang on Ascend on dual-GPU nodes

Through internal testing, we have confirmed that the hang issue only occurs on Ascend dual-GPU (nextgen) nodes. We’re still unsure why...

Read more
1 year 2 months ago 1 year 2 months ago
Python version mismatch in Jupyter + Spark instance Software Resolved
(workaround)

You may encounter the following error message when running a Spark instance using a custom kernel in the Jupyter + Spark app:

25/04/25 10:49:01 WARN TaskSetManager:...
Read more
1 year 1 month ago 7 months 3 days ago
ptrace Disabled so Debuggers and Tracers Do Not Work Unresolved

ptrace has been disabled globally across all OSC systems to mitigate a newly identified Linux kernel vulnerability. This disables most functionality of debuggers,...

Read more
1 month 3 weeks ago 2 weeks 2 days ago
Project space giving errors "No space left on device" filesystem Resolved

11/01/2016 11:52AM Update: This issue has been fixed. 

We have become aware of a problem with the Project storage space that gives errors "No space left on device". The...

Read more
9 years 8 months ago 9 years 8 months ago
Proj13 file system difficulties filesystem Resolved

We are currently experiencing difficulties with the servers for the filesystem mounted at /nfs/proj13.

12 years 9 months ago 12 years 9 months ago

Pages