Known issues

Unresolved known issues

Known issue with an Unresolved Resolution state is an active problem under investigation; a temporary workaround may be available.

Resolved known issues

A known issue with a Resolved (workaround) Resolution state is an ongoing problem; a permanent workaround is available which may include using different software or hardware.

A known issue with Resolved Resolution state has been corrected.

Known Issues

Title Category Resolutionsort descending Description Posted Updated
Podman storage error due to outdated configuration file Software Resolved
(workaround)

If you experience a storage error such as:

 write /var/tmp/storage388772891/1: no space left on device.

you may have an outdated ~/....

Read more
1 year 1 week ago 1 year 1 week ago
All HPC systems are available Login Problems, Operations, Outage Resolved

8/24/16 3:57PM: All HPC systems are availalbe including:

  • Oakley cluster for general access
  • Ruby cluster for restricted access
  • Owens cluster for...
Read more
9 years 11 months ago 9 years 11 months ago
ORCA Bind to CORE Failure Software Resolved
(workaround)

The default CPU binding for ORCA jobs can fail sporadically.  The failure is almost immediate and produces a cryptic error message, e.g.:

...
Read more
3 years 3 months ago 1 year 3 months ago
Brief disruption of GPFS on 8/28/2013 filesystem Resolved

On the morning August 28th, 2013 we will briefly disrupt the GPFS filesystem to reboot servers. This is necessary to upgrade the GPFS system. The in-place upgrade should only briefly interrupt...

Read more
12 years 11 months ago 12 years 11 months ago
Rolling reboot of login nodes of clusters at 7:00AM Dec 19, 2017 login Resolved

We will have rolling reboot of login nodes of clusters at 7:00AM Dec 19, 2017 for GPFS version upgrade. It is supposed to be completed in a short period of time. f you encounter any login issues,...

Read more
8 years 7 months ago 8 years 7 months ago
cuMemHostRegister Fails with CUDA_ERROR_INVALID_VALUE on RHEL 9.6 Ascend, Cardinal, GPU, system software Resolved

After upgrading the operating system to RHEL 9.6 during the scheduled downtime on May 12, 2026,  applications utilizing UCX (...

Read more
2 months 2 weeks ago 1 month 2 weeks ago
Intermittent issue with connecting to batch server Batch, Owens Resolved

Updated on June 18, 2018, at 3:15 PM:

This issue has been fixed. 

Posted on June 18, 2018, at 12:30 PM:

We've been having intermittent...

Read more
8 years 1 month ago 8 years 1 month ago
GPFS errors on compute nodes filesystem Resolved

We've seen an increase in transient problems that result in compute nodes losing access to the GPFS file systems for ~5 minutes.

Any jobs running on these nodes accessing files on GPFS may...

Read more
5 years 8 months ago 4 years 8 months ago
PyTorch hangs on dual-gpu node on Ascend Ascend, GPU Resolved
(workaround)

PyTorch can hang on Ascend on dual-GPU nodes

Through internal testing, we have confirmed that the hang issue only occurs on Ascend dual-GPU (nextgen) nodes. We’re still unsure why...

Read more
1 year 3 months ago 1 year 3 months ago
Armstrong inaccessible Resolved

Update: 2PM March 12th: Armstrong is back up and running.  Please notify oschelp@osc.edu of any lingering issues.


As of 10AM Thursday March 12th...

Read more
11 years 5 months ago 11 years 5 months ago

Pages