A rolling reboot is scheduled to remediate a ptrace vulnerability and safely restore debugger functionalit

Known issues

Unresolved known issues

Known issue with an Unresolved Resolution state is an active problem under investigation; a temporary workaround may be available.

Resolved known issues

A known issue with a Resolved (workaround) Resolution state is an ongoing problem; a permanent workaround is available which may include using different software or hardware.

A known issue with Resolved Resolution state has been corrected.

Known Issues

Title Category Resolutionsort descending Description Posted Updated
Application Errors client portal Resolved

When beginning a major or discovery-level application for resources at OSC, you are asked for a required justification on the Additional Documents page. However, there is no mechanism for you to...

Read more
7 years 2 months ago 6 years 11 months ago
Apptainer sandbox fails on GPFS due to stale NFS file handles Software Resolved

Users may encounter errors when attempting to run a sandbox on GPFS-mounted directories such as /fs/scratch or /fs/ess. The error...

Read more
11 months 2 weeks ago 7 months 3 weeks ago
Segmentation fault from openmpi/1.10-hpcx and 2.0-hpcx on Owens Owens, Software Resolved

We have found that recent MPI jobs using openmpi/1.10-hpcx and openmpi/2.0-hpcx on Owens may complete or hang until the job is killed, but receive segmentation fault. Some applications might be ...

Read more
6 years 11 months ago 6 years 11 months ago
Email Issues client portal Resolved

OSU is having ongoing periodic problems with Microsoft (their mail hosting provider) severely delaying outbound email. There is no solution being offered and no timeline for getting it resolved....

6 years 9 months ago 6 years 5 months ago
Rolling reboot of all clusters, starting from Wednesday morning, April 19, 2017 Batch, Maintenance, Owens, Ruby Resolved

1:40PM 4/27/2017 Update: Rolling reboots are completed. 

3:10PM 4/18/2017 Update: Rolling reboots on Owens have started to address GPFS errors occured...

Read more
9 years 2 months ago 9 years 2 months ago
Slurm database repair on 01/25/2024 Outage Resolved

We have scheduled a Slurm database repair, which is planned to start at 8:30 am US/Eastern on Thursday, January 25, 2024. During the repair, Slurm database will be offline; running jobs and...

Read more
2 years 5 months ago 2 years 5 months ago
Lustre is still offline. HPC systems back up Maintenance Resolved

Day One of the scheduled downtime has been completed, and HPC operations have resumed. As planned, Lustre work will extend into Day Two. Jobs using /fs/lustre or $PFSDIR cannot run until this work...

Read more
12 years 1 week ago 12 years 4 days ago
MVAPICH2 build of CP2K 6.1 Pitzer Resolved

We have found some types of CP2K jobs would fail or have poor performance using cp2k.popt and cp2k.psmp from MVAPICH2 build (gnu/4.8.5 mvapich2/2.3). This version will be removed on December 15th...

Read more
5 years 7 months ago 5 years 4 months ago
GPU Memory Not Released Causing OOM in Subsequent Jobs GPU Resolved
(workaround)

We have noticed that GPU memory is not being properly released in some jobs, causing subsequent jobs on the same nodes to run out of memory (OOM). We are currently working on a resolution...

Read more
9 months 3 weeks ago 4 months 5 days ago
Replacement of Owens Ethernet switches from Dec 14, 2018 Network, Owens Resolved

Updated on Jan 16, 2019, at 09:20 AM:

The replacement is done except for the three switches including the login nodes of Owens. We posted another notice for more...

Read more
7 years 10 months ago 7 years 6 months ago

Pages