Known issues

Unresolved known issues

Known issue with an Unresolved Resolution state is an active problem under investigation; a temporary workaround may be available.

Resolved known issues

A known issue with a Resolved (workaround) Resolution state is an ongoing problem; a permanent workaround is available which may include using different software or hardware.

A known issue with Resolved Resolution state has been corrected.

Known Issues

Title Category Resolutionsort descending Description Posted Updated
Storage Problems - GPFS services (Project / Scratch) filesystem Resolved

Updates at 15:52 March 11, 2020:

The issue with the Project file system that causes deletes of file system snapshots to fail has now been resolved. OSC Project file system...

Read more
6 years 5 months ago 6 years 5 months ago
Lustre is still offline. HPC systems back up Maintenance Resolved

Day One of the scheduled downtime has been completed, and HPC operations have resumed. As planned, Lustre work will extend into Day Two. Jobs using /fs/lustre or $PFSDIR cannot run until this work...

Read more
12 years 1 month ago 12 years 1 month ago
OnDemand unresponsive login Resolved

Some of the login nodes on Owens and Pitzer are in bad states. User can't log into OnDemand. And scratch is unresponsive sometimes. We are working on this issue. We will update when we have more...

Read more
6 years 2 months ago 6 years 2 months ago
Replacement of Owens Ethernet switches from Dec 14, 2018 Network, Owens Resolved

Updated on Jan 16, 2019, at 09:20 AM:

The replacement is done except for the three switches including the login nodes of Owens. We posted another notice for more...

Read more
7 years 11 months ago 7 years 7 months ago
- --gpus-per-task is not working Batch Resolved

Updated: This is fixed. 

Original Post:

After the recent Slurm upgrade, the option --gpus-per-task is currently not functioning as...

Read more
1 year 7 months ago 1 year 7 months ago
Problems with Project Space (/nfs/gpfs) filesystem Resolved

(9/8/15 14:21 Eastern) Project space appears to be back to normal operation. We are running some tests to verify that the problem is fully resolved.


As of early afternoon, Sept. 8,...

Read more
10 years 11 months ago 10 years 11 months ago
Problems with home directory servers filesystem Resolved

We had several auto-reboots with our home directory servers, starting from around 11pm August 30. Jobs might be impacted. The systems are working properly now.  

... Read more
4 years 11 months ago 4 years 11 months ago
Segmentation fault from openmpi/1.10-hpcx and 2.0-hpcx on Owens Owens, Software Resolved

We have found that recent MPI jobs using openmpi/1.10-hpcx and openmpi/2.0-hpcx on Owens may complete or hang until the job is killed, but receive segmentation fault. Some applications might be ...

Read more
7 years 3 weeks ago 7 years 2 days ago
Poor performance with hybrid MPI+OpenMPI jobs and more than 4 MPI Tasks on multiple nodes Software Resolved
(workaround)

RELION versions prior to 5 may exhibit suboptimal performance in hybrid MPI+OpenMP jobs when the number of MPI tasks exceeds four across multiple nodes.

Workaround...

Read more
1 year 1 week ago 1 year 1 week ago
Rolling reboot of all clusters, starting from Wednesday morning, April 19, 2017 Batch, Maintenance, Owens, Ruby Resolved

1:40PM 4/27/2017 Update: Rolling reboots are completed. 

3:10PM 4/18/2017 Update: Rolling reboots on Owens have started to address GPFS errors occured...

Read more
9 years 4 months ago 9 years 3 months ago

Pages