Known issues

Unresolved known issues

Known issue with an Unresolved Resolution state is an active problem under investigation; a temporary workaround may be available.

Resolved known issues

A known issue with a Resolved (workaround) Resolution state is an ongoing problem; a permanent workaround is available which may include using different software or hardware.

A known issue with Resolved Resolution state has been corrected.

Known Issues

Title Category Resolutionsort descending Description Posted Updated
Problems with home directory servers filesystem Resolved

We had several auto-reboots with our home directory servers, starting from around 11pm August 30. Jobs might be impacted. The systems are working properly now.  

... Read more
5 years 2 weeks ago 5 years 2 weeks ago
Segmentation fault from openmpi/1.10-hpcx and 2.0-hpcx on Owens Owens, Software Resolved

We have found that recent MPI jobs using openmpi/1.10-hpcx and openmpi/2.0-hpcx on Owens may complete or hang until the job is killed, but receive segmentation fault. Some applications might be ...

Read more
7 years 1 month ago 7 years 1 month ago
Poor performance with hybrid MPI+OpenMPI jobs and more than 4 MPI Tasks on multiple nodes Software Resolved
(workaround)

RELION versions prior to 5 may exhibit suboptimal performance in hybrid MPI+OpenMP jobs when the number of MPI tasks exceeds four across multiple nodes.

Workaround...

Read more
1 year 1 month ago 1 year 1 month ago
Rolling reboot of all clusters, starting from Wednesday morning, April 19, 2017 Batch, Maintenance, Owens, Ruby Resolved

1:40PM 4/27/2017 Update: Rolling reboots are completed. 

3:10PM 4/18/2017 Update: Rolling reboots on Owens have started to address GPFS errors occured...

Read more
9 years 5 months ago 9 years 4 months ago
Intermittent home directory performance issues filesystem Resolved

Users may experience performance issues in home directory. It is recommended to use temporary directory ($TMPDIR, or scratch) or project storage to minimize the impact on...

Read more
3 years 2 months ago 3 years 1 month ago
GPU Memory Not Released Causing OOM in Subsequent Jobs GPU Resolved
(workaround)

We have noticed that GPU memory is not being properly released in some jobs, causing subsequent jobs on the same nodes to run out of memory (OOM). We are currently working on a resolution...

Read more
11 months 3 weeks ago 6 months 1 week ago
Lustre is still offline. HPC systems back up Maintenance Resolved

Day One of the scheduled downtime has been completed, and HPC operations have resumed. As planned, Lustre work will extend into Day Two. Jobs using /fs/lustre or $PFSDIR cannot run until this work...

Read more
12 years 2 months ago 12 years 2 months ago
Scheduling suspended Batch Resolved

We have temporarily suspended scheduling due to some problems with the parallel scratch file system.

11 years 11 months ago 11 years 11 months ago
Client Portal inaccessible client portal Resolved

All users are currently unable to login to OSC client portal.

When attempting to login to OSC...

Read more
5 years 8 months ago 5 years 8 months ago
NFS outage on Thursday Jan 17 from 7am to 8am filesystem Resolved

Update:

This work has been canceled and will be done during downtime on Feb. 5. 

Original Post:

On Thursday, January 17th from 7 am to 8 am OSC will have a GPFS...

Read more
7 years 8 months ago 7 years 8 months ago

Pages