Known issues: Difference between revisions

Revision as of 18:32, 25 July 2017

Other languages:

English
français

Intro[edit]

Please report issues to support@computecanada.ca

Shared issues[edit]

The CC slurm configuration preferentially encourages whole-node jobs. Users should if possible request whole-nodes rather than per-core resources. See Job Scheduling - Whole Node Scheduling (Patrick Mann 20:15, 17 July 2017 (UTC))
- Cpu and Gpu backfill partitions have been created on both clusters. If a job is submitted with <24hr runtime, it will be automatically entered into the cluster-wide backfill partition. This partition has a low priority, but will allow increased utilization of the cluster by serial jobs. (Nathan Wielenga)
Quotas on /project are all 1 TB. The Storage National team is working on a project/RAC based schema. Fortunately Lustre have announced group-based quotas but that will need installation. (Patrick Mann 20:12, 17 July 2017 (UTC))
SLURM epilog does not fully clean up processes from ended jobs, especially if the job did not exit normally. (Greg Newby) Fri Jul 14 19:32:48 UTC 2017)
The SLURM 'sinfo' command yields different resource-type detail on graham and cedar. (Greg Newby) 16:05, 23 June 2017 (UTC))
The status page at http://status.computecanada.ca/ is not updated automatically yet, so does not necessarily show correct, current status.
"Nearline" capabilities are not yet available (see https://docs.computecanada.ca/wiki/National_Data_Cyberinfrastructure for a brief description of the intended functionality)
- Update July 17: still not working. If you need your nearline RAC2017 quota then please ask CC support. (Patrick Mann 20:45, 17 July 2017 (UTC))
Operations will occasionally time out with a message like "Socket timed out on send/recv operation" or "Unable to contact slurm controller (connect failure)". As a temporary workaround, attempt to resubmit your jobs/commands, they should go through in a few seconds. (Nathan Wielenga) 08:50, 18 July 2017 (MDT))
- Should be resolved after a VHD migration to a new backend for slurmctl. (NW)
Auto-creation of project directories such as /project/$USER was an interim solution. Soon there will be /project/gid where gid is the project group identifier. This will be symlinked to /project/projects/pname where pname is the "friendly" project (RAPI) name). And then, /project/gid/$USER can be where user subdirectories for that project will live. Note that quotas in /project are project-based, not user-based. (Greg Newby) Thu Jul 20 00:45:00 UTC 2017)

Cedar only[edit]

Graham only[edit]

Graham has returned to service after a maintenance outage, which included rearrangement of /project. /scratch is unavailable, but will be returned to service as soon as possible.

@@ Line 21: / Line 21: @@
 == Graham only == <!--T:4-->
-* '''Graham is currently in a maintenance outage (Tuesday, 2017-07-25).'''
+* Graham has returned to service after a maintenance outage, which included rearrangement of /project.  /scratch is unavailable, but will be returned to service as soon as possible.
 == Other issues == <!--T:5-->
 </translate>

Known issues: Difference between revisions

Revision as of 18:32, 25 July 2017

Contents

Intro[edit]

Shared issues[edit]

Cedar only[edit]

Graham only[edit]

Other issues[edit]

Navigation menu

Known issues: Difference between revisions

Revision as of 18:32, 25 July 2017

Intro[edit]

Shared issues[edit]

Cedar only[edit]

Graham only[edit]

Other issues[edit]

Navigation menu

Search