National systems: Difference between revisions

From Alliance Doc
Jump to navigation Jump to search
No edit summary
 
(92 intermediate revisions by 11 users not shown)
Line 1: Line 1:
<table border="5" style="border-color:red;width:100%">
<languages />
<tr><td>Sharing: Public</td><td align="center" style="font-size:18px">'''New Systems'''</td><td align="right">Author:CC Migration team</td></tr>
<translate>
</table>


==Compute==
==Compute clusters== <!--T:1-->


===Overview===
<!--T:3-->
A ''general-purpose'' cluster is designed to support a wide variety of types of jobs, and is composed of a mixture of different nodes.  We broadly classify the nodes as:
* ''base'' nodes, containing typically about 4GB of memory per core;
* ''large-memory'' nodes, containing typically more than 8GB memory per core;
* ''GPU'' nodes, which contain [https://en.wikipedia.org/wiki/Graphics_processing_unit graphic processing units].


GP2 and GP3 are almost identical systems with some minor differences in the actual mix of large memory, small memory and GPU nodes.
<!--T:17-->
The ''large parallel'' cluster [[Niagara]] is designed to support multi-node parallel jobs requiring more than 1000 CPU cores, although jobs as small as a single node are also supported there.  Niagara is composed of nodes of a uniform design, with an interconnect optimized for large jobs.


{| class="wikitable sortable"
<!--T:18-->
|-
All clusters have large, high-performance storage attached.  For details about storage, memory, CPU model and count, GPU model and count, and the number of nodes at each site, please click on the cluster name in the table below.
! Name !! Description !! Approximate Capacity !! Availability
|-
| [[CC-Cloud Resources]] (GP1) || Cloud || 7,640 cores || In production (integrated with west.cloud)
|-
| [[GP2]] || heterogeneous, general-purpose cluster (Serial and small parallel jobs)<br />Small cloud partition || 25,000 cores || April, 2017
|-
| [[GP3]] || heterogeneous, general-purpose cluster (Serial and small parallel jobs)<br />Small cloud partition || 25,000 cores || April, 2017
|-
| [[LP]] || a cluster designed for large parallel jobs || 60,000 cores || Late 2017
|}


Note that GP1, GP2 and LP will all have large, high-performance attached storage.
===List of compute clusters=== <!--T:14-->


===GP2 (SFU)===
<!--T:15-->
 
{| class="wikitable"
The GP2 system evaluation is not yet completed (as of early November 2016).  Anticipated specifications, based on SFU's RFP and bids received, include the following.  This information is '''not guaranteed''' and might not be complete.  It is provided for planning purposes.
 
{| class="wikitable sortable"
|-
|-
| '''Parallel filesystem:''' || Approximately 4PB usable capacity for temporary (<code>/scratch</code>) storage.  Aggregate performance of approximately 40GB/s.  Available to all nodes.
! Name and link !! Type !! Sub-systems !! Status
|-
|-
|'''External persistent storage:'''
| [[Béluga/en|Béluga]]
||
| General-purpose
10PB or more of storage for persistent (<code>/project</code> and other) storage. Available to compute nodes, but not designed for parallel workloads.
|
* beluga-compute
* beluga-gpu
* beluga-storage
| In production
|-
|-
|'''High performance interconnect:'''
| [[Cedar|Cedar]]
||
| General-purpose
Low-latency high-performance fabric connecting all nodes and temporary storage. The design of GP2 is to support multiple simultaneous parallel jobs of at least 1024 cores in a fully non-blocking manner. Jobs larger than 1024 cores would be less well-suited for the topology.
|
|}
* cedar-compute
 
* cedar-gpu
====Node types and characteristics:====
* cedar-storage
 
| In production
{| class="wikitable sortable"
|-
|-
| "Base" compute nodes: || Over 500 nodes || 128GB of memory.
| [[Graham|Graham]]
| General-purpose
|
* graham-compute
* graham-gpu
* graham-storage
| In production
|-
|-
| "Large" compute nodes: || Over 100 nodes || 256GB of memory.
| [[Narval/en|Narval]]
| General-purpose
|
* narval-compute
* narval-gpu
* narval-storage
| In production
|-
|-
| "Bigmem512" and "Bigmem1500" nodes || 24 nodes || each with 512GB and 1.5TB of memory respectively.
| [[Niagara|Niagara]]
| Large parallel
|
* niagara-compute
* niagara-storage
* hpss-storage
| In production
|}
|}


All of the above nodes will have 16 cores/socket (32 cores/node), with an anticipated frequency of 2.1Ghz.
==Cloud - Infrastructure as a Service== <!--T:16-->
 
Our cloud systems are offering an Infrastructure as a Service (IaaS) based on OpenStack.
{| class="wikitable sortable"
|-
| "GPU base" nodes: || Over 100 nodes with 4 GPUs each. || 12 cores/socket (24 cores/node) with an anticipated frequency of 2.2Ghz.  GPUs on a dual PCI root.
|-
| "GPU large" nodes.  || Approximately 30 nodes || same configuration as "GPU base," but a single PCI root.
|-
| "Bigmem3000" nodes || 4 nodes || each with 3TB of memory. These are 4-socket nodes with 8 cores/socket.
|}
 
All of the above nodes will have local (on-node) storage.
 
====Login, Gateway and Data Transfer nodes====
 
Compute Canada is not currently able to disclose the specific GPU model or specifications.  The RFP used NVIDIA K80 as a baseline specification.
 
The total GP2 system is anticipated to have over 25,000 cores, 900 nodes, and 500 GPUs.
 
The delivery and installation schedule is not yet known, but the procurement team has confidence that the system will be in production for the start of the allocations year on April 1, 2017.
 
===GP3 (Waterloo)===
 
P3 system evaluation is not yet completed (as of early November 2016).  Anticipated specifications, based on Waterloo's RFP and bids received, include the following.  This information is NOT GUARANTEED and might not be complete.  It is provided for planning purposes.
 
The parallel filesystem, interconnects and external persistent storage will be the same as GP2.
 
====Node types and characteristics:====


{| class="wikitable sortable"
<!--T:4-->
{| class="wikitable"
|-
|-
| "Base" compute nodes: || Over 500 nodes ||  128GB of memory.
! Name and link !! Sub-systems !! Description !! Status
|-
|-
| "Large" compute nodes: || Over 100 nodes || 256GB of memory.
| [[Cloud_resources#Arbutus_cloud|Arbutus cloud]]
|
* arbutus-compute-cloud
* arbutus-persistent-cloud
* arbutus-dcache
|
* VCPU, VGPU, RAM
* Local ephemeral disk
* Volume and snapshot storage
* Shared filesystem storage (backed up)
* Object storage
* Floating IPs
* dCache storage
| In production
|-
|-
| "Bigmem512" nodes|| TBD || each with 512GB
| [[Cloud_resources#B.C3.A9luga_cloud|Béluga cloud]]
|
* beluga-compute-cloud
* beluga-persistent-cloud
|
* VCPU, RAM
* Local ephemeral disk
* Volume and snapshot storage
* Floating IPs
| In production
|-
|-
| "Bigmem3000" nodes || TBD || 3 TB
| [[Cloud_resources#Cedar_cloud|Cedar cloud]]
|}
|
 
* cedar-persistent-cloud
====GPU nodes====
* cedar-compute-cloud
 
|
All of the above nodes will have local (on-node) storage.
* VCPU, RAM
 
* Local ephemeral disk
====Login, Gateway and Data Transfer Nodes====
* Volume and snapshot storage
 
* Floating IPs
Compute Canada is not currently able to disclose the specific GPU or CPU model or specifications.  The RFP used NVIDIA K80 as a baseline specification.
| In production
 
The total GP3 system is anticipated to have over 20,000 cores.
 
The delivery and installation schedule is not yet known.
 
==National Data Cyberinfrastructure (NDC)==
 
{| class="wikitable sortable"
|-
|-
! Type !! Location !! Capacity !! Availability !! Comments
| [[Cloud_resources#Graham_cloud|Graham cloud]]
|-
|
| '''Nearline (Long-term) <br />File Storage''' ||
* graham-persistent-cloud
SFU and Waterloo
|
* NDC-SFU
* VCPU, RAM
* NDC-Waterloo
* Local ephemeral disk
|| 15 PB each || Late Autumn 2016 ||
* Volume and snapshot storage
* Tape backup
* Floating IPs
* '''Silo replacement'''
| In production
* Aiming for redundant tape backup across sites
|-
| '''Object Store'''
||
All 4 sites
* NDC-Object
||
Small to start (A few PB usable)
||
Late 2017
||
* Fully distributed, redundant object storage
* Accessible anywhere
* Allows for redundant, high availability architectures
* S3 and File interfaces
* New service aimed at experimental and observational data
|-
| '''Special Purpose''' || All 4 sites || ~3.5 PB || Specialized migration plans ||
* Special purpose for Atlas, LHC, SNO+, CANFAR
* dCache and other customized configurations
|}
|}


Note that due to Silo decommissioning Dec.31/2016 it will be necessary to provide interim storage while the NDC is developed. See [[Migration2016:Silo]] for details.
</translate>
 
[[Category:Migration2016]]

Latest revision as of 14:23, 8 April 2022

Other languages:

Compute clusters

A general-purpose cluster is designed to support a wide variety of types of jobs, and is composed of a mixture of different nodes. We broadly classify the nodes as:

  • base nodes, containing typically about 4GB of memory per core;
  • large-memory nodes, containing typically more than 8GB memory per core;
  • GPU nodes, which contain graphic processing units.

The large parallel cluster Niagara is designed to support multi-node parallel jobs requiring more than 1000 CPU cores, although jobs as small as a single node are also supported there. Niagara is composed of nodes of a uniform design, with an interconnect optimized for large jobs.

All clusters have large, high-performance storage attached. For details about storage, memory, CPU model and count, GPU model and count, and the number of nodes at each site, please click on the cluster name in the table below.

List of compute clusters

Name and link Type Sub-systems Status
Béluga General-purpose
  • beluga-compute
  • beluga-gpu
  • beluga-storage
In production
Cedar General-purpose
  • cedar-compute
  • cedar-gpu
  • cedar-storage
In production
Graham General-purpose
  • graham-compute
  • graham-gpu
  • graham-storage
In production
Narval General-purpose
  • narval-compute
  • narval-gpu
  • narval-storage
In production
Niagara Large parallel
  • niagara-compute
  • niagara-storage
  • hpss-storage
In production

Cloud - Infrastructure as a Service

Our cloud systems are offering an Infrastructure as a Service (IaaS) based on OpenStack.

Name and link Sub-systems Description Status
Arbutus cloud
  • arbutus-compute-cloud
  • arbutus-persistent-cloud
  • arbutus-dcache
  • VCPU, VGPU, RAM
  • Local ephemeral disk
  • Volume and snapshot storage
  • Shared filesystem storage (backed up)
  • Object storage
  • Floating IPs
  • dCache storage
In production
Béluga cloud
  • beluga-compute-cloud
  • beluga-persistent-cloud
  • VCPU, RAM
  • Local ephemeral disk
  • Volume and snapshot storage
  • Floating IPs
In production
Cedar cloud
  • cedar-persistent-cloud
  • cedar-compute-cloud
  • VCPU, RAM
  • Local ephemeral disk
  • Volume and snapshot storage
  • Floating IPs
In production
Graham cloud
  • graham-persistent-cloud
  • VCPU, RAM
  • Local ephemeral disk
  • Volume and snapshot storage
  • Floating IPs
In production