Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

SERC Resources

SDSS-CC Overview

Stanford Research Computing’s Sherlock HPC platform is SDSS’s principal compute resource. As detailed below, this includes the public partitoins, available to all Sherlock users (which includes all PI gropus at Stanford) and the serc partition, which is shared by Stanford Doerr School of Sustainability (SDSS) research groups. Where Sherlock is not an appropriate, or insufficient platform, SDSS-CC may support alternative options, including:

Sherlock partitions

Sherlock is configured in the “condo” model, which means that -- in addition to a handful of special purpose public partitions (see below), Sherlock is divided into many (small) PI owned partitions. PI groups who buy into the model have exclusive access to resource in their partition and are also granted access to machines in the owners partition (see below).

This differs from the partitioning in larger, National Labs scale computing centers, where partitions are likely defined by hardware type and access to resrouces is controlled by Accounting and some form of pay- or apply for- service. For example, a user might be granted 10k CPU hours on a system; their use is tracked by SLURM (or a different job controller software) Accounting.

The various hardware configurations in Sherlock can significantly affect job performance and resource availabiltiy, so it is important to understand what resources are available, what are the performance characteristics of those resources, and how to request those resoruces to optimize job performance.

Public Partitions

SERC Partition

Overview

SERC is a large partition, shared by SDSS researchers. In short, SERC constitutes SDSS’s principal compute platform and includes a variety of resources to facilitate different types of computation. Modes of compute include interactive Jupyter Notebooks, multi-node MPI simuations, single task (node) serial or thread-based (OpenMP, etc.)

General Compute

SH3_CBASE, SH3_CBASE.1 (104, 96)

SH4_CBASE (128)

Performance Compute

SH4_CPERF (16)

High Capacity Nodes

SH3_CPERF (8)

SH4_CSCALE (4)

GPUs:

SERC presently includes 88 NVIDIA A100 and 8 H100 GPUs, as part of the SH04 acquisition. Another 8 V100 devices will be decomissioned with SH02, sometime in 2025. Note that these machines are very expensive and in very high demand, so it is important that GPU jobs exercise optimized worflows and codes that use these valuable resources efficiently and effectively.

In particular IO performance can often be significantly improved and memory requirements dramatically reduced (or at least better controlled and regulated) by pre-processing input data into HDF5, NetCDF, or similar data formats. This often can be accomplished in a few short paragraphs of code, can reduce memory requirements to easily -- and predictably, fit within the `128/256 GB/GPU’ system RAM on SERC’s A/H100 machines, and improve compute performance by a factor of betwen 10 and 100. Examples in this documentation include HDF5 image examples, HDF5 I/O performance, and DICOM and HDF5.

SH3_G4TF64 (2/8)

SH3_G8TF64 (10/80)

SH4_G8TF64 (1/8)

Selecting hardware

SERC CPU slots

Specific hardware configurations can be requested using the --constraint SLURM directive. When selecting hardware, consider not only the compute requirements, but what resources are available in the partition. As shown in the figure to the right -- which shows the availability of CPU combinations for SERC SH03+SH04 (circa 2025), as a function of CPUs requested, small changes in a resource request can have a large impact on the number of resources available to fulfill the request.

For example:

To see a list of labels available to use with --constraint, run the command, sh_node_feat -p serc

Single Task Jobs:

For Single task (most) jobs, hardware selection is nominally of second order importance -- as compared to MPI jobs (see below). For most single task jobs, SLURM math will be more important than clock speed or memory bandwidth -- in other words, in most cases, the most flexible request will be the fastest path to results. Even if a job is assigned slower hardware, it will wait less time in the queue. Nonetheless, for single task jobs, hardware selection should take into account:

MPI Jobs

It is important that MPI jobs run on homogeneous hardware configurations. Firstly -- assuming the job domains are well balanced, an MPI job with a heterogeneous configuration will always wait on its slowest MPI rank. Additionally, it is possible that different buffer sizes, etc. on the various machines can cause data exchange errors or slowdowns resulting from other mistmatches betweenthe HW and operating systems. Considerations to tak into account for MPI jobs include:

Note the square brackets [ ] tell SLURM to fulfill the request with “all of this, or that, or the other,” so any one of those classes might be used but all selected nodes will satisfy the same requirement -- in this case, all machines will be of the same “class.”

GPU Jobs

GPU hardware can be specified by a handful of constraints, depending on the specific requirement. Again, looking at the output from sh_node_feat -p serc, the features GPU_GEN:VLT, GPU_GEN:AMP, and GPU_GEN:HPR can be used to select V100 (to be retired with SH02 hardware), A100, or H100 GPUs, respectively. Similarly, GPU memory may be used as a discriminator, eg. --constraint=GPU_MEM:80GB. Some examples of constraint specifications include:

Note also that SLURM might, by default, fulfill a request for multiple GPUs (eg, --gpus=4) on multiple tasks and nodes. Unless multi-node parallelization is explicitly enabled, most codes and frameworks will only parallelize across GPUs (or CPUs) on one machine (shared memory or thread based parallelization), so multiple GPU allocations will be wasted. Unless a code is explicitly known to handle multi-node parallelization (eg. MPI), multi-GPU jobs should include the --ntasks=1 SLURM directive (resource request).