Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Choosing a Storage Space on Sherlock

Sherlock offers several different storage spaces, each with different capacity, speed, and retention tradeoffs. Picking the right one can make a big difference in how fast your data-heavy pipelines run. This tutorial walks through the available options and how to choose between them.

The Storage Space Zoo

Sherlock’s available storage options are:

$HOME

15GB

Private to your user. Be careful not to fill it!

$GROUP_HOME

1TB of shared space for your group

$OAK

“Cheap and deep”

For large files and backups

$SCRATCH

100TB for each user

Purges files older than 90 days

$GROUP_SCRATCH

100TB of shared scratch space

Purges files older than 90 days

$L_SCRATCH

Scratch local to the active compute node, deletes when job ends

Size varies from 100s of GB to a few TB

An Analogy: HPC as Cooking

It can help to think about these spaces the way you’d think about storing and preparing food:

How Do I Choose Which Space Is Right for Me?

You’ll likely need to use more than one of these spaces together. A few questions can help you decide:

  1. What is your dataset size?

    • Small (a few GBs) → $HOME

    • Medium (dozens of GBs) → $GROUP_HOME

    • Large (TBs) → keep going

  2. Is long-term storage needed?

    • Yes → $OAK

    • No → keep going

  3. Do you have large files, or many small files?

    • Large files → $OAK

    • Many small files → keep going

  4. Do you need to share these files with others?

    • Yes → $GROUP_SCRATCH

    • No → keep going

  5. Do you need ultra-low latency?

    • Yes → $L_SCRATCH

    • No → $SCRATCH

Use Case 1: One Large File

Example: machine learning with the MOSAIKS dataset — a 4.7TB single-file table of satellite image encodings (4,005 columns, 146M rows), sourced from Redivis. Training requires reading the file into memory and distributing the workload across many CPUs or nodes.

Storage solutions:

Use Case 2: Many Small Files

Example: an AI pipeline using pre-tiled Landsat imagery — 9.8TB spread across roughly 105,000 files of about 100MB each. The pipeline runs in Python on GPUs with PyTorch and CUDA, and every file needs to be transferred into CPU memory before being loaded and unloaded from the GPU in batches.

Storage solutions:

Getting Help