Expanse GPU

Expanse GPU has 56 nodes with four NVIDIA V100s (32 GB each) linked by NVLink. Suits multi-GPU AI/ML training and other workloads that scale across GPUs.

Submitting Jobs

 

GPU nodes are allocated as a separate resource. The GPU nodes can be accessed via either the "gpu" or the "gpu-shared" partitions.

#SBATCH -p gpu

or

#SBATCH -p gpu-shared

 

When users request 1 GPU, in gpu-shared partition, by default they will also receive, 1 CPU, and 1G memory.  

For more information and example job scripts see Expanse Running Jobs.

Queue specifications

Metrics updated 2026-09-10

Name Purpose CPU cores / node GPUs / node Num nodes Node RAM Max wallclock Wait time
30-day trend
Wall time
30-day trend
Number of jobs run
30 days
gpu Used for exclusive access to the GPU nodes Xeon Gold 6248 (40 cores) 4 NVIDIA V100 SXM2 (32 GB vRAM) 56 384 GB 48h
gpu wait time: average 23.1 hours, range 4.4 to 68.2 hours over 30 days
gpu wall time: average 13.8 hours, range 0.4 to 15.8 hours over 30 days, wall-time limit 48h
480
gpu-debug Priority access to gpu-shared nodes set aside for testing of jobs with short wall time and limited resources; max two gpus per job Xeon Gold 6248 (40 cores) 4 NVIDIA V100 SXM2 (32 GB vRAM) 2 384 GB 30m
gpu-debug wait time: average 0.0 hours, range 0 to 0 hours over 30 days
gpu-debug wall time: average 0.1 hours, range 0 to 0.3 hours over 30 days, wall-time limit 30m
1,881
gpu-shared Non-refundable discounted jobs to run on unallocated nodes that can be pre-empted by higher priority queues Xeon Gold 6248 (40 cores) 4 NVIDIA V100 SXM2 (32 GB vRAM) 56 384 GB 48h
gpu-shared wait time: average 0.8 hours, range 0 to 16.9 hours over 30 days
gpu-shared wall time: average 1.5 hours, range 0.3 to 9.4 hours over 30 days, wall-time limit 48h
41,727
nairr-gpu Expanse AI resource using compute with Nvidia GPUs 2 Intel Sapphire Road processors (36 cores) 4 Nvidia H100 (80 GB vRAM) 34 1 TB 48h N/A N/A N/A

Software

No software usage data is currently reported for Expanse GPU in XDMoD.

SEE ALL SOFTWARE AVAILABLE ON EXPANSE GPU


Storage

Expanse has a small backed-up Home for source code and configuration, a 12 PB Lustre parallel file system for active scratch and project data, and fast node-local SSD that exists only for the duration of a job. There is no archival tier and the Lustre file system is not backed up, so anything you cannot lose has to be copied off Expanse yourself.

SDSC limits each user to 2 million files under /expanse/lustre/scratch. If your workflow needs a very large number of small files, contact support first - it loads the metadata server and can affect the whole system. Node-local scratch is a different size on each node type; the Notes column gives the figures.

Jobs that need Lustre must ask for it explicitly with #SBATCH --constraint="lustre", which can be combined with other constraints. A job that needs Lustre and omits the constraint is scheduled on a node without it and will fail. See [Expanse Storage] for the full description.

File System

Directory Path Quota Purge Backup Notes
Scratch Lustre /expanse/lustre/scratch 10 TB 90 days after allocation expiration. Not backed up​ This is not an archival file system, it is not backed up, and will be purged according to purge policy.
Scratch GPU Node /scratch/$USER/job_$SLURM_JOB_ID 1600 GB Users only have access to these SSDs during job execution at the local file system path to the gpu node.
Home /home 100 GB 8 week rolling backup The home directory is limited in space and should be used only for source code storage. Jobs should never be run from the home file system, as it is not set up for high performance throughput.

External Storage

Expanse has no archival tier of its own. Project space is purged 90 days after the allocation expires and Lustre is never backed up, so move anything you need to keep to an external archive before your allocation ends.

To see which projects you can charge to and how much of your allocation is left, run expanse-client on a login node. See [Expanse Account Management] for its syntax.


File Transfer

Use Globus for anything large or made of many files - it retries and resumes on its own. Expanse has no separate data transfer node, so command-line transfers with scp, sftp or rsync go to login.expanse.sdsc.edu, and SDSC asks you not to use the login nodes for large or numerous transfers. See [Expanse Data Movement] for the Globus collections and their mount points.

Supported Methods Data Transfer Node / Globus Collection Notes
GLOBUS | RECOMMENDED SDSC HPC - Expanse Lustre Globus

Datasets

Name Description
OceanTopography

OpenTopography provides efficient, user-friendly access to high-resolution topography data, processing tools, and resources to advance understanding of the Earth's surface, vegetation, and built environment.

OpenAltimetry

OpenAltimetry is a web based data visualization and discovery tool for exploring surface elevation profiles over time using satellite altimetry data from NASA's ICESat and ICESat-2 missions.

OpenForest4D

OpenForest4D is a web-based platform that leverages multi-source remote sensing data and artificial intelligence to generate on-demand, research-grade estimates of forest structure and above-ground biomass in four dimensions for global forest monitoring.