REPACSS GPU

RP account needed

REPACSS GPU is the GPU partition of REPACSS, a compute cluster with 8 nodes, each with 4 NVIDIA H100 NVL GPUs (94 GB), dual Intel Xeon Gold 6448Y CPUs, and 512 GB RAM. The cluster is power-aware and runs on variable energy. It is well suited to deep-learning training and inference and other GPU-offloaded code.

Submitting Jobs Documentation

You can run jobs at different sizes and durations on REPACSS GPU. The following lists the different queues that you can submit to, describing how many nodes you get, how long you can run, the type of resources you get, and the average wait time.

Jobs run through Slurm; the login node is for editing, compiling, and staging only. GPU jobs go to the h100 partition and must request GPUs explicitly, e.g. sbatch -p h100 --gres=gpu:nvidia_h100_nvl:2 job.sh or interactive -p h100 --gres=… for an interactive session (REPACSS recommends the interactive wrapper over salloc). Load CUDA only after you land on a GPU node; module load cuda fails on the login node. Max wall time isn't published; run sinfo to check current limits.

For more detail, please visit the [Job Basics] page.

Queue specifications Documentation

Queue CPU cores / node GPUs / node Num nodes Node RAM
h100
All GPU jobs run here; this is REPACSS GPU's only partition.
2x Intel Xeon Gold 6448Y (64 cores) 4 NVIDIA H100 NVL (94 GB vRAM) 8 512 GB

Software Documentation

No software usage data is currently reported for REPACSS GPU in XDMoD.

SEE ALL SOFTWARE AVAILABLE ON REPACSS GPU


Storage Documentation

REPACSS storage is allocated per group (shared by project members): $HOME for scripts and config, $SCRATCH for active job I/O (not for long-term storage), and $WORK for results to keep. Each group shares 9 TB across the three areas. Check your group's usage from a login node with quota -s.

For more information, visit the [Introduction to REPACSS: A Beginner’s Guide].

File System Documentation

Directory Path Quota Purge Backup Notes
Home $HOME See notes Persistent personal storage for user scripts and configuration files. Quota varies by allocation; check your group's usage with df -h /mnt/$(id -gn)
Scratch $SCRATCH See notes High-performance temporary storage space subject to periodic purging. Quota varies by allocation; check your group's usage with df -h /mnt/$(id -gn)
Work $WORK See notes Long-term storage for research outputs and work purposes. Quota varies by allocation; check your group's usage with df -h /mnt/$(id -gn)
Shared Area /mnt/SHARED-AREA See notes Quota varies by allocation.
Shared Scratch /mnt/SHARED-SCRATCH See notes Quota varies by allocation.

File Transfer Documentation

Globus Connect is the recommended way to move data in and out of REPACSS, through the collection named REPACSS. Do not use scp, sftp, rsync, or direct connections to the login nodes. For step by step instructions, see [REPACSS File Transfer with Globus].

Supported Methods Data Transfer Node / Globus Collection Notes
GLOBUS CONNECT | RECOMMENDED REPACSS Transfer files in the Globus web app