# REAL HPC documentation

> Plain Markdown copy of the How to page.

[Open the website view](/cluster/#docs)

## 0. Setup

### SSH access to real

On your laptop, use an existing SSH key or create one at an unused path:

```bash
ssh-keygen -t ed25519 -f ~/.ssh/id_ed25519_real
```

Send the public key (.pub) to the cluster admin and ask for an account on the real front node. Keep the private key on your laptop. Add your assigned username and key path to ~/.ssh/config:

```bash
Host real
    HostName real3.itu.dk
    User USER
    IdentityFile ~/.ssh/id_ed25519_real
```

### Backup of /home and extra storage

Get an ITU HPC account and authorize an SSH public key on it through ITU HPC support. Keep its private key in your laptop’s ~/.ssh folder. Download and run the setup script on your laptop:

```bash
scp real:/opt/real-cluster-source/bin/setup-itu-backup.sh ./setup-itu-backup.sh
bash ./setup-itu-backup.sh
```

Enter real, your ITU username, and the path to your ITU private key when prompted. The script creates a separate restricted backup key; your ITU login key stays on your laptop.

**Extra storage:** Your ITU HPC account’s home is not only for the backup: you can keep all your data there. It is on a shared 321 TB disk with no per-user limit.

## 1. SSH into real

```bash
ssh -A real
```

### Get your code onto real

Forward your laptop’s SSH agent, then clone or pull from GitHub. This uses your loaded laptop key without copying its private key to real. Only forward your agent to a machine you trust. (Other options: a read-only deploy key for one repository, or gh auth login on real.)

```bash
git clone git@github.com:OWNER/REPOSITORY.git
cd REPOSITORY
```

## 2. Submit a job

A job script asks Slurm for resources, then runs your commands. main is the default for submitted background jobs and allows up to 24 hours. You do not need to name it in the submission command or job script. interactive is only for the documented Bash shell and allows up to 1 hour.

Every job must state --time; a job without it is refused. Set it to what the job needs plus some margin: the job then starts sooner, because Slurm can fit it into gaps and shorter requests get a priority bonus. A job that runs past its --time is stopped. Slurm accepts --time=0-4 (4 hours), --time=1-0 (a day), --time=240 (240 minutes), or --time=04:00:00; note that --time=4:00 means 4 minutes. Memory defaults to 8 GiB per requested CPU and is enforced. GPU jobs get four CPUs per GPU unless you request another amount. Interactive jobs are limited to 4 GPUs per user.

### job.sh: save this file in your project

```bash
#!/bin/bash
#SBATCH --gpus=1
#SBATCH --time=00:05:00
# A private Slack message from REAL HPC when this job ends:
#SBATCH --mail-type=END
#SBATCH --mail-user=slack

set -euo pipefail
uv run --locked python script.py
```

set -euo pipefail stops the script when a command fails, an unset variable is used, or a command in a pipeline fails. Replace uv run --locked python script.py with the command that runs your code. uv is installed on real and the compute machines. Jobs currently have outbound internet, so uv can download locked packages; downloads use job time. Use nvidia-smi when you only want to check the GPU assigned to the job.

### Terminal: type these commands after saving job.sh

```bash
cd ~/YOUR_PROJECT
sbatch job.sh
```

Where to submit from, and the trade-offs of each way, are in Shared vs scratch below.

[Current machines](/cluster/machines.md) · [Cluster policy](/cluster/policy.md)

## Example job scripts

- [Download gpu-check-job.sh](/cluster/job-examples/gpu-check-job.sh): one GPU, checks the assigned GPU.
- [Download one-gpu-uv-job.sh](/cluster/job-examples/one-gpu-uv-job.sh): one GPU, runs a uv Python script.
- [Download four-gpu-torchrun-job.sh](/cluster/job-examples/four-gpu-torchrun-job.sh): four GPUs, runs PyTorch distributed training.
- [Download cpu-memory-job.sh](/cluster/job-examples/cpu-memory-job.sh): one GPU with explicit CPU and RAM requests.
- [Download interactive-notebook.sh](/cluster/job-examples/interactive-notebook.sh): Jupyter or VS Code in an interactive shell (run it inside srun --partition=interactive ...).

## Do's and don'ts

**Do:** run installs (uv sync, pip install, conda) and Jupyter or VS Code sessions inside a job: sbatch, or an interactive shell with interactive-notebook.sh. The same goes for large downloads, into your home or scratch. Keep Jupyter and VS Code protected by their token, as the script does.

**Don't:** run them on real. real is one machine everyone shares (it also runs Slurm, the homes, and the website); heavy or long-running processes there slow down everyone else. Don't keep services of your own, crontabs, or lingering, or open ports on real either: run your work as Slurm jobs.

## Batch vs interactive

### Batch job

Save your commands in a job script and submit it with sbatch (see Shared vs scratch for where to submit from). Slurm starts it in the main queue when the requested resources are free, for up to 24 hours. You can log out while it waits or runs; its output goes to the file named by --output.

When to use: training runs, sweeps, and anything that runs without you watching it.

```bash
cd ~/YOUR_PROJECT/.worktrees/RUN
sbatch job.sh
squeue --me
```

### Interactive shell

This is the only command form accepted by the interactive partition. You can change its resource amounts and time. Slurm waits for those resources, selects the compute node, and opens Bash there. Exit the shell to release the resources. Direct SSH to compute nodes remains administrator-only.

When to use: debugging, short tests, and checking that your code and environment work on a GPU before you submit a long batch job.

```bash
srun --partition=interactive --gpus=1 --cpus-per-task=4 --time=01:00:00 --pty bash -l
```

## Shared vs scratch

Shared mode works on your project files in your home; scratch mode works on a private copy on the compute machine. The tabs explain each way and its trade-offs.

### Shared mode

#### With a git worktree

Make a separate git worktree for each run and submit from it with sbatch. The job runs that fixed copy of the code, so you can keep editing and pulling in your main checkout. The worktree holds the last commit only, so commit your changes first, and add .worktrees/ and logs/ to .gitignore. Write results to an absolute path in your home outside the worktree, such as $HOME/YOUR_PROJECT/results: they are saved as the job writes them, and a job stopped at its time limit loses only the work in progress. Cost: each worktree is a full copy of the code in your home with its own .venv, and the job reads both over the network. Building a new worktree’s .venv is quick: uv links packages from its cache in your home (~/.cache/uv) instead of copying them. Remove the worktree when the job has ended; git worktree remove refuses if it holds files you have not committed.

When to use: training runs and sweeps, especially long ones or many at once while you keep changing the code.

```bash
cd ~/YOUR_PROJECT
mkdir -p logs
RUN=.worktrees/$(date +%Y%m%d-%H%M%S)
git worktree add --detach "$RUN"
(cd "$RUN" && sbatch --output="$HOME/YOUR_PROJECT/logs/job-%j.log" job.sh)

# After the job has ended:
git worktree remove "$RUN"
```

#### Plain

Use this when you want the job to read and write the shared project directly. There is no project copy or output copy-back. A git pull, branch switch, or file edit can change code used by processes that start later in a pending or running job.

When to use: jobs that spend their time on CPU or GPU work and read or write few files, when you will not edit the project while the job waits or runs.

```bash
cd ~/YOUR_PROJECT
sbatch job.sh
```

#### With local scratch

Keep code and results in your home, and put large datasets and temporary files on the compute machine’s local disk. Add these lines to job.sh. stage-dataset --private copies a dataset folder from your home (path relative to your home) to local scratch once, and later jobs on the same machine reuse it. Files in TMPDIR are not copied back; the last line deletes them.

When to use: large datasets that are read many times, or jobs that write many temporary files, where reading and writing over the shared home would slow the job.

```bash
export TMPDIR=/scratch/$USER/tmp/$SLURM_JOB_ID
mkdir -p "$TMPDIR"
DATA=$(stage-dataset --private datasets/MY_DATASET)
uv run --locked python train.py --data "$DATA" --output "$HOME/YOUR_PROJECT/results/$SLURM_JOB_ID"
rm -rf "$TMPDIR"
```

### Scratch mode

#### Plain

Use this when you want a private project copy while the job runs. cluster-submit copies Git-tracked files when the job starts, including uncommitted edits. Changes made while the job waits may be copied; later changes do not affect the running copy. Add #CLUSTER include= for other inputs and #CLUSTER copy-back= for results to return, including after an application failure. Git commands such as git rev-parse do not work in the scratch copy, which has no .git folder. --mode is required (--mode=shared is the same as plain sbatch). sbatch options go before the script in the --name=value form, for example cluster-submit --mode=scratch --output=$HOME/YOUR_PROJECT/logs/job-%j.log job.sh.

When to use: projects with many small files or heavy reading and writing inside the project folder, or when the job should use a fixed copy of the project including uncommitted edits.

```bash
cd ~/YOUR_PROJECT
cluster-submit --mode=scratch job.sh
```

#### Results in home

Run from a private scratch copy, but write results and checkpoints to an absolute path in your home. They are saved as the job writes them, so they do not depend on copy-back at the end of the job, and a time limit or crash loses only the work in progress. These paths need no #CLUSTER copy-back= line.

When to use: long scratch-mode runs whose checkpoints you cannot afford to lose.

```bash
#!/bin/bash
#SBATCH --gpus=1
#SBATCH --time=12:00:00
#SBATCH --output=training-%j.log
# A private Slack message from REAL HPC when this job ends:
#SBATCH --mail-type=END
#SBATCH --mail-user=slack
set -euo pipefail
uv run --locked python train.py --output "$HOME/YOUR_PROJECT/results/$SLURM_JOB_ID"

# Submit with:
# cd ~/YOUR_PROJECT
# cluster-submit --mode=scratch job.sh
```

#### Extra inputs and outputs

For scratch mode. Add these lines to the job script. Outputs return to the same relative path inside the original project. Choose a different output directory for each run. Non-Git projects need explicit inputs.

```bash
#CLUSTER include=datasets/small-test
#CLUSTER copy-back=results/
```

#### Recover partial results

For scratch mode. After the application exits, even with an error, the helper copies the declared outputs back. If that fails, or the job is stopped before it finishes (a time limit or cancellation gives it 5 minutes), the files stay in the job’s scratch folder on the compute machine until 7 days after the job ended, then they are deleted. The job log shows a ready-to-paste sbatch line when the job starts, and again at the end whenever the scratch copy is kept. Run it on real after the job has ended to copy them back to your project. Node loss can still lose scratch files, so write important checkpoints to an absolute path in your home during the job.

```bash
grep -A1 "copy them back" training-JOB_ID.log
```

## Other settings

### Inspect your jobs

squeue shows queued and running jobs. sacct includes completed jobs. sstat shows CPU and memory use so far. An overlapping job step can inspect the assigned GPUs. seff is not installed; use sacct and sstat. Terminal output files and application checkpoints are separate outputs.

```bash
squeue --me
sacct --starttime today --format=JobID,JobName,State,Elapsed,AllocTRES
sstat -j JOB_ID.batch --format=JobID,AveCPU,MaxRSS
srun --jobid=JOB_ID --overlap nvidia-smi
less training-JOB_ID.log
```

### Home space

Your home on real has no soft limit. Above 300 GB you see a notice when you log in to real, listing your largest folders and marking git repos and worktrees. From 400 GB your new jobs wait in the queue (reason AssocMaxJobsLimit) until your home is below 400 GB again; running jobs continue. 500 GB is the hard limit: above it you cannot write. Usage is checked every hour. Old worktrees and results you no longer need are the usual places to free space.

```bash
du -sh ~/* ~/.[!.]* 2>/dev/null | sort -h | tail -5
git worktree list
squeue --me --format="%.10i %.20j %.8T %R"
```

### GPU software

You cannot change the driver: it is set per machine and decides the newest CUDA version your programs can use. The shared CUDA toolkit is in /usr/local/cuda; its compiler is /usr/local/cuda/bin/nvcc. The NVIDIA HPC SDK compilers (nvc++, nvc, nvfortran) are not on your PATH, so call them by their full path. For another CUDA version, add it to your project environment with uv: PyTorch and JAX bring their own CUDA libraries, and the nvidia-cuda-nvcc package gives another nvcc. Never use a CUDA version newer than the driver runs. Slurm gives each job only the GPUs it asked for.

Current versions on each machine: [machines.md](/cluster/machines.md)

```bash
# Inside a job or an interactive shell on a compute machine:
nvidia-smi
/usr/local/cuda/bin/nvcc --version
ls /opt/nvidia/hpc_sdk/Linux_x86_64/
```

### Containers (Apptainer)

Apptainer runs containers in your jobs. It works the same way as Docker and Singularity, and is compatible with Docker images: it runs them straight from a registry (docker://...), or you can pull an image once into a local .sif file and reuse it. It is the successor of Singularity: the singularity command also works, and existing .sif images run unchanged. Add --nv to use GPUs; the container sees only your job’s GPUs. Differences from Docker: there is no docker command and no daemon, the container runs as you (no root inside), and your home folder is available inside it. In jobs, downloaded images are cached in /scratch/$USER/cache/apptainer (removed after 14 days unused, like the other caches); on real the cache is ~/.apptainer/cache in your home.

```bash
#!/bin/bash
#SBATCH --partition=main
#SBATCH --gpus=1
#SBATCH --time=00:15:00
# A private Slack message from REAL HPC when this job ends:
#SBATCH --mail-type=END
#SBATCH --mail-user=slack
apptainer exec --nv docker://nvidia/cuda:13.0.0-base-ubuntu22.04 nvidia-smi -L

# To keep a local image file and reuse it:
# apptainer pull pytorch.sif docker://pytorch/pytorch
# apptainer exec --nv pytorch.sif python train.py
```

### Caches

Caches are set up for you in every job. Hugging Face, Torch, pip, and other downloads are cached on the compute machine in /scratch/$USER/cache (XDG_CACHE_HOME). uv packages are cached in ~/.cache/uv in your home, so a .venv in your repo is filled by linking, not copying. Scratch mode keeps the uv cache on scratch. To change this, set XDG_CACHE_HOME, HF_HOME, or UV_CACHE_DIR in your job script. A cache folder on scratch with nothing used for 14 days is deleted; the next job downloads again.

```bash
# Inside a job:
echo "$XDG_CACHE_HOME" "$UV_CACHE_DIR"
du -sh /scratch/$USER/cache
```

### Slack notifications

Add these two lines to your job script to get a private Slack message from the REAL HPC app when your job ends (it is not an email): how it ended, how long it ran, and where its log is. It needs your Slack member ID on the cluster; ask the admin. Leave them out if you do not want the message.

```bash
# A private Slack message from REAL HPC when this job ends:
#SBATCH --mail-type=END
#SBATCH --mail-user=slack
```

## Jobs examples

### One GPU

Submit with sbatch (see Shared vs scratch). Main is selected by default. The program must accept --output, or change that argument to match your program. Results go to an absolute path in your home, so they are saved as the job writes them.

```bash
#!/bin/bash
#SBATCH --partition=main
#SBATCH --gpus=1
#SBATCH --time=02:00:00
# A private Slack message from REAL HPC when this job ends:
#SBATCH --mail-type=END
#SBATCH --mail-user=slack
set -euo pipefail
uv run --locked python train.py --output "$HOME/YOUR_PROJECT/results/$SLURM_JOB_ID"
```

### Four GPUs

For training code that already supports four GPU workers. Requesting four GPUs does not make single-GPU code use them. Submit with sbatch (see Shared vs scratch).

```bash
#!/bin/bash
#SBATCH --partition=main
#SBATCH --gpus=4
#SBATCH --cpus-per-task=16
#SBATCH --time=02:00:00
#SBATCH --output=training-%j.log
# A private Slack message from REAL HPC when this job ends:
#SBATCH --mail-type=END
#SBATCH --mail-user=slack
set -euo pipefail
uv run --locked torchrun --nproc-per-node=4 train.py
```

### More CPUs and RAM

Request 16 CPUs and 192 GiB of RAM for one GPU. These requests fit vapor; it has 32 logical CPUs. Submit with sbatch (see Shared vs scratch).

```bash
#!/bin/bash
#SBATCH --partition=main
#SBATCH --gpus=1
#SBATCH --cpus-per-task=16
#SBATCH --mem=192G
#SBATCH --time=02:00:00
#SBATCH --output=training-%j.log
# A private Slack message from REAL HPC when this job ends:
#SBATCH --mail-type=END
#SBATCH --mail-user=slack
set -euo pipefail
uv run --locked python train.py
```
