User Tools

Site Tools


en:first_steps:tutorial

Tutorial

The goal of this tutorial is to guide you through your first steps on the Eden cluster. After completing it, you will be able to:

  • log in to Eden,
  • check available resources,
  • run your first job,
  • check its status,
  • stop a job.

Logging in

You can access the Eden system via SSH.

From the faculty network

When connected to the faculty network, you can connect directly:

ssh EDEN_LOGIN@eden.mini.pw.edu.pl

Replace EDEN_LOGIN with your Eden cluster login.

From outside the faculty network

From outside the faculty network, you can connect indirectly using the faculty server as a jump host.

Use the -J option:

ssh -J MINI_LOGIN@ssh.mini.pw.edu.pl EDEN_LOGIN@eden.mini.pw.edu.pl

In this case, you first authenticate to the faculty server, and the connection is then automatically forwarded to the Eden cluster.

Replace:

  • MINI_LOGIN with your faculty login,
  • EDEN_LOGIN with your Eden cluster login.

After a successful login, you should see the Eden welcome screen:

Eden welcome screen after logging in

The Eden machine you log in to via SSH is only a login node. Do not run computations on it!

After logging in, you have access to a Linux terminal and to the Slurm workload manager, which is used to run jobs on the cluster.

Basic commands

Useful commands:

  • sfree — displays currently available resources on the cluster.
  • pestat — shows the status of compute nodes and information about who is currently using the resources.
  • squeue — displays the entire job queue.
  • squeue -u EDEN_LOGIN — displays only your jobs.
  • sinfo — displays information about available partitions (queues) and node status.
  • sshare -l — displays information about user shares (FairShare) and priorities.

Stopping a job

If you want to stop one of your jobs, first find its JOBID:

squeue -u EDEN_LOGIN

Then run:

scancel JOBID

Running interactive jobs

An interactive session is useful for short tests, experimentation, and checking your code directly on a compute node.

First, let us see which partitions are available:

sinfo

Example output:

Example output of the sinfo command

We can see, among others, the student partition.

We can also check which resources are currently available:

sfree

Example output:

Example output of the sfree command

We can see that both stud-1 and stud-2 have free CPUs. Let us start a session requesting:

  • 1 CPU,
  • 4 GB of RAM,
  • a maximum running time of 30 minutes.

Run:

srun -p student -A GROUP_NAME --cpus-per-task=1 --mem=4G --time=00:30:00 --pty bash

Replace GROUP_NAME with the name of your group.

Once Slurm allocates the requested resources, we can check which node we are on:

hostname

Example interactive session:

Starting an interactive session and checking the node name

We can see that we were assigned stud-1. Notice that the terminal prompt has also changed from @eden to @stud-1. This means that we are now on a compute node.

We can now run computations, test our code, start Python, etc.

When we finish working, we leave the session with:

exit

If your SSH connection is interrupted, computations running in the interactive session will also be terminated.

You can avoid this by using tools such as tmux or screen.

What do the individual options mean?

  • srun — a command used to run jobs. In this configuration, it starts an interactive session.
  • -p student — selects the partition (queue) named student.
  • -A GROUP_NAME — assigns the job to a particular accounting group or project.
  • –cpus-per-task=1 — requests one CPU core.
  • –gres=gpu:NUM_OF_GPUS — requests access to a specified number of GPUs, for example –gres=gpu:1.
  • –mem=MEM_IN_MBYTES — reserves RAM for the entire job, for example –mem=4000M or –mem=4G.
  • –time=TIME — sets the maximum running time of the job in the DD-HH:MM:SS format, for example –time=00:30:00. The job will be terminated after this time limit is reached.
  • –pty bash — creates an interactive pseudo-terminal running the Bash shell.

Running jobs in the background

For longer computations, you should not use an interactive session. A better solution is to submit jobs using sbatch.

The job description and requested resources are specified in a text file.

The file consists of:

  • Slurm directives — lines beginning with #SBATCH,
  • commands that should be executed.

The first line should specify the interpreter, for example:

#!/usr/bin/env bash

Step-by-step example

1. Preparing the program

Let us create a simple Python program:

echo 'print("Hello Eden")' > hello.py

2. Preparing the Slurm script

Create a file called hello.sh with the following contents:

#!/usr/bin/env bash
#SBATCH --partition student
#SBATCH --account=GROUP_NAME
#SBATCH --cpus-per-task=1
#SBATCH --gres=gpu:0
#SBATCH --mem=4G
#SBATCH --time=00:10:00
#SBATCH --job-name=hello_test
#SBATCH --output=slurm_logs/hello_test-%j.log
 
python3 hello.py

Replace GROUP_NAME with the name of your group.

The %j placeholder in the output filename will automatically be replaced with the unique JOBID of your job.

3. Creating the log directory

Before submitting the job, make sure that the log directory exists.

Slurm will not create it automatically.

mkdir -p slurm_logs

4. Submitting the job

Submit the job to the queue:

sbatch hello.sh

Slurm will return a message similar to:

Submitted batch job 1752516

Example:

Submitting a job using sbatch

The number at the end is the job's JOBID.

We can check whether the job is currently in the queue:

squeue -u EDEN_LOGIN

5. Checking the result

After the job finishes, its output can be found in the slurm_logs directory.

For example, if the job's JOBID is 12345:

cat slurm_logs/hello_test-12345.log

You should see:

Hello Eden

Example for a job with JOBID equal to 1752516:

Reading the job output from the log file

A job that we can see in the queue

The previous program finishes so quickly that we may not have enough time to see it using squeue.

Let us therefore modify hello.py:

import time
 
print("Hello Eden")
time.sleep(30)
print("Hello Eden after sleep")

Submit the job again:

sbatch hello.sh

and immediately check the queue:

squeue -u EDEN_LOGIN

Example output:

A running job visible in the Slurm queue

If the job has already started, we should be able to see it in the queue for approximately 30 seconds.

In the ST column, the value R means that the job is currently running (Running).

After the job finishes, running:

squeue -u EDEN_LOGIN

again should no longer show it.

What is Slurm?

Slurm is a cluster management and job scheduling system widely used in high-performance computing (HPC) environments and on supercomputers.

Its main responsibilities include:

  • allocating computing resources to users,
  • managing the job queue,
  • running jobs on available compute nodes,
  • releasing resources after computations finish.

A typical workflow is as follows:

  1. The user prepares an sbatch script specifying the required resources and commands to execute, and then submits the job to the queue.
  2. Slurm checks whether the requested resources are available.
  3. When suitable resources become available, Slurm reserves them and starts the job.
  4. After the computation finishes, the resources are released and can be allocated to another job.

Official documentation:

Slurm documentation

Next steps

TODO


This tutorial was prepared based on materials created by Tymon Tumialis. We would like to thank him for preparing the instructions, examples, and graphical materials.

en/first_steps/tutorial.txt · Last modified: by pchojecki