====== Tutorial ======
The goal of this tutorial is to guide you through your first steps on the Eden cluster. After completing it, you will be able to:
* log in to Eden,
* check available resources,
* run your first job,
* check its status,
* stop a job.
===== Logging in =====
You can access the Eden system via SSH.
==== From the faculty network ====
When connected to the faculty network, you can connect directly:
ssh EDEN_LOGIN@eden.mini.pw.edu.pl
Replace ''EDEN_LOGIN'' with your Eden cluster login.
==== From outside the faculty network ====
From outside the faculty network, you can connect indirectly using the faculty server as a //jump host//.
Use the ''-J'' option:
ssh -J MINI_LOGIN@ssh.mini.pw.edu.pl EDEN_LOGIN@eden.mini.pw.edu.pl
In this case, you first authenticate to the faculty server, and the connection is then automatically forwarded to the Eden cluster.
Replace:
* ''MINI_LOGIN'' with your faculty login,
* ''EDEN_LOGIN'' with your Eden cluster login.
After a successful login, you should see the Eden welcome screen:
{{:pl:first_steps:tutorial_1.png?900|Eden welcome screen after logging in}}
**The Eden machine you log in to via SSH is only a login node. Do not run computations on it!**
After logging in, you have access to a Linux terminal and to the Slurm workload manager, which is used to run jobs on the cluster.
===== Basic commands =====
Useful commands:
* ''sfree'' — displays currently available resources on the cluster.
* ''pestat'' — shows the status of compute nodes and information about who is currently using the resources.
* ''squeue'' — displays the entire job queue.
* ''squeue -u EDEN_LOGIN'' — displays only your jobs.
* ''sinfo'' — displays information about available partitions (queues) and node status.
* ''sshare -l'' — displays information about user shares (FairShare) and priorities.
==== Stopping a job ====
If you want to stop one of your jobs, first find its ''JOBID'':
squeue -u EDEN_LOGIN
Then run:
scancel JOBID
===== Running interactive jobs =====
An interactive session is useful for short tests, experimentation, and checking your code directly on a compute node.
First, let us see which partitions are available:
sinfo
Example output:
{{:pl:first_steps:tutorial_2.png?900|Example output of the sinfo command}}
We can see, among others, the ''student'' partition.
We can also check which resources are currently available:
sfree
Example output:
{{:pl:first_steps:tutorial_3.png?900|Example output of the sfree command}}
We can see that both ''stud-1'' and ''stud-2'' have free CPUs. Let us start a session requesting:
* 1 CPU,
* 4 GB of RAM,
* a maximum running time of 30 minutes.
Run:
srun -p student -A GROUP_NAME --cpus-per-task=1 --mem=4G --time=00:30:00 --pty bash
Replace ''GROUP_NAME'' with the name of your group.
Once Slurm allocates the requested resources, we can check which node we are on:
hostname
Example interactive session:
{{:pl:first_steps:tutorial_4.png?1000|Starting an interactive session and checking the node name}}
We can see that we were assigned ''stud-1''. Notice that the terminal prompt has also changed from ''@eden'' to ''@stud-1''. This means that we are now on a compute node.
We can now run computations, test our code, start Python, etc.
When we finish working, we leave the session with:
exit
**If your SSH connection is interrupted, computations running in the interactive session will also be terminated.**
You can avoid this by using tools such as ''tmux'' or ''screen''.
==== What do the individual options mean? ====
* ''srun'' — a command used to run jobs. In this configuration, it starts an interactive session.
* ''-p student'' — selects the partition (queue) named ''student''.
* ''-A GROUP_NAME'' — assigns the job to a particular accounting group or project.
* ''--cpus-per-task=1'' — requests one CPU core.
* ''--gres=gpu:NUM_OF_GPUS'' — requests access to a specified number of GPUs, for example ''--gres=gpu:1''.
* ''--mem=MEM_IN_MBYTES'' — reserves RAM for the entire job, for example ''--mem=4000M'' or ''--mem=4G''.
* ''--time=TIME'' — sets the maximum running time of the job in the ''DD-HH:MM:SS'' format, for example ''--time=00:30:00''. The job will be terminated after this time limit is reached.
* ''--pty bash'' — creates an interactive pseudo-terminal running the Bash shell.
===== Running jobs in the background =====
For longer computations, you should not use an interactive session. A better solution is to submit jobs using ''sbatch''.
The job description and requested resources are specified in a text file.
The file consists of:
* Slurm directives — lines beginning with ''#SBATCH'',
* commands that should be executed.
The first line should specify the interpreter, for example:
#!/usr/bin/env bash
==== Step-by-step example ====
=== 1. Preparing the program ===
Let us create a simple Python program:
echo 'print("Hello Eden")' > hello.py
=== 2. Preparing the Slurm script ===
Create a file called ''hello.sh'' with the following contents:
#!/usr/bin/env bash
#SBATCH --partition student
#SBATCH --account=GROUP_NAME
#SBATCH --cpus-per-task=1
#SBATCH --gres=gpu:0
#SBATCH --mem=4G
#SBATCH --time=00:10:00
#SBATCH --job-name=hello_test
#SBATCH --output=slurm_logs/hello_test-%j.log
python3 hello.py
Replace ''GROUP_NAME'' with the name of your group.
The ''%j'' placeholder in the output filename will automatically be replaced with the unique ''JOBID'' of your job.
=== 3. Creating the log directory ===
Before submitting the job, make sure that the log directory exists.
Slurm will not create it automatically.
mkdir -p slurm_logs
=== 4. Submitting the job ===
Submit the job to the queue:
sbatch hello.sh
Slurm will return a message similar to:
Submitted batch job 1752516
Example:
{{:pl:first_steps:tutorial_5.png?700|Submitting a job using sbatch}}
The number at the end is the job's ''JOBID''.
We can check whether the job is currently in the queue:
squeue -u EDEN_LOGIN
=== 5. Checking the result ===
After the job finishes, its output can be found in the ''slurm_logs'' directory.
For example, if the job's ''JOBID'' is ''12345'':
cat slurm_logs/hello_test-12345.log
You should see:
Hello Eden
Example for a job with ''JOBID'' equal to ''1752516'':
{{:pl:first_steps:tutorial_6.png?1000|Reading the job output from the log file}}
==== A job that we can see in the queue ====
The previous program finishes so quickly that we may not have enough time to see it using ''squeue''.
Let us therefore modify ''hello.py'':
import time
print("Hello Eden")
time.sleep(30)
print("Hello Eden after sleep")
Submit the job again:
sbatch hello.sh
and immediately check the queue:
squeue -u EDEN_LOGIN
Example output:
{{:pl:first_steps:tutorial_8.png?1000|A running job visible in the Slurm queue}}
If the job has already started, we should be able to see it in the queue for approximately 30 seconds.
In the ''ST'' column, the value ''R'' means that the job is currently running (//Running//).
After the job finishes, running:
squeue -u EDEN_LOGIN
again should no longer show it.
===== What is Slurm? =====
Slurm is a cluster management and job scheduling system widely used in high-performance computing (HPC) environments and on supercomputers.
Its main responsibilities include:
* allocating computing resources to users,
* managing the job queue,
* running jobs on available compute nodes,
* releasing resources after computations finish.
A typical workflow is as follows:
- The user prepares an ''sbatch'' script specifying the required resources and commands to execute, and then submits the job to the queue.
- Slurm checks whether the requested resources are available.
- When suitable resources become available, Slurm reserves them and starts the job.
- After the computation finishes, the resources are released and can be allocated to another job.
Official documentation:
[[https://slurm.schedmd.com/documentation.html|Slurm documentation]]
===== Next steps =====
TODO
----
//This tutorial was prepared based on materials created by **Tymon Tumialis**. We would like to thank him for preparing the instructions, examples, and graphical materials.//