User Tools

Site Tools


en:first_steps:tutorial

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
en:first_steps:tutorial [2026/09/27 13:58] – created pchojeckien:first_steps:tutorial [2026/09/30 01:58] (current) – pchojecki
Line 1: Line 1:
 ====== Tutorial ====== ====== Tutorial ======
  
-Celem tej instrukcji jest wykonanie pierwszych kroków na klastrze Eden. Po jej przejściu:+The goal of this tutorial is to guide you through your first steps on the Eden cluster. After completing it, you will be able to:
  
-  * zalogujesz się na Eden, +  * log in to Eden, 
-  * sprawdzisz dostępne zasoby, +  * check available resources, 
-  * uruchomisz pierwsze zadanie, +  * run your first job, 
-  * sprawdzisz jego stan, +  * check its status, 
-  * zakończysz zadanie.+  * stop a job.
  
-===== Logowanie =====+===== Logging in =====
  
-Logowanie do systemu Eden jest możliwe przez SSH.+You can access the Eden system via SSH.
  
-==== Z sieci wydziałowej ====+==== From the faculty network ====
  
-Będąc w sieci wydziałowej, możesz połączyć się bezpośrednio:+When connected to the faculty network, you can connect directly:
  
 <code bash> <code bash>
-ssh LOGIN_EDEN@eden.mini.pw.edu.pl+ssh EDEN_LOGIN@eden.mini.pw.edu.pl
 </code> </code>
  
-Zastąp ''LOGIN_EDEN'' swoim loginem na klastrze Eden.+Replace ''EDEN_LOGIN'' with your Eden cluster login.
  
-==== Spoza sieci wydziałowej ====+==== From outside the faculty network ====
  
-Spoza sieci wydziałowej możesz połączyć się pośrednio, wykorzystując serwer wydziałowy jako tzw. //jump host//.+From outside the faculty network, you can connect indirectly using the faculty server as a //jump host//.
  
-Służy do tego flaga ''-J'':+Use the ''-J'' option:
  
 <code bash> <code bash>
-ssh -J LOGIN_MINI@ssh.mini.pw.edu.pl LOGIN_EDEN@eden.mini.pw.edu.pl+ssh -J MINI_LOGIN@ssh.mini.pw.edu.pl EDEN_LOGIN@eden.mini.pw.edu.pl
 </code> </code>
  
-W tym przypadku najpierw następuje autoryzacja na serwerze wydziałowym, a następnie ruch jest automatycznie przekierowywany na klaster Eden.+In this case, you first authenticate to the faculty server, and the connection is then automatically forwarded to the Eden cluster.
  
-Zastąp:+Replace:
  
-  * ''LOGIN_MINI'' swoim loginem wydziałowym, +  * ''MINI_LOGIN'' with your faculty login, 
-  * ''LOGIN_EDEN'' swoim loginem na klastrze Eden.+  * ''EDEN_LOGIN'' with your Eden cluster login. 
 + 
 +After a successful login, you should see the Eden welcome screen: 
 + 
 +{{:pl:first_steps:tutorial_1.png?900|Eden welcome screen after logging in}}
  
 <WRAP center round important 80%> <WRAP center round important 80%>
-**Eden, na który logujesz się przez SSH, jest tylko węzłem dostępowym. Nie wykonujemy na nim obliczeń!**+**The Eden machine you log in to via SSH is only a login node. Do not run computations on it!**
 </WRAP> </WRAP>
  
-Po zalogowaniu mamy dostęp do terminala z systemem Linux oraz do systemu kolejkowego Slurm, za pomocą którego uruchamiamy zadania na klastrze.+After logging in, you have access to a Linux terminal and to the Slurm workload manager, which is used to run jobs on the cluster.
  
-===== Podstawowe komendy =====+===== Basic commands =====
  
-Przydatne polecenia:+Useful commands:
  
-  * ''sfree'' — wyświetla aktualnie wolne zasoby na klastrze. +  * ''sfree'' — displays currently available resources on the cluster. 
-  * ''pestat'' — pokazuje status węzłów i informacje o tym, kto aktualnie używa zasobów. +  * ''pestat'' — shows the status of compute nodes and information about who is currently using the resources. 
-  * ''squeue'' — wyświetla całą kolejkę zadań. +  * ''squeue'' — displays the entire job queue. 
-  * ''squeue -u LOGIN_EDEN'' — pokazuje tylko Twoje zadania. +  * ''squeue -u EDEN_LOGIN'' — displays only your jobs. 
-  * ''sinfo'' — wyświetla informacje o dostępnych partycjach (kolejkach) i stanie węzłów. +  * ''sinfo'' — displays information about available partitions (queues) and node status. 
-  * ''sshare -l'' — wyświetla informacje o udziałach (FairShare) i priorytetach użytkowników.+  * ''sshare -l'' — displays information about user shares (FairShare) and priorities.
  
-==== Zatrzymanie zadania ====+==== Stopping a job ====
  
-Jeżeli chcesz przerwać swoje zadanie, najpierw sprawdź jego identyfikator ''JOBID'':+If you want to stop one of your jobs, first find its ''JOBID'':
  
 <code bash> <code bash>
-squeue -u LOGIN_EDEN+squeue -u EDEN_LOGIN
 </code> </code>
  
-Następnie użyj:+Then run:
  
 <code bash> <code bash>
Line 71: Line 75:
 </code> </code>
  
-===== Uruchamianie zadań interaktywnych =====+===== Running interactive jobs =====
  
-Sesja interaktywna jest wygodna do krótkich testów, eksperymentowania oraz sprawdzania kodu na węźle obliczeniowym.+An interactive session is useful for short tests, experimentation, and checking your code directly on a compute node.
  
-Najpierw zobaczmy, jakie kolejki są dostępne:+First, let us see which partitions are available:
  
 <code bash> <code bash>
Line 81: Line 85:
 </code> </code>
  
-Widzimy m.in. kolejkę ''student''.+Example output:
  
-Możemy też sprawdzić, jakie zasoby są obecnie wolne:+{{:pl:first_steps:tutorial_2.png?900|Example output of the sinfo command}} 
 + 
 +We can see, among others, the ''student'' partition. 
 + 
 +We can also check which resources are currently available:
  
 <code bash> <code bash>
Line 89: Line 97:
 </code> </code>
  
-O, widzimy że zarówno węzeł ''stud-1'', jak i ''stud-2'' mają wolne CPU. Uruchommy sesję z:+Example output: 
 + 
 +{{:pl:first_steps:tutorial_3.png?900|Example output of the sfree command}} 
 + 
 +We can see that both ''stud-1'' and ''stud-2'' have free CPUs. Let us start a session requesting:
  
   * 1 CPU,   * 1 CPU,
-  * 4 GB RAM, +  * 4 GB of RAM, 
-  * maksymalnym czasem pracy 30 minut.+  * a maximum running time of 30 minutes.
  
-Uruchamiamy:+Run:
  
 <code bash> <code bash>
Line 101: Line 113:
 </code> </code>
  
-Zastąp ''GROUP_NAME'' nazwą swojej grupy.+Replace ''GROUP_NAME'' with the name of your group.
  
-Gdy system przyzna nam zasoby, możemy sprawdzić, na jakim węźle się znajdujemy:+Once Slurm allocates the requested resources, we can check which node we are on:
  
 <code bash> <code bash>
Line 109: Line 121:
 </code> </code>
  
-Widzimy, że dostaliśmy ''stud-1''. Oznacza to, że jesteśmy już na węźle obliczeniowym. Możemy tutaj uruchamiać obliczenia, testować kod, uruchomić Pythona itd.+Example interactive session:
  
-Po zakończeniu pracy wychodzimy z sesji:+{{:pl:first_steps:tutorial_4.png?1000|Starting an interactive session and checking the node name}} 
 + 
 +We can see that we were assigned ''stud-1''. Notice that the terminal prompt has also changed from ''@eden'' to ''@stud-1''. This means that we are now on a compute node. 
 + 
 +We can now run computations, test our code, start Python, etc. 
 + 
 +When we finish working, we leave the session with:
  
 <code bash> <code bash>
Line 118: Line 136:
  
 <WRAP center round important 80%> <WRAP center round important 80%>
-**Jeżeli przerwie się połączenie SSH, obliczenia uruchomione w sesji interaktywnej również zostaną przerwane.**+**If your SSH connection is interrupted, computations running in the interactive session will also be terminated.**
  
-Można temu zapobiec, korzystając np. z ''tmux'' albo ''screen''.+You can avoid this by using tools such as ''tmux'' or ''screen''.
 </WRAP> </WRAP>
  
-==== Co oznaczają poszczególne opcje? ====+==== What do the individual options mean? ====
  
-  * ''srun'' — narzędzie służące do uruchamiania zadań. W tej konfiguracji pozwala na pracę interaktywną. +  * ''srun'' — a command used to run jobs. In this configuration, it starts an interactive session. 
-  * ''-p student'' — wybiera partycję (kolejkę) o nazwie ''student''. +  * ''-p student'' — selects the partition (queue) named ''student''. 
-  * ''-A GROUP_NAME'' — przypisuje zadanie do konkretnej grupy rozliczeniowej lub projektu. +  * ''-A GROUP_NAME'' — assigns the job to a particular accounting group or project. 
-  * ''--cpus-per-task=1'' — żąda jednego rdzenia CPU. +  * ''--cpus-per-task=1'' — requests one CPU core. 
-  * ''--gres=gpu:NUM_OF_GPUS'' — żąda dostępu do określonej liczby GPU, np. ''--gres=gpu:1'' +  * ''--gres=gpu:NUM_OF_GPUS'' — requests access to a specified number of GPUs, for example ''--gres=gpu:1''. 
-  * ''--mem=MEM_IN_MBYTES'' — rezerwuje pamięć RAM dla całego zadania, np. ''--mem=4000M'', lub ''--mem=4G'' +  * ''--mem=MEM_IN_MBYTES'' — reserves RAM for the entire job, for example ''--mem=4000M'' or ''--mem=4G''. 
-  * ''--time=TIME'' — ustala maksymalny czas trwania zadania w formacie ''DD-HH:MM:SS'', np. ''--time=00:30:00''. Po upływie tego czasu zadanie zostanie zakończone. +  * ''--time=TIME'' — sets the maximum running time of the job in the ''DD-HH:MM:SS'' format, for example ''--time=00:30:00''. The job will be terminated after this time limit is reached. 
-  * ''--pty bash'' — tworzy interaktywny pseudoterminal z powłoką Bash.+  * ''--pty bash'' — creates an interactive pseudo-terminal running the Bash shell.
  
-===== Zadania wykonywane w tle =====+===== Running jobs in the background =====
  
-Do dłuższych obliczeń nie należy używać sesji interaktywnej. Lepszym rozwiązaniem są zadania wysyłane za pomocą ''sbatch''.+For longer computations, you should not use an interactive session. A better solution is to submit jobs using ''sbatch''.
  
-Opis zadania oraz wymagane zasoby zapisujemy w pliku tekstowym.+The job description and requested resources are specified in a text file.
  
-Plik składa się z:+The file consists of:
  
-  * dyrektyw Slurma — linii zaczynających się od ''#SBATCH'', +  * Slurm directives — lines beginning with ''#SBATCH'', 
-  * poleceń, które mają zostać wykonane.+  * commands that should be executed.
  
-Pierwsza linia powinna wskazywać interpreter, np.:+The first line should specify the interpreter, for example:
  
 <code bash> <code bash>
Line 151: Line 169:
 </code> </code>
  
-==== Przykład krok po kroku ====+==== Step-by-step example ====
  
-=== 1. Przygotowanie programu ===+=== 1. Preparing the program ===
  
-Utwórzmy prosty program w Pythonie:+Let us create a simple Python program:
  
 <code bash> <code bash>
Line 161: Line 179:
 </code> </code>
  
-=== 2. Przygotowanie skryptu Slurm ===+=== 2. Preparing the Slurm script ===
  
-Tworzymy plik ''hello.sh'' o następującej zawartości:+Create a file called ''hello.sh'' with the following contents:
  
 <code bash> <code bash>
Line 179: Line 197:
 </code> </code>
  
-Zastąp ''GROUP_NAME'' nazwą swojej grupy.+Replace ''GROUP_NAME'' with the name of your group.
  
-Symbol ''%j'' w nazwie pliku zostanie automatycznie zastąpiony numerem ''JOBID'' danego zadania.+The ''%j'' placeholder in the output filename will automatically be replaced with the unique ''JOBID'' of your job.
  
-=== 3. Utworzenie katalogu na logi ===+=== 3. Creating the log directory ===
  
-Przed wysłaniem zadania upewnij się, że katalog na logi istnieje.+Before submitting the job, make sure that the log directory exists.
  
-Slurm nie utworzy go automatycznie.+Slurm will not create it automatically.
  
 <code bash> <code bash>
Line 193: Line 211:
 </code> </code>
  
-=== 4. Wysłanie zadania ===+=== 4. Submitting the job ===
  
-Wysyłamy zadanie do kolejki:+Submit the job to the queue:
  
 <code bash> <code bash>
Line 201: Line 219:
 </code> </code>
  
-Slurm zwróci komunikat podobny do:+Slurm will return a message similar to:
  
 <code> <code>
Line 207: Line 225:
 </code> </code>
  
-Liczba na końcu to ''JOBID'' zadania.+Example:
  
-Możemy sprawdzić, czy zadanie znajduje się w kolejce:+{{:pl:first_steps:tutorial_5.png?700|Submitting a job using sbatch}} 
 + 
 +The number at the end is the job's ''JOBID''. 
 + 
 +We can check whether the job is currently in the queue:
  
 <code bash> <code bash>
-squeue -u LOGIN_EDEN+squeue -u EDEN_LOGIN
 </code> </code>
  
-=== 5. Sprawdzenie wyniku ===+=== 5. Checking the result ===
  
-Po zakończeniu zadania wynik znajdziemy w katalogu ''slurm_logs''.+After the job finishes, its output can be found in the ''slurm_logs'' directory.
  
-Przykładowo, jeżeli ''JOBID'' zadania wynosi ''12345'':+For example, if the job's ''JOBID'' is ''12345'':
  
 <code bash> <code bash>
Line 225: Line 247:
 </code> </code>
  
-Powinniśmy zobaczyć:+You should see:
  
 <code> <code>
Line 231: Line 253:
 </code> </code>
  
-==== Zadanie, które możemy zobaczyć w kolejce ====+Example for a job with ''JOBID'' equal to ''1752516'':
  
-Poprzedni program wykonuje się tak szybko, że możemy nie zdążyć zobaczyć go za pomocą ''squeue''.+{{:pl:first_steps:tutorial_6.png?1000|Reading the job output from the log file}}
  
-Zmieńmy więc ''hello.py'' na:+==== A job that we can see in the queue ==== 
 + 
 +The previous program finishes so quickly that we may not have enough time to see it using ''squeue''. 
 + 
 +Let us therefore modify ''hello.py'':
  
 <code python> <code python>
Line 245: Line 271:
 </code> </code>
  
-Ponownie uruchamiamy:+Submit the job again:
  
 <code bash> <code bash>
Line 251: Line 277:
 </code> </code>
  
-i od razu sprawdzamy kolejkę:+and immediately check the queue:
  
 <code bash> <code bash>
-squeue -u LOGIN_EDEN+squeue -u EDEN_LOGIN
 </code> </code>
  
-Przez około 30 sekund powinniśmy zobaczyć nasze zadanie na liście.+Example output:
  
-Po jego zakończeniu ponowne wykonanie:+{{:pl:first_steps:tutorial_8.png?1000|A running job visible in the Slurm queue}} 
 + 
 +If the job has already started, we should be able to see it in the queue for approximately 30 seconds. 
 + 
 +In the ''ST'' column, the value ''R'' means that the job is currently running (//Running//). 
 + 
 +After the job finishes, running:
  
 <code bash> <code bash>
-squeue -u LOGIN_EDEN+squeue -u EDEN_LOGIN
 </code> </code>
  
-nie powinno już go pokazywać.+again should no longer show it.
  
-===== Czym jest Slurm? =====+===== What is Slurm? =====
  
-Slurm jest systemem zarządzania klastrami i planowania zadań powszechnie wykorzystywanym w środowiskach HPC i na superkomputerach.+Slurm is a cluster management and job scheduling system widely used in high-performance computing (HPC) environments and on supercomputers.
  
-Jego głównym zadaniem jest:+Its main responsibilities include:
  
-  * przydzielanie użytkownikom zasobów obliczeniowych, +  * allocating computing resources to users, 
-  * zarządzanie kolejką zadań, +  * managing the job queue, 
-  * uruchamianie zadań na dostępnych węzłach, +  * running jobs on available compute nodes, 
-  * zwalnianie zasobów po zakończeniu obliczeń.+  * releasing resources after computations finish.
  
-Typowy przebieg pracy wygląda następująco:+A typical workflow is as follows:
  
-  - Użytkownik przygotowuje skrypt ''sbatch'', w którym określa potrzebne zasoby i polecenia do wykonania, a następnie wysyła zadanie do kolejki. +  - The user prepares an ''sbatch'' script specifying the required resources and commands to execute, and then submits the job to the queue. 
-  - Slurm sprawdza dostępność zasobów. +  - Slurm checks whether the requested resources are available. 
-  - Gdy odpowiednie zasoby są dostępne, Slurm je rezerwuje i uruchamia zadanie. +  - When suitable resources become available, Slurm reserves them and starts the job. 
-  - Po zakończeniu obliczeń zasoby zostają zwolnione i mogą zostać przydzielone kolejnemu zadaniu.+  - After the computation finishes, the resources are released and can be allocated to another job.
  
-Oficjalna dokumentacja:+Official documentation:
  
-[[https://slurm.schedmd.com/documentation.html|Dokumentacja Slurm]]+[[https://slurm.schedmd.com/documentation.html|Slurm documentation]]
  
-===== Dalsze kroki =====+===== Next steps =====
  
 TODO TODO
 +
 +----
 +
 +//This tutorial was prepared based on materials created by **Tymon Tumialis**. We would like to thank him for preparing the instructions, examples, and graphical materials.//
en/first_steps/tutorial.1790535530.txt.gz · Last modified: by pchojecki