Interactive & Parallel Jobs#
Not every job is a batch script. Use spur run (Slurm srun) to run a
command directly as a job or job step, spur alloc (Slurm salloc) to get
an interactive shell on allocated resources, and spur attach (Slurm
sattach) to connect to a running job’s input and output. This page covers
all three and how to request GPUs for them.
Run a Command — srun#
spur run (Slurm srun) runs a command as a job and streams its output
live. It operates in two modes:
Standalone — when run outside an allocation, it submits a new job, waits for the allocation, streams output from the node agent, and exits with the job’s exit code. Pressing Ctrl-C cancels the job. After the allocation is granted, standalone
srun(including--pty) resolves the job owner from the controller for step, keepalive, and cancel RPCs, the same waysrundoes inside ansallocshell viaSPUR_JOB_USER.Step mode — when run inside an existing allocation (
SPUR_JOB_IDis set, as undersallocor in a batch script), it creates a job step against the parent allocation instead of submitting a new job.
Inside an allocation, a bare srun inherits the allocation’s size
(--ntasks, --cpus-per-task, nodes, partition, account, QOS), so it runs
at the allocation’s scale unless you override on the command line.
A step is named after the command line that created it. That name is stored
capped at 256 bytes, so a longer command line is truncated on a character
boundary and ends in ... where scontrol show step prints it. The step
still runs the full command — only the displayed name is shortened.
Common options:
Long |
Short |
Description |
|---|---|---|
|
|
Number of nodes. Standalone default |
|
|
Number of tasks. Standalone default |
|
|
CPUs per task. Default |
|
Generic resources, e.g. |
|
|
|
GPU shorthand, folded into |
|
Per-task GPU binding: |
|
|
|
Partition to run in. |
|
|
Wall-clock limit. |
|
Allocate a pseudo-terminal (use with an interactive shell). Inside an
existing allocation this runs as an interactive step: a single task on
one node, started in the job’s environment and working directory
rather than your shell’s. |
|
|
|
Prefix each output line with its task rank. |
|
|
File for stdout. |
|
|
File for stderr. |
|
Target a running job (use with |
|
|
Share resources with the targeted job’s existing steps. |
Examples:
srun -N2 -n16 --gres=gpu:8 hostname # parallel command across 2 nodes
spur run --pty bash # interactive shell on an allocated node
srun -n4 python train.py
srun --jobid 1024 --overlap rocm-smi # run a command inside a running job
Note
--jobid requires --overlap; --jobid alone is an error.
--input/-i is ignored in step mode.
Interactive Allocation — salloc#
spur alloc (Slurm salloc) requests an interactive allocation, waits for
it to start (up to 300 seconds), then spawns your $SHELL with the allocation
environment exported (SPUR_JOB_ID, SPUR_JOB_USER, SPUR_NODELIST,
SPUR_NNODES, SPUR_NTASKS, SPUR_CPUS_PER_TASK, the
partition/account/QOS variables, and their SLURM_* twins). When you exit the
shell, the allocation is released. Ctrl-C cancels it.
When JWT authentication is enabled, salloc also passes $SPUR_AUTH_TOKEN
(or ~/.spur/token) into the allocation shell so step commands can
authenticate to the controller. With [auth] plugin = "spur", the shell does
not inherit a bearer; each command mints a fresh credential from the local
socket. SPUR_JOB_USER records the job owner bound at submit time (for
example the JWT subject); srun inside the shell uses it when step RPCs run
without a token.
Inside that shell, srun runs as a job step. With no step-level -N or
-w, it uses the allocation’s nodes and inherits the allocation’s task count
(SPUR_NTASKS / SLURM_NTASKS). For buffered steps, -N/--nodes
limits the step to that many allocated nodes without dropping that inherited
task count, and -w/--nodelist selects an exact subset of allocated
nodes. A step cannot request more nodes than its allocation.
--pty does not follow those node-selection rules: it still ignores
--nodes (one task on one node) and, with -w, uses only the first listed
name.
The -N/-n defaults of 1 apply when srun submits a new job, not
when it runs as a step inside salloc/sbatch. Common options: --nodes/
-N, --ntasks/-n, --cpus-per-task/-c, --mem, --time/
-t, --gres, --gpus/-G, --partition/-p, --constraint/
-C, --nodelist/-w, --exclude/-x, --reservation, and
--exclusive.
Examples:
salloc -N1 --gres=gpu:2 -t 2:00:00
spur alloc --partition gpu --exclusive
Attach to a Running Job — sattach#
spur attach (Slurm sattach) connects to a running job’s I/O. The
positional argument is JOB_ID or JOB_ID.STEP_ID; the step-id component is
accepted, but only the job id is used to route the attach. By default it opens a
full interactive raw-mode terminal attached through the node agent. Attach RPCs
use the job owner from SPUR_JOB_USER (when set), the controller job record, or
the local login name, matching step-mode srun under JWT authentication.
Option |
Description |
|---|---|
|
Which stream to attach to. Default |
|
Stream output one-way instead of an interactive terminal. |
sattach 1024
sattach 1024 --output-only --output stderr
Surviving an Agent Restart#
Work is supervised by a process that outlives the node agent, so an
administrator restarting or upgrading spurd does not kill it. An allocation
stays held, a running srun step keeps running and still reports its exit
status. A terminal is the exception: the allocation and its supervisors survive,
but a terminal already open inside it ends, and you reconnect by running
srun --pty again. A step that finishes while the agent is down is covered
too: its supervisor records the exit durably, that record is kept for ten
minutes after the step settles, and the restarted agent answers the waiting
caller with the exit and the step’s output instead of reporting the step as
lost. Any number of restarts inside that window are covered; a caller that only
comes back after it is still told the step is unknown. A session that nothing
can ever settle — one whose supervisor was killed before it recorded an exit —
is kept for a day and then swept.
One case is not covered: a step given its own --container-image. It ends
when the agent stops, and leaves its unpacked rootfs behind. A standalone
srun --pty that allocates its own job loses the job as well: srun
cancels the job it created as soon as its connection drops. Take an allocation
first and run srun --pty inside it, so a restart costs you the terminal
rather than the work:
salloc -N1
srun --pty bash # inside the allocation shell
Requesting GPUs#
srun, salloc, and batch jobs all request GPUs the same way:
srun --gres=gpu:2 python infer.py # 2 GPUs of any type
srun --gres=gpu:mi300x:8 python train.py # 8 GPUs of a specific type
srun -G 4 python infer.py # -G shorthand
salloc --gres=gpu:2 -t 1:00:00
When GPUs are allocated, the job sees them through several variables, each set to the allocated device ordinals so GPU runtimes find the right devices:
Variable |
Meaning |
|---|---|
|
Allocated GPU ordinals. |
|
Allocated GPU ordinals. |
|
Allocated GPU ordinals. |
|
Allocated GPU ordinals (Spur-native). |
Under srun, --gpu-bind adjusts the visible set per task —
--gpu-bind=closest binds each task to its nearest GPU, and
map_gpu:.../mask_gpu:... set an explicit per-task mapping or mask.