Kubernetes Deployment#
Deploy Spur on an existing Kubernetes cluster. The controller runs as a StatefulSet with Raft consensus, and compute nodes are managed by the spur-k8s-operator.
Prerequisites#
Kubernetes cluster with
kubectlconfigured
Build and load container images:
# Build
docker build --target runtime -t spur:<tag> .
# Load onto each node (if not using a registry)
docker save spur:<tag> -o spur.tar
# SCP to each node, then:
sudo ctr -n k8s.io images import spur.tar
Components#
spurctld — Controller. Runs as a StatefulSet with Raft consensus for high availability. Handles accounting (backed by PostgreSQL via
accounting.database_url) and serves the Slurm-compatible REST API on port 6820.spurd — Node agent. Runs on each compute node (DaemonSet or Deployment).
spur-k8s-operator — Watches
SpurJobcustom resources and submits them to the controller.
Example manifests for production-style deployment live in examples/k8s/.
Deploy#
Note
Before applying, review the manifests and update namespaces, image names/tags, resource limits, and storage classes to match your environment. Ensure the --controller argument in spurd.yaml includes the http:// scheme (e.g. http://spurctld.spur.svc.cluster.local:6817).
Apply manifests in order:
kubectl apply -f examples/k8s/namespace.yaml
kubectl apply -f examples/k8s/rbac.yaml
kubectl apply -f examples/k8s/spurjob-crd.yaml
spurctld reads spur.conf from a Secret, because the file carries the
accounting database password. Fill in database_url in
examples/k8s/spur.conf and build the Secret from it before you apply the
controller:
kubectl create secret generic spur-config \
--from-file=spur.conf=examples/k8s/spur.conf -n spur
kubectl apply -f examples/k8s/spurctld.yaml
kubectl apply -f examples/k8s/spurd.yaml
kubectl apply -f examples/k8s/operator.yaml
kubectl apply -f examples/k8s/pdb.yaml
Configuration#
examples/k8s/spur.conf is the controller configuration. The controllers
run with the Secret spur-config that you build from it. The whole file goes
into the Secret, not only the password: the configuration loader reads one TOML
file and has no separate source for accounting.database_url or
auth.jwt_key. A cluster with no accounting and no auth.jwt_key can hold
the same file in a ConfigMap instead, built with
kubectl create configmap --from-file. The file sets:
cluster_name = "spur-k8s"
[controller]
peers = [
"spurctld-0.spurctld.spur.svc.cluster.local:6821",
"spurctld-1.spurctld.spur.svc.cluster.local:6821",
"spurctld-2.spurctld.spur.svc.cluster.local:6821",
]
[scheduler]
interval_secs = 2
plugin = "backfill"
[[partitions]]
name = "default"
state = "UP"
default = true
Raft peers use StatefulSet DNS names. The node ID is auto-derived from each pod’s
position in peers by matching the pod hostname (e.g. spurctld-0) against
each entry’s host part, so controller.node_id never needs to be set. Each
pod’s hostname must correspond to its own peers entry.
Resolution precedence is: explicit controller.node_id -> position in
peers -> hostname ordinal. The resolved id must fall within
1..=len(peers); if a pod’s hostname matches no entry (or matches more than
one), the controller fails fast at startup rather than joining with a wrong ID.
Adjust partition definitions to match your cluster hardware. Once the controller
is running, scontrol reconfigure applies many sections live, while others
need a controller or agent restart — see
the configuration reference for the per-field breakdown.
reconfigure runs on the Raft leader only — followers keep their startup
config until restarted, at which point they re-read this same Secret and
converge. To roll all controllers onto an updated Secret, restart the
StatefulSet pods.
Submitting Jobs#
Jobs are submitted as SpurJob custom resources:
apiVersion: spur.amd.com/v1alpha1
kind: SpurJob
metadata:
name: training-run
spec:
script: |
#!/bin/bash
#SBATCH --job-name=train
#SBATCH -N 2
#SBATCH --gres=gpu:8
torchrun --nnodes=2 train.py
Apply with kubectl:
kubectl apply -f job.yaml
The operator watches SpurJob resources, submits them to the controller, and updates status fields as the job progresses.
Operator connection to the controller#
The operator opens each channel to spurctld with these bounds:
A connect timeout of 10 s. A Service with no ready endpoint drops the connection attempt, and the kernel retries it for more than two minutes. The operator stops earlier and connects again when the controller is ready.
HTTP/2 keepalive pings every 10 s. A ping with no answer in 5 s closes the channel.
A bound of 30 s on each request.
The job controller and the node watcher each keep one channel open. After a transport error they
drop the channel and open a new one on the next call. A refusal from the controller, for example
NOT_FOUND, is an answer and keeps the channel. A request that reaches the 30 s bound is a
transport error and is retried.
The node watcher registers each Kubernetes node with the controller. A transport error during a registration restarts the node watcher, which lists every node again. A refusal that is not a transport error, for example a missing admission token, is written to the log. The node watcher tries the registration again at the next event of that node.
Authenticating the operator agent surface#
The operator serves a virtual-agent gRPC surface on --listen (port 6818) that carries a
cluster-wide pod-create privilege, so reaching it must not be enough to ask the operator to run
work. Authentication mirrors the cluster [auth] mode:
--auth-mode disableddoes not authenticate callers at all; treat the port as an administrative boundary if you use this.--auth-mode permissive(default) verifies a credential when one is presented and otherwise logs and allows — the migration default.--auth-mode requiredrejects every call that carries no credential. It refuses to start without a key.--jwt-key/SPUR_JWT_KEYis the cluster’s signing key —[auth] jwt_key, or the contents of the file[auth] jwt_key_filepoints at — that the operator verifies credentials against; source it from a Secret. Inpermissivemode with no key, a controller that does present a credential is rejected (there is no key to verify it), so set the key before controllers start sending one.
See the commented --auth-mode / SPUR_JWT_KEY lines in examples/k8s/operator.yaml.
Enforcing account quotas#
--enable-quota turns on a reconciler that projects each SPUR account’s GrpTRES allocation
into a per-account ResourceQuota/LimitRange. An account with no GrpTRES set gets a
closed quota (pods: 0) rather than an uncapped one, so pods submitted under it are rejected
until an allocation is granted:
sacctmgr modify account name=myaccount set grptres=cpu=16,mem=32768,gres/gpu=8
Verify#
# All pods running
kubectl get pods -n spur
# Controller logs (check Raft leader election)
kubectl logs statefulset/spurctld -n spur
# Node registration
kubectl exec -n spur spurctld-0 -- spur nodes