Upgrading Spur#
This page covers upgrading Spur binaries, whether on a single host or across a whole
cluster. It describes the spur self-update self-updater, re-running install.sh,
and the Ansible playbooks that upgrade a running cluster with or without an outage.
Note
There are two upgrade scopes: upgrading the binaries on a single host and
upgrading a whole cluster. There is no in-process hot-swap that preserves
running jobs inside a single daemon process. spur self-update only swaps the
binaries on disk — it does not restart the daemons. The drain-aware, jobs-preserving
cluster path is the Ansible rolling_upgrade.yml playbook (see below).
Single-Host: spur self-update#
spur self-update downloads the latest release and replaces the spur, spurctld,
and spurd binaries on the current host. It is a single-host convenience: it does not
restart daemons and gives no drain or quorum protection, so it is not a substitute for the
cluster playbooks below.
Check whether an update is available:
spur version --check
update available: 0.3.0 → v0.3.1
Run `spur self-update` to install.
Install the update:
spur self-update
spur update is an alias for spur self-update. Add --nightly to pull from the
nightly channel instead of the latest stable release:
spur self-update --nightly
The updater downloads the release tarball, verifies it against its published SHA256
checksum, then replaces each binary atomically: the current binary is renamed to
<name>.spur-old as a backup, the new binary is copied into place, and the backups are
deleted on success. If any copy fails, the .spur-old backup is restored. The install
directory is auto-detected as wherever the running spur binary already lives.
After a successful update the CLI prints:
Updated spur to v0.3.1
Note: Restart running daemons (spurctld, spurd) to use the new version.
Warning
spur self-update never restarts daemons. Running spurctld and spurd
processes keep executing the old binary until you restart them yourself:
sudo systemctl restart spurctld spurd
The [update] config block#
The optional [update] block in spur.conf controls the daemon startup update check.
Its fields are:
Field |
Default |
Effect |
|---|---|---|
|
|
Check the GitHub releases API when the daemon starts and log if an update exists. |
|
|
Download and install an available update automatically. Even when |
|
|
Release channel to check: |
|
|
Directory for the update-check cache (1-hour TTL). |
Note
Even with auto_update = true, a new binary on disk does not take effect until the
daemon is restarted. The config block applies to spurctld; spurd never
auto-installs updates.
See Configuration Reference (spur.conf) for the full [update] field reference.
Single-Host: Re-running install.sh#
Re-running the one-line installer upgrades an existing single-host install in place. It
downloads the requested release, verifies its checksum, and copies the binaries over the
install directory (default ~/.local/bin):
curl -fsSL https://raw.githubusercontent.com/ROCm/spur/main/install.sh | bash
Pass nightly or a pinned vX.Y.Z to select a specific release, or set
INSTALL_DIR to install elsewhere:
curl -fsSL https://raw.githubusercontent.com/ROCm/spur/main/install.sh | bash -s -- nightly
curl -fsSL https://raw.githubusercontent.com/ROCm/spur/main/install.sh | bash -s -- v0.3.1
Like spur self-update, install.sh does not manage systemd units or restart
daemons. Restart spurctld and spurd yourself after re-installing.
Cluster Upgrades with Ansible#
For a multi-node cluster, the Ansible toolkit is the recommended upgrade path. Two
playbooks are supported; both reuse the same install, config, and health-check roles as
deploy.yml, so their behavior stays consistent.
Rebuild all three binaries from the same source tree together — they share a Raft write-ahead-log schema, and mixing binaries from different builds can leave a controller unable to parse a log written by a differently-versioned peer:
cargo build --release -p spur-cli -p spurctld -p spurd
Binaries roll out by content, not version string: Ansible compares checksums, so an unchanged re-run is a near no-op.
Full convergence (deploy.yml)#
deploy.yml is the simplest upgrade: it re-installs binaries and restarts every daemon
on every host in one play. This causes a brief cluster-wide outage — in-flight jobs are
disrupted, and in an HA setup all controllers bounce together, briefly losing the Raft
leader. It is non-destructive by default (state is preserved). Use it for topology changes
or when a short blip is acceptable:
ansible-playbook playbooks/deploy.yml -i inventory/hosts.ini -e spur_binary_src=/path/to/target/release
Note
Spur 0.3.0 has no online Raft membership change. Adding, removing, or reordering a
controller fails early unless you also pass -e spur_wipe_state=true (a Raft
reinit that wipes state). Compute agents are not Raft members, so they can be added or
removed freely without a wipe.
Rolling upgrade (rolling_upgrade.yml)#
rolling_upgrade.yml is the seamless, no-full-outage path: it upgrades one host at a
time so the cluster keeps scheduling and running jobs throughout.
ansible-playbook playbooks/rolling_upgrade.yml -i inventory/hosts.ini -e spur_binary_src=/path/to/target/release
The playbook proceeds in order:
Guard rails. Abort if
spur_wipe_state=true(never wipe Raft mid-upgrade), ifspur_transport=wireguard(not yet supported by this playbook), or if the cluster is not already healthy (spur nodesmust return success).Upgrade controllers one at a time (
serial: 1, no failures tolerated), preserving Raft quorum. Each controller’s binaries are force-reinstalled and the daemon restarted; the “wait for Raft leader” step is the health gate before moving to the next controller. The existingspur.confis preserved unless you pass-e spur_overwrite_conf=true.Drain and upgrade agents in batches. For each agent: drain it from the controller (
spur node drain <node> --reason "ansible rolling upgrade"), wait until its state isDRAINEDorDOWN(running jobs finish first — drain never force-kills), swap the binary, restartspurd, wait for the node to re-register, then resume it (scontrol update NodeName=<node> State=RESUME).Verify. Submit a real test job to confirm the upgraded cluster schedules work.
The rolling upgrade is controlled with these -e flags:
Flag |
Default |
Effect |
|---|---|---|
|
(unset) |
Directory of pre-built |
|
|
Agents upgraded per batch. Controllers are always upgraded one at a time regardless. |
|
|
Skip agents unreachable over SSH instead of aborting. |
|
|
Leave a still-busy node on its current binary and continue, rather than aborting. |
|
|
Kill running jobs and containers on a busy node and upgrade it anyway. Affected jobs
are marked |
|
|
Re-render |
Note
A larger spur_rolling_batch_size upgrades faster but drains more capacity at once. A
single-controller cluster still has a short outage while its own controller restarts —
true zero-downtime requires an HA quorum of 3 or more controllers.
Safe Upgrade Order#
Follow this order for any cluster upgrade:
Rebuild all three binaries together from the same source tree — they share a Raft WAL schema and must stay version-matched.
Upgrade controllers before agents. Both playbooks do this automatically, one controller at a time to preserve quorum.
Drain agents before swapping binaries. The rolling playbook drains automatically; a running job blocks the swap unless you force it.
Never wipe state during an upgrade. Keep the default
spur_wipe_state=false. Wiping resets the Raft job-id counter and destroys accounting history.HA membership is fixed at init in Spur 0.3.0. You can roll new binaries onto the existing controller set freely, but changing which hosts are controllers requires
deploy.yml -e spur_wipe_state=true.