Download

Get Expanse 1.1.9.

One hybrid ISO installs a node on x86_64 hardware. Build it from the flake, or write a release image to a USB stick, then follow the quickstart.

Latest release

Expanse 1.1.9

Released 2026-09-30 · leader elections spare healthy workloads, and VMs spread across nodes
expanse-1.1.9-x86_64-linux.iso
Hybrid UEFI + BIOS installer image · about 1.7 GB · needs a 4 GB USB stick
SHA-256 1411d36c045d04a7f9e6692ae655595018b17d91e65c4111e541dcc3cfcb93ae

Release images are attached to the v1.1.9 release on GitHub. Check the download against the release's SHA256SUMS with sha256sum -c SHA256SUMS in the same folder. Building from source gives you the same release: the flake asserts the ISO is named expanse-1.1.9-* and that the installer reports 1.1.9.

Requirements

The design target is a decade-old laptop, not a rack.

2
CPU cores
2 GB
RAM
20 GB
disk (2+ for a mirror)
4 GB
USB stick for the installer
  • Architectures: x86_64, booting via UEFI or legacy BIOS. An aarch64 image (UEFI only) builds from the flake but has not been tested.
  • Cluster sizes: one node works; three voters (or two plus a witness) survive losing one.
  • Licence: Expanse is free software under the Apache License 2.0; the source is on GitHub.
  • Hardware status: automated coverage is on QEMU/KVM. Laptop, mini-PC, server and ARM SBC targets are listed in the hardware matrix as untested.
Build from source

Reproducible by design.

Expanse is a Nix flake. Building the ISO needs Nix with flakes enabled; the dev shell brings Go, golangci-lint, protoc and QEMU for everything else.

  • nix build .#iso produces the installer; nix build .#expanse just the CLI and agent.
  • nix flake check runs lint, unit tests and the NixOS VM tests (KVM required).
  • The release version lives in one place, nix/version.nix; the ISO name and expanse version both read it.
# the 1.1.9 installer ISO, straight from the release tag
$ nix build github:team-expanse/expanse/v1.1.9#iso
$ ls result/iso/
expanse-1.1.9-x86_64-linux.iso

# or from a checkout
$ nix develop          # go, golangci-lint, protoc, qemu, …
$ make build && make test && make lint
$ nix build .#iso -L
$ sudo dd if=result/iso/*.iso of=/dev/sdX bs=4M \
    status=progress oflag=direct
Changelog

Release notes

Versions follow Semantic Versioning. History before 1.0.0 is not reconstructed.

1.1.9Leader elections spare healthy workloads, and VMs spread across nodes2026-09-30

Fixed

  • Killing the raft leader could restart VM and other workloads on healthy nodes. While the cluster elected a new leader, a node's read of its desired state failed as unavailable, and the reconciler took the failed read for an empty desired state and stopped everything it ran; a VM workload then cold booted again once the new leader was up. A failed read now skips the round, and a desired-state entry that fails to decode keeps its workload running.
  • A leader election could still restart a VM on a healthy node, through its disk. Reading a volume's record during the election failed, the failure was reported as "not found", and the component that writes each node's desired state dropped the VM's disk mount and changed its spec. The storage controller had the same blind spot and could request a second volume under an existing name. Read failures now keep their real kind, and both components skip the round instead of acting on a partial view.
  • Workloads and their disks piled onto the same few nodes. Every replicated disk went to the same nodes (the lowest node IDs), a VM has to run where its disk is, and the least-loaded score saw no node capacity so it never counted. Disks now go to the nodes holding the fewest replicas, the scheduler knows each node's capacity, and workloads placed together see each other's load, so VMs spread evenly across the cluster.

Added

  • A cluster test with VM workloads that serve HTTP from replicated disks, watched from outside the cluster, measuring forming, node-loss failover, rejoin and leader-loss failover; it runs at 6 nodes and at 12, each VM pinned to its own physical cores. At 6 nodes a node loss costs a workload about 105 seconds (node declared lost, then a nested-VM cold boot) with no acknowledged write lost, and a leader loss costs nothing.
1.1.8Installs about 2.5 times faster, copied off the ISO2026-09-29

Changed

  • Installs are about 2.5 times faster: the install step takes about two minutes in QEMU, down from five, and downloads 18 MiB instead of 192 MiB. The installed node now builds the installer's own expanse, so it is copied off the ISO rather than compiled; the ISO carries a reference node for each disk layout and the tools that disko and the node-specific parts build with; and nixos-install copies NixOS's small local-only derivations off the ISO instead of rebuilding them. The ISO grows from 1.47 GB to about 1.65 GB.
  • The tty1 host console drops its Node ID row when the ID is only the hostname again (a clustered node) or a bare UUID.

Fixed

  • Installed nodes reported their commit as installer; they now report the commit the ISO was built from.
1.1.7Apache-2.0 licence, bundled third-party notices, quiet clock checks at boot2026-09-28

Added

  • Expanse is licensed under the Apache License, Version 2.0: LICENSE and NOTICE at the top of the source tree, and meta.license on the Nix package.
  • The expanse package ships LICENSE, NOTICE, a THIRD-PARTY.md listing every vendored Go module and its version, and each module's own licence, under share/licenses/expanse. Installed nodes and the installer link them at /run/current-system/sw/share/licenses/expanse. Earlier ISOs shipped the binaries without the notices their MIT and BSD modules require.
  • scripts/release.sh publishes a release: it builds and checks the ISO, stages SHA256SUMS, the changelog notes and a release.json manifest, pushes the tag, creates the GitHub release with the ISO attached and runs the website's tools/sync-release hook. See docs/RELEASING.md.

Fixed

  • Nodes read DEGRADED for about two minutes after each boot, long enough to fire ExpanseNodeDegraded: the clock-sync check judged chrony's last correction, which is the clock step chrony makes at boot, instead of how far the clock is off now (System time in chronyc tracking).
1.1.6Installer facelift, tty1 host console, quiet boots2026-09-28

Changed

  • Installer TUI has a proper console layout: a framed, centred page that fits the terminal (80×25 upwards, redrawn on resize), a "Step n of 6" indicator, a disk table with model, size, type and current contents, a reverse-video selection cursor, a red destructive list on review, a real progress view (stage, progress bar, elapsed time and a scrolling log tail) and a done screen with the node's addresses and next steps. Colours are the VT's 16; box drawing falls back to ASCII with EXPANSE_TUI_ASCII=1. Keys and flow are unchanged.

Added

  • expanse console: a read-only host information screen, like ESXi's DCUI, shown on tty1 in place of a login (expanse-console.service; tty2+ and the serial console keep their logins). Version, hostname and node ID, addresses and web UI URL, cluster name, role and quorum, health, CPU, memory, disks, md mirror state (a degraded array in red) and uptime, refreshing every 3 s. Health comes from the agent's heartbeat, so the console never re-runs checks. --once prints one screen to stdout.

Fixed

  • Every node reported UNHEALTHY (and fired ExpanseNodeUnhealthy) for a minute after each boot while chrony synchronised. For the first 10 minutes after boot an unsynchronised clock now reads "unknown".
  • The clock-sync check's large-offset warning (> 100 ms) could never fire.
1.1.5A polished web UI2026-09-27

Changed

  • The web UI has a proper app shell and design system: sidebar, header with cluster name and quorum/health, user menu, light/dark toggle, breadcrumbs and a footer naming the serving node and version; status pills, styled tables, cards, forms with inline validation, toasts and <dialog> confirmations. Responsive to phone width, keyboard accessible, hand-written CSS, no build step, no CDN, self-only Content-Security-Policy.
  • / is a live dashboard: nodes, quorum, blocks by phase, volumes by state, firing alerts, the current generation, a "needs attention" list and recent events.

Added

  • Pages for Nodes, Generations (history, detail, diff of any two, rollback behind a confirm), Events (a live, filtered log of store writes) and Settings (admin password, OIDC status, UI CA download).

Fixed

  • Plain form posts from a browser were rejected with "CSRF token mismatch".
  • A node's reported health read "unknown" permanently: the agent's PATH lacked chronyc.
  • The web UI footer read "Expanse dev" on installed nodes.
1.1.4Mirror layout boots from either disk2026-09-27

Changed

  • The ESP and the system partition are now md RAID1 arrays across the first two disks (btrfs on the md array). Verified in QEMU: install, boot with each disk removed, and a replacement disk rebuilt and booted alone. Existing mirror installs keep the old layout until reinstalled.
  • On an md ESP, systemd-boot is installed with relaxed ESP checks and without an NVRAM entry; each disk boots through \EFI\BOOT\BOOTX64.EFI.
  • Mirror nodes ship sgdisk and log md events to the journal; the storage docs carry the disk-replacement runbook.
  • expanse doctor storage reads the md array under the system btrfs: PASS for a healthy mirror, WARN when degraded.
1.1.3Install from the ISO works end to end2026-09-27

Fixed

  • Verified in QEMU through the interactive installer, nixos-install, first boot, a one-node cluster running a block, and a reboot that wipes root and keeps the cluster.
  • Installer TUI key handling, dropped SSH keys, missing tools; the ISO's <nixpkgs> needed flakes; installed nodes had no hardware configuration and did not run the agent; the impermanence rollback raced the system disk; cluster init now refuses while expansed runs.
  • The web UI could not be opened in any browser (SSL_ERROR_NO_CYPHER_OVERLAP): it now has its own ECDSA P-256 CA; import ui-ca.pem.
  • An installed node's firewall blocked the web UI (8443) and metrics (7447) from other machines.
  • The ISO and installed nodes log to the serial console too, with a login there.

Known issues

  • Binary-backed block types (nginx, redis, …) need their program on the node's PATH, which an installed node does not ship; util/echo runs.
1.1.2The installer reports its real version2026-09-26

Fixed

  • expanse version on the installer and on installed nodes reported 0.0.1 (installer). The release version lives in nix/version.nix, and the ISO is named expanse-<version>-<system>.iso.
1.1.1Faster replica rebuilds under load2026-09-26

Changed

  • DRBD resources set c-min-rate 4M: a new or rebuilt replica on a busy volume syncs at 4 MiB/s or more instead of being throttled toward 250 KiB/s. Existing volumes pick it up on upgrade.
1.1.0Single-node clusters that grow2026-09-26

Added

  • cluster init --expect 1 forms a working cluster on one machine; default volumes, blocks with storage and VIPs run there.
  • Volumes grow to their replication target automatically as nodes join, one fully synced replica at a time; DRBD quorum turns on live at three replicas.
  • A new UnderReplicated volume state, shown as "1 of 3 (no redundancy)"; volume list shows volumes waiting for placement with the reason.

Changed

  • A volume or block storage entry without an explicit replication takes its storage class's target and is placed on as many eligible nodes as exist. An explicit --replication N still waits for N nodes.
1.0.0First release2026-09-26

Added

  • Installer ISO with TUI and unattended modes, declarative btrfs+LVM layout, impermanent root.
  • Clustering: mDNS discovery, HMAC join tokens, a Raft control plane over mTLS, generations with rollback, cordon/drain/remove, witness nodes.
  • DRBD-replicated volumes on LVM thin pools: create, resize, snapshot/restore, online repair, split-brain recovery.
  • Blocks: a scheduler for systemd-supervised workloads with rolling updates, anti-affinity, singleton and daemonset placement.
  • Postgres, iSCSI and QEMU/KVM blocks with failover.
  • WireGuard mesh, per-node nftables, L4/L7 load balancing with VIP failover.
  • Web UI with OIDC login. restic backup and restore of cluster config and volume data. Prometheus metrics and alert rules. Cluster CA with zero-downtime rotation; age-encrypted secrets. Rolling upgrades, one node at a time, with zero acked-write loss.
  • Every feature has NixOS VM-test coverage of real node kills, partitions and restarts.
Known issues

Read before you install.

Carried in the changelog as of 1.1.9. Everything under 1.0.0 still applies unless a later release says otherwise.

Open in 1.1.9

  • Binary-backed block types on installed nodes. nginx, redis and similar types need their program on the node's PATH, which an installed node does not yet ship. util/echo runs.
  • Web UI VIP. The UI takes one address from the external VIP pool; size the pool one larger than the block VIPs you need.
  • Idle control-plane CPU with a volume attached has not been measured on real hardware; without volumes, bare metal measured 2.5% of one core on the Raft leader. With a volume, the leader is projected at about 4%.
  • SMB/NFS file shares are paused: Samba deploys, but failover has an open bug.
  • The on-prem LLM subsystem is paused (no accelerator hardware to validate on).
  • TPM sealing of secrets is deferred (no TPM hardware to validate on).
  • Mirror installs made before 1.1.4 keep the old layout (ESP on the first disk only) until reinstalled.

Ready to boot it?

The quickstart covers the USB stick, the installer, the first boot, forming the cluster and opening the web UI.