1.1.9 Leader elections spare healthy workloads, and VMs spread across nodes

The cluster OS for the machines you already have.

Expanse turns ordinary hardware — the design target is a decade-old laptop — into a node of a self-healing cluster: replicated storage, scheduled services, VIP failover and a live web console. Built on NixOS from proven parts, not hand-rolled distributed systems.

2 cores · 2 GB RAM · 20 GB disk Guided or unattended install Starts on one node, grows as nodes join
Left: the live web console every node serves on :8443. Right: the read-only host console every installed node shows on tty1. Illustrative mock-ups rendered in HTML, not screenshots.
Adopted, not reinvented
NixOS Raft DRBD 9 LVM thin btrfs md RAID1 WireGuard nftables systemd restic Prometheus QEMU/KVM
What you get

Everything a small fleet needs, in one image.

One bootable ISO installs a node. Nodes find each other, join with single-use tokens, and share one strongly consistent control plane. Everything below ships in 1.1.9 and is exercised by NixOS VM tests that kill nodes, partition networks and pull disks.

Install from a USB stick

A hybrid UEFI/BIOS ISO with a full-screen installer TUI or a one-file unattended mode. Declarative btrfs + LVM layout, an impermanent root that is wiped on every boot, and a two-disk mirror layout that boots from either disk alone.

Installer details

A control plane that keeps quorum

Raft-replicated state over mTLS, linearizable reads by default, single-writer leases as the anti-split-brain primitive, witness nodes for 2+1 topologies, and a read-only degraded mode when a node cannot reach a leader.

Clustering

Replicated volumes

DRBD 9 on LVM thin pools: create, resize, snapshot and restore, online repair and verified split-brain recovery. A volume starts on one node and grows to its replication target as nodes join, one fully synced replica at a time.

Storage

Blocks: services as YAML

Declare a service, its replicas, resources and placement; the scheduler runs it as hardened systemd units. Rolling updates, anti-affinity, singleton and daemonset placement, cross-node logs, and explain when something will not place.

Blocks

Mesh, VIPs and load balancing

A WireGuard overlay between every node, lease-fenced VIPs that move with their service, L4 and L7 balancing, cluster DNS for <block>.<ns>.expanse.local, and a default-deny nftables firewall per node.

Networking

Generations and rolling upgrades

Every fleet-wide configuration change is a generation: diff any two, roll back behind a confirm. Upgrade one node at a time with no quorum loss; a switch watchdog rolls a node back automatically if a switch goes wrong.

Upgrades

A live web console

Every node serves the UI on :8443: dashboard, cluster, nodes, health, blocks, volumes, generations and events, updating over Server-Sent Events. Password login plus optional OIDC single sign-on with a fail-closed allow-list.

Web UI and host console

Health, metrics and alerts

A Prometheus /metrics endpoint on every node, a shipped alert-rule file and Grafana dashboard, and an in-UI health page that derives the same critical alerts with no external tooling at all.

Observability

Secure by construction

A per-cluster CA with zero-downtime rotation, HMAC single-use join tokens, mTLS on every internal port, age-encrypted secrets, and a browser-facing UI CA that is separate from the cluster's own.

Security
How it works

Install. Form a cluster. Manage.

One installed node is already a usable cluster. Each step below is the real command, not a sketch.

Install a node

Write the ISO to a USB stick and boot the target. The installer picks disks, network and an SSH key, shows what it will destroy, then installs. Headless? SSH in and run it there.

$ nix build github:team-expanse/expanse/v1.1.9#iso
$ sudo dd if=result/iso/*.iso of=/dev/sdX bs=4M
# boot the target from the stick, then:
$ expanse install --tui

Form a cluster

Bootstrap the first node as the sole voter. It prints a single-use join command; run it on each further node. Volumes and blocks gain replicas automatically as nodes join.

$ systemctl stop expansed
$ expanse cluster init --expect 1
cluster initialized: 5c0b… (name "expanse")
Join command (single-use, 15m):
  expanse cluster join --address 10.0.0.11:7446 \
    --token expanse-join-…
$ systemctl start expansed

Manage it

Open the web console on any node, or use the CLI: apply a block, create a volume, watch health. Everything the UI does, the CLI does too — it is one control plane with two ways in.

$ expanse ctl block apply -f web.yaml
created default/web
$ expanse ctl volume create pg --size 10Gi
$ expanse cluster status
quorum:    3/2   # 3 voters, 2 needed
One CLI, one control plane

Operate the whole cluster from any node.

Writes are forwarded to the Raft leader, reads are linearizable by default, and every action the web console offers is a CLI verb underneath. When a block will not place, explain tells you why each node was rejected.

  • Declarative and idempotent. block apply is a no-op when nothing changed and a rolling update when something did.
  • Cross-node logs. block logs default/web --replica 1 streams from whichever node runs it.
  • Doctors for storage and network. PASS/WARN/FAIL checks with a remediation hint on every row.
$ expanse cluster status
cluster:   expanse (5c0b3e2a-…)
version:   1.1.9
generation: 12
leader:    n1
quorum:    3/2
nodes:     3

  ID   ROLE    STATE   RAFT             API
  n1   voter   leader  10.0.0.11:7444   10.0.0.11:7443
  n2   voter   voter   10.0.0.12:7444   10.0.0.12:7443
  n3   voter   voter   10.0.0.13:7444   10.0.0.13:7443

$ expanse ctl block apply -f web.yaml --dry-run
# validated, nothing written
$ expanse ctl block apply -f web.yaml
created default/web
$ expanse ctl block explain web
web: 3/3 placed (phase RUNNING)

$ expanse ctl volume list
ID      NAME    SIZE   STATE     REPLICAS  NODES (PRIMARY)
vol-…   pgdata  10Gi   healthy   3/3       n1,n2,n3 (n1)

$ expanse doctor storage --vg expanse
CHECK           STAT  DETAIL
drbd-module     PASS  loaded, version 9.2.14
volume-group    PASS  expanse: 41.3% free of 943718400000 bytes
thin-pools      PASS  1 pool(s) healthy: pool
system-mirror   PASS  md md127 RAID1 across 2 devices
drbd-resources  PASS  1 resource(s) healthy
110+
NixOS VM tests that kill nodes, partition networks and pull disks — not just unit tests.
0
Acknowledged writes lost across leader failover, partitions and rolling upgrades in those tests.
≤15s
VIP failover budget after a hard power-off of the holder, measured in VM tests from an external client (typically ~10 s).
~2.5%
Of one core for the idle Raft leader: three containers on one bare-metal host, no volumes; followers 1.1–1.2%.
Why Expanse

Built the boring way, on purpose.

The design rule is simple: adopt a proven component for the data plane, build only the control plane that ties them together. That keeps the whole system small enough to reason about — and to test with real failures.

Expanse is

  • A host OS and cluster manager in one image. There is nothing to install on top; the node boots into the agent, the web console and the tty1 host console.
  • Sized for the machines you have. 2 cores, 2 GB RAM and a 20 GB disk is the design target, and a single node is a valid cluster from day one.
  • Strongly consistent. One Raft log holds cluster state, leases, generations and even join-token consumption, so two nodes can never both believe they own something.
  • Reproducible. Every node is a NixOS configuration; an upgrade is a new generation with a boot entry to fall back to.
  • Tested with real failures. Every feature has VM-test coverage of node kills, partitions and restarts, plus a chaos suite with a split-brain invariant checker.

Expanse is not

  • Kubernetes. Blocks are systemd-supervised units described in YAML, not pods. There is no Kubernetes API, no container runtime and no operator ecosystem.
  • A hypervisor first. VMs are one block type (QEMU/KVM on a replicated disk). Failover is a cold restart elsewhere; there is no live migration.
  • A file server yet. SMB/NFS shares are paused: the Samba block deploys but its failover has an open bug. Use volumes, Postgres, iSCSI or VMs today.
  • TPM-sealed. Secrets are age-encrypted, keyed from the cluster secret. TPM sealing is a named, deferred feature, not a shipped one.
  • Validated on every box. Automated coverage is on QEMU/KVM; the laptop, mini-PC, server and ARM targets in the hardware matrix are still marked untested.
ConcernAdopted componentWhat Expanse adds
Host OS and upgradesNixOS, systemd-bootGenerations, rollback, a switch watchdog, an impermanent root
Cluster statehashicorp/raft, memberlistLeases, generations, join protocol, mTLS CA with rotation
VolumesDRBD 9, LVM thin, btrfs, md RAID1Placement, repair, split-brain resolution, grow-as-you-join
Servicessystemd units with sandboxingA scheduler with rolling updates, anti-affinity, singleton leases
NetworkWireGuard, nftables, miekg/dnsMesh reconciliation, lease-fenced VIPs, L4/L7 balancing, cluster DNS
Backup and metricsrestic, Prometheus, GrafanaConsistent snapshot paths, one-command restore, shipped rules and dashboard
How Expanse is made

A joint human–AI project.

Expanse's code, tests and documentation, and this website, were built by a human maintainer working with AI coding assistants. The AI wrote much of the code and prose. The maintainer directs the work and decides what ships.

  • It is in the history. Most commits in both repositories credit an AI co-author in a Co-Authored-By trailer. Read the commit log.
  • Claims come with evidence. What this site says is drawn from the project's tests, docs and changelog, and what has not been verified is marked as such. Real hardware, for example, is still untested.
  • Judge it on its record. The VM tests exercise failures, not just the happy path, but automated tests are no substitute for your own evaluation. Read the known issues before trusting Expanse with data you care about.

Boot a node tonight. Add the second one whenever.

The quickstart takes a blank machine to a one-node cluster with a web console, and shows how each further node joins with a single command.