Every capability below is in the release, documented, and covered by NixOS VM tests that exercise real node kills, partitions and restarts. The last section lists what is deliberately not built yet, so you are never guessing.
From a blank machine to a node with one USB stick.
The x86_64 ISO is a hybrid image that boots under UEFI and legacy BIOS. A full-screen installer starts automatically; headless machines can be installed over SSH with the same TUI or with a one-file unattended config.
Six steps. Welcome (hardware summary), Disks, Network, SSH key (paste one or gh:<username>), Review with a red list of what will be destroyed, Progress with a stage, bar and log tail.
Refuses to surprise you. Disks holding data are marked CONTAINS DATA; unattended installs will not touch them without --force, and --dry-run prints every command without running one.
A layout built for recovery. btrfs subvolumes @root @nix @persist @log; the root is wiped on every reboot and only /persist survives. The rest of every disk becomes the LVM volume group expanse.
Mirror layout on two or more disks. The ESP and the system partition are md RAID1 across the first two disks, so the node boots from either disk alone; a replacement disk rebuilds without a reinstall.
Idempotent. Boot the stick again and re-run: same config, same system. Node identity lives in /persist and survives.
expanse-install.yaml · unattended install
version: 1
disks:
layout: auto # auto | single | mirrordevices: [/dev/sda] # empty = use every disknetwork:
mode: dhcp
ssh:
authorized_keys:
- ssh-ed25519 AAAA...timezone: UTC
One Raft log, mutually authenticated, from the first node.
Cluster state, leases and generations live in one Raft-replicated store over mTLS. cluster init bootstraps a single voter; every further node joins with a single-use token and grows the same cluster.
Join tokens that cannot be replayed. HMAC-SHA256, TTL-bounded, consumed inside the Raft log — of two simultaneous joins exactly one wins. The joiner proves possession of its own Ed25519 key before a certificate is issued.
Linearizable by default. Reads take one quorum round trip; stale reads are opt-in. Writes from any node are forwarded to the leader and fail fast (about 2 s) against a deposed one.
Leases with a guard band. 15 s TTL, renewed every 5 s; a takeover cannot happen until at least 8 s after a partitioned holder has stopped. That holds on clock rate, never on clock offset.
Degrades honestly. A node that cannot reach a leader goes read-only and freezes reconciliation instead of guessing.
Witness nodes and node lifecycle. A witness is a full voter that never runs workloads (2 + 1 topologies). Cordon, drain and remove are interlocked; removal transfers leadership first.
mDNS discovery.expanse cluster discover lists clusters and unjoined nodes on the L2 network; cluster join --discover finds the join endpoint itself.
leader n1, then a new node n2
# on the leader, with its agent stopped:$ expanse cluster token create --ttl 15m --uses 1expanse-join-…# on the new node, before its agent starts:$ expanse cluster join --address 10.0.0.11:7446 \
--token expanse-join-…joined cluster 5c0b3e2a-… as n2 (10.0.0.12:7444)$ systemctl start expansed$ expanse cluster statusquorum: 2/2nodes: 2
API / gRPC
7443 (mTLS)
Raft
7444
Gossip
7445
Join
7446
Storage
Replicated volumes that grow with the cluster.
Volumes are DRBD 9 resources (protocol C) on LVM thin volumes, placed by the leader. Replication is a target, not a precondition: a volume lands on one node today and gains replicas as nodes join, one fully synced replica at a time.
Full lifecycle. Create, grow, snapshot and restore, verify and resync a scrubbed replica, retire a dead node's copy, or hand-pick the surviving side of a split-brain with volume diverged --choose.
Honest states.UnderReplicated means every replica is healthy but there are fewer than the target — the CLI and UI show "1 of 3 (no redundancy)" rather than a green tick.
Quorum switches on live. DRBD quorum turns on at three replicas without interrupting I/O, verified under continuous acknowledged writes.
Rebuilds that finish. A resync floor of 4 MiB/s keeps a busy volume from starving its new replica; a 256 MiB volume under a heavy fsync writer gained a replica in 33 s on VMs.
A doctor for it.expanse doctor storage checks the DRBD module, VG headroom, thin-pool data and metadata, the system mirror and every DRBD resource, with a remediation hint on anything short of PASS.
the stack, per node
GPT ─┬─ ESP (FAT32) → /boot
├─ btrfs partition → @root @nix @persist @log
└─ LVM PV → VG "expanse" → thin pool → thin LVs
└─ DRBD 9 (protocol C, quorum at ≥3)
└─ filesystem or raw (block replica)
$ expanse ctl volume create data --size 20Gi$ expanse ctl volume listID NAME SIZE STATE REPLICAS NODES (PRIMARY)vol-… data 20Gi underreplicated 1/3 n1 (n1) no redundancy# … two nodes join …vol-… data 20Gi healthy 3/3 n1,n2,n3 (n1)$ expanse ctl volume resize data --size 40Gi$ expanse ctl volume snapshot data --name before-migration
Blocks
Services as YAML, run as hardened systemd units.
A block declares a catalog type, replicas, resources, placement and networking. The leader's scheduler filters and scores nodes; each node's agent renders the unit and supervises it. Health flows back through the store and the block is promoted to RUNNING when every replica is.
Rolling updates that keep serving.maxUnavailable, maxSurge and minReadySeconds are honoured; the rolling-update VM test asserts zero failed health probes during a roll. A mid-roll crash resumes deterministically.
Placement you can reason about. Node anti-affinity, even spread, required capabilities (for example kvm), daemonsets, and singletons fenced by a Raft lease so exactly one instance ever runs.
Self-healing. A node silent past the 30 s grace has its placements retired and rescheduled under the same rules.
Sandboxed. Units run with DynamicUser, ProtectSystem=strict, PrivateTmp and friends; config is validated against the type's JSON Schema before anything is written.
Eleven shipped types.util/echo, web/nginx, web/static-site, web/whoami, db/redis, db/postgres, monitor/node-exporter, ai/ollama, share/smb, iscsi/target, vm/instance. A new type is four files plus one flake entry.
A mesh, VIPs that move, and a firewall that never reloads.
Every node joins a WireGuard overlay reconciled from the store — a key rotation converges in one pass. Services are reached through VIPs held under a lease, balanced at L4 or L7, and named by cluster DNS.
Lease-fenced VIPs. The holder announces the address with gratuitous ARP; losing the lease removes the address before the listener stops. A replacement takes over within 15 s of a hard power-off, typically about 10 s.
L4 and L7 balancing. TCP splicing round-robin with drain-on-removal, and an HTTP reverse proxy routing by <block>.<ns>.expanse.local and path prefix with idempotent retries.
Cluster DNS. An authoritative server on each node's overlay address: block names to VIPs, replica names to addresses, SRV for declared ports; scale-ups visible within 5 s.
Default-deny nftables. One inet expanse table per node with dynamic sets for peers, VIPs and block ports. After bootstrap it is only ever updated element by element, so conntrack is never dropped.
expanse doctor network. Twelve live checks — interfaces, overlay peers and path MTU, the port matrix, exactly-one VIP holder, ARP consistency, DNS, firewall, conntrack, time sync.
address plan
External clients
│
▼ (VIP 192.168.1.100 — ARP announced by the lease holder)
┌───────────────────────────────────────────────┐
│ Physical LAN 192.168.1.0/24 │
│ n1 .11 n2 .12 n3 .13 │
└───────────────────────────────────────────────┘
│ WireGuard overlay 10.42.0.0/16 (exp0, MTU 1420)
▼ Service VIPs 10.43.0.0/16 (+ a LAN VIP pool you declare)
VIP takeover
≤ 15 s, ~10 s typical
L7 proxy (loopback)
≥ 20k req/s, +p99 ≤ 1 ms
DNS p99 (loopback)
≤ 1 ms
Generations and rolling upgrades
Every change is a generation. Every node can roll back.
Fleet-wide desired state is versioned inside the Raft store: hash-addressed, append-only, kept for the last 50 generations or 30 days. Rollback restores an earlier generation as a new one, so history never loses a step.
Diff anything.generation diff 11 12 in the CLI, or the Generations page in the web UI: added, removed and changed keys, with rollback behind a confirm dialog.
Upgrades are ordinary NixOS switches. The agent builds the flake attribute and runs the same activation script nixos-rebuild switch uses, restarting only what changed and installing a boot entry.
A watchdog for bad switches. If a pending switch is still unconfirmed after 10 minutes, the node rolls back to its previous generation and reboots. A bad remote change cannot brick a node.
One node at a time, no gap. Confirm quorum and every replica UpToDate, drain any volume primary, switch, wait for the agent to rejoin, then the next. A surviving node's reads and writes keep working throughout — proven with a background write loop, not a spot check.
Scope stated plainly. The proof covers builds that share the replicated log's wire version, which has not changed since it was introduced; a future bump needs its own dual-version plan.
upgrade one node
# first: quorum reads 3/2 and every volume is UpToDate$ expanse cluster status$ expanse ctl resource apply - \
--socket /run/expanse/agent.sock <<'EOF'nix-config:node: type: nix-config flake: path:/etc/nixos attr: nixosConfigurations.expanse-node.config.system.build.toplevel switch_mode: switchEOF# wait for expansed to return and quorum to read 3/2,# then the next node$ expanse ctl generation list$ expanse ctl generation diff 11 12$ expanse ctl generation rollback 11# becomes generation 13
Web UI and host console
Two consoles, no agents to install.
Every node serves the web UI on :8443 as soon as its agent starts, and shows a read-only host console on tty1 in place of a login prompt — the same idea as a hypervisor's local console, for a cluster node.
Live pages. Dashboard, Cluster, Nodes, Health & alerts, Blocks (deploy, scale, delete, per-replica logs), Volumes (create, snapshot, resize), Generations (diff, rollback), Events and Settings, updating over Server-Sent Events, never by polling.
Built to work offline. Embedded templates, hand-written CSS, vendored htmx, no build step, and a self-only Content-Security-Policy. Light and dark themes, keyboard accessible, usable on a phone.
Sessions that survive failover. Sessions and CSRF tokens live in the replicated store, so a login keeps working when the UI's VIP moves to another node.
One admin, never a default password. The first agent to open its store generates a random password, hashes it with argon2id and logs the plaintext exactly once. Reset locally with expanse ctl admin reset-password.
Optional OIDC single sign-on. Additive to password login, gated by a required allow-list of verified e-mail addresses, issuing the same store-backed session.
The tty1 host console. Version, hostname and node ID, every address and the web UI URL, cluster name, role and quorum, health, CPU, memory, disks and md mirror state, refreshed every 3 s. It never offers a shell or shows a secret; Alt+F2 is a login shell, and expanse console --once prints the same screen over SSH.
root@n1 — ssh
$ expanse console --onceEXPANSE 1.1.9 2026-09-28 17:42:10 UTC┌─ n1 ─────────────────────────────────────────────────────┐│ Hostname n1 Uptime 3d 4h 12m ││ Health HEALTHY │├─ Management ─────────────────────────────────────────────┤│ eno1 10.0.0.11 ││ Web UI https://10.0.0.11:8443 │├─ Cluster ────────────────────────────────────────────────┤│ Cluster expanse Role leader ││ Quorum 3/2 (3 nodes) │└──────────────────────────────────────────────────────────┘# lost the admin password? reset it locally, shown once:$ expanse ctl admin reset-password \
--socket /run/expanse/agent.sock# trust the UI's own CA in your browser:$ scp root@n1:/persist/expanse/ca/ui-ca.pem .
Observability
The same health signal, everywhere you look.
Expanse does not ship its own metrics database. Every node exports Prometheus metrics for node, resource, volume and quorum health; the web UI's Health page derives the same critical alerts in-process, so the two can never disagree.
Authenticated /metrics. A dedicated listener on :7447 behind the cluster CA and a bearer token you mint with expanse ctl metrics set-token.
Shipped alert rules. Seven critical rules that fire on the first true evaluation — node unhealthy or unreachable, resource unhealthy, volume failed or read-only, quorum without a leader or degraded — and four sustained-window warnings.
A shipped Grafana dashboard with six panels, verified rendering real data through Grafana's own query API.
Within the 30 s budget. A node failure is visible as a firing alert in the UI and Grafana inside 30 s; a measured run at a 2 s scrape cadence closed the pipeline in 11.5–11.7 s. Scrape at 5 s or faster.
Routing is yours. Alertmanager receivers (e-mail, webhooks, pagers) are deliberately left to the operator.
mTLS everywhere, a CA you can rotate without downtime.
Each cluster has its own Ed25519 CA; every node holds a certificate for internal mTLS that renews itself every six hours. Rotation is just "renew onto a new CA": both roots are trusted during the transition and nothing restarts.
Three commands.ca rotate makes the new root primary, ca status lists nodes still pending, ca complete retires the old root — and refuses until every node has renewed, so it is safe to run early.
Nothing joins without a token. Single-use, expiring join tokens; self-signed client certificates are rejected with unknown ca; zero cleartext on the join path, verified in the VM tests.
Two CAs on purpose. Browsers reject Ed25519 TLS, so the web UI has its own ECDSA P-256 CA, created once per cluster. Node-to-node traffic and metrics keep the cluster CA.
Secrets at rest. Key material is age-encrypted, keyed by HKDF over the cluster secret. TPM sealing is a named, deferred feature: no TPM hardware has been available to validate it.
Firewall by default. Cluster ports are reachable only over the overlay; the web UI and metrics ports are open but TLS-authenticated.
rotate the cluster CA
$ systemctl stop expansed$ expanse cluster ca rotate && systemctl start expansedCA rotation started: the new CA is now primary; the old CA stays trusted.$ expanse cluster ca statusrotating: primary CA fingerprint sha256:…pending nodes (not yet renewed): [n3]# … after every node's renewal loop has ticked …every node has renewed onto the new CA; safe to run `cluster ca complete`$ systemctl stop expansed$ expanse cluster ca complete && systemctl start expansedCA rotation complete: outgoing CA retired.
Backup and disaster recovery
restic underneath, one command to bring a node back.
Expanse does not write a backup engine. Each node's cluster state — identity, CA, node TLS, the Raft log and the generation history — sits under /persist/expanse, which restic backs up to any S3-compatible destination. Volume bytes are captured through thin snapshots.
One repository per node.expanse cluster restore pulls exactly that node's state back and the daemon rejoins under its original identity, with no re-init or re-join step.
Whole-cluster loss is covered. Restore every node and start every daemon: quorum reforms cold from the restored logs and the reconciler converges the fleet back to the captured desired state — proven end to end in a VM test.
Volume data as a flat image.expanse ctl volume snapshot takes an LVM thin snapshot; restic and dd move the bytes out and back onto the live DRBD device, checksum-equal.
Scheduling is yours. Backups run from cron, a systemd timer or by hand; an automatic cadence is a named, deferred extension point.
back up and restore a node
$ export RESTIC_REPOSITORY=s3:https://s3.example.net/backups/n1$ export RESTIC_PASSWORD=…$ restic init# once, per node$ restic backup /persist/expanse# on whatever cadence fits# the node is rebuilt from the ISO; bring its state back:$ expanse cluster restorerestored /persist/expanse from the latest backup$ systemctl start expansed# rejoins as the same node
Stateful workloads
Postgres, iSCSI and virtual machines, each with failover.
Three shipped block types put real state on the cluster. Each one reuses the same primitives — replicated volumes, leases, VIPs and singleton scheduling — rather than a bespoke mechanism.
db/postgres
Each replica runs on its own node with its own volume; PostgreSQL streaming replication is the redundancy. A lease-gated controller picks the primary, standbys clone from it with pg_basebackup, and on primary loss a survivor is promoted live with pg_promote() — no restart. Clients use one VIP routed to the primary only.
iscsi/target
One raw DRBD-backed volume exported as a single LUN through LIO. The block is a singleton colocated with the volume's primary and exposed at a stable VIP; when its node fails the VIP moves with the new primary and the initiator's own session recovery reconnects to the same portal address.
vm/instance
A QEMU/KVM guest on a raw replicated disk, scheduled only to nodes with hardware virtualization, with a MAC pinned from the block's name so it comes back at the same address. Node loss means a cold boot elsewhere — the guest's filesystem was verified intact under sustained fsync load through the kill. There is no live migration.
Not built yet
What Expanse does not do today.
Each item below is recorded in the project's own documentation as paused, deferred or a known issue. None of them is quietly missing.
Paused, deferred or known
SMB/NFS file shares. Paused: the Samba block deploys, but its failover has an open bug.
On-prem LLM subsystem. Paused; no accelerator hardware has been available to validate against.
TPM sealing of secrets. Deferred; secrets stay age-encrypted.
VM live migration and an interactive VM console. Deferred; failover is a cold restart and the guest console is the unit's journal.
Alertmanager routing and backup scheduling. Left to the operator by design.
Binary-backed block types on installed nodes. Types such as nginx and redis need their program on the node's PATH, which an installed node does not yet ship; util/echo runs.
Real-hardware validation. Automated tests run on QEMU/KVM. The old-laptop, mini-PC, server and ARM targets in the hardware matrix are marked untested, and idle CPU with a replicated volume attached has not been measured on bare metal.
See it on your own hardware.
The quickstart goes from a USB stick to a one-node cluster with a web console, then adds nodes.