Getting started

From a blank machine to a running cluster.

This walkthrough assumes no familiarity with Nix or btrfs. It takes one machine to an installed node, forms a one-node cluster, opens the web console, and then adds more nodes.

What you need

  • Any machine with at least 2 CPU cores, 2 GB RAM and a 20 GB disk. A ten-year-old laptop is the design target. x86_64 is the tested architecture; aarch64 images build from the flake but are untested.
  • A USB stick of at least 4 GB (the ISO is about 1.7 GB).
  • An SSH public key. Optional, but strongly recommended: the installer warns when none is given.
  • A machine with Nix to build the ISO, or a downloaded release image (see Download).

Installing Expanse destroys all data on the disks you select. After installation the root filesystem is wiped on every reboot; only /persist survives. That is the point, but back things up first.

1. Make the installer USB

Build the ISO from the flake (or use a downloaded release image), then write it to the stick. Replace /dev/sdX with your USB device — check with lsblk first.

$ nix build github:team-expanse/expanse/v1.1.9#iso
$ sudo dd if=result/iso/*.iso of=/dev/sdX bs=4M status=progress oflag=direct

The x86_64 image is hybrid: it boots under UEFI and legacy BIOS alike.

2. Boot the installer

Boot the target machine from the stick. A full-screen installer starts automatically on the console.

Headless machine? The console prints a random root password, and the installer advertises itself over mDNS as _expanse-installer._tcp. Log in over SSH and start the same installer there:

$ ssh root@<ip>
$ expanse install --tui

3. Walk through the installer

The installer is a framed, six-step console page that fits any terminal from 80×25 up. The step indicator shows where you are.

  1. Welcome — a hardware summary.
  2. Disks — select target disks. One disk gives a single layout (a btrfs system partition plus the remainder as LVM). Two or more give a mirror: the first two disks each carry an ESP and a system partition mirrored as md RAID1 so either disk alone boots, and every disk's remainder joins the LVM data pool. Disks holding data are marked CONTAINS DATA and require explicit confirmation.
  3. Network — DHCP (recommended) or static.
  4. SSH key — paste a public key, or type gh:<username> to fetch your keys from GitHub.
  5. Review — the full plan, with a red list of everything that will be destroyed. Type INSTALL to proceed; the install starts on the last letter, with no Enter needed.
  6. Progress — the current stage, a progress bar, elapsed time and a scrolling tail of the log (the full log is in /tmp/expanse-install.log).

When it finishes you get your node ID (a UUID), the machine's addresses and the next steps: the web UI URL, where the admin password is logged, and how to form a one-node cluster. Reboot and remove the stick.

Unattended install (no TUI)

Write a config file and run the installer against it. Use --dry-run first: it prints every command that would run and touches nothing.

expanse-install.yaml
version: 1
disks:
  layout: auto          # auto | single | mirror
  devices: [/dev/sda]   # empty = use every disk
network:
  mode: dhcp
ssh:
  authorized_keys:
    - ssh-ed25519 AAAA...
timezone: UTC
$ expanse install --config /tmp/expanse-install.yaml --dry-run
$ expanse install --config /tmp/expanse-install.yaml --force
  • Without --force the installer refuses to touch a disk that contains data and exits without modifying it.
  • --target-flake REF installs from a specific flake instead of the one embedded in the ISO.

4. First boot

The node boots to multi-user.target, generates its identity (a UUID and an Ed25519 keypair) in /persist/expanse/identity, and comes up on the network via DHCP and mDNS.

The screen on tty1 shows the host console: the Expanse version, hostname and node ID, every address with the web UI URL, the cluster name, role and quorum (or "not in a cluster yet"), the node's health, CPU, memory, disks and md mirror state, and uptime. It refreshes every few seconds and is read-only — it never offers a shell or shows a secret. Press Alt+F2 for a login shell on tty2, or use the serial console.

Log in with your SSH key and verify:

$ ssh root@expanse-<node-id-prefix>.local
$ expanse version
$ expanse node info
$ btrfs subvolume list /    # @root @nix @persist @log
$ vgs expanse                # the LVM data pool volumes will live in

The volume group exists but is otherwise empty. Replicated volumes (the thin pool and DRBD) are an explicit opt-in, covered in step 8.

5. Form a one-node cluster

One installed node is already a usable cluster. Form it with the agent stopped — cluster init refuses while expansed runs, since the running agent would never join the new cluster.

$ systemctl stop expansed
$ expanse cluster init --expect 1
cluster initialized: 5c0b3e2a-… (name "expanse")
data dir:    /persist/expanse
raft addr:   10.0.0.11:7444

Join command (single-use, 15m):
  expanse cluster join --address 10.0.0.11:7446 --token expanse-join-…
$ systemctl start expansed
$ expanse cluster status
cluster:   expanse (5c0b3e2a-…)
version:   1.1.9
generation: 1
leader:    n1
quorum:    1/1
nodes:     1

Keep the printed join command: it is what a second node runs, and it is valid for 15 minutes. Volumes and blocks run on this node now, with no redundancy until more nodes join.

6. Open the web UI

Every node serves the UI on port 8443 over TLS as soon as its agent starts. Point a browser at https://<node-ip>:8443/.

There is one account, admin. Its initial password is random, generated the first time the agent opens its store, and logged exactly once in the agent's journal:

$ journalctl -u expansed | grep "initial web UI admin password"
… generated initial web UI admin password — save it now, it will not be shown again  password=…

If it has scrolled past, reset it instead of searching: expanse ctl admin reset-password --socket /run/expanse/agent.sock prints a new one, once.

The certificate is issued by the cluster's own web UI CA. Browsers show it as untrusted until you import /persist/expanse/ca/ui-ca.pem (not ca.pem, which is the cluster's Ed25519 CA and rejected by browsers) as a trusted authority:

$ scp root@<node>:/persist/expanse/ca/ui-ca.pem .
# Firefox: Settings → Privacy & Security → Certificates → Authorities → Import
# curl:    --cacert ui-ca.pem

7. Add nodes

Install each further machine the same way. Then, on the new node with its agent not yet started, run the join command the leader printed (or mint a fresh token on the leader with expanse cluster token create, with that node's agent stopped):

# on the new node
$ systemctl stop expansed
$ expanse cluster join --address 10.0.0.11:7446 --token expanse-join-…
joined cluster 5c0b3e2a-… as n2 (10.0.0.12:7444)
$ systemctl start expansed

# from any node
$ expanse cluster status
quorum:    3/2    # three voters, two needed: the fully healthy reading

Nodes on the same L2 network can also find the join endpoint themselves: expanse cluster discover lists clusters and unjoined nodes, and cluster join --discover skips --address. Raft needs three voters — or two plus a --role witness node that never runs workloads — to survive the loss of one.

8. Volumes and blocks

Enable replicated volumes

Replicated volumes need the DRBD kernel module and a thin pool, so they are an explicit opt-in. Add to the node's NixOS configuration (/persist/etc/nixos) and rebuild, then create the pool once per node — its size is a capacity decision, so it is not inferred:

configuration.nix
expanse.agent.storageVG = "expanse";
expanse.agent.storagePool = "pool";   # empty = thick volumes, no snapshots
$ lvcreate --type thin-pool -l 95%FREE -n pool expanse   # leave headroom
$ expanse ctl volume create data --size 10Gi
$ expanse ctl volume list
$ expanse doctor storage --vg expanse

A volume without an explicit --replication takes its class's target (3 by default) and is placed on as many eligible nodes as exist; it reads underreplicated until the target is met and gains replicas automatically as nodes join.

Deploy a block

Blocks are services described in YAML and run as systemd units. util/echo is the shipped test workload and runs on a fresh install; other types are listed with expanse ctl catalog list.

echo.yaml
apiVersion: expanse.io/v1
kind: Block
metadata:
  name: echo
  namespace: default
spec:
  type: util/echo
  replicas: 1
  resources:
    requests:
      cpu: 100m
      memory: 64Mi
  config:
    port: 18080
    body: "hello from expanse"
$ expanse ctl block apply -f echo.yaml --dry-run   # validate only
$ expanse ctl block apply -f echo.yaml
created default/echo
$ expanse ctl block explain echo
echo: 1/1 placed (phase RUNNING)
$ expanse ctl block logs default/echo

The same deploy, scale and delete actions are on the Blocks page of the web UI, with schema errors shown inline.

Troubleshooting

SymptomFix
"refusing to wipe non-empty disk"That is intentional. Add --force only if you mean it.
Install fails at partitionCheck lsblk; the target may be the USB stick itself.
Root not wiped on rebootThe @root-blank snapshot is missing — expanse doctor storage flags it. Re-install to restore it; never hand-edit the subvolume layout.
Clock warnings in logsOld CMOS battery; chrony fixes the time after a few steps. Health reads "unknown" rather than unhealthy for the first 10 minutes after boot while it synchronises.
Want to see what changed on rebootjournalctl -u expanse-impermanence-check
Browser says the UI's certificate is invalidImport ui-ca.pem, not ca.pem (see step 6).

Re-installing is safe: boot the stick again and re-run. The installer is idempotent, and node identity survives in /persist unless you destroy the system partition yourself.