Cluster

Turn many servers into one panel: roles, placement, load balancers with health checks, floating IPs, replicated sites and live account migration.

In preview cluster version 0.1.0

Actions

44 actions, callable from the panel, the command palette and the API as POST /api/v1/a/<id>. Internal actions used between modules are not listed.

ActionWhat it doesRiskPreview
cluster.node.list List nodes with roles, health and load (read) low No dry run
cluster.node.get Get one node (hardware, accounts, pools, VIPs) (read) low No dry run
cluster.node.refresh Re-read a node's hardware and addresses low No dry run
cluster.node.label Set zone, labels, account cap and weight low No dry run
cluster.node.cordon Exclude a node from new placements medium No dry run
cluster.node.uncordon Allow new placements on a node again medium No dry run
cluster.node.roles Set the roles of a node high
cluster.node.drain Drain a node: cordon it and migrate its accounts away high
cluster.node.undrain Return a drained node to service medium No dry run
cluster.node.remove Remove a node from the cluster critical
cluster.topology.get Nodes, roles, load balancers and floating IPs as a graph (read) low No dry run
cluster.firewall.requirements Ports the cluster needs open between nodes and to the internet (read) low No dry run
cluster.placement.list List placement policies (read) low No dry run
cluster.placement.set Create or change a placement policy medium No dry run
cluster.placement.delete Delete a placement policy medium No dry run
cluster.placement.simulate Show where a new account would land and why (read) low No dry run
cluster.placement.assign Pick (and record) the node for a new account medium No dry run
cluster.placement.get Where an account lives and why (read) low No dry run
cluster.lb.list List load-balancer pools (read) low No dry run
cluster.lb.get Get a pool with live backend state (read) low No dry run
cluster.lb.create Create a load-balancer pool high
cluster.lb.update Change a load-balancer pool high
cluster.lb.delete Delete a load-balancer pool high
cluster.lb.apply Validate, diff and activate load-balancer configuration high
cluster.lb.backend Add, remove, enable, disable, drain or re-weight a backend medium
cluster.lb.stats Live traffic and health per backend (read) low No dry run
cluster.lb.history Backend up/down history of a pool (read) low No dry run
cluster.vip.list List floating IPs with MASTER/BACKUP state (read) low No dry run
cluster.vip.create Create a floating IP (VRRP) between nodes high
cluster.vip.update Change a floating IP high
cluster.vip.delete Delete a floating IP high
cluster.vip.state Which node holds a floating IP right now (read) low No dry run
cluster.vip.failover Move a floating IP to another node high
cluster.migrate.start Move an account to another node high
cluster.migrate.cancel Cancel a running or queued migration medium No dry run
cluster.migrate.purge_source Remove the old copy a finished move left on the source node high
cluster.migrate.retry Retry a failed migration high No dry run
cluster.migrate.get Get a migration with its steps (read) low No dry run
cluster.migrate.list List migrations (read) low No dry run
cluster.sync.create Replicate a site directory to other nodes high
cluster.sync.list List replicated directories with their lag (read) low No dry run
cluster.sync.run Replicate now low No dry run
cluster.sync.delete Stop replicating a directory medium No dry run
cluster.reconcile Re-apply load balancers and floating IPs to nodes low No dry run

Permissions and limits

Permissions

  • cluster.node.view View nodes, placement and topology
  • cluster.node.manage Label, cordon and assign roles to nodes
  • cluster.node.drain Drain and undrain nodes
  • cluster.node.remove Remove nodes
  • cluster.placement.manage Manage placement policies
  • cluster.lb.view View load balancers
  • cluster.lb.manage Manage load balancers
  • cluster.vip.manage Manage floating IPs
  • cluster.migrate Migrate accounts between nodes
  • cluster.storage.manage Manage replicated site directories

Plan limits

  • cluster.lb_pools Load-balancer pools
  • cluster.lb_backends Backends per pool
  • cluster.vips Floating IPs
  • cluster.parallel_migrations Parallel migrations

Engineering notes

Generated from modules/cluster/docs.md at build d90e9e2. These are the notes the engineers keep next to the code: precise, technical, and honest about what is not done yet.

Spec docs/specs/cluster.md. Desired state lives in schema m_cluster (node_meta, placement_policies, placements, lb_pools, vips, migrations, sync_links, lb_events); the node list itself (id, roles, status) stays in core and is read with core.node.list / written with core.node.update|retire. Enrolment tokens stay core.node.enrol_token.create.

What exists

  • Nodes: cluster.node.list|get|refresh|label|cordon|uncordon|roles|drain|undrain|remove, cluster.topology.get. refresh runs cluster.node.inventory (hostname, OS, kernel, CPU, RAM, load, disks, private/public IPs, haproxy/keepalived/nginx active) and caches it; roles previews packages (lb -> haproxy+keepalived+psmisc) and installs them through the allow-listed system.apt path. control can neither be given nor taken away. Removal refuses while accounts, pools, VIPs or sync links remain (force needs confirm=<node id>).
  • Placement: policies (least_loaded, spread, pack, pinned; role, plan, labels, zone, priority). simulate shows every candidate and why it was rejected (cordoned, draining, cap, zone, label, no role, offline). assign records the decision and its reason (cluster.placement.get). accounts.created records where an account landed; accounts.deleted forgets it. accounts.create without a node_id asks the internal cluster.placement.decide and records the decision as the placement reason; an explicit node_id wins.
  • Load balancers (lbgen renders, the agent validates): HAProxy on lb nodes; a pool is one document (hostnames, port, http/tcp, algorithm, health check with expected status or body regex and interval/rise/fall, sticky cookie-insert / prefix / source / header, TLS terminate / re-encrypt / passthrough with redirect + HTTP/2, maxconn, per-IP connection cap and request rate, allow/deny CIDRs, PROXY protocol to backends, maintenance 503 page, backend weight / backup / drain / disabled). Pools on the same port share a frontend and route by Host. Every user value is validated against a strict grammar before it reaches the file; there is no free-text snippet. create|update|delete|backend|apply all go: render per LB node -> haproxy -c on a scratch copy -> diff -> certs + atomic swap -> graceful reload -> proof that a new worker pid answers (a failed reload leaves the old worker running and systemd still says "Reloaded") -> otherwise the previous generation is restored and the call fails with apply_failed. The pool row is stored only after every LB node accepted it; nodes already changed are put back if a later one fails. Dry run returns a Plan with the real diff per node. Backend drain/disable/remove first set the server to DRAIN over the runtime socket and wait for its sessions to end (pool drain_timeout_sec, max 120 s). cluster.lb.poll (1 min) records UP/DOWN transitions and emits cluster.lb.backend_down|up. Certificates: tls.cert_id uses the PEM the ssl module already deployed on that LB node (/var/lib/respirecloud/ssl/certs/<id>/{fullchain,privkey}.pem, keys never travel over the bus); without it the agent makes a stable self-signed bundle so the pool works at once.
  • Floating IPs: keepalived VRRP with unicast peers (works where multicast is blocked), per-node priority, nopreempt unless preempt, a pgrep -x haproxy track script (-30 priority when haproxy is gone), PASS auth generated and never returned or shown in diffs. state reports MASTER/BACKUP/STOPPED per node by looking for the address on the interface; failover stops keepalived on the holder until the target owns the address, then starts it again as standby.
  • Replicated site directories (cluster.sync.*): one-way pull from the source node to targets over the same pinned one-shot TLS stream migrations use, newer files only (tar --newer-mtime), run by cluster.sync.tick every minute and on demand. Chosen over lsyncd / rsync-over-ssh because it needs no SSH keys, no extra daemon and no secret shared between nodes. Consequence: lag is interval_sec (min 60) + up to 1 min, deletes are not propagated, and writes must go to the source node. The zero-lag alternative (shared SeaweedFS filer mount from storage) is not wired.
  • Migration / drain (cluster.migrate.*, cluster.node.drain): steps preflight -> dump (databases listed by databases.db.list; mariadb/mysql via mysqldump, postgres via pg_dump, local default socket) -> transfer (source opens a TLS listener on 38400-38499 with a single-use token and a pinned fingerprint, the target pulls; the target home must be empty) -> restore -> verify (file count, bytes and a digest of sorted path+size must match the live source) -> switch -> cleanup (staging removed; the source copy is kept, marked by the agent so a move back replaces it and cluster.migrate.purge_source can remove it). Failure or cancel before the switch removes what was imported on the target; the source is never modified. A drain cordons the node, plans one migration per account with the placement policy (counting accounts already planned), runs them (2 at a time) and flips the node to drained when none are left. Dry run lists the moves with size estimates.
  • Firewall: cluster.firewall.requirements lists the ports (4222, 38400-38499, VRRP proto 112, every pool frontend port) for the security module to open; this module does not touch the firewall.

Verified in the lab (PROFILE=cluster, 2026-10-10, HAProxy 2.8.16, keepalived on Ubuntu 24.04)

See docs/tasks/w8/w8-01-cluster-lb.handover.md for the command log: LB pool over cp1+web2 on both nodes, kill a backend -> traffic stays on the other with zero failures, 4 reloads under load with no failed request, sticky cookie, TLS terminate + h2 + redirect, port clash -> automatic rollback, VIP between two LB nodes (hard failure of the master: traffic back in 2.8 s), manual failover, incremental site sync, migration web2 -> cp1 with verified identical copy, rollback on failure and on cancel, drain with cordon.

w9-02 follow-ups (landed)

  • accounts.set_node {id, node_id} (internal, system only) moves the row, ensures the user on the target and emits accounts.moved. Migration steps are now preflight, dump, transfer, restore, verify, switch, follow, cleanup. follow calls the idempotent handlers web|ssl|dns|mail|databases.on_account_moved in that order (they are also subscribed to the event). A failing follower leaves the migration awaiting_switch with the reason, emits cluster.migration.follow_failed, and the tick retries every ~90 s; after the switch the copy is never rolled back or cancelled. web re-renders vhosts on the new node then removes them from the old; ssl copies chain+key node to node (cluster.cert.mirror), installs for the web targets and removes the old copy; dns rewrites A/AAAA holding the old node address (records pointing elsewhere, e.g. a pooled zone, stay); databases records the new node and re-applies db/users/grants; mail only reports (mailboxes live on mail servers). Placement follows via cluster.on_account_moved; drain completes when no account is left.
  • Pooled sites: a pool spec has site_id; web.sites.set_nodes renders the vhost on every backend node (table m_web.site_nodes) and removes it when a backend leaves; the site's files must exist on each node (use cluster.sync.*).
  • LB certificates: for terminate/reencrypt pools ssl.cert.lb_status finds the site's active certificate and ssl.cert.lb_deploy copies it to each LB node (cert_lb_nodes); renewals re-copy and emit ssl.cert.lb_deployed, which makes cluster reload HAProxy. No certificate yet = self-signed fallback.
  • Firewall: cluster.firewall.requirements also returns per-node needs with source addresses; security.firewall.cluster applies them as rc-cluster: allow rules through the normal apply/auto-rollback path (only on nodes that already have a panel-managed firewall).

w10-fix-14a

  • Move followers are called in order with ordered:true (web renders on the new node, ssl, dns re-points, mail, databases) and a last finalize pass removes the vhosts from the old node after a 12 s grace; the plain accounts.moved event no longer re-points DNS or removes old vhosts.
  • Zone records that name "the server" use the node the account lives on (dns accountIPs), not the DNS server's node.
  • cluster.node.drain refuses the control node; cluster.node.roles refuses dropping a role the node is serving.
  • The agent deletes a home on rollback only when this migration created it (markers/imported-<id>).

w10-fix-16

  • Pooling prepares the node (cluster/010): web.sites.set_nodes (called by lb.create|update|delete) first asks accounts (accounts.node.prepare, internal to web) to create the account's Linux user on each new backend node with the same uid/gid as on its home node (system.user.ensure takes optional uid/gid; an existing user with other ids or ids taken by another name are refused). The dry run of an LB call lists these steps. When the account's last site leaves a node accounts.node.release removes the user there again together with its home (only a replica of the home node's files; a leftover home would make a later move to that node refuse). Delete the cluster.sync link of that site first. The home node is never touched.
  • Switch grace follows the DNS TTL (cluster/005): the DNS follower answers with the longest TTL of the records it re-pointed; the old node keeps its vhosts for that TTL + 2 s, at least 12 s and at most 10 minutes (moveGraceMax). A resolver that holds an address longer than 10 minutes can still reach the old node afterwards.
  • cluster.migrate.purge_source also drops the old node's database copies (cluster/004): the account's databases (databases.db.list) are dropped on the source together with the kept home (agent-checked kept-copy marker, strict name grammar); the dry run lists each one. (Database users: see w10-fix-18.)

Known gaps / contract requests

  • A VRRP need is expressed as "allow any protocol from the peer" because the firewall rule model has no raw IP-protocol match (broader than needed).
  • Pooled site files are not replicated automatically; create a cluster.sync link for the docroot.
  • cluster.parallel_migrations / cluster.lb_pools limit keys are declared but not enforced (admin-only module); the migration semaphore is fixed at 2.
  • DB replication helpers (MariaDB primary/replica, Postgres streaming), WireGuard overlay, rolling agent upgrade, SSH enrolment, HTTP/3, OCSP stapling, CrowdSec bouncer, UDP pools, DNS-TTL lowering and mail/DNS repointing during migration are not built.
  • Stats are polled per call (no history in stats yet); lb_events rows are per LB node.

w10-fix-18

  • One uid/gid per account, cluster-wide (cluster/013): accounts allocates the number once (counter id_counter in its schema, starting above every recorded uid and at least 10000), stores it as accounts.uid/gid and passes it to every system.user.ensure for the account (create, accounts.node.prepare, the move's import via ClusterImport.UID/GID). The agent refuses an id that belongs to another name ("is already used by"); a new account then takes the next number. Accounts from before keep their node's id (gid recorded as the uid); a node where that id is taken refuses the pool/move with both numbers in the message.
  • Moves resume (cluster/011): every step is recorded; after a restart the module requeues running moves (OnStart, and cluster.migrate.tick for a stalled one) and the move continues from the recorded step. Before it does, reconcile asks accounts.get where the account lives: on the target = every step up to the switch counts as done; a step that was running runs again (an interrupted transfer first clears the half-imported home, which the agent removes only for the migration that imported it). rollback and the failure path never remove anything on the node the account lives on, and when that cannot be determined the move is paused (status paused): cluster.migrate.retry continues it from the recorded step, cluster.migrate.cancel rolls it back when it never switched (after the switch the copy on the old node stays; moving back is a new move). A failing step after the switch also pauses instead of failing.
  • Database users move with their credentials (cluster/012): the dump step asks the source agent (cluster.migrate.dbdump with users) to write each user's authentication string (MariaDB: native hash per host; MySQL: native or caching_sha2; PostgreSQL: the SCRAM verifier) into the migration's staging directory; it travels in the pinned stream with the dumps and the restore step creates the users from it on the target before the switch. No plain password is read, and nothing goes through the control plane or the bus. The databases follower then finds the users and re-applies hosts, TLS, connection limits and grants. cluster.migrate.purge_source drops the old node's database users together with the databases (agent-checked kept-copy marker). A user whose password is not a SCRAM verifier (md5) or whose plugin cannot be carried (e.g. unix_socket) stops the move at the dump step with a message naming the user.

w10-fix-19

  • PostgreSQL moves keep owners and grants (cluster/012): the dump keeps owners and ACLs; the target has <db>_own/_rw/_ro first, and after the restore the database's standard grants and default privileges are re-applied. What the app could do on the source it can do on the target (tables, sequences, views).
  • Records follow the service (cluster/016): a move re-points the account's web names (apex, www, app names) at the new node. Name-server glue (ns1/ns2 and any NS target inside the zone) stays on the DNS server's node. Mail names (MX targets inside the zone, mail, smtp, imap, pop, autoconfig, autodiscover) stay where the mailboxes live. Mail does not move with an account: mailbox data and the mail domain stay on their mail server and DNS keeps pointing there, until a mailbox move exists. The dry run of cluster.migrate.start says so.
  • Moving onto a pool replica (cluster/018): if a pool of the account's site lists the target as a backend, the replica home there is replaced by the moved copy (same account, same uid/gid); any other data on the target is still refused. cluster.node.drain on a node that is already draining is the retry (accounts with a failed move get a new one; the answer lists blocked accounts with the reason); cluster.node.undrain cancels the drain.
  • HAProxy applies (cluster/014): a port another process holds is refused before anything is written, naming the port and the process; the old configuration keeps serving. A node whose last reload failed ("Reload failed!") is restarted on its last good configuration by the next apply (and by the rollback).
  • A failed pool create/update releases what it prepared (cluster/015): the account's user and home on the nodes it had added are removed again.
  • The switch grace survives a restart (cluster/017): the end of the DNS-TTL wait is stored in the move row (follow_until).