Actions
44 actions, callable from the panel, the command palette and the API as
POST /api/v1/a/<id>. Internal actions used between modules are not listed.
| Action | What it does | Risk | Preview |
|---|---|---|---|
cluster.node.list | List nodes with roles, health and load (read) | low | No dry run |
cluster.node.get | Get one node (hardware, accounts, pools, VIPs) (read) | low | No dry run |
cluster.node.refresh | Re-read a node's hardware and addresses | low | No dry run |
cluster.node.label | Set zone, labels, account cap and weight | low | No dry run |
cluster.node.cordon | Exclude a node from new placements | medium | No dry run |
cluster.node.uncordon | Allow new placements on a node again | medium | No dry run |
cluster.node.roles | Set the roles of a node | high | |
cluster.node.drain | Drain a node: cordon it and migrate its accounts away | high | |
cluster.node.undrain | Return a drained node to service | medium | No dry run |
cluster.node.remove | Remove a node from the cluster | critical | |
cluster.topology.get | Nodes, roles, load balancers and floating IPs as a graph (read) | low | No dry run |
cluster.firewall.requirements | Ports the cluster needs open between nodes and to the internet (read) | low | No dry run |
cluster.placement.list | List placement policies (read) | low | No dry run |
cluster.placement.set | Create or change a placement policy | medium | No dry run |
cluster.placement.delete | Delete a placement policy | medium | No dry run |
cluster.placement.simulate | Show where a new account would land and why (read) | low | No dry run |
cluster.placement.assign | Pick (and record) the node for a new account | medium | No dry run |
cluster.placement.get | Where an account lives and why (read) | low | No dry run |
cluster.lb.list | List load-balancer pools (read) | low | No dry run |
cluster.lb.get | Get a pool with live backend state (read) | low | No dry run |
cluster.lb.create | Create a load-balancer pool | high | |
cluster.lb.update | Change a load-balancer pool | high | |
cluster.lb.delete | Delete a load-balancer pool | high | |
cluster.lb.apply | Validate, diff and activate load-balancer configuration | high | |
cluster.lb.backend | Add, remove, enable, disable, drain or re-weight a backend | medium | |
cluster.lb.stats | Live traffic and health per backend (read) | low | No dry run |
cluster.lb.history | Backend up/down history of a pool (read) | low | No dry run |
cluster.vip.list | List floating IPs with MASTER/BACKUP state (read) | low | No dry run |
cluster.vip.create | Create a floating IP (VRRP) between nodes | high | |
cluster.vip.update | Change a floating IP | high | |
cluster.vip.delete | Delete a floating IP | high | |
cluster.vip.state | Which node holds a floating IP right now (read) | low | No dry run |
cluster.vip.failover | Move a floating IP to another node | high | |
cluster.migrate.start | Move an account to another node | high | |
cluster.migrate.cancel | Cancel a running or queued migration | medium | No dry run |
cluster.migrate.purge_source | Remove the old copy a finished move left on the source node | high | |
cluster.migrate.retry | Retry a failed migration | high | No dry run |
cluster.migrate.get | Get a migration with its steps (read) | low | No dry run |
cluster.migrate.list | List migrations (read) | low | No dry run |
cluster.sync.create | Replicate a site directory to other nodes | high | |
cluster.sync.list | List replicated directories with their lag (read) | low | No dry run |
cluster.sync.run | Replicate now | low | No dry run |
cluster.sync.delete | Stop replicating a directory | medium | No dry run |
cluster.reconcile | Re-apply load balancers and floating IPs to nodes | low | No dry run |
Permissions and limits
Permissions
cluster.node.viewView nodes, placement and topologycluster.node.manageLabel, cordon and assign roles to nodescluster.node.drainDrain and undrain nodescluster.node.removeRemove nodescluster.placement.manageManage placement policiescluster.lb.viewView load balancerscluster.lb.manageManage load balancerscluster.vip.manageManage floating IPscluster.migrateMigrate accounts between nodescluster.storage.manageManage replicated site directories
Plan limits
cluster.lb_poolsLoad-balancer poolscluster.lb_backendsBackends per poolcluster.vipsFloating IPscluster.parallel_migrationsParallel migrations
Engineering notes
Generated from modules/cluster/docs.md at build d90e9e2. These are the notes the engineers keep
next to the code: precise, technical, and honest about what is not done yet.
Spec docs/specs/cluster.md. Desired state lives in schema m_cluster (node_meta, placement_policies, placements, lb_pools,
vips, migrations, sync_links, lb_events); the node list itself (id, roles, status) stays in core and is read with
core.node.list / written with core.node.update|retire. Enrolment tokens stay core.node.enrol_token.create.
What exists
- Nodes:
cluster.node.list|get|refresh|label|cordon|uncordon|roles|drain|undrain|remove,cluster.topology.get.refreshrunscluster.node.inventory(hostname, OS, kernel, CPU, RAM, load, disks, private/public IPs, haproxy/keepalived/nginx active) and caches it;rolespreviews packages (lb-> haproxy+keepalived+psmisc) and installs them through the allow-listedsystem.aptpath.controlcan neither be given nor taken away. Removal refuses while accounts, pools, VIPs or sync links remain (force needsconfirm=<node id>). - Placement: policies (
least_loaded,spread,pack,pinned; role, plan, labels, zone, priority).simulateshows every candidate and why it was rejected (cordoned, draining, cap, zone, label, no role, offline).assignrecords the decision and its reason (cluster.placement.get).accounts.createdrecords where an account landed;accounts.deletedforgets it.accounts.createwithout anode_idasks the internalcluster.placement.decideand records the decision as the placement reason; an explicitnode_idwins. - Load balancers (
lbgenrenders, the agent validates): HAProxy onlbnodes; a pool is one document (hostnames, port, http/tcp, algorithm, health check with expected status or body regex and interval/rise/fall, sticky cookie-insert / prefix / source / header, TLS terminate / re-encrypt / passthrough with redirect + HTTP/2, maxconn, per-IP connection cap and request rate, allow/deny CIDRs, PROXY protocol to backends, maintenance 503 page, backend weight / backup / drain / disabled). Pools on the same port share a frontend and route by Host. Every user value is validated against a strict grammar before it reaches the file; there is no free-text snippet.create|update|delete|backend|applyall go: render per LB node ->haproxy -con a scratch copy -> diff -> certs + atomic swap -> graceful reload -> proof that a new worker pid answers (a failed reload leaves the old worker running and systemd still says "Reloaded") -> otherwise the previous generation is restored and the call fails withapply_failed. The pool row is stored only after every LB node accepted it; nodes already changed are put back if a later one fails. Dry run returns a Plan with the real diff per node. Backenddrain/disable/removefirst set the server to DRAIN over the runtime socket and wait for its sessions to end (pooldrain_timeout_sec, max 120 s).cluster.lb.poll(1 min) records UP/DOWN transitions and emitscluster.lb.backend_down|up. Certificates:tls.cert_iduses the PEM thesslmodule already deployed on that LB node (/var/lib/respirecloud/ssl/certs/<id>/{fullchain,privkey}.pem, keys never travel over the bus); without it the agent makes a stable self-signed bundle so the pool works at once. - Floating IPs: keepalived VRRP with unicast peers (works where multicast is blocked), per-node priority,
nopreemptunlesspreempt, apgrep -x haproxytrack script (-30 priority when haproxy is gone), PASS auth generated and never returned or shown in diffs.statereports MASTER/BACKUP/STOPPED per node by looking for the address on the interface;failoverstops keepalived on the holder until the target owns the address, then starts it again as standby. - Replicated site directories (
cluster.sync.*): one-way pull from the source node to targets over the same pinned one-shot TLS stream migrations use, newer files only (tar --newer-mtime), run bycluster.sync.tickevery minute and on demand. Chosen over lsyncd / rsync-over-ssh because it needs no SSH keys, no extra daemon and no secret shared between nodes. Consequence: lag isinterval_sec(min 60) + up to 1 min, deletes are not propagated, and writes must go to the source node. The zero-lag alternative (shared SeaweedFS filer mount fromstorage) is not wired. - Migration / drain (
cluster.migrate.*,cluster.node.drain): steps preflight -> dump (databases listed bydatabases.db.list; mariadb/mysql via mysqldump, postgres via pg_dump, local default socket) -> transfer (source opens a TLS listener on 38400-38499 with a single-use token and a pinned fingerprint, the target pulls; the target home must be empty) -> restore -> verify (file count, bytes and a digest of sorted path+size must match the live source) -> switch -> cleanup (staging removed; the source copy is kept, marked by the agent so a move back replaces it andcluster.migrate.purge_sourcecan remove it). Failure or cancel before the switch removes what was imported on the target; the source is never modified. A drain cordons the node, plans one migration per account with the placement policy (counting accounts already planned), runs them (2 at a time) and flips the node todrainedwhen none are left. Dry run lists the moves with size estimates. - Firewall:
cluster.firewall.requirementslists the ports (4222, 38400-38499, VRRP proto 112, every pool frontend port) for thesecuritymodule to open; this module does not touch the firewall.
Verified in the lab (PROFILE=cluster, 2026-10-10, HAProxy 2.8.16, keepalived on Ubuntu 24.04)
See docs/tasks/w8/w8-01-cluster-lb.handover.md for the command log: LB pool over cp1+web2 on both nodes, kill a backend -> traffic stays
on the other with zero failures, 4 reloads under load with no failed request, sticky cookie, TLS terminate + h2 + redirect, port clash ->
automatic rollback, VIP between two LB nodes (hard failure of the master: traffic back in 2.8 s), manual failover, incremental site sync,
migration web2 -> cp1 with verified identical copy, rollback on failure and on cancel, drain with cordon.
w9-02 follow-ups (landed)
accounts.set_node {id, node_id}(internal, system only) moves the row, ensures the user on the target and emitsaccounts.moved. Migration steps are now preflight, dump, transfer, restore, verify, switch, follow, cleanup.followcalls the idempotent handlersweb|ssl|dns|mail|databases.on_account_movedin that order (they are also subscribed to the event). A failing follower leaves the migrationawaiting_switchwith the reason, emitscluster.migration.follow_failed, and the tick retries every ~90 s; after the switch the copy is never rolled back or cancelled. web re-renders vhosts on the new node then removes them from the old; ssl copies chain+key node to node (cluster.cert.mirror), installs for the web targets and removes the old copy; dns rewrites A/AAAA holding the old node address (records pointing elsewhere, e.g. a pooled zone, stay); databases records the new node and re-applies db/users/grants; mail only reports (mailboxes live on mail servers). Placement follows viacluster.on_account_moved; drain completes when no account is left.- Pooled sites: a pool spec has
site_id;web.sites.set_nodesrenders the vhost on every backend node (tablem_web.site_nodes) and removes it when a backend leaves; the site's files must exist on each node (usecluster.sync.*). - LB certificates: for terminate/reencrypt pools
ssl.cert.lb_statusfinds the site's active certificate andssl.cert.lb_deploycopies it to each LB node (cert_lb_nodes); renewals re-copy and emitssl.cert.lb_deployed, which makes cluster reload HAProxy. No certificate yet = self-signed fallback. - Firewall:
cluster.firewall.requirementsalso returns per-node needs with source addresses;security.firewall.clusterapplies them asrc-cluster:allow rules through the normal apply/auto-rollback path (only on nodes that already have a panel-managed firewall).
w10-fix-14a
- Move followers are called in order with
ordered:true(web renders on the new node, ssl, dns re-points, mail, databases) and a lastfinalizepass removes the vhosts from the old node after a 12 s grace; the plainaccounts.movedevent no longer re-points DNS or removes old vhosts. - Zone records that name "the server" use the node the account lives on (dns
accountIPs), not the DNS server's node. cluster.node.drainrefuses the control node;cluster.node.rolesrefuses dropping a role the node is serving.- The agent deletes a home on rollback only when this migration created it (
markers/imported-<id>).
w10-fix-16
- Pooling prepares the node (cluster/010):
web.sites.set_nodes(called bylb.create|update|delete) first asks accounts (accounts.node.prepare, internal to web) to create the account's Linux user on each new backend node with the same uid/gid as on its home node (system.user.ensuretakes optionaluid/gid; an existing user with other ids or ids taken by another name are refused). The dry run of an LB call lists these steps. When the account's last site leaves a nodeaccounts.node.releaseremoves the user there again together with its home (only a replica of the home node's files; a leftover home would make a later move to that node refuse). Delete thecluster.synclink of that site first. The home node is never touched. - Switch grace follows the DNS TTL (cluster/005): the DNS follower answers with the longest TTL of the records it re-pointed; the old node keeps its vhosts for
that TTL + 2 s, at least 12 s and at most 10 minutes (
moveGraceMax). A resolver that holds an address longer than 10 minutes can still reach the old node afterwards. cluster.migrate.purge_sourcealso drops the old node's database copies (cluster/004): the account's databases (databases.db.list) are dropped on the source together with the kept home (agent-checked kept-copy marker, strict name grammar); the dry run lists each one. (Database users: see w10-fix-18.)
Known gaps / contract requests
- A VRRP need is expressed as "allow any protocol from the peer" because the firewall rule model has no raw IP-protocol match (broader than needed).
- Pooled site files are not replicated automatically; create a
cluster.synclink for the docroot. cluster.parallel_migrations/cluster.lb_poolslimit keys are declared but not enforced (admin-only module); the migration semaphore is fixed at 2.- DB replication helpers (MariaDB primary/replica, Postgres streaming), WireGuard overlay, rolling agent upgrade, SSH enrolment, HTTP/3, OCSP stapling, CrowdSec bouncer, UDP pools, DNS-TTL lowering and mail/DNS repointing during migration are not built.
- Stats are polled per call (no history in
statsyet);lb_eventsrows are per LB node.
w10-fix-18
- One uid/gid per account, cluster-wide (cluster/013):
accountsallocates the number once (counterid_counterin its schema, starting above every recorded uid and at least 10000), stores it asaccounts.uid/gidand passes it to everysystem.user.ensurefor the account (create,accounts.node.prepare, the move's import viaClusterImport.UID/GID). The agent refuses an id that belongs to another name ("is already used by"); a new account then takes the next number. Accounts from before keep their node's id (gid recorded as the uid); a node where that id is taken refuses the pool/move with both numbers in the message. - Moves resume (cluster/011): every step is recorded; after a restart the module requeues running moves (
OnStart, andcluster.migrate.tickfor a stalled one) and the move continues from the recorded step. Before it does,reconcileasksaccounts.getwhere the account lives: on the target = every step up to the switch counts as done; a step that was running runs again (an interrupted transfer first clears the half-imported home, which the agent removes only for the migration that imported it).rollbackand the failure path never remove anything on the node the account lives on, and when that cannot be determined the move is paused (statuspaused):cluster.migrate.retrycontinues it from the recorded step,cluster.migrate.cancelrolls it back when it never switched (after the switch the copy on the old node stays; moving back is a new move). A failing step after the switch also pauses instead of failing. - Database users move with their credentials (cluster/012): the dump step asks the source agent (
cluster.migrate.dbdumpwithusers) to write each user's authentication string (MariaDB: native hash per host; MySQL: native or caching_sha2; PostgreSQL: the SCRAM verifier) into the migration's staging directory; it travels in the pinned stream with the dumps and the restore step creates the users from it on the target before the switch. No plain password is read, and nothing goes through the control plane or the bus. The databases follower then finds the users and re-applies hosts, TLS, connection limits and grants.cluster.migrate.purge_sourcedrops the old node's database users together with the databases (agent-checked kept-copy marker). A user whose password is not a SCRAM verifier (md5) or whose plugin cannot be carried (e.g. unix_socket) stops the move at the dump step with a message naming the user.
w10-fix-19
- PostgreSQL moves keep owners and grants (cluster/012): the dump keeps owners and ACLs; the target has
<db>_own/_rw/_rofirst, and after the restore the database's standard grants and default privileges are re-applied. What the app could do on the source it can do on the target (tables, sequences, views). - Records follow the service (cluster/016): a move re-points the account's web names (apex, www, app names) at the new node. Name-server glue (ns1/ns2 and any NS target inside the zone) stays on the DNS server's node. Mail names (MX targets inside the zone, mail, smtp, imap, pop, autoconfig, autodiscover) stay where the mailboxes live. Mail does not move with an account: mailbox data and the mail domain stay on their mail server and DNS keeps pointing there, until a mailbox move exists. The dry run of
cluster.migrate.startsays so. - Moving onto a pool replica (cluster/018): if a pool of the account's site lists the target as a backend, the replica home there is replaced by the moved copy (same account, same uid/gid); any other data on the target is still refused.
cluster.node.drainon a node that is already draining is the retry (accounts with a failed move get a new one; the answer listsblockedaccounts with the reason);cluster.node.undraincancels the drain. - HAProxy applies (cluster/014): a port another process holds is refused before anything is written, naming the port and the process; the old configuration keeps serving. A node whose last reload failed ("Reload failed!") is restarted on its last good configuration by the next apply (and by the rollback).
- A failed pool create/update releases what it prepared (cluster/015): the account's user and home on the nodes it had added are removed again.
- The switch grace survives a restart (cluster/017): the end of the DNS-TTL wait is stored in the move row (
follow_until).