Stats

Fleet, node and per-account usage metrics, scoped queries and alerts.

In preview stats version 0.1.0 Part of Stats and logs

Actions

15 actions, callable from the panel, the command palette and the API as POST /api/v1/a/<id>. Internal actions used between modules are not listed.

ActionWhat it doesRiskPreview
stats.overview.get Admin overview (read) low No dry run
stats.node.get Node metrics (read) low No dry run
stats.usage.get Account usage (read) low No dry run
stats.top.get Top consumers (read) low No dry run
stats.query Metric query (scoped) (read) low No dry run
stats.alert.rule.list List alert rules (read) low No dry run
stats.alert.rule.create Create an alert rule medium No dry run
stats.alert.rule.update Update an alert rule medium No dry run
stats.alert.rule.delete Delete an alert rule medium No dry run
stats.alert.list List alerts (read) low No dry run
stats.alert.ack Acknowledge an alert low No dry run
stats.alert.resolve Resolve an alert by hand low No dry run
stats.retention.get Read retention and store disk use (read) low No dry run
stats.retention.set Set metrics retention high No dry run
stats.reconcile Re-apply the metrics store and collectors low No dry run

Permissions and limits

Permissions

  • stats.overview.view See the fleet overview
  • stats.node.view See node metrics
  • stats.usage.view See account usage
  • stats.query Run scoped metric queries
  • stats.alert.view See alerts and rules
  • stats.alert.manage Create and change alert rules
  • stats.alert.respond Acknowledge and resolve alerts
  • stats.admin Metrics retention and collectors

Engineering notes

Generated from modules/stats/docs.md at build d90e9e2. These are the notes the engineers keep next to the code: precise, technical, and honest about what is not done yet.

Metrics for the fleet, nodes and every account, stored in VictoriaMetrics on the control node.

How it works

  • Collect: the plugin asks every node's agent for a snapshot (stats.collect) every 15 s and writes it to VictoriaMetrics (/api/v1/import/prometheus, loopback). The agent only samples (internal/agent/metrics.go: /proc, statfs, systemctl list-units, cgroup v2 of every rc-acct-<user>.slice); there is no node_exporter. This is control-plane-pulled rather than agent-pushed because agents have no route to the store and the op bus already is the authenticated channel. The plugin adds node, and for rc_account_* replaces user with account_id + username.
  • Store: stats.vm.install installs the pinned VictoriaMetrics release (version and SHA-256 compiled into the agent, URL built from them, cached under /var/lib/respirecloud/cache so offline labs can preload the tarball) as rc-victoria-metrics.service (DynamicUser, loopback :8428, -retentionPeriod). The collector self-heals: if the store does not answer it re-runs the install (at most once a minute) and emits stats.store.down once.
  • Series: rc_node_* (cpu seconds by mode, load, memory, swap, disk/inodes per mount, net per device, PSI, uptime), rc_node_up (written by the plugin, 0 when the agent does not answer), rc_service_up{unit}, rc_account_* (cpu usage seconds, memory, pids, io bytes, OOM kills, CPU throttle count, cpu/memory limits).
  • Scoping (AC-stats-28): stats.query, stats.usage.get, stats.top.get pass VictoriaMetrics' extra_filters[] with {account_id="<id>"} (user), {account_id=~"id1|id2"} (reseller: accounts.list as the caller), nothing (admin). VictoriaMetrics forces the filter onto every selector, so no expression can read another account's series. The filter is derived from the authenticated actor; account ids are validated as UUIDs. Node/fleet metrics have no account_id label, so tenants cannot see them.
  • Alerts: threshold rules (expr returning a vector + op + threshold + for_seconds + severity), evaluated every minute (stats.alert.evaluate, internal schedule). pending -> firing after for_seconds -> resolved when the series stops matching. A store outage skips the round (never resolves). Events stats.alert.fired|resolved|acked. Six default rules (node down, service down, disk > 85%, inodes > 90%, memory pressure, load) are seeded on first evaluation.
  • Retention: stats.retention.set stores the days and re-runs the install (rewrites the unit, restarts the store). Default 90 days; VictoriaMetrics OSS has no rollups, so the spec's tiered 15 d / 90 d / 13 months is not implemented.

Actions

stats.overview.get, stats.node.get, stats.usage.get, stats.top.get, stats.query, stats.alert.rule.{list,create,update,delete}, stats.alert.{list,ack,resolve}, stats.retention.{get,set}, stats.reconcile, internal stats.alert.evaluate.

Verified (lab w4-05)

See docs/tasks/w4/w4-05-stats-logs.handover.md.

Not built (spec items left)

Channels (email/webhook/Slack/push), routing/escalation, silences, uptime probes, topology, capacity forecast, site analytics from access logs, reports, saved dashboard layouts, per-key limit meters beyond cgroups, threshold notifications at 80/95/100%.