Nothing happens
without a trace.
Telemetry streamed from the clusters themselves, and the durable trail behind it: metrics history, the activity audit, shipped logs and archived indices.
Telemetry for the whole estate, on one screen, at the resolution you ask for.
Live now,
and why it matters.
Instance distribution, resource metrics, the heaviest consumers and the state of every node, refreshed from the clusters themselves rather than from a cache.
The running-to-stopped ratio, kept separate for containers and virtual machines rather than lumped into one number.
CPU, memory, storage and network I/O, each with minimum, average and maximum overlaid on the series.
Five-minute intervals through to twenty-four hours, with zoom in and out on the chart itself.
Sortable highest to lowest, so the instance eating a node is the first row rather than something you hunt for.
Capacity per pool across distributed backends, so a pool approaching its limit is visible before it fills.
Health, connectivity and resource allocation for every registered node, as cards rather than a buried table.
What is happening right now, and what is wrong.
Every asynchronous cluster task still running, with its precise timestamp and execution state, from an open console session to an instance being provisioned.
Critical, major, warning and informational events in one categorised list, surfacing an offline node, a missing network controller or a failing storage mount.
The alarm statistics matrix on the monitoring page synchronises with this list, so a severity count and the warnings themselves never disagree.
A warning raised by a cluster links straight to the resource it names, rather than leaving you to work out which instance it meant.
From a live metric to an archived index, in one path.
A number on screen now,
and a record of it later.
CPU, memory, disk and network streamed live, straight from the cluster.
A time-series chart you can zoom and pan, the instances consuming the most resources, node connection state, and alerts broken down by severity.
A background collector writes historical metrics to PostgreSQL, kept for a retention window you configure, and pruned by a periodic task.
Every action recorded with the actor and the result, alongside session tracking for every terminal and console session.
A sidecar tails container output and ships it to OpenSearch, with index lifecycle management applying age-based retention policies.
Create, list, restore, delete and verify archived indices from the dashboard. Admin only, and the endpoints are not mounted unless the feature is enabled.
Access logs, audit logs and an overview, as dashboards you query directly rather than a log file you grep.