Get alerted when a service crosses a threshold
Turn on, tune and acknowledge the alert rules of a service. Covers the four rules (volume usage, memory, restart loops and CPU), when an alert opens and resolves, the canvas pill and Alerts panel, and who is notified.
A full volume or a process that keeps hitting its memory limit takes a service down without a deploy to blame. Alert rules watch the numbers the cluster already reports and tell you before that happens.
Before you begin
- Reading alerts and rules needs the
viewerrole. Changing a rule and acknowledging an alert needdeployeror higher. - The service has a healthy deployment. A rule measures what a running release reports; a service that is stopped or never started has nothing to measure.
- Memory and CPU rules measure usage against a limit. Set a memory limit, and a CPU limit for the CPU rule, on every process of the service. Without one, the rule shows Not evaluated and says so.
- To be notified outside the console, add a channel under Settings > Integrations. See Get notified about events.
The four rules
Every service has the same four rules. Three are on from the start and the CPU rule is off.
| Rule | Opens when | Default | Resolves when |
|---|---|---|---|
| Volume usage | Any volume of the service is fuller than the threshold. | On, above 85% | The volume is below the threshold minus 5 points. |
| Memory usage | Memory stays above the threshold, as a share of the memory limit, for 5 minutes. | On, above 90% | Memory is below the threshold minus 5 points. |
| Restart loop | Containers restart at least the threshold number of times within 10 minutes. | On, 3 restarts | No container restarted for 10 minutes. |
| CPU usage | CPU stays above the threshold, as a share of the CPU limit, for 10 minutes. | Off, above 90% | CPU is below the threshold minus 5 points. |
A percentage threshold takes a whole number from 10 to 99. A restart loop takes 2 to 100: one restart is a restart, not a loop.
The 5-point gap keeps a reading that hovers at the threshold from opening and closing an alert every minute. Between the threshold and the gap, an open alert stays open and keeps its latest value.
A usage alert opens only when every one of the last 5 (or 10) minutes is above the threshold. A minute with no sample counts against it: a gap in the data is never read as a breach, and it is never read as a recovery either.
Alerts are critical when a volume is 95% full or more, memory is 97% of its limit or more, or a service is in a restart loop. Other alerts are warnings.
Change a rule
Select the service on the canvas, then select Settings. Scroll to Alerts.
Use the switch on a rule's row to turn it on or off. Turning a rule off closes its open alerts.
Type a new number in the rule's box and press Enter or move out of the box. A number outside the allowed range is refused with the range.
Select Reset to return a changed rule to its default. The button appears only on rules you changed.
Each row ends with what the last check measured, such as Now 62%, or why it measured nothing, such as Not evaluated: the running release does not set a memory limit on every process…. A rule that is on and has not been checked yet reads Not evaluated yet.
A rule belongs to the service and applies in every environment it runs in. Changing it needs no deployment and is recorded in Activity.
What "not evaluated" means
Nebula never fills in a number it did not measure. A rule is not evaluated, and opens or resolves nothing, when:
- the cluster has not reported the service's pods, or its newest report is older than 3 minutes;
- memory or CPU has no limit to be measured against;
- a volume shares the node's disk, as the default K3s storage class
local-pathdoes, so it has no usage of its own and its volume row reads Usage not reported; - a volume has not reported for 10 minutes, or the service has no volumes;
- the control plane restarted and has not yet received a report from the cluster.
An open alert whose rule cannot be measured stays open. After 30 minutes without a measurement it closes as not reported instead of recovered, and says so.
Restarts are counted from the control plane's own checks, one a minute. A restart before the first check of a pod is not counted, so a restart loop is reported after the control plane has watched it, not retroactively.
See alerts on the canvas
Three places show an environment's alerts, and all read the same list, refreshed every 30 seconds.
- The card pill. A service with open alerts carries a pill on the left of its card's top edge, reading 1 alert or N alerts. It is red when the worst alert is critical and amber otherwise. Hovering shows "Worst critical" or "Worst warning". Selecting the pill opens the service's Settings tab at Alerts. An acknowledged alert no longer counts toward the pill. A card shows at most three pills; the rest fold into a +N pill whose hover text lists them. The alert pill comes first, then Update, Skipped and the staged-change pill (New, Edited or Removed), so it is the last to fold.
- The attention chip. While any alert in the environment is open, a chip reading N alerts sits in the row of status chips at the top of the canvas, with a dot that is red or amber by the worst severity. Select it to open the Alerts panel. It disappears when every alert is acknowledged or resolved.
- The Alerts panel. The panel lists every open and acknowledged alert of the environment, critical first and the newest first within a severity. Its address carries
?panel=alerts, so a copied link reopens it. Each row shows what is wrong, then the service name, Opened with the time, and Acknowledged by your teammate's name once someone has acknowledged it. Select the service name to open that service's Settings tab at Alerts. With nothing past a threshold, the panel says so.
You can also press Ctrl K (⌘K on a Mac) and run Open alerts. The command is listed while the environment has an open or acknowledged alert. See Run commands from the command palette.
Acknowledge an alert
An alert has three states: open, acknowledged and resolved. Acknowledging tells your team someone has seen it. The alert stays listed and measured, and it resolves like any other.
Open the Alerts panel from the attention chip, or open the service's Settings tab and find the alert under Alerts. Open alerts are listed above the rules.
Select Acknowledge on the alert. The button reads Acknowledging… while the request runs. The row then reads Acknowledged by your name, and the card pill and the attention chip drop it from their count.
If the rule recovers and is breached again later, that is a new alert with its own notification.
Who is notified
When an alert opens and when it resolves, Nebula writes the inbox of the organization's admins and owners. In a production environment it also posts to Slack channels subscribed to Alerts and opens a PagerDuty incident, which the resolution closes. Alerts in other environments reach the inbox only. A re-run of the check never notifies twice for the same alert.
A resolved alert says recovered. An alert closed because its rule was switched off, or because it was not reported for 30 minutes, says that instead and does not claim a recovery.
API
| Operation | Call |
|---|---|
listEnvironmentAlerts | GET /api/v1/environments/ENVIRONMENT_ID/alerts, with state=resolved for the alerts that resolved in the last 7 days |
acknowledgeAlert | POST /api/v1/alerts/ALERT_ID/acknowledge |
listServiceAlertRules | GET /api/v1/services/SERVICE_ID/alert-rules?environmentId=ENVIRONMENT_ID |
setServiceAlertRule | PUT /api/v1/services/SERVICE_ID/alert-rules/RULE with {"enabled": true, "threshold": 80} |
resetServiceAlertRule | DELETE /api/v1/services/SERVICE_ID/alert-rules/RULE |
RULE is volume_usage, memory_usage, restart_loop or cpu_usage. An alert resolved at most 30 days ago stays readable; older ones are deleted.
curl "BASE_URL/api/v1/environments/ENVIRONMENT_ID/alerts" \
-H "Authorization: Bearer NEBULA_TOKEN"Replace BASE_URL, ENVIRONMENT_ID and NEBULA_TOKEN with your installation's address, the environment's id and an API token.
{
"items": [
{
"id": "6d1c0e7a-2f43-4e0b-9a5d-3c8b7f2e1a90",
"serviceName": "api",
"rule": "volume_usage",
"subject": "data",
"state": "open",
"severity": "warning",
"value": 91.4,
"threshold": 85,
"unit": "percent",
"summary": "Volume data of api is 91% full"
}
]
}Verify
Fill a test volume past its threshold, or set a low threshold such as 10 on a volume that holds data. The alert appears under Alerts and in the inbox of an admin within a few minutes of the cluster reporting the new usage. Raise the threshold or delete data: once the volume is 5 points below the threshold, the alert resolves and the inbox gets a second message.
Next steps
- Get notified about events covers the inbox and channels.
- Read logs and metrics shows the numbers the rules measure.
- Add and back up a volume covers growing a volume that is filling up.
Open a shell and browse files in an instance
Open a terminal in a running instance from the browser, and list, download and upload its files. Covers who may do it, protected environments, the debug container fallback, limits and what is recorded.
Run commands from the command palette
Jump to a service, run its deploy, restart and copy commands, find a setting by name, compare environments, open alerts, generate a domain or open a shell, and deploy a repository by pasting its URL, all from the keyboard.