Deployment states
Every state a deployment can be in and every blocker that can name why it is not moving, with what each means and what to do about it.
A deployment is one attempt to make a release active in an environment. It carries a state, and when it is not moving or has failed, a blocker and a message. The console shows them on the service's deployment card, and the API returns them in the state, blocker and message fields of a deployment. For the model behind them, see Deployments.
State machine
flowchart LR
approval(["awaiting-approval"]) -->|approved| pending
approval -->|rejected| rejected
approval -->|"expired after 24 h"| cancelled
approval -->|"a newer deployment"| superseded
pending --> admitted --> releasing --> starting --> qualifying --> activating --> draining --> healthy
admitted -.->|"no release command"| starting
moving["pending to draining"] --> failed
moving --> cancelled
moving --> supersededA deployment starts in awaiting-approval in a protected environment, and in pending everywhere else. The diagram shows the usual route. When a release uses a volume that accepts downtime, the previous release is stopped in a draining stage that runs before starting, because it must give up the volume first.
The seven states from pending to draining are the stages of a deployment that is moving. A deployment reaches one of five final states: healthy, failed, rejected, cancelled or superseded. It never leaves a final state.
States
| State | Console label | Meaning | What to do |
|---|---|---|---|
awaiting-approval | Waiting for approval | A protected environment holds the deployment until a person decides. The approval expires after 24 hours. | Approve or reject it. See Approvals and deploy freezes. |
pending | Deploying: Pending | The deployment is queued. It waits here while it is held behind another deployment, and the agent has not started it yet. | If it stays here, read its message. A held deployment shows Waiting and says what it waits for. |
admitted | Deploying: Admitted | The cluster has enough CPU and memory for the new replicas. The check runs once per deployment. | Nothing. If there is not enough room, the deployment fails with the capacity blocker. |
releasing | Deploying: Release command | The release command is running as a one-off job against the new release. A release without a release command skips this stage. | Wait. If the command fails, the deployment fails with release-command. |
starting | Deploying: Starting | The agent creates the release's workloads and services. | Wait. A container that cannot start fails the deployment with image-pull, startup or scheduling. |
qualifying | Deploying: Qualifying | Every replica of the primary process must stay ready for its stabilization time, and its readiness check must pass through the service. The stage ends after 5 minutes. | Wait. A replica that never stays ready fails the deployment with readiness. |
activating | Deploying: Activating | Traffic switches to the new release in one step. | Wait. If the switch keeps failing for 2 minutes, the deployment fails with route-activation. |
draining | Deploying: Draining | The previous release keeps running for the process's drain time so a rollback stays cheap, then stops. The new release is watched during this window. | Wait. A release that becomes unhealthy here sends traffic back to the previous release and fails with readiness. |
healthy | Healthy | The release is active. It is the release that serves traffic. | Nothing. |
failed | Failed: BLOCKER | The deployment ended without becoming healthy. The previous release keeps serving. | Read the blocker and message. See Blockers. |
rejected | Rejected | A person rejected the approval. | Change what was rejected and deploy again. |
cancelled | Cancelled | A person cancelled the deployment, or its approval expired (message: Approval expired after 24 h). | Deploy again if you still want the change. |
superseded | Superseded, or the message that names the newer release | A newer deployment of the same service replaced this one while it was waiting for approval or still moving. Being replaced is not a failure. | Nothing. Follow the newer deployment. |
An environment that has never deployed a service shows Not deployed.
Blockers
A blocker names the reason a deployment is not progressing. The message beside it says what happened, what state that leaves and what to do. A deployment that fails because of a blocker leaves the previous release serving.
The agent in the cluster reports most blockers, and a deployment that carries one has the state failed. The control plane sets five itself. It sets external-dependency when it ends a deployment as failed. It sets the other four on a deployment that is still moving, and clears each one when its cause is gone.
| Blocker | Set by | Meaning | What to do |
|---|---|---|---|
image-pull | Agent | A node could not pull the release's image: the registry refused the credentials, the image does not exist, or the registry was unreachable. The message says which. | Check the image name and tag and the registry credentials in the service's source settings, then deploy again. See Git providers and registries. |
startup | Agent | A container did not stay running: it crashed (the message gives the exit code), ran out of memory, has no start command in the image, or was rejected before it started. | Open the service's runtime output to read the last log lines. Fix the cause, or raise the memory limit, then deploy again. See Processes. |
scheduling | Agent | No node can run a pod of the release, usually because of placement rules, taints or a lack of room. | Check node labels and taints, or add a compatible node. |
capacity | Agent | The cluster does not have the CPU or memory the new replicas request. The message names the deficit. This is checked in the admitted stage, before anything starts. | Lower the process's resource request, or add a node with more capacity. |
release-command | Agent | The release command exited with an error or ran past its timeout. The message quotes the end of its output. | Fix the command and deploy again. See Release commands and files. |
pvc-attachment | Agent | A volume's claim stayed unbound for more than 2 minutes, or a pod waited that long for its claim. | Check the cluster's storage class and the volume's size. See Volumes. |
readiness | Agent | The replicas did not stay ready within 5 minutes, the readiness check kept failing, or the release became unhealthy during the drain window. | Check the process's readiness path and port, and its logs. See Processes. |
route-activation | Agent | Switching traffic to the release kept failing for 2 minutes. The previous release keeps serving. | Check that the cluster's API server is healthy, then deploy again. |
database-unhealthy | Agent | A database service did not become healthy within 15 minutes of starting (2 hours when it is a point-in-time restore). The message gives CloudNativePG's last report. | Read the message. For a restore, choose a moment at or before the newest activity on the original database. See Databases. |
external-dependency | Control plane | A service this one is set to deploy after ended without becoming healthy, so this deployment never started. Message: Not deployed: <service> failed to deploy (r-<n>). .... | Fix the service it waits for. Then deploy this one again. See Deploy order. |
configuration | Control plane | The control plane cannot compile the service's desired state, for example because of an invalid reference. The message gives the reason. | Fix the variable, reference or setting the message names. The deployment continues by itself. |
delivery | Control plane | The deployment is past its stage's usual time and the cause is between the control plane and the cluster: the environment is not bound to a cluster, the agent is offline, the agent is behind, or it has the state and has not started the release. | Read the message. Bind the environment to a cluster, or bring the agent back. The deployment continues by itself. See Clusters and agents. |
stalled | Control plane | A stage ran past its usual time and the agent reports no specific failure. | Look at the release's pods and the agent's log, as the message says. |
waiting | Control plane | The deployment is held in pending. It is not a fault. Messages: Waiting for <service> to finish deploying (r-<n>) (deploy order) or Waiting for r-<n> to finish taking over traffic (an older release is still moving traffic). | Wait. The deployment starts when the one it waits for finishes. |
source | None | Accepted by the API. This version never reports it. | |
policy | None | Accepted by the API. This version never reports it. |
configuration, delivery, stalled and waiting do not make a deployment failed by themselves.
How long a stage may take
The control plane starts to explain a stuck stage once it has run longer than its usual time. It checks every 15 seconds. It calls an agent offline when no heartbeat has arrived for 1 minute.
| Stage | Usual time before the control plane explains the delay |
|---|---|
pending | 1 minute |
admitted | 3 minutes |
releasing | 30 minutes |
starting | 10 minutes |
qualifying | 7 minutes |
activating | 4 minutes |
draining | 20 minutes |
The agent has shorter limits of its own, and its specific failure comes first when it has one: qualifying 5 minutes, activating 2 minutes, an unbound volume claim 2 minutes. A database service gets 15 minutes to start. awaiting-approval has no limit but the approval's 24 hours, because it waits for a person.
Related
- Troubleshoot deployments, by the message you see.
- Roll back and promote.
- Limits.
Variable expressions
The syntax of ${{ }} expressions in variable values, what each form produces, when it is evaluated, where it is allowed, and the messages a bad one produces.
Limits
Every hard limit in NebulaCtrl 0.39.0, with its exact value and what is refused or clamped when you reach it. Covers API requests, streams, rate limits, sizes and counts in configuration files, retention periods and timeouts.