Deployments
How a release becomes active in an environment. The deployment states, what ends a deployment, and how supersede, rollback, promotion and deploy order work.
A release is an immutable image digest plus the configuration frozen when it was created, numbered per service per environment (r-1, r-2). A deployment is one attempt to make a release active in an environment. It moves through a visible state machine instead of stalling silently, and it always ends in one of five final states.
A deployment has a kind: deploy, rollback or promote.
The state machine
stateDiagram-v2
state "awaiting-approval" as awaiting
[*] --> awaiting: protected environment
[*] --> pending: any other environment
awaiting --> pending: approved
awaiting --> rejected: rejected
awaiting --> cancelled: cancelled or expired
pending --> admitted
admitted --> releasing: release has a release command
admitted --> starting
releasing --> starting
starting --> qualifying
qualifying --> activating
activating --> draining
draining --> healthy
healthy --> [*]
A deployment waits for an approval only in a protected environment. After that it walks eight stages in order, and releasing appears only when the release has a release command. From any stage before healthy it can also end as failed, cancelled or superseded.
| State | Meaning |
|---|---|
awaiting-approval | A deployment in a protected environment, waiting for someone to decide its approval. Nothing reaches the cluster. The console reads Waiting for approval. |
pending | Created and waiting to be admitted, or held behind a dependency that is still deploying. |
admitted | Admitted onto the environment's cluster. |
releasing | The release command runs with the new release's image, before any process starts. The console calls this stage Release command. |
starting | The release's processes are starting. |
qualifying | The agent checks the new processes from inside the cluster before they take traffic. |
activating | The environment's stable entry point switches traffic to the new release. |
draining | The previous release stops taking traffic and shuts down. |
healthy | Final. The release is active. |
failed | Final. The deployment stopped on a blocker. |
rejected | Final. Its approval was rejected. |
cancelled | Final. Someone cancelled it, or its approval expired. |
superseded | Final. A newer deployment of the same service in the same environment took over first. |
What each stage checks
- Qualifying. Every replica of the primary process must stay Ready for the process's stabilization period. For an
httpprocess the agent also sends aGETto the readiness path,/by default, through the release's own service, and the response must be below 500. For aworkerwith a port it opens a TCP connection. A release with no long-running process, such as cron or job only, qualifies at once. A process that isn't ready within 5 minutes fails with the blockerreadiness. - Releasing. A failing or timed-out release command fails the deployment with the blocker
release-command. The previous release keeps serving. - Activating. Traffic switches only when the agent can repoint the stable service. If the Kubernetes API keeps refusing the change for 2 minutes, the deployment fails with
route-activation, and the previous release keeps serving. - Draining. A release that touches a volume and has the downtime acknowledged drains the previous release before
starting, because the old pod must release the volume first.
When a deployment fails
A failed deployment names a blocker: the reason it stopped, such as image-pull, capacity, readiness or release-command. The previous release keeps serving, because traffic only moves at activating. The control plane also sets a blocker on a deployment that's still open and clears it when the cause is gone: configuration when the service can't be compiled, delivery when the cluster's agent is offline, and stalled when a stage ran past its normal time. The deployment states reference lists every blocker with its fix, and Troubleshooting deployments starts from the message you see.
A newer deployment replaces the one in flight
Deploying, redeploying, rolling back or promoting while an earlier deployment of the same service is still in progress replaces it. What happens depends on how far the earlier one got:
- It has not taken traffic yet (
pendingthroughqualifying). It ends assupersededwith a message such asSuperseded by r-7. Its pods are removed, and the response to the new request lists what it ended. You don't need to cancel it first. The release that was serving keeps serving until the new one is healthy. - It is taking over traffic (
activatingordraining). It finishes, because stopping it halfway could leave nothing serving. The new deployment starts right behind it and showsWaiting for r-6 to finish taking over traffic.
A deployment waiting for approval stops nothing until it's approved. Approving it supersedes what's in flight at that moment. Rolling back or promoting also ends the service's queued and running builds in that environment, and a build that finished before the rollback doesn't create its release afterwards.
Rollback, promotion and redeploy
- A rollback makes a previously active release active again. It's a deployment like any other and passes through the same approval, unless the organization's Rollbacks skip approval policy is on. The 3 most recent retired releases of a service stay available for rollback.
- A promotion creates a new release in another environment from an existing release's image, then deploys it there.
- A redeploy creates a new release from the image that's active now. It needs an environment that has had a healthy release.
What waits for what
A deployment waits for the services it depends on. A service depends on another service or database in the project when:
- It references it. A variable that reads
${{ api.KEY }}or${{ db.DATABASE_URL }}, or a database connection. A${{ host(api) }}is only an address and doesn't order deploys. - It says so. The service's Deploy after list, set in its settings in the console or with
[deploy] depends_onin itsnebula.toml. The file wins and shows read-only in the console. A cycle, an unknown slug or the service itself is refused when the list is set.
While a dependency has a deployment in flight in the same environment that started first, the dependent deployment stays pending with the blocker waiting and the message Waiting for db to finish deploying (r-4). Nothing reaches the cluster for it, and its current release keeps serving. It starts as soon as the dependency is healthy. A dependency that isn't being deployed costs nothing, so most deploys never wait.
If the dependency ends failed, cancelled or rejected, the held deployment ends failed: Not deployed: db failed to deploy (r-4). api keeps running r-7. Deploy api again once db is healthy. Nothing that went healthy is rolled back, and services that don't depend on the failed one carry on.
A change set applies in dependency order: databases first, then along the same edges. See Deploy order.
Concurrency
A project can cap how many deployments run at once with Deploy concurrency in the project's settings. It's unlimited by default. A deployment past the limit is refused, not queued, with an error that says the project is already running its maximum number of concurrent deployments. A deployment held behind a dependency still counts toward the limit.
Related
- Change sets and approvals: who decides a production deployment.
- Rollbacks and promotions: the console steps.
- Deployment states: every state and blocker.
Projects and environments
How NebulaCtrl organizes work. Organizations own projects, projects hold environments and services, and each service runs as processes from a source.
Change sets and approvals
How edits to an environment are staged, reviewed and applied together, and how production changes wait for an approval, a deploy freeze or a concurrency limit.