Skip to content
NebulaCtrldocs

Deployments

Find the message a refused, waiting, failed or stalled deployment shows, and fix its cause: approvals, deploy order, freezes, capacity, image pulls, startup and readiness.

A deployment is refused when you start it, waits for a person or another deployment, or fails or stalls on the cluster. A failed deployment ends with a blocker and the previous release keeps serving. Find the message in the alert titled Blocked on in the deployment's details, in the error toast did not deploy, or in the API detail field.

Approvals and waiting

Waiting for approval, but no one else in this organization can approve it: your policy forbids approving your own deploys and you are its only admin. Invite another admin, or turn off "Approvers can't approve their own requests" in Settings > Security & policy.

Cause: The deployment's message shows this in a protected environment when the organization forbids self-approval and you are the only member who can decide approvals. The deployment stays at Waiting for approval.

Fix: Add another admin or owner, or turn off Approvers can't approve their own requests under Change policy in Settings > Security & policy.

Verify: The other admin sees the request under Approvals and can select Approve. See Approvals and deploy freezes.

you requested approval <id>, and your organization's policy forbids approving your own request; it is still pending — ask another admin or owner to decide it

Cause: The API answers 403 with this detail, and the console shows the toast The approval was not recorded, when you approve a request you made while self-approval is forbidden.

Fix: Ask another admin or owner to approve it.

Verify: After their decision the deployment leaves Waiting for approval and moves to Deploying: Pending.

Approval expired after 24 h

Cause: The deployment's message shows this when nobody decided its approval within 24 hours. The deployment ends as Cancelled. Approving after the deadline is refused with approval <id> expired at <time> and is still pending until expiry processing rejects it; request the production change again for a new approval.

Fix: Deploy again to create a new approval.

Verify: The new deployment shows Waiting for approval, and the approval's detail states when it expires.

Superseded by r-<n>

Cause: The deployment's state label shows this message. A newer deployment of the same service replaced this one while it waited for approval or was still moving. Being replaced is not a failure.

Fix: None. Follow the newer deployment.

Verify: The row reads Superseded by followed by the newer release, and that deployment continues.

Waiting for <service> to finish deploying (r-<n>)

Cause: The deployment's details show this when the service depends on the named service, through a variable reference, a link or its Deploy after list, and that service has a deployment in flight. Nothing is sent to the cluster yet and the current release keeps running.

Fix: Wait. To stop the wait, cancel the dependency's deployment with Cancel deployment, or remove it from Deploy after. See Deploy order.

Verify: When the dependency is healthy the message disappears and the deployment moves through its stages.

Not deployed: <service> failed to deploy (r-<n>). <other> keeps running r-<m>. Deploy <other> again once <service> is healthy.

Cause: The deployment's message shows this when a dependency it waited for failed. The wording is was cancelled or was rejected when the dependency ended that way. The blocker is external-dependency. Nothing was sent to the cluster.

Fix: Fix and deploy the dependency first, then deploy this service again.

Verify: The dependency shows Healthy and the new deployment carries no Not deployed message.

Waiting for r-<n> to finish taking over traffic

Cause: The deployment's details show this when an older deployment of the same service is activating or draining. The newer deployment waits in pending. This is not a fault.

Fix: None. The deployment starts when the older one finishes.

Verify: The message clears and the stage changes from Deploying: Pending.

Refused when you start a deployment

<n> of <limit> already running; wait for one to finish or raise the limit in project settings

Cause: The API answers 409 with this detail, and the toast did not deploy shows it, when the project has as many deployments in flight as its Deploy concurrency allows. Deployments of other services count. This service's own earlier deployments do not, because the new one replaces them. The deployment is refused, not queued.

Fix: Wait for a deployment to finish, or raise Deploy concurrency in Project settings.

Verify: Deploy again; the deployment starts at Deploying: Pending.

production is in a deploy freeze until <day> <time> <zone>; rollbacks are still allowed, or change the freeze windows in the organization's policy

Cause: The API answers 403 with this detail when a deployment, an approval or a change set for a production environment arrives during a freeze window.

Fix: Wait for the window to end, roll back instead, or edit the windows under Weekend deploy freeze in Change policy.

Verify: Outside the window the deployment is accepted.

runAsUser must be set on process "<process>": <image> runs as root by default and the agent enforces runAsNonRoot, which kubelet would otherwise refuse at deploy time; set runAsUser: 0 to allow it, or use an image with a non-root default user

Cause: The API answers 422 with this detail when the image runs as root and no process sets Run as user. For an image with a named user the message says runs as the named user "<user>", which Kubernetes cannot verify as non-root, so nothing was deployed.

Fix: Set a numeric USER in the Dockerfile, or set Run as user on the process. 0 allows root.

Verify: The deployment passes Deploying: Admitted.

process "<process>" mounts volume "<volume>": activating this release will stop the previous release first; retry with downtimeAccepted to proceed

Cause: The API answers 409 with this detail when a release touches a volume. The console opens a confirmation titled Deploy with the service name and with downtime?.

Fix: Select Deploy with downtime, or send "downtimeAccepted": true over the API.

Verify: The Draining stage detail reads stopping previous release <id> before starting, and the deployment continues.

<service> is staged in <person>'s change set for <environment> and hasn't been committed: deploy it by committing that change set (Review → Deploy), or remove it from the change set first.

Cause: The API answers 409 with this detail when you deploy or build a service that exists only in an uncommitted change set.

Fix: Select Review on the change set, then Deploy, or remove the service from it.

Verify: The service deploys as part of the change set.

this cluster's agent can't mount files yet: re-run the installer to update the agent and its Release CRD, then deploy again

Cause: The API answers 409 with this detail when the service has file mounts and the cluster's agent does not report that capability. A release command gets this cluster's agent can't run release commands yet: re-run the installer to update its Release CRD, then deploy again.

Fix: Update the agent. If the cluster is old, re-apply its manifest. See Agent updates.

Verify: Deploy again; the refusal does not return.

environment has no active release to re-release

Cause: The API answers 409 with this detail when you select Redeploy in an environment where no release ever reached Healthy.

Fix: Deploy the service once with a new release.

Verify: The deployment reaches Healthy, and Redeploy works afterward.

Failed on the cluster

Needs <amount> more than the cluster can allocate

Cause: The alert Blocked on capacity shows this. The agent checks once, in the admitted stage, against the ready, schedulable nodes. The message names the missing CPU or memory.

Fix: Lower the process's CPU or memory request, or add a node. The alert links to Open settings and add a node. See Connect a cluster.

Verify: Deploy again; the deployment passes Deploying: Admitted.

pod <name> cannot be scheduled: <reason>

Cause: The alert Blocked on scheduling shows this when the scheduler reports a pod unschedulable. A pod that waits only for its volume claim is not counted here.

Fix: Check node labels and taints, or add a compatible node.

Verify: The pod starts and the deployment reaches Deploying: Qualifying.

The registry refused to give <service> its image <image>: the credentials are missing or wrong, or the image does not exist

Cause: The alert Blocked on image-pull shows this the first time a node reports a failed pull. Every message of this kind ends with Any previous release keeps serving. and the kubelet's own reason.

Fix: Check the Registry credential and the image name in the service's Source settings. See Git providers and registries.

Verify: Deploy again; the deployment passes Deploying: Starting.

<service>'s image <image> was not found in the registry

Cause: The alert Blocked on image-pull shows this when the registry answers that the image or tag does not exist.

Fix: Correct the image name or tag in Source, or rebuild the image.

Verify: Deploy again; the pull succeeds.

The node could not pull <service>'s image <image>: the registry was unreachable or answered with an error

Cause: The alert Blocked on image-pull shows this for any other pull failure, usually a temporary one.

Fix: Deploy again in a few minutes. If it keeps failing, check the registry address.

Verify: The deployment passes Deploying: Starting.

<service>'s image runs as root, and NebulaCtrl runs every container as a non-root user

Cause: The alert Blocked on startup shows this when the kubelet refuses a root image that the deploy-time check could not inspect. A named user gives sets its user to <user>, and Kubernetes cannot verify that a name is not root.

Fix: Add a numeric USER to the Dockerfile, or set Run as user on the process. 0 allows root.

Verify: Deploy again; the pod starts.

<service>'s container "<container>" exited with code 127: its start command was not found in the image

Cause: The alert Blocked on startup shows this when the container keeps crashing. Exit code 126 reads its start command exists but cannot be executed.

Fix: Check the process's command and args, and the file's permissions for 126.

Verify: Deploy again; the deployment passes Deploying: Qualifying.

<service>'s container "<container>" was killed for running out of memory

Cause: The alert Blocked on startup shows this with (limit <n> MiB), or on the node when the process has no limit.

Fix: Raise the process's Default memory limit, or reduce the app's memory use.

Verify: Deploy again; the container stays running.

<service>'s container "<container>" keeps crashing on start (last exit code <n>)

Cause: The alert Blocked on startup shows this for any other crash loop. Without a previous exit the message reads keeps crashing right after it starts.

Fix: Read Runtime output in the deployment's logs, fix the cause, and deploy again.

Verify: The new deployment passes Deploying: Qualifying.

<service>'s container needs the key <key> from <Secret or ConfigMap> <name>, which does not exist

Cause: The alert Blocked on startup shows this when a variable or linked resource points at something deleted. A missing object reads references the <kind> "<name>", which does not exist.

Fix: Check the service's variables and linked resources, then deploy again.

Verify: The container starts.

readiness probe never returned a healthy status within 5 minutes; last response was <status>

Cause: The alert Blocked on readiness shows this. The agent requests the process's Readiness path (/ when unset) through the release's own service, and a status of 500 or above is unhealthy. Other forms are readiness probe never succeeded within 5 minutes: <error>, TCP readiness dial never succeeded within 5 minutes: <error> for a worker with a port, and replicas never stayed ready continuously within 5 minutes when too few replicas stayed ready for the stabilization time.

Fix: Make the path answer below 500, and check that the process port matches the port the app listens on. Read the logs for restarts. See Processes.

Verify: Deploy again; the deployment passes Deploying: Qualifying.

release has only <n> of <m> replicas ready during the drain window

Cause: The alert Blocked on readiness shows this after traffic switched. A container that restarted 3 times reads a container has restarted <n> times during the drain window. The agent returns traffic to the previous release.

Fix: Read the logs, fix the cause, and deploy again.

Verify: The details read r-<n> is still serving; nothing changed for users.

release command exited <code>

Cause: The alert Blocked on release-command shows this followed by the last 20 lines of output. Secret values of 6 or more characters show as [redacted]. A timeout reads release command timed out after <duration>.

Fix: Fix the command, or raise Timeout (seconds) under Release command, then deploy again. See Release commands and files.

Verify: The stage Release command completes.

persistentvolumeclaim <name> has stayed <phase> for over 2m0s

Cause: The alert Blocked on pvc-attachment shows this when a volume's claim stays unbound. A pod waiting that long reads pod <name> has waited over 2m0s for its volume claim: <reason>.

Fix: Check the cluster's storage class and the volume size. See Volumes.

Verify: Deploy again; the claim binds.

switching traffic to <release> did not complete within 2m0s (<error>); ...

Cause: The alert Blocked on route-activation shows this when the cluster's API server keeps refusing the traffic switch. The previous release keeps serving.

Fix: Check that the cluster's API server is healthy, then deploy again.

Verify: The deployment passes Deploying: Activating.

Held or stalled by the control plane

Release <n> of <service> has been in the <stage> stage for <time>, longer than the usual <time>

Cause: The label Deploying with the stage name stays and the alert Blocked on stalled shows this when a stage runs past its usual time and the agent reports no failure. The message quotes the agent's last report. The deployment can still finish.

Fix: Run kubectl get pods --all-namespaces and kubectl -n nebula-system logs deploy/nebula-agent, or cancel with Cancel deployment.

Verify: The message clears when the stage moves on.

Release <n> of <service> has not started after <time>

Cause: The alert Blocked on delivery shows this when a pending deployment waits longer than 1 minute. A clause after the colon names the cause, as the next entries show.

Fix: Read the clause and follow its entry. The deployment continues by itself once the cause is gone.

Verify: The blocker clears.

its environment is not bound to a cluster, so nothing can run it

Cause: The alert Blocked on delivery shows this when the environment has no cluster.

Fix: Bind a cluster in the environment's settings with Bind cluster.

Verify: The deployment starts.

the agent of cluster <name> is offline (<age>)

Cause: The alert Blocked on delivery shows this when the agent sent no heartbeat for 1 minute, or never did. The previous release keeps serving.

Fix: Run kubectl -n nebula-system get pods and check that the cluster reaches the control plane. See Clusters and agents.

Verify: The cluster shows as connected and the deployment continues.

cluster <name>'s agent has applied revision <a> of <b>, so it has not yet received or applied this change

Cause: The alert Blocked on delivery shows this. The control plane sends the state again once a minute. If the agent cannot apply it, an organization admin sees why under Clusters. When the agent has the state but starts nothing, the message reads cluster <name>'s agent has the current desired state (revision <n>) but has not reported this release starting.

Fix: Wait one minute, then check the agent's log with kubectl -n nebula-system logs deploy/nebula-agent.

Verify: The applied revision reaches <b> and the stage moves past Deploying: Pending.

<service> can't be deployed: <reason>

Cause: The alert Blocked on configuration shows this when the control plane cannot compile the service, for example a variable that points at a deleted service. Running releases keep serving. Environment <name> can't be deployed means the whole environment.

Fix: Fix what the reason names, such as a variable, then deploy again.

Verify: The message clears when the service compiles.

Cluster <name> can't be sent new changes: <reason>

Cause: The alert Blocked on configuration shows this when one project's configuration blocks the whole cluster's state. A deployment of another project sees Another project on <cluster> can't be deployed right now, so this deployment is waiting, and one past its stage time sees the control plane cannot build cluster <name>'s desired state right now, so nothing new is delivered to it. Those readers cannot see the reason.

Fix: Fix the reason. An organization admin reads it under Clusters. Delivery resumes within a minute.

Verify: The message clears on its own.

On this page

Approvals and waitingWaiting for approval, but no one else in this organization can approve it: your policy forbids approving your own deploys and you are its only admin. Invite another admin, or turn off "Approvers can't approve their own requests" in Settings > Security & policy.you requested approval <id>, and your organization's policy forbids approving your own request; it is still pending — ask another admin or owner to decide itApproval expired after 24 hSuperseded by r-<n>Waiting for <service> to finish deploying (r-<n>)Not deployed: <service> failed to deploy (r-<n>). <other> keeps running r-<m>. Deploy <other> again once <service> is healthy.Waiting for r-<n> to finish taking over trafficRefused when you start a deployment<n> of <limit> already running; wait for one to finish or raise the limit in project settingsproduction is in a deploy freeze until <day> <time> <zone>; rollbacks are still allowed, or change the freeze windows in the organization's policyrunAsUser must be set on process "<process>": <image> runs as root by default and the agent enforces runAsNonRoot, which kubelet would otherwise refuse at deploy time; set runAsUser: 0 to allow it, or use an image with a non-root default userprocess "<process>" mounts volume "<volume>": activating this release will stop the previous release first; retry with downtimeAccepted to proceed<service> is staged in <person>'s change set for <environment> and hasn't been committed: deploy it by committing that change set (Review → Deploy), or remove it from the change set first.this cluster's agent can't mount files yet: re-run the installer to update the agent and its Release CRD, then deploy againenvironment has no active release to re-releaseFailed on the clusterNeeds <amount> more than the cluster can allocatepod <name> cannot be scheduled: <reason>The registry refused to give <service> its image <image>: the credentials are missing or wrong, or the image does not exist<service>'s image <image> was not found in the registryThe node could not pull <service>'s image <image>: the registry was unreachable or answered with an error<service>'s image runs as root, and NebulaCtrl runs every container as a non-root user<service>'s container "<container>" exited with code 127: its start command was not found in the image<service>'s container "<container>" was killed for running out of memory<service>'s container "<container>" keeps crashing on start (last exit code <n>)<service>'s container needs the key <key> from <Secret or ConfigMap> <name>, which does not existreadiness probe never returned a healthy status within 5 minutes; last response was <status>release has only <n> of <m> replicas ready during the drain windowrelease command exited <code>persistentvolumeclaim <name> has stayed <phase> for over 2m0sswitching traffic to <release> did not complete within 2m0s (<error>); ...Held or stalled by the control planeRelease <n> of <service> has been in the <stage> stage for <time>, longer than the usual <time>Release <n> of <service> has not started after <time>its environment is not bound to a cluster, so nothing can run itthe agent of cluster <name> is offline (<age>)cluster <name>'s agent has applied revision <a> of <b>, so it has not yet received or applied this change<service> can't be deployed: <reason>Cluster <name> can't be sent new changes: <reason>