Skip to content
NebulaCtrldocs
Guides

Update a cluster agent

Update the agent that runs in a cluster from the console, roll it back, and re-apply the install manifest when an older cluster needs new permissions.

A control plane update does not touch your clusters. Each cluster's agent updates on its own, from the console, when you start it. Workloads keep running during the update.

Before you begin

  • You need the admin or owner role, and a session or a token without project limits.
  • The cluster must be connected.
  • No deployment or build may be running on the cluster. Deployments that wait for approval do not count.
  • The control plane offers an agent image newer than the one the cluster runs. The default is ghcr.io/nebulactrl/nebula-agent at the control plane's own version. The --agent-image server flag overrides it. A tagged image must be 0.34.1 or newer, and not newer than the control plane.

Update the agent

  1. Open Clusters and select the cluster.
  2. In the header, select Update agent. The button appears when an update is available. You can also use Update now under Settings > Agent. In the cluster list, Update agent in the context menu opens the cluster's overview, where the header button is.
  3. Watch Settings > Agent. The Version row reads Updating to TAG… until the new agent reports in.

Settings > Updates lists every cluster under cluster agents with its node count and agent version. A cluster whose agent runs a different release from the control plane carries an update badge. That comparison is made only when the control plane is a release build and the agent reports a release version. After a successful control plane update, the same page shows "Cluster agents are updated separately. These still run an older version." and a row for each cluster that trails, with an Update agent button. The rows stay while any agent trails.

To start the update from the API:

curl -s -X POST \
  -H "Authorization: Bearer $NEBULA_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{}' \
  https://nebula.example.com/api/v1/clusters/CLUSTER_ID/agent/update

An empty body uses the control plane's agent image. Send {"image": "IMAGE"} to choose one. NEBULA_TOKEN is an API token and CLUSTER_ID the cluster's id.

What the update does

  1. The control plane records the target image and sends the update command.
  2. A second component, the access broker, patches the agent's Deployment to the new image. The agent cannot patch its own Deployment. The broker can change only that image, and only to ghcr.io/nebulactrl/nebula-agent.
  3. Kubernetes rolls the Deployment out one extra pod at a time with no pod unavailable. The old agent keeps running until the new one can reach the control plane.
  4. The update is Healthy when the new agent reports the target image. The pod's image is what counts, not the Deployment's.
  5. If the new agent has not reported within the deadline, the update is closed as Rolled back with the new agent did not report in within 10 minutes, and the control plane starts a rollback to the previous image by itself. The deadline is the --agent-update-deadline server flag and defaults to 10 minutes. A rollback that also times out ends Failed.

An update can also fail to start. The console shows The agent update did not start or Could not start the agent update, and the current agent keeps running. The usual refusals:

MessageWhat to do
an agent update is already running on this clusterWait for it to finish.
the cluster must be connected before its agent can be updatedBring the agent online first.
wait for the deployments and builds on this cluster to finish, then update the agentWait, or cancel the build or deployment.
this cluster already runs IMAGENothing to update.

Roll back to an earlier agent

  1. Open Settings > Agent on the cluster.
  2. In Agent history, find the attempt you want to undo. The table lists Attempt, From → to, Kind, Status and Started.
  3. Select Roll back, then Roll back again in the confirmation.

What an update changes in the cluster

  • The agent's image. Nothing else in the Deployment changes.
  • Four custom resource definitions. The agent applies builds.nebula.dev, domains.nebula.dev, releases.nebula.dev and workloads.nebula.dev each time it starts, so an update brings newer fields with it. Its cluster role lets it patch and update only those four definitions.
  • No RBAC. An update never changes roles, role bindings or namespaces. The broker gives the agent access to each environment's namespace. Those permissions come from the install manifest.

Re-apply the manifest on an older cluster

An agent installed before it gained permission to update the definitions cannot apply them. It logs one warning and keeps running. The warning begins:

cannot update the nebula.dev CustomResourceDefinitions: the agent's ClusterRole predates this version and may not update the nebula.dev CustomResourceDefinitions, so the cluster keeps the ones it has and features that need newer fields (such as file mounts) stay unavailable

Builds and deployments that need a newer field are then refused with a message that tells you to re-apply the agent manifest. One example begins this cluster's build policy, Build CRD or agent predates Railpack builds. To re-apply it:

  1. Open the cluster's Settings > Agent. Beside Install credentials, select Regenerate, then confirm. The previous token stops working, and the cluster reads Awaiting agent until an agent reports in with the new token.
  2. Copy the Install command and the One-time token. The token is shown once and expires in one hour.
  3. Run the command as root on a server node of that cluster. Paste the token at the silent NebulaCtrl token: prompt.

The script is safe to run again with a newly issued token.

Verify

  • The Version row shows the new version and latest.
  • Agent history lists the attempt as Healthy.
  • The cluster is connected and its workloads are unchanged.

Next steps

On this page