Skip to content
NebulaCtrldocs

Clusters and agents

Find the message from a failed cluster install, a disconnected or degraded cluster, or a rolled-back agent update, and fix its cause.

Use this page when connecting a cluster stops, a cluster reads Disconnected or Degraded, or an agent update does not finish. Each entry starts with the message you see.

install token is invalid, expired or already used

Cause: The install dialog shows this on the failed step, and the install command prints it, when the one-time token is spent or older than one hour. Variants name a join token, enrollment token or registration token. A stale script link answers this link has expired or was already used; create a new one from the cluster page.

Fix: Open the cluster, select Settings, and select Regenerate beside Install credentials. Run the new Install command and paste the new One-time token at the NebulaCtrl token: prompt. The prompt does not echo; press Enter after you paste. Do not create a second cluster.

Verify: The cluster reads Awaiting agent, then Healthy once the agent reports in.

port <port> is held by <holder>; stop and disable it, then run this command again

Cause: The Checking the server step fails on the first install of a server when another process listens on port 80, 443 or 6443.

Fix: Stop and disable the service the message names, then run the command again.

Verify: ss -ltn 'sport = :80' prints no listener, and Checking the server completes.

firewalld is active; NebulaCtrl manages ufw

Cause: The Checking the server step refuses a host where firewalld runs, because the installer configures ufw.

Fix: Stop and disable firewalld, then run the command again.

Verify: systemctl is-active firewalld prints inactive.

NebulaCtrl installs on Ubuntu 22.04 or newer and Debian 12 or newer; this is <name>

Cause: The Checking the server step refuses an older operating system. A host that is not x86_64 or aarch64 is refused with NebulaCtrl supports x86_64 and aarch64 servers; this is <architecture>.

Fix: Run the install command on a supported host.

Verify: Checking the server completes and Installing packages starts.

this server needs at least <count> CPUs to run as a <role> node; it has <count>

Cause: The Checking the server step enforces a floor by role. A server needs 2 CPUs and about 1.9 GB of memory; an agent node needs 1 CPU and about 480 MB. A memory shortfall reads this server needs at least <kB> kB of memory to run as a <role> node; it has <kB> kB.

Fix: Resize the machine or use a larger one.

Verify: nproc and grep MemTotal /proc/meminfo meet the floor, and the step completes.

this server already runs K3s for another cluster; run /usr/local/bin/k3s-uninstall.sh first if you mean to reuse it

Cause: The Checking the server step finds a K3s configuration that belongs to another cluster, or a K3s binary with no configuration.

Fix: To reuse the machine, run /usr/local/bin/k3s-uninstall.sh on it, then run the install command again. Otherwise use another machine.

Verify: Checking the server completes.

<command> exited <status>

Cause: The install dialog shows this on the failed step when a command fails without a more specific message. Steps run in this order: preflight, packages, sysctl, mesh, join, k3s, cnpg (servers only), firewall, ssh, fail2ban, updates and agent. The high-availability script reports etcd.

Fix: Read /var/log/nebula-install.log on the machine; the installer also sends its last 40 lines to the control plane. Remove the cause of the named command, select Regenerate beside Install credentials, and run the command again. The installer is safe to run again.

Verify: The failed step completes on the next run.

could not join the WireGuard mesh: <reason>

Cause: The Joining the WireGuard mesh step prints the control plane's answer. Reasons include registration token is invalid, expired or was minted for a different purpose; request a fresh install command, these install credentials already admitted a machine to the mesh and the control plane at <url> could not be reached.

Fix: For a token reason, regenerate the install credentials. For the last reason, open the control plane's address from the machine and fix DNS, TLS or the firewall in between. See networking.

Verify: Joining the WireGuard mesh completes.

a machine named <name> is already on this cluster's mesh

Cause: The mesh admits one machine per hostname, and you are reinstalling a machine that is still registered.

Fix: Select Remove from mesh on the node, or change the hostname, then run a fresh join command. Removal retires the machine's mesh key for good.

Verify: Joining the WireGuard mesh completes and the node appears in the cluster.

cannot reach the cluster's server at <host>:6443 over the WireGuard mesh after 120 seconds

Cause: A joining node waited two minutes for the cluster's server during the k3s step.

Fix: Check that the server is online and that systemctl status nebula-mesh-sync.timer reports active on it. Then run a fresh join command.

Verify: The node appears in the cluster and reads Ready.

<cluster> is disconnected

Cause: The cluster page shows this alert when the control plane received no heartbeat for 60 seconds; the agent sends one every 15 seconds. Drain, Uncordon and other node commands then wait or are refused with this cluster was last seen at <time>, so it cannot be reached.

Fix: On a server node, run k3s kubectl -n nebula-system logs deploy/nebula-agent. The lines stream disconnected and control plane reachability probe failed mean the agent cannot reach the control plane; fix DNS, TLS or the firewall between them. The line agent token rotation requested; identity cleared, restart the agent with a new enrollment token means you must regenerate the install credentials.

Verify: The cluster reads Healthy.

<node> is NotReady

Cause: The cluster reads Degraded when a connected cluster has a node whose kubelet stopped reporting Ready.

Fix: On that node, read systemctl status k3s for a server or k3s-agent for a worker, and systemctl status wg-quick@nebula0 for the mesh.

Verify: The cluster returns to Healthy.

the new agent did not report in within 10 minutes

Cause: The agent update history shows Rolled back when the new agent did not connect before the deadline, 10 minutes by default. The control plane then starts an automatic rollback. A rollback that also times out reads Failed.

Fix: Read k3s kubectl -n nebula-system describe deploy nebula-agent for an image pull or readiness failure. Then select Update agent again. See agent updates.

Verify: The history row reads Healthy.

wait for the deployments and builds on this cluster to finish, then update the agent

Cause: Update agent is refused with this API detail while a deployment or build reconciles on the cluster. Related refusals: the cluster must be connected before its agent can be updated and an agent update is already running on this cluster.

Fix: Wait for the work to finish, reconnect the cluster, or wait for the running update. Then select Update agent again.

Verify: A toast reads Agent update started.

cannot update the nebula.dev CustomResourceDefinitions

Cause: The agent logs this warning at start when its role predates the permission to update its own definitions. It keeps running, but features that need newer fields, such as file mounts, stay unavailable, and a build or deployment that needs one is refused with a message to re-apply the agent manifest.

Fix: In the cluster's Settings, select Regenerate beside Install credentials, then run the new Install command on the first server.

Verify: The warning is absent from the agent log after the next restart.

<address> presented a host key that does not match the fingerprint you confirmed

Cause: Adding a server over SSH pins the fingerprint you confirm with Fingerprint matches — install, and the machine now presents another key. Related refusals: <host>:<port> did not complete an ssh handshake and refused to sign in as <user> with this organization's key.

Fix: If you reinstalled the machine, start a new install and confirm the new fingerprint. Otherwise treat the connection as intercepted. For the sign-in refusal, add the public key from the dialog to ~<user>/.ssh/authorized_keys.

Verify: The dialog moves to joining and the node appears.

<count> environments still bound to this cluster; unbind each one (or move it to another cluster) before removing it

Cause: Remove cluster is refused with a 409 while an environment runs on the cluster. The message reads 1 environment for one. A cluster that fronts domains for edge clusters is refused with this cluster serves <count> domains of other clusters as their edge; switch them under Domains → Serve through, then remove this cluster, which reads 1 domain for one.

Fix: Move or unbind each environment, change Serve through on the listed domains, then remove the cluster again. Workloads keep running and K3s stays installed.

Verify: DELETE /api/v1/clusters/<id> answers 204.

the tailscale operator deployment did not become ready within 180 seconds

Cause: The script behind Connect, in the Tailnet section of the cluster's Settings, waits 180 seconds for the operator. no tailscale OAuth client was supplied means the client ID or secret was empty.

Fix: On the server, run kubectl -n tailscale get pods and kubectl -n tailscale logs deploy/operator, correct the OAuth client, then run the command again. See Tailscale.

Verify: The script ends with Done. This cluster is connected to your tailnet.

On this page