Skip to content
NebulaCtrldocs
Guides

Read logs and metrics

Tail and search a service's logs, read build and deploy logs, and read the Metrics tab and the numbers on each canvas card. Also covers opting a worker's own metrics in.

When a service misbehaves, its logs say what happened and its metrics say how much. This page covers each place the console shows them and the API call behind it.

Before you begin

  • Reading logs, metrics, instances and builds needs the viewer role or higher. Opting a worker's metrics in needs deployer.
  • The service has a running deployment. A service that never started has no lines and no points.

Tail a service's logs

Select the service on the canvas, then select Logs. With a service selected, pressing L opens the same tab. The tab shows the newest 200 retained lines, newest at the bottom, and a Filter lines… box that matches text. Under the lines you see live · tailing N instances and the environment's slug.

Choose all, warn+ or error to set the minimum level. warn+ includes errors. A line with no recognizable level shows only under all.

Select Pause to hold new lines aside. The status reads paused · N lines buffered. Select Resume to append them.

Each line shows its UTC time, level, instance (such as web-0) and text. A divider in the tail marks each release's activation, so a spike lines up with what shipped.

A live tail closes after 15 minutes, and a credential holds at most 4 open streams. The console reconnects on its own, waiting from 1 to 30 seconds between attempts. The HTTP API page lists every stream limit.

Search logs

The environment view searches every service at once and filters by instance. It has no button; add panel=logs to the project's address, for example ?env=ENVIRONMENT_SLUG&panel=logs.

Open the project with panel=logs. A Logs panel opens: "Every service in ENVIRONMENT_SLUG, newest at the bottom. Kept for 7 days." Columns are Time, Service, Pod and Message.

Type in Search log text. After a short pause, the panel reloads with matching lines, highlights each match and counts them.

Pick an instance in All pods and a window in Time range: Last 15m, Last 1h (the default), Last 6h, Last 24h or Last 7d. Set the level with all, warn+ or error.

Select Load older lines to page back, 200 lines at a time. Select Download to save the lines on screen as a .log file.

A search matches whole words, ignoring case. A line must contain every word you type, in any order. Part of a word does not match. Lines that arrive live are filtered by case-insensitive substring instead.

A range that reaches past retention ends in a marker: log lines before TIMESTAMP are outside the 7-day retention window and were never stored. NebulaCtrl keeps 7 whole UTC days of lines and drops older days hourly.

The same search is listEnvironmentLogs:

curl -G "BASE_URL/api/v1/environments/ENVIRONMENT_ID/logs" \
  -H "Authorization: Bearer NEBULA_TOKEN" \
  --data-urlencode "q=connection refused" \
  --data-urlencode "level=error" \
  --data-urlencode "from=FROM_TIMESTAMP" \
  --data-urlencode "limit=50"

q, serviceId, podName, from, to and level are optional. Results come newest first, 50 per page by default and 200 at most. Follow nextCursor for the next page.

Read a build's or a deployment's logs

Select the service, then Deployments, then the deployment. The deployment's card opens with the tabs Build logs, Deploy logs and Details. Build logs appears only for a deployment that came from a build.

Select Build logs for the build output, grouped by step, or Deploy logs for the output of that release's instances.

Type in Search N lines to filter, or use Wrap, Follow, Copy and Download.

A build's log is kept for 7 days. The card fetches at most 5,000 build lines and says so when it stops. Through the API, listBuildLogs pages a build's log oldest first (GET /api/v1/builds/{id}/logs, 200 lines by default, 1,000 at most), and streamBuildLogs tails it (GET /api/v1/builds/{id}/logs/stream). See Builds.

Read service metrics

Select the service, then Metrics. The tab shows the cards CPU, Memory, Requests, p95 latency and Error rate, refreshed every 30 seconds.

Choose 1h, 24h or 7d. The note under the cards reads "One point per 1 min", "One point per 15 min" or "One point per 120 min".

CardWhat it is
CPU, MemoryUsage summed over every pod of the service. With a limit, the card draws it as a line and shows "N% of limit".
RequestsRequests per minute through the cluster's proxy.
p95 latency95th-percentile response time. A bucket with too few requests has no point.
Error rateShare of requests answered 5xx. A bucket with no requests has no point.

"N% of limit" turns amber from 70% and red from 85%. The limit is the sum of the limits of the http and worker processes across their replicas. If any of them has no limit, the card says "no limit set".

A card with no points reads — and, in place of the chart, not reported followed by a reason. It was not measured; it does not mean zero. For CPU and Memory the reason is "the cluster reports no CPU or memory metrics" or "no samples in this window". For Requests, p95 latency and Error rate it is "the cluster reports no traffic metrics" or "no requests in this window". A stopped service shows No metrics while stopped above the cards, and the cards disappear when the window holds no points. A service whose last deployment failed with no earlier release live shows Metrics unavailable and an Open logs button the same way.

The API operations are getServiceMetricsSummary (GET /api/v1/services/{id}/metrics/summary, windows 1h, 24h and 7d; the console uses it), getServiceMetrics (GET /api/v1/services/{id}/metrics) and getServiceTraffic (GET /api/v1/services/{id}/traffic), both with windows 1h, 6h and 24h. Each takes environmentId. Points are kept for 7 days, like logs.

Read canvas metrics

Each service card on the canvas shows three numbers for the trailing 30 minutes, in 60-second buckets (up to 30 points). The agent samples every 15 seconds and the console reads every 30 seconds. A database service's template decides the card. Any other service follows its first process by name.

CardShows
http processCPU, p95, RPM
worker processCPU, Mem, Restarts. Once opted in and scraped: CPU, Queue, Jobs/min.
PostgreSQLConns, QPS (writes per second), Lag (replication lag), writer and reader pips
ValkeyMem, Ops/s, Hit (cache hit rate)
cron processLast run, Next, 7 runs (N/M ok)
MySQL, MongoDB, job, proxyno numbers

CPU shows a share of the limit when one is known, and millicores otherwise. RPM and p95 cover the whole 30 minutes. p95 needs at least 10 requests in that window.

A reading turns amber, and an amber dot appears beside the service's status, from 90% of its warning level. Hover the dot for the list, such as "CPU 92% of 1 vCPU · near limit".

ReadingWarning level
p95300 ms
Queue60
Conns160
Lag1 s
CPU, Memthe service's limit

Nothing else turns amber. With no limit known, hovering CPU or Mem says "CPU limit not reported" or "Memory limit not reported" and the reading never turns amber. A limit is known only when every http and worker process of the service has both a CPU and a memory limit.

An RPM or p95 slot with no reading says why on hover. A service with only tailnet domains, or with no public domain, never passes through the proxy where traffic is measured, and a cluster without a Traefik metrics endpoint reports no traffic. See Domains and Tailscale.

A PostgreSQL card needs CloudNativePG on the cluster. Its pips show the instances CloudNativePG last reported, each lit when ready. A database silent for 2 minutes shows none. A Valkey card needs nothing: the agent reads its own INFO.

Opt a worker's metrics in

A worker shows its own numbers (Queue and Jobs/min) after you opt in and the agent scrapes at least one point for either.

Select the service, then Settings. Under Processes, select Edit… on the worker.

Turn on Scrape app-exported metrics for the canvas card. Metrics port and Metrics path appear.

Enter the port. Leave Metrics path empty for /metrics. Select Save process. A toast reads "Saved WORKER_NAME".

If nebula.toml sets the key, the switch and its fields are disabled and a note names the file that sets them. Remove the key from the file to edit it here.

A deployment freezes process settings, so the card changes after the next deployment of the service.

Your process exposes Prometheus text on that port and path:

NameTypeCard slot
nebula_queue_depthgaugeQueue
nebula_jobs_processed_totalcounterJobs/min

Expose either or both. The agent reads the first series of each name, whatever type it declares. It scrapes one pod, the lowest-named, every 15 seconds. A scrape that fails is skipped silently. That covers a timeout over 3 seconds, a refused connection, a non-200 answer, a body over 256 KiB, unparseable text and a redirect, which the agent never follows. A counter that resets on restart skips one point.

The process needs a port, its own or Metrics port. Metrics apply only to a worker. The path must start with / and cannot start with // or contain @, ? or #. The card follows the service's first process by name, so a worker that is not first shows no app numbers. Turning the option off returns the card to Mem and Restarts once the last scraped points leave the card's 30-minute window.

Verify

  • Logs: new lines appear and the status reads live.
  • Metrics: the cards show values, or a reason after not reported.
  • Opt-in: after a deployment and a minute or two, the worker card shows Queue or Jobs/min. If it still shows Mem, open the worker's logs and check that the endpoint answers with Prometheus text.

Next steps

On this page