TurboPanel Docs
Deployment

Deployment Troubleshooting

This guide helps resolve common issues when deploying or running TurboPanel in production. For security-related issues, see Security. For setup verification, see Control plane and Daemon setup.

Introduction

Production issues typically fall into TLS/Caddy misconfiguration, instance socket permissions, daemon WebSocket connectivity, Postgres socket access, or Ansible playbook failures. The self-hosted browser entrypoint is HTTPS on port 8443. Co-located development uses that same listener.

PortWhen
8443Always. Control plane HTTPS. Every name is https://<host>:8443.
80Only during Let's Encrypt issuance or renewal.
443Hosting only.

Hosting Caddy also binds 80 while hosted sites are deployed. A control plane with no hosted sites leaves 80 closed except during that issuance or renewal window. See Hostnames and TLS.

Port conflicts (8443 / 80 / 443)

Problem: Control-plane Caddy cannot bind :8443. During Let's Encrypt issuance or renewal, another process holds port 80.

Solutions:

  1. For :8443, stop or move the process that holds it. The managed Caddyfile binds :8443 literally and does not read CADDY_PORT (only the co-located development overlay does), so the port cannot be changed
  2. When the apply reports port 80 is held by <process> (for example port 80 is held by nginx), stop that process or move it off port 80, then apply again. Hosting Caddy is the expected holder when hosted sites are already deployed
  3. Check listeners: ss -tlnp | grep -E '8443|80|443'
  4. Port 443 is hosting Caddy. Leave it for hosted sites

UI or API not loading

Problem: Browser shows errors at https://<host>:8443.

Solutions:

  1. Check the control plane: systemctl status turbopanel-instance
  2. Check Caddy: systemctl status turbopanel-caddy
  3. Test health over HTTPS: curl -k https://localhost:8443/api/health (omit -k when the leaf is publicly trusted). Trust the Platform CA for a self-hosted name.
  4. Logs: journalctl -u turbopanel-instance -u turbopanel-caddy -f
  5. Static UI: ensure the UI export exists in /opt/turbopanel/share/ui if TURBOPANEL_UI_MODE=static
  6. Dev UI: ensure turbopanel-ui is running when TURBOPANEL_UI_MODE=dev

Daemon shows offline but service is running

Problem: UI badge is Offline (or Update fails) while systemctl status turbopaneld is active and the Cell panel shows recent activity.

Cause: The server list uses a Postgres presence projection that can drift from live WebSocket state after a failed upgrade or reconnect.

Solution: Re-run the release installer — see Refresh a stuck daemon.

Daemon not connecting

Problem: Remote server daemon fails to register.

Solutions:

  1. Verify the control plane URL: TURBOPANEL_INSTANCE_URL=https://<host>:8443. The daemon refuses an http:// control-plane URL.
  2. Test from the server over HTTPS: curl -k https://<host>:8443/api/health (omit -k when the origin is publicly trusted). LAN names trust the Platform CA, or use -k only for this bootstrap check.
  3. Confirm the Platform CA: compare /etc/turbopanel/instance-ca.pem with GET /api/daemon/v1/instance/ca (bundle, current CA first). A 404 on a Let's Encrypt or uploaded name means that name uses the system trust store, or the uploaded issuer for a private pair — do not install instance-ca.pem for that name. An unlisted name on :8443 still receives the Platform CA bundle.
  4. Firewall: allow 8443 from remote daemons to the control plane. Allow 80 from the internet only while Let's Encrypt is issuing or renewing. 443 is hosting. The installer removes ufw, firewalld and iptables-persistent from every host it installs on — TurboPanel owns the host firewall — so a distribution firewall cannot be what blocks 8443 after a completed install. If curl https://127.0.0.1:8443/api/health answers locally but not remotely, look upstream: a cloud security group, a provider firewall, or a NAT in front of the host.
  5. Logs: journalctl -u turbopaneld -f

Stale platform CA (tls-trust parked)

Problem: Daemon log shows tls-trust parked (dialed host, CA path, fingerprint) and reconnects only every 5 minutes to 1 hour — not a silent 30 s loop.

Cause: The host still trusts an old platform CA after the control plane rotated (or the leaf SAN no longer matches the dialed hostname). Control-plane identity (JWKS JWT) is unchanged; only the transport anchor is stale.

Solutions:

  1. Prefer the overlap path: on the control plane, keep the old CA in ca-bundle.pem and enqueue server.tls.trust.reconcile while the existing WSS session is still valid.
  2. If the session is already dead, confirm the new CA's fingerprint out of band and re-run the installer with --instance-ca <pem>. Without it, run.sh fetches with the existing pin; when that cannot connect it retries once against the system trust store only and installs a PEM only if it validates the live leaf. It never retries with verification off, and when it cannot verify, it keeps the existing CA and prints both fingerprints. --insecure-tls makes the bootstrap fetch unverified; the run still stops unless the fetched CA validates the live leaf.
  3. If verification fails, keep the existing CA and fix the control plane URL / SAN list rather than blindly replacing the file.

Postgres connection failures (self-hosted)

Problem: The control plane cannot connect to Postgres.

Solutions:

  1. Check Postgres container: docker ps | grep postgres
  2. Verify Unix socket under /var/run/turbopanel/postgres/
  3. Confirm TURBOPANEL_DATABASE_URL in turbopanel-instance.service — a full Postgres URL, typically injected by the instance-launch role
  4. Socket directory permissions: /run/turbopanel should be 2770 tp:tp

Install wizard / PAM failures (Deno)

Problem: Install bootstrap rejects host credentials.

Solutions:

  1. Install pamtester on the host
  2. Confirm tpctrl user sudoers allows pamtester login * authenticate
  3. Use a host account in sudo / wheel / admin, or root

UI Update stuck on "update already in progress"

Problem: After clicking Update, Ansible runs but the daemon keeps the same PID for a long time. A second update attempt returns update already in progress.

Cause: The UI update path uses run.sh --no-start (does not stop the daemon during reconcile). The daemon then restarts itself with sudo -n systemctl enable and sudo -n systemctl restart --no-block, as the tp user. If that restart is not accepted, the old process keeps running and its in-memory lock blocks further updates.

Check logs on the server:

Terminal
sudo tail -30 /var/log/turbopanel/daemon.err.log
sudo grep -i 'systemctl\|sudo\|update' /var/log/turbopanel/daemon.log | tail -20

Immediate recovery on the server:

Terminal
sudo systemctl restart turbopaneld

Or reconcile by re-running run.sh — see Refresh a stuck daemon.

Stuck or failed platform upgrades

Problem: Admin → Updates stays on Updating… or Needs attention, or a server row is needs_attention.

What the statuses mean:

StatusWhat to do
waitingThe daemon is offline. It is dispatched when it reconnects.
failed / rolled_backThe step ended. A rolled-back step is retried once; after that it becomes needs_attention.
needs_attentionRetries are exhausted (3 dispatches with no progress for 15 minutes, or a rollback that was already retried). Retry on that row, or fix the host and retry.
Run failedThe control-plane step failed. Servers stay gated until a new run gets this host's daemon and the control plane onto the target.

Pre-flight codes (preflight_disk, preflight_manifest, preflight_trust, preflight_backup, preflight_in_progress) are explained on Refresh a stuck daemon. A control-plane health failure that restored the previous build reports health_timeout, health_mismatch, or restart_failed. recovery_required means that automatic restore did not bring the control plane back.

Restore from the automatic backup. On the control-plane host, as root, use the recovery command from the pre-flight sheet (the upgrade id is in the path):

Terminal
sudo -n /opt/turbopanel/share/orchestration/scripts/tp-orchestrate playbook -i localhost, -c local \
  -e turbopanel_upgrade_id=<upgrade id> \
  instance-rollback.yml

The dump, the /etc/turbopanel tarball, and meta.json are under /backup/control-plane/<upgrade id>/ (TURBOPANEL_BACKUP_DIR when that is set). The playbook restores the .prev builds and restores the database only when the migration fingerprint changed. The newest three attempts are kept.

If the co-located daemon is too old to take that backup (managed-upgrade-v1 missing), update only the daemon with a pinned manifest — Upgrade and rollback. That command does not install the control plane.

Ansible / upgrade failures

Problem: daemon orchestration fails (or, on a contributor dev host only, the console's Upgrade System action — managed installs have no such button; they re-run the installer).

Solutions:

  1. Check daemon logs — on managed hosts Ansible runs as the tp user; on co-located dev it runs as the current dev user
  2. Dirty git checkouts block upgrade; commit or stash changes in $HOME source repos (~/turbopaneld, ~/turbopanel, ~/ui, ~/website)

What to point a monitor at

/api/health is not an outage signal. It is a static identity payload — licence, version, the commit the build came from — and it never touches the database. It answers 200 with Postgres stopped, which is exactly when an operator most wants it to say something. A game day on a dev stack confirmed it: the control-plane database was killed and /api/health returned 200 throughout.

Monitor /api/daemon/v1/readiness instead. It reads the database, so it is the endpoint that can fail, and its three answers are distinguishable:

ResponseMeaning
200 {"ok":true,"ready":true}Serving
503 {"ok":true,"ready":false,"needsInstall":true}Up, but the install wizard has not created the first organization and superadmin yet
503 {"ok":false,"ready":false,"error":"database unavailable"}The database is down, unreachable or mid-failover

Alert on the third. The second is a normal state for a freshly installed control plane, and co-located daemons already poll this endpoint and wait it out.

Getting paged when a server goes dark

The control plane watches its servers on a sweep and marks a daemon offline when it stops answering. Without somewhere to send that, it is a log line nobody reads at 3am. Point it at an incoming webhook:

Terminal
curl -X PUT https://<host>:8443/api/admin/v1/settings/alert-webhook \
  -H 'content-type: application/json' \
  --cookie "$SESSION" \
  -d '{"url":"https://hooks.slack.com/services/T000/B000/XXXX"}'

{"url": null} clears it. Two alerts go out:

AlertWhen
server.offlineOne daemon stopped answering and was marked offline
fleet.mass_disconnectA single sweep lost at least three hosts and either half the connected servers or ten hosts outright — a shared cause (a bad daemon release, a broken ingress, a cell outage) rather than N coincidences. It is sent before the per-server alerts it explains

The body carries a text field, which Slack, Mattermost, Rocket.Chat and Discord all render, plus kind and detail for anything that parses JSON.

Two things to know about the URL:

  • It is a credential. In every common webhook scheme the path is the secret. It is stored sealed, the same way an SMTP password is, and reading the setting back returns only the origin and whether one is configured — never the path. If you lose it, set a new one; there is no way to read the old one out.
  • It has to be https, with no credentials in the URL. That is the whole rule — a receiver on your LAN (an Alertmanager next to the control plane) is fine. The forge URLs keep a stricter, public-only gate because those fetches carry the App's credentials; this one carries nothing and its response goes nowhere.

A delivery that fails is logged and dropped. It never delays or fails the sweep — an unreachable Slack must not be the reason a dead host stays marked online.

What recovers on its own

FaultRecovery
Postgres container crashesDocker's unless-stopped policy restarts it; the control plane's connection pool reconnects. No operator action, and turbopanel-instance is not restarted — if systemd bounced the unit, something else is wrong
RabbitMQ container crashesSame, and the command consumer reconnects with capped backoff, retrying until the broker is back
Control plane process exitsRestart=always in the unit, with a 2s delay

One trap worth knowing before you test any of this by hand: docker kill <name> is not a crash simulation. Docker records it as a manual stop, and a restart policy does not fire for a manual stop — the container stays down until you run docker start. To simulate a real crash, signal the container's process directly:

Terminal
sudo kill -9 "$(docker inspect -f '{{.State.Pid}}' turbopanel-database)"

(docker exec <name> kill -9 1 does nothing at all — PID 1 of a PID namespace ignores SIGKILL sent from inside that namespace.)

Contributors can run the whole exercise, with these assertions checked for them, from the dev checkout:

Terminal
./scripts/game-day.sh --all --include-operator-stop

It refuses to run outside a development guest, bounds every wait, and exits non-zero if the stack does not come back.

General debugging

TaskCommand
Control plane healthcurl -k https://localhost:8443/api/health (omit -k when the leaf is publicly trusted; trust the Platform CA for a self-hosted name)
Control plane readinesscurl -k https://localhost:8443/api/daemon/v1/readiness — the probe that reads the database; see What to point a monitor at
Client statuscurl -k https://localhost:8443/api/client/v1/status (omit -k when the leaf is publicly trusted)
Service statussystemctl status turbopanel-instance turbopanel-caddy turbopaneld
Follow logsjournalctl -u turbopanel-instance -u turbopanel-caddy -u turbopaneld -f
Daemon WS pathwss://<host>:8443/ws/daemon/v1
Edit on GitHub

Last updated on

On this page