Deployment Troubleshooting
This guide helps resolve common issues when deploying or running TurboPanel in production. For security-related issues, see Security. For setup verification, see Control plane and Daemon setup.
Introduction
Production issues typically fall into TLS/Caddy misconfiguration, instance socket permissions, daemon WebSocket connectivity, Postgres socket access, or Ansible playbook failures. The self-hosted browser entrypoint is HTTPS on port 8443. Co-located development uses that same listener.
| Port | When |
|---|---|
| 8443 | Always. Control plane HTTPS. Every name is https://<host>:8443. |
| 80 | Only during Let's Encrypt issuance or renewal. |
| 443 | Hosting only. |
Hosting Caddy also binds 80 while hosted sites are deployed. A control plane with no hosted sites leaves 80 closed except during that issuance or renewal window. See Hostnames and TLS.
Port conflicts (8443 / 80 / 443)
Problem: Control-plane Caddy cannot bind :8443. During Let's Encrypt issuance or renewal, another process holds port 80.
Solutions:
- For
:8443, stop or move the process that holds it. The managed Caddyfile binds:8443literally and does not readCADDY_PORT(only the co-located development overlay does), so the port cannot be changed - When the apply reports
port 80 is held by <process>(for exampleport 80 is held by nginx), stop that process or move it off port 80, then apply again. Hosting Caddy is the expected holder when hosted sites are already deployed - Check listeners:
ss -tlnp | grep -E '8443|80|443' - Port 443 is hosting Caddy. Leave it for hosted sites
UI or API not loading
Problem: Browser shows errors at https://<host>:8443.
Solutions:
- Check the control plane:
systemctl status turbopanel-instance - Check Caddy:
systemctl status turbopanel-caddy - Test health over HTTPS:
curl -k https://localhost:8443/api/health(omit-kwhen the leaf is publicly trusted). Trust the Platform CA for a self-hosted name. - Logs:
journalctl -u turbopanel-instance -u turbopanel-caddy -f - Static UI: ensure the UI export exists in
/opt/turbopanel/share/uiifTURBOPANEL_UI_MODE=static - Dev UI: ensure
turbopanel-uiis running whenTURBOPANEL_UI_MODE=dev
Daemon shows offline but service is running
Problem: UI badge is Offline (or Update fails) while systemctl status turbopaneld is active and the Cell panel shows recent activity.
Cause: The server list uses a Postgres presence projection that can drift from live WebSocket state after a failed upgrade or reconnect.
Solution: Re-run the release installer — see Refresh a stuck daemon.
Daemon not connecting
Problem: Remote server daemon fails to register.
Solutions:
- Verify the control plane URL:
TURBOPANEL_INSTANCE_URL=https://<host>:8443. The daemon refuses anhttp://control-plane URL. - Test from the server over HTTPS:
curl -k https://<host>:8443/api/health(omit-kwhen the origin is publicly trusted). LAN names trust the Platform CA, or use-konly for this bootstrap check. - Confirm the Platform CA: compare
/etc/turbopanel/instance-ca.pemwithGET /api/daemon/v1/instance/ca(bundle, current CA first). A 404 on a Let's Encrypt or uploaded name means that name uses the system trust store, or the uploaded issuer for a private pair — do not installinstance-ca.pemfor that name. An unlisted name on:8443still receives the Platform CA bundle. - Firewall: allow 8443 from remote daemons to the control plane. Allow 80 from the internet only while Let's Encrypt is issuing or renewing. 443 is hosting. The installer removes
ufw,firewalldandiptables-persistentfrom every host it installs on — TurboPanel owns the host firewall — so a distribution firewall cannot be what blocks 8443 after a completed install. Ifcurl https://127.0.0.1:8443/api/healthanswers locally but not remotely, look upstream: a cloud security group, a provider firewall, or a NAT in front of the host. - Logs:
journalctl -u turbopaneld -f
Stale platform CA (tls-trust parked)
Problem: Daemon log shows tls-trust parked (dialed host, CA path, fingerprint) and reconnects only every 5 minutes to 1 hour — not a silent 30 s loop.
Cause: The host still trusts an old platform CA after the control plane rotated (or the leaf SAN no longer matches the dialed hostname). Control-plane identity (JWKS JWT) is unchanged; only the transport anchor is stale.
Solutions:
- Prefer the overlap path: on the control plane, keep the old CA in
ca-bundle.pemand enqueueserver.tls.trust.reconcilewhile the existing WSS session is still valid. - If the session is already dead, confirm the new CA's fingerprint out of band and re-run the installer with
--instance-ca <pem>. Without it,run.shfetches with the existing pin; when that cannot connect it retries once against the system trust store only and installs a PEM only if it validates the live leaf. It never retries with verification off, and when it cannot verify, it keeps the existing CA and prints both fingerprints.--insecure-tlsmakes the bootstrap fetch unverified; the run still stops unless the fetched CA validates the live leaf. - If verification fails, keep the existing CA and fix the control plane URL / SAN list rather than blindly replacing the file.
Postgres connection failures (self-hosted)
Problem: The control plane cannot connect to Postgres.
Solutions:
- Check Postgres container:
docker ps | grep postgres - Verify Unix socket under
/var/run/turbopanel/postgres/ - Confirm
TURBOPANEL_DATABASE_URLinturbopanel-instance.service— a full Postgres URL, typically injected by theinstance-launchrole - Socket directory permissions:
/run/turbopanelshould be2770 tp:tp
Install wizard / PAM failures (Deno)
Problem: Install bootstrap rejects host credentials.
Solutions:
- Install
pamtesteron the host - Confirm
tpctrluser sudoers allowspamtester login * authenticate - Use a host account in
sudo/wheel/admin, orroot
UI Update stuck on "update already in progress"
Problem: After clicking Update, Ansible runs but the daemon keeps the same
PID for a long time. A second update attempt returns update already in progress.
Cause: The UI update path uses run.sh --no-start (does not stop the daemon
during reconcile). The daemon then restarts itself with sudo -n systemctl enable
and sudo -n systemctl restart --no-block, as the tp user. If that restart
is not accepted, the old process keeps running and its in-memory lock blocks further
updates.
Check logs on the server:
sudo tail -30 /var/log/turbopanel/daemon.err.log
sudo grep -i 'systemctl\|sudo\|update' /var/log/turbopanel/daemon.log | tail -20Immediate recovery on the server:
sudo systemctl restart turbopaneldOr reconcile by re-running run.sh — see Refresh a stuck daemon.
Stuck or failed platform upgrades
Problem: Admin → Updates stays on Updating… or Needs attention, or a server row is needs_attention.
What the statuses mean:
| Status | What to do |
|---|---|
waiting | The daemon is offline. It is dispatched when it reconnects. |
failed / rolled_back | The step ended. A rolled-back step is retried once; after that it becomes needs_attention. |
needs_attention | Retries are exhausted (3 dispatches with no progress for 15 minutes, or a rollback that was already retried). Retry on that row, or fix the host and retry. |
Run failed | The control-plane step failed. Servers stay gated until a new run gets this host's daemon and the control plane onto the target. |
Pre-flight codes (preflight_disk, preflight_manifest, preflight_trust, preflight_backup, preflight_in_progress) are explained on Refresh a stuck daemon. A control-plane health failure that restored the previous build reports health_timeout, health_mismatch, or restart_failed. recovery_required means that automatic restore did not bring the control plane back.
Restore from the automatic backup. On the control-plane host, as root, use the recovery command from the pre-flight sheet (the upgrade id is in the path):
sudo -n /opt/turbopanel/share/orchestration/scripts/tp-orchestrate playbook -i localhost, -c local \
-e turbopanel_upgrade_id=<upgrade id> \
instance-rollback.ymlThe dump, the /etc/turbopanel tarball, and meta.json are under /backup/control-plane/<upgrade id>/ (TURBOPANEL_BACKUP_DIR when that is set). The playbook restores the .prev builds and restores the database only when the migration fingerprint changed. The newest three attempts are kept.
If the co-located daemon is too old to take that backup (managed-upgrade-v1 missing), update only the daemon with a pinned manifest — Upgrade and rollback. That command does not install the control plane.
Ansible / upgrade failures
Problem: daemon orchestration fails (or, on a contributor dev host only, the console's Upgrade System action — managed installs have no such button; they re-run the installer).
Solutions:
- Check daemon logs — on managed hosts Ansible runs as the
tpuser; on co-located dev it runs as the current dev user - Dirty git checkouts block upgrade; commit or stash changes in
$HOMEsource repos (~/turbopaneld,~/turbopanel,~/ui,~/website)
What to point a monitor at
/api/health is not an outage signal. It is a static identity payload — licence, version, the commit the build came from — and it never touches the database. It answers 200 with Postgres stopped, which is exactly when an operator most wants it to say something. A game day on a dev stack confirmed it: the control-plane database was killed and /api/health returned 200 throughout.
Monitor /api/daemon/v1/readiness instead. It reads the database, so it is the endpoint that can fail, and its three answers are distinguishable:
| Response | Meaning |
|---|---|
200 {"ok":true,"ready":true} | Serving |
503 {"ok":true,"ready":false,"needsInstall":true} | Up, but the install wizard has not created the first organization and superadmin yet |
503 {"ok":false,"ready":false,"error":"database unavailable"} | The database is down, unreachable or mid-failover |
Alert on the third. The second is a normal state for a freshly installed control plane, and co-located daemons already poll this endpoint and wait it out.
Getting paged when a server goes dark
The control plane watches its servers on a sweep and marks a daemon offline when it stops answering. Without somewhere to send that, it is a log line nobody reads at 3am. Point it at an incoming webhook:
curl -X PUT https://<host>:8443/api/admin/v1/settings/alert-webhook \
-H 'content-type: application/json' \
--cookie "$SESSION" \
-d '{"url":"https://hooks.slack.com/services/T000/B000/XXXX"}'{"url": null} clears it. Two alerts go out:
| Alert | When |
|---|---|
server.offline | One daemon stopped answering and was marked offline |
fleet.mass_disconnect | A single sweep lost at least three hosts and either half the connected servers or ten hosts outright — a shared cause (a bad daemon release, a broken ingress, a cell outage) rather than N coincidences. It is sent before the per-server alerts it explains |
The body carries a text field, which Slack, Mattermost, Rocket.Chat and Discord all render, plus kind and detail for anything that parses JSON.
Two things to know about the URL:
- It is a credential. In every common webhook scheme the path is the secret. It is stored sealed, the same way an SMTP password is, and reading the setting back returns only the origin and whether one is configured — never the path. If you lose it, set a new one; there is no way to read the old one out.
- It has to be
https, with no credentials in the URL. That is the whole rule — a receiver on your LAN (an Alertmanager next to the control plane) is fine. The forge URLs keep a stricter, public-only gate because those fetches carry the App's credentials; this one carries nothing and its response goes nowhere.
A delivery that fails is logged and dropped. It never delays or fails the sweep — an unreachable Slack must not be the reason a dead host stays marked online.
What recovers on its own
| Fault | Recovery |
|---|---|
| Postgres container crashes | Docker's unless-stopped policy restarts it; the control plane's connection pool reconnects. No operator action, and turbopanel-instance is not restarted — if systemd bounced the unit, something else is wrong |
| RabbitMQ container crashes | Same, and the command consumer reconnects with capped backoff, retrying until the broker is back |
| Control plane process exits | Restart=always in the unit, with a 2s delay |
One trap worth knowing before you test any of this by hand: docker kill <name> is not a crash simulation. Docker records it as a manual stop, and a restart policy does not fire for a manual stop — the container stays down until you run docker start. To simulate a real crash, signal the container's process directly:
sudo kill -9 "$(docker inspect -f '{{.State.Pid}}' turbopanel-database)"(docker exec <name> kill -9 1 does nothing at all — PID 1 of a PID namespace ignores SIGKILL sent from inside that namespace.)
Contributors can run the whole exercise, with these assertions checked for them, from the dev checkout:
./scripts/game-day.sh --all --include-operator-stopIt refuses to run outside a development guest, bounds every wait, and exits non-zero if the stack does not come back.
General debugging
| Task | Command |
|---|---|
| Control plane health | curl -k https://localhost:8443/api/health (omit -k when the leaf is publicly trusted; trust the Platform CA for a self-hosted name) |
| Control plane readiness | curl -k https://localhost:8443/api/daemon/v1/readiness — the probe that reads the database; see What to point a monitor at |
| Client status | curl -k https://localhost:8443/api/client/v1/status (omit -k when the leaf is publicly trusted) |
| Service status | systemctl status turbopanel-instance turbopanel-caddy turbopaneld |
| Follow logs | journalctl -u turbopanel-instance -u turbopanel-caddy -u turbopaneld -f |
| Daemon WS path | wss://<host>:8443/ws/daemon/v1 |
Related documentation
- Security — TLS and authentication
- Control plane — Service layout
- Daemon setup — Node installer
- Upgrade and rollback — managed order, backups, recovery command
- Dev console troubleshooting — Local dev issues
Last updated on