Skip to content

5. Troubleshooting

Symptom Cause Resolution
Installer refuses the chosen disk, or a destructive step is blocked The disk (or a partition on it) is currently mounted Re-check with lsblk -f and pick the correct spare device — the installer's mount check is deliberate and will not be bypassed
"Could not determine disk-by-id path" during the cluster role detect-disk-id.sh was not run on the peer node first, or the value was mistyped/reordered Run detect-disk-id.sh on BOTH nodes before starting either real install; keep disk order identical on both sides
Cluster installer refuses to run on an already-deployed cluster Intentional data-safety guard — the live configuration already contains a VIP resource To rebuild from scratch (destroys all data): pcs cluster stop --all; pcs cluster destroy --all, then re-run secondary, then primary
Cluster primary step fails with "DRBD ... never connected" / "not in a fresh state" Only the primary was (re-)run after a partial failure Re-run BOTH nodes, secondary first, then primary
SSH access refused after several failed login attempts fail2ban's sshd jail has banned the source IP Wait for the ban to expire (default 10 minutes), or from a still-reachable session: fail2ban-client set sshd unbanip <ip>
Cannot SSH in as root after install Remote root login is disabled by design at the end of installation Use the console, or the local "support" operator account, or the GUI
Browser shows a certificate warning at https://<host>:8443/ The GUI uses a self-signed certificate until a real one is uploaded Expected — proceed past the warning, or upload a CA certificate under Configuration → TLS certificate
Standalone node's dashboard shows cluster/quorum/DRBD error panels The dashboard code has cluster-first defaults; those panels genuinely don't apply to a standalone node Cosmetic only, safe to ignore — confirmed non-fatal, every panel fails independently
Firewall page shows "No firewall apply has been run yet" The default ruleset has not been applied on this node yet Configuration → Firewall → Apply, once per node
Package install fails during an offline/unattended run ("no installation candidate") A stray active deb cdrom: line in /etc/apt/sources.list breaks apt-get update on that node Comment out the cdrom line in /etc/apt/sources.list and retry; current ISO builds no longer ship with this issue
"sudo: unable to resolve host \<name>" warnings during install The node has no DNS configured yet at that point in the install Cosmetic only — does not affect the outcome, resolves itself once DNS is configured
Config backup or NTP/timezone change appears to have no effect Change was saved on one node only; these are per-node settings Confirm the same value on BOTH cluster nodes — Network/NTP/timezone settings are intentionally not auto-replicated
Need to redeploy just the GUI without a full reinstall — Run installer/install-gui.sh [--config FILE] on the node (needs /etc/nascore/cluster.conf, already written by the role)