Five findings from a CodeAnt pass over the SSH host-key work:
- fail CLOSED in verify_transport_host_key: if get_remote_server_key()
raises we now reject instead of silently proceeding to auth unverified.
Every caller already wraps the call in its connect try/except, so the
raise is handled exactly like a changed key.
- port-aware pinning: a non-standard SSH port is keyed as [host]:port in
known_hosts (port 22 stays the bare host). Previously a non-22 host was
treated as unknown on every connect and never actually pinned. ssh_pool
and xhm now pass their real port.
- IPv6 host parsing in remove_host_keys: split(':')[0] truncated IPv6
addresses at the first colon; a dedicated token parser handles bare and
bracketed IPv6, bracketed IPv4:port, and comma host-lists.
- standalone WS server: serialize known_hosts writes with a lock so
concurrent TOFU saves can't corrupt the trust file.
- standalone WS server: warn loudly instead of silently falling back to a
cwd-relative known_hosts path when PEGAPROX_SSH_KNOWN_HOSTS is unset.
Unit-tested (fail-closed / port pin+reject / IPv6 removal) and live E2E'd
against real ESXi (pin, accept, downgrade-guard + same-type reject-on-change).
.ssh_ws_server.py re-synced from the embedded string in vms.py.
Follow-up to the TOFU host-key work: removing a cluster only deleted its DB rows,
leaving the pinned SSH host keys in known_hosts. Re-adding the SAME running node
worked (key still matches), but re-adding a node that had been REINSTALLED in the
meantime presented a new host key and tripped reject-on-change — the reconnect
failed until someone hand-edited known_hosts.
delete_cluster now collects the cluster host + node IPs (manager host, the
node-IP cache, and get_nodes() where reachable) and calls the new
ssh_security.remove_host_keys() before stopping the manager, so a later re-add
re-pins the current key via TOFU regardless of a reinstall.
Verified end-to-end against a live PVE node: add -> pin, simulate reinstall
(stale key blocks connect), delete -> pin removed, re-add -> fresh TOFU connects;
and the baseline remove+re-add of a live (non-reinstalled) node stays green.
Follow-up to the paramiko TOFU fix: the adversarial review found a parallel fleet
of subprocess ssh/scp commands still using StrictHostKeyChecking=no (accept ANY
host key → MitM), including the HA-fence hub->node paths (poweroff + fencing) and
node command execution.
Add cli_hostkey_opts() to ssh_security: returns (accept-new|yes, known_hosts_path)
— accept-new rejects a CHANGED key while allowing a genuinely new host, upgraded to
yes under strict mode, pinned to the same known_hosts the paramiko paths use.
Migrated all 8 manager.py subprocess callsites (_ha_fence_node poweroff,
_ssh_run_command x3, HA sshpass ssh x2) to StrictHostKeyChecking=accept-new +
UserKnownHostsFile=<shared known_hosts>. The node->node scp (executed ON the source
PVE node) uses accept-new against that node's own known_hosts rather than pinning
the hub file.
Live-verified against a real PVE node: the migrated subprocess ssh args connect +
pin the key, a changed key is rejected (rc 255), and strict mode rejects an unknown
host (rc 255).
Still open (deliberately NOT touched here — they run against ESXi/XCP-ng and can't
be safely live-tested overnight, and one carries the migration path just proven
end-to-end): 14 subprocess callsites in core/v2p.py (incl. _esxi_ssh, which sends
the ESXi password) and 3 in core/xhm.py. Same accept-new + shared-known_hosts
treatment needed, but each requires an ESXi/XCP-ng retest before shipping.
Aikido flagged 10 CRITICAL 'disabled SSH host key verification' findings: every
paramiko connection used AutoAddPolicy/WarningPolicy, which accept an unknown host
key silently — an attacker between the hub and a node (PVE/PBS/ESXi/XCP-ng/storage)
could impersonate it and capture credentials and commands. Worse, the historical
'TOFU' never actually worked: it wrote known_hosts to pegaprox/config/... which does
not exist at runtime, so no key was ever persisted and a changed key was never
rejected — effectively blind-trust on every connection.
New central module pegaprox/utils/ssh_security.py implements trust-on-first-use with
reject-on-change and an opt-in strict mode (PEGAPROX_SSH_STRICT_HOST_KEYS):
- apply_host_key_policy(): loads known_hosts + a custom verifying MissingHostKeyPolicy
for client.connect paths.
- verify_transport_host_key(): keyboard-interactive auth runs over a manually-built
paramiko.Transport, which BYPASSES the SSHClient policy entirely — so these paths
had no verification at all. Now the server key is checked (before any credential is
sent) on every Transport path; a known host that presents a changed key OR an
unpinned key type is rejected (blocks key-change + keytype-downgrade MitM).
- known_hosts now resolves to the REAL runtime config dir (<cwd>/config), shared by
every path including the system-ssh sshpass fallback (StrictHostKeyChecking
accept-new, upgraded to yes under strict mode) and the standalone terminal WS
subprocess (inline policy, path handed in via env).
Migrated every paramiko callsite: utils/ssh.py (3 clients + 2 KI transports + sshpass
fallback), utils/ssh_pool.py, utils/vnc_tunnel.py (client; the local forward's
server-side Transport is intentionally left as-is), core/xhm.py, core/pbs.py,
core/xcpng.py, core/manager.py, api/storage.py, api/vms.py (6 client paths, each now
persisting; the remove-node cleanup deliberately does not pin a departing node) and
the embedded ssh-ws server_script (inline TOFU).
Strictly E2E-tested against a live PVE node (192.168.1.2): TOFU first-connect pins
the key; second connect verifies; a changed key (same keytype) is rejected; an
unpinned-keytype downgrade is rejected; strict mode rejects unknown / allows known;
recovery after strict-off; the app boots with the new code; the standalone WS script
starts clean (no NameError). Adversarial-reviewed.
Scope: this covers the paramiko paths (the Aikido criticals). A separate follow-up
hardens the subprocess ssh/scp fleet (StrictHostKeyChecking=no / UserKnownHostsFile=
/dev/null in v2p + HA-fence + sshfs paths), which shares the same known_hosts file.