226 Commits

Author SHA1 Message Date
MrMasterbay
18d03f0476 feat(update): install-method-aware self-update (apt / docker / source)
The in-app updater and update.sh are a git-tree + pip file update that only fits
the source/deploy.sh layout. On an apt/dpkg install the correct path is apt
upgrade (deps are system python3-* packages; a pip run diverges from dpkg and
can't lift the dpkg-owned crypto libs -> fail-closed TLS on restart); on Docker
a file update is discarded at the next image pull.

- settings.py: _detect_install_method() (docker via /.dockerenv + cgroup, apt via
  dpkg -S, else source); perform_pegaprox_update refuses on apt/docker with the
  right guidance + copy-paste command (409, allow_managed override); check-update
  reports install_method / in_app_update_supported / managed_update_{hint,command}.
- update.sh: same detection + guard before touching anything (--force override).
- settings_modal.js: the Install button becomes a package-manager / image hint
  with a copyable command on apt/docker; performUpdate handles the 409.
- tests: 5 guard/reporting tests.
2026-08-11 19:02:44 +02:00
mkellermann97
ae7a7275d7 harden self-update: don't restart into a failed dependency upgrade
Both the in-app updater (api/settings.py) and update.sh now run a crypto/TLS
preflight in the target interpreter (the venv) before bouncing the service. Our
TLS startup is fail-closed (#633), so if a release's dependency bump didn't land
a loadable cryptography/pyOpenSSL pair, a restart takes the service down with no
way back up. On preflight failure we keep the current process running and tell
the operator to install deps + restart manually.

- api/settings.py: gate the restart thread on the preflight; the response now
  carries deps_ok / restarting=false + a clear message instead of a green check.
- update.sh: same preflight before the systemctl restart; exit with the fix
  steps instead of bouncing into a service that won't come back.
2026-08-11 08:20:08 +02:00
MrMasterbay
b57f454736 security(ssrf): pin webhook target to resolved IP + gate custom-template image_url at wget sink
- webhooks._guard_url now returns (ok, reason, url_to_use); public http(s) targets are
  IP-pinned via resolve_and_pin_url so a DNS rebind between the guard check and the POST
  can't swing the request to an internal host. allow_private opt-in skips the pin so an
  internal ntfy/Gotify keeps working. Both send_to_channel / _post_ntfy callers use the
  pinned url. (Aikido #469089218)
- templates_lib deploy: re-run the SSRF guard on tpl['image_url'] right before the node
  wgets it. add_custom_template validates on entry, but built-in catalog and DB-stored
  URLs reached the sink unchecked. (Aikido #469089270)
- tests: webhook-guard 3-tuple contract + metadata/loopback block + allow_private no-pin.
2026-08-10 08:53:20 +02:00
MrMasterbay
d26f044ccd security(authz): tighten backup-job target gate + guard unpooled-VM int()
- _authz_backup_targets: a non-admin backup.schedule holder now needs an
  explicit include-list of VMs they own. all=1 / pool / exclude-mode /
  selMode=all|exclude|pool and an *empty* selection are admin-only — PVE
  treats several of those as "every VM", and the load->edit->save round-trip
  of an admin job re-hits the gate on PUT so a scoped user can't retune it.
- get_vms_without_pool: a single non-numeric vmid no longer int()-throws and
  500s the whole unpooled-VM listing; skip the malformed row instead.
- 2 regression tests (empty/exclude selection + PUT all=1 denied).
2026-08-10 08:50:15 +02:00
MrMasterbay
4a702cd742 security: fix web-push IPv6 SSRF bypass + persist-credentials on CI checkouts (Aikido)
- push.py _is_internal_or_metadata_host (#469089273): stopped splitting the host on
  ':' (which mangled every IPv6 literal, '::1' -> '', bypassing the block) and unwrap
  IPv4-mapped IPv6 (::ffff:a.b.c.d) so metadata/private checks apply. +6 unit tests.
- docker.yml + release-images.yml (#348463387/#348463388): persist-credentials:false on
  the actions/checkout steps that never push.
2026-08-10 08:36:27 +02:00
MrMasterbay
ebc876c3e7 security(authz): per-VM authorization on 5 BOLA routes (Aikido)
The cluster-reach check (check_cluster_access) grants access via the pool/VM-ACL
fallback and expects a downstream per-VM gate — these 5 routes lacked it:
- clusters.py cancel_task (#469089252): parse the task's VMID from the UPID and
  require user_can_access_vm(vm.stop) — no cross-pool/tenant task cancellation.
- storage.py create/update backup job (#469089226): authorize every submitted
  VMID (vm.backup); cluster-wide (all=1) / pool jobs are admin-only.
- static_files.py vms-without-pool (#469089182): filter to accessible VMs.
- search.py cluster tags (#469089237): count only accessible VMs' tags.
- xhm.py migration detail/list (#469089253): require source-VM access, matching
  the plan/start gate.
Admins pass user_can_access_vm unchanged. +5 regression tests.
2026-08-10 08:31:47 +02:00
MrMasterbay
3e2879a681 security(authz): object-level gates on VMware + PBS write/delete/diagnose (Aikido BOLA)
The read routes already call check_vmware_access / check_pbs_access, but the
mutating ones were guarded only by @require_auth(perms=[...]):
- vmware.py update/delete/diagnose (#469089245/#207/#204): a vmware.config holder
  could take over, delete, or probe another tenant's ESXi server by id.
- pbs.py update/delete (#469089284/#244) and auto-storage (#469089213): a
  pbs.config holder could overwrite/delete another tenant's PBS record, or push
  its stored credentials onto an arbitrary cluster's storage config.
Each now calls check_* first (admins + unlinked servers unaffected); auto-storage
also check_cluster_access on every target cluster. +6 regression tests.
2026-08-10 08:18:08 +02:00
MrMasterbay
a07324be82 fix: audit follow-ups — token-scoped rate reads, i18n relative-time, arm64 CI gate, BMC tests
Low-severity items from the Testing-branch audit:
- power/cost list_rates: scope by the token's floored effective_role, not the
  owner's account role, so an admin-owned but scoped token can't enumerate every
  cluster's rates (mirrors the upsert admin check).
- i18n relative-time: 'timeAgo' was always appended ('5s vor' in German). Split
  into locale-ordered timeAgoSec/timeAgoMin templates for all 8 locales
  (de/fr/es/pt lead, en/zh/ko/it trail).
- release-images CI: arm64 builds under qemu emulation; a failed arm64 leg no
  longer blocks the amd64 release (continue-on-error on the arm64 matrix legs;
  publish already attaches whatever artifacts exist).
- tests: BMC in-band read key->agent->password fallback order (#609) now has
  direct coverage (3 tests).

(OIDC cross-source collision finding intentionally left alone — OIDC is hands-off.)
2026-08-09 21:14:51 +02:00
MrMasterbay
f335057d0a security(authz): scope cost-rate reads too (Aikido IDOR follow-up)
Same gap as the power-rate fix: GET /api/cost/rates listed every cluster's
cost rates to any authed user, and GET /api/cost/rates/<id> had no tenant gate.
Scope the list via get_user_clusters (keep __default__) and add the
check_cluster_access gate to the single-cluster read, mirroring api/power.py.
+ 3 regression tests.
2026-08-09 20:13:48 +02:00
MrMasterbay
965ec2ce6e security(authz): close 2 Aikido code-audit gaps (portal reboot perm + power-rates IDOR)
- client_portal _vm_power: reboot was gated on vm.start, so a start-only
  portal user could reboot (stop+start) a guest. Now gates on vm.restart,
  matching the main VM API and the scheduler. (Aikido #469089251)
- api/power.py list_rates (GET /api/power/rates): returned every cluster's
  power/cost rates to any authed user. Now scoped via get_user_clusters
  (admins / default tenant see all), always keeping the shared __default__
  fallback row. (Aikido #469089241)
- tests: tests/test_aikido_batch3.py (5) — scoped vs admin list + the
  reboot/start/shutdown permission mapping.
2026-08-09 20:00:23 +02:00
MrMasterbay
978f0a9701 Merge main into Testing: OIDC username-collision fix (#651, @rmalchow)
Brings rmalchow s #651 (keep the full preferred_username so bob@corp.com and
bob@partner.com no longer collapse onto one account) onto Testing. It had been
merged to main by mistake (a silent gh retarget failure), so merge main back in
to land it where our work actually ships and keep Testing a descendant of main.
2026-08-09 14:23:13 +02:00
MrMasterbay
63ce5de69d fix(v2p): refresh a stale ESXi session before the migration plan/start
On a long-lived server the ESXi REST/CIS session expires, and unlike every other
vmware route (get_vmware_vms, single-VM GET, snapshots, performance, ... all call
mgr.ensure_connected() first) the migration-plan and migrate-start handlers went
straight to get_vm(). get_vm() then hit /api/vcenter/vm/<id> on the dead session and
returned {error: VM not found}, so the wizard reported the VM missing and the
migration never started — i.e. "VMware migration stopped working" after the box had
been up a while, even though the VM list (which does reconnect) still showed it.

Add mgr.ensure_connected() at the top of get_vmware_migration_plan and
start_vmware_migration, mirroring the other routes. Reproduced + confirmed live
against a real ESXi 8.0.3 host: plan 400d on the stale session, then returned the
full plan once the session was refreshed; a full 90GB VM migrate then ran end-to-end.
2026-08-07 18:20:12 +02:00
MrMasterbay
77bfe71d98 security: enforce authz/validation gaps from the Aikido Testing-branch pentest (batch 2/2)
Second half of the adversarially-verified findings (batch 1 = 35d4078). Each fixed,
code-reviewed and covered by tests/test_aikido_batch2.py (17 new, full suite 401 green).

- power: only a global admin (effective_role) may overwrite the shared __default__
  power-rate row — a cluster.config holder edits only its own cluster
- metrics exporter: /api/metrics requires an admin-role token, not any valid token
  (it emits cluster-wide, cross-tenant infra gauges)
- portal: build_authz_user in _vm_power so a token's effective_role is honoured;
  invalidate the user's other sessions on portal password change
- cluster-groups: treat a global (tenant_id NULL) group as admin-only for the
  delete + balance-now writes, matching the earlier update fix
- vm-tags: reject a non-numeric vmid before the global DELETE+rewrite, and roll back
  save_vm_tags on error so a mid-loop failure can't persist a partial table wipe
- datacenter/multipath: allowlist the path_selector policy, and reject non-member
  nodes before SSH (no more `_get_node_ip(node) or node` fallback to a raw hostname)
- storage: pin http download-url fetches to the validated IP (DNS-rebind); https is
  left as the hostname since the node's TLS cert check already defeats a rebind
- multi-sdn: advertise an in-flight span's zone/controller so a concurrent purge
  can't tear down infra a create is still building (TOCTOU)
- ws-token validate: enforce node.shell for the standalone node SSH shell path
  (shell=node) — the VM termproxy path is unaffected
- SSE: scope vmware_vms / vmware_vm_detail to the server's linked_clusters instead
  of broadcasting guest_info/performance to every client; scope the portal audit
  task feed to the cluster it happened on (portal writers now set cluster=)
- LDAP: authoritative re-sync — rebuild LDAP-sourced perms/tenant_permissions from
  the current group mapping instead of only unioning them in, so group removal revokes
2026-08-07 17:31:57 +02:00
MrMasterbay
35d40781bd security: enforce authz gaps from the Aikido Testing-branch pentest (batch 1/2)
Adversarially verified all 45 AI-pentest findings against the current code; this batch
fixes the 10 HIGH + 2 MED confirmed REAL (17 were by-design/known admin-wide debt, 2
already fixed, 1 false positive). Full in-process test suite is green — the one case that
failed was a test that encoded the very BMC credential-exfil this now guards; corrected it
and added the negative case.

- rbac.py: user_can_access_vmware_vm honours effective_role (token-scoped), like its Proxmox
  twin, so an admin-owned viewer-scoped API token no longer gets the full-admin VMware bypass.
- nodes.py / vmware.py: reject a masked password paired with a CHANGED host - the stored BMC /
  VMware secret can no longer be shipped to a caller-chosen host (credential exfiltration).
- schedules.py: create/update enforce the same per-VM ACL as the live action (build_authz_user
  + user_can_access_vm), not just cluster reachability.
- pbs.py: restore (overwrite/test) enforces per-VM authorization on the destination VMID;
  backup-diff now requires pbs.datastore.view instead of plain pbs.view.
- users.py: create/update cap delegated permissions and custom roles to what the caller holds
  (a tenant admin.users delegate can no longer mint accounts with admin.* perms via the
  permissions[] list or the custom-role tier fallback); update_user rejects a foreign-tenant
  custom role; update_tenant is tenant-scoped like get_tenant_quota.
- groups.py: a global (tenant_id NULL) cluster group is admin-only for writes.
2026-08-07 15:31:13 +02:00
MrMasterbay
96a43f22a6 fix(migrations): route the ESXi + cross-hypervisor migration lists to their handlers (#654)
The @bp.route + @require_auth decorators for GET /api/vmware/migrations and
/api/xhm/migrations were attached to the _migration_reachable / _xhm_reachable IDOR
helpers instead of the list handlers — a copy-paste slip when those NS-Jul-2026 helpers
were inserted between the decorators and the intended functions. Flask therefore called
the helper with no args (TypeError: missing 't') and returned 500, so ESXi/XCP-ng
migration failed immediately and list_vmware_migrations / xhm_list were never routed.

Move the decorators back onto the real list handlers; the reachability/IDOR filter is
unchanged (still applied inside the handlers). Verified via url_map: /api/vmware/migrations
now maps to vmware.list_vmware_migrations and /api/xhm/migrations to xhm.xhm_list.

Confirmed live in #654 (miketate1985) and flagged by the CodeAnt full-repo sweep
(vmware.py:1130, xhm.py:194).
2026-08-06 21:30:06 +02:00
MrMasterbay
2ca4804872 fix(update): scheduled rolling updates evacuate local disks by default (#630)
Unattended scheduled updates now default allow_local_disks=True instead of only
following the cluster balance_local_disks flag — nobody is watching a 3am run, so
pausing it because one local-disk VM cannot live-migrate is the worst outcome. On
shared-storage clusters this is a no-op; a per-schedule override is honoured if set.

Also surface allow_local_disks + reboot_timeout in the manual run log summary so a
paused evacuation can be diagnosed from the log (was: only evacuation_timeout shown).
2026-08-06 08:22:16 +02:00
ruben
9bb78ebb7a oidc(#486): bind the legacy-username fallback to the same subject
The back-compat lookup added in the previous commit adopted an existing
account whenever its key matched the local part of the derived username. That
re-opened the very collision the commit set out to close, just one step later:
bob@partner.com logging in found the 'bob' row that bob@corp.com had
provisioned pre-change and inherited its role and tenant.

A legacy row is now only reused when it is provably the same identity -- its
auth_source is an OIDC-family one and its stored oidc_sub equals the incoming
sub claim. Local and LDAP rows are never adopted. A sub is only unique within
an issuer and no issuer is recorded on the user row, so installs that repoint
at a different IdP still have to rekey by hand; likewise a pre-change account
that never completed an OIDC login has no oidc_sub and now needs a manual
rekey rather than being silently adopted.

'+' joins the allowed char set. sanitize_username() permits it, and dropping
it merged identities the same way dropping '@' did -- a+b@example.com derived
onto a real ab@example.com account.

oidc_derive_username() takes the already-loaded users table. Both callers in
the login path had one in hand and were loading it a second time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:36:45 +02:00
ruben
3f00e828ec oidc(#486): keep the full preferred_username instead of truncating at '@'
Username derivation split preferred_username at '@', so bob@corp.com and
bob@partner.com both provisioned onto a single 'bob' account and whoever
logged in second inherited the first one's role and tenant.

Keep the whole claim. '@' also has to join the allowed char set, or dropping
the split just yields 'bobexample.com'. Nothing downstream constrains it:
sanitize_username() already permits '@', LDAP provisioning already produces
such names, username is an unconstrained TEXT PRIMARY KEY, and no username
reaches a filesystem path, shell command or interpolated SQL.

Existing accounts keep their current key via a back-compat lookup, so no
install loses permissions and nobody is locked out under auto_create_users:
false. Installs that already hold a truncated 'bob' therefore keep the
collision -- closing that needs a rekey migration, left out deliberately.

The derivation was duplicated in the auto_create_users pre-check, with a
comment recording that a previous divergence caused 403s. Both call sites now
share one helper.

Adds the first test coverage for OIDC username derivation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:09:14 +02:00
mkellermann97
bfb10272aa fix(maintenance): honour local-disk evacuation on the maintenance-mode paths (#629)
The anti-affinity fix threaded --with-local-disks through the balancer, but the
manual maintenance-mode toggle (api/vms.py) and the scheduled-update path
(api/schedules.py) still called enter_maintenance_mode() without allow_local_disks,
so local-disk VMs bailed out of a maintenance evacuation even with "Balance VMs
with Local Disks" enabled. Pass the same cluster balance_local_disks flag through
both call sites (the rolling-update path in settings.py already did). Off by
default, so no behaviour change unless the setting is on.
2026-08-05 09:42:35 +02:00
mkellermann97
62c3024c49 feat(console): SPICE console alongside noVNC (downloadable virt-viewer .vv)
Adds a SPICE option next to the VNC console for QEMU VMs — the same model PVEs
own web UI uses: a downloadable virt-viewer .vv connection file that opens in
remote-viewer, giving full SPICE (audio / USB redirection / multi-monitor) that
an in-browser client cannot.

- core/manager.py get_spice_ticket() gains a `proxy` arg (defaults to the cluster
  API host) so remote-viewer tunnels the SPICE stream through the PVE host´s
  pveproxy (:3128) — works behind a single public IP / when the node is not
  directly reachable.
- api/vms.py: new GET .../<vmid>/spice endpoint, same authz as the VNC console
  (vm.console + per-VM ACL). Calls spiceproxy and assembles the [virt-viewer] .vv
  (multi-line CA single-line-escaped), served as an application/x-virt-viewer
  attachment. Returns 409 with a clear hint when the VM has no SPICE display.
- dashboard.js: a "SPICE" action (qemu-only, running) that downloads the .vv.

Verified: backend py_compile + route registration + .vv-format unit test, and a
LIVE end-to-end run against PVE 9.2.3 — spun up a throwaway qxl VM, confirmed
spiceproxy returned a real SPICE endpoint (tls-port + password + CA, proxy via
the PVE host), our .vv came out well-formed, then destroyed the VM. Frontend
build clean. i18n uses English fallbacks for now (SPICE label is language-neutral).
2026-08-04 11:53:11 +02:00
mkellermann97
d54e6e0c32 fix(console): keep the VNC websocket alive through a transient hub stall
The VNC ws server runs as a monkey-patched greenlet on the shared gevent hub, so
a brief hub stall stops its asyncio loop from answering a keepalive pong in time —
and that same stall is what makes the browser´s SSE watchdog reconnect. With
ping_timeout=10 the console got dropped as a keepalive failure right when SSE
renewed, which read as "the VNC connection drops whenever SSE refreshes". Loosened
the keepalive to ping_interval=30 / ping_timeout=60 (both the primary and the IPv6
fallback listener) so ordinary hub jitter no longer kills a healthy session; TCP
and the client´s own reconnect still catch a genuinely dead peer. The anti-idle
ping that #92 added stays.
2026-08-03 17:18:36 +02:00
mkellermann97
6c945697e7 fix(console): stop intermittent "No access to cluster" on node SSH consoles
The ws-token validate resolved the caller by reloading the WHOLE users table
(load_users(), which decrypts every user´s TOTP). Under gevent that read can
transiently fail on SQLite/WAL contention and swallow the error to {}, at which
point get_user_clusters() no longer sees the admin/all-access sentinel and treats
the caller as a default-tenant viewer — so opening a node SSH console 403´d with
"No access to cluster <id>" every so often (and, on an unscoped default tenant,
the opposite: a fail-open allow). Resolve the token´s user by its indexed row
(get_db().get_user) and treat an unresolvable identity as a retryable 401, never
a cluster denial. A genuinely unauthorised user still resolves to a real dict and
still gets a real 403, so the cross-cluster BOLA check is unchanged.
2026-08-03 17:18:36 +02:00
mkellermann97
7b6bbc5e46 fix(repl): keep the intermediate clone on the source storage (#641)
Snapshot replication does an intermediate full-clone of the guest on the
SOURCE node, then migrates that clone to the target node with
targetstorage=target_storage. But the source clone was ALSO being handed
target_storage, which names a storage on the destination (or the "local-lvm"
fallback when none is picked). On a cross-node job that storage need not exist
on the source, so the clone died with e.g. "storage ´local-lvm´ is not
available on node ´pve-03´" — breaking replication for any guest not sitting
on a like-named storage (BTRFS setups especially).

Only pin the clone onto target_storage when it is the final replica, i.e. a
same-node run. Cross-node, the clone stays on the source volume´s own storage
(omitting storage makes PVE keep each volume where it is) and the existing
migrate step re-homes it to target_storage on the destination. Reported by
@ripperrd; owes a live 2-node retest.
2026-08-03 11:42:26 +02:00
mkellermann97
1309008e80 multi-sdn(#612): harden cross-cluster EVPN after a deep review
A close review of the #612 feature (backend + frontend + an in-process API run)
came back sound — auth matrix, per-member BOLA, input validation and SSRF surface
are all correct — but turned up a handful of low-severity gaps. Fixed:

Backend (api/multi_sdn.py):
- edit / reconcile / scan did a read-modify-write of per_cluster_status WITHOUT
  the per-vid lock and REPLACED it from a pre-fan-out snapshot, so a concurrent
  add/remove-member could lose a member entry (and edit could revert a concurrent
  membership change in desired_state). New _merge_status_write() takes the per-vid
  lock, re-reads inside it, MERGES the fresh results into the stored map, drops
  departed members, and preserves the current member_clusters. Self-healing before,
  consistent now.
- scan was gated on node.view but it overwrites the shared drift snapshot everyone
  sees; gated it like every other span writer (sdn.manage + admin.settings) + the
  member-access check.
- create takes a name-scoped lock + re-checks for an existing same-name record
  before INSERT, so two admins racing the same name can no longer end up with two
  aggregate records for one span (name IS the SDN vnet id; the fan-out is idempotent).

Frontend (vm_modals.js + dashboard.js):
- the whole EVPN feature + every mutation control was shown to node.view-only
  users who then 403 on everything; added a canManage prop (isAdmin || sdn.manage
  && admin.settings) that hides Create / Edit / Re-apply / Reconcile / Scan /
  Forget / Purge / add- and remove-member, leaving the read-only list. mcevpnReadOnly
  note in all 7 locales.
- the validate preview rendered a member whose pre-flight SDN read FAILED as green
  "ok" while the banner said "issues found"; added a p.error red branch.
- the auto-reconcile toggle reused the preposition t("on")/t("off") (no "off" key
  existed) → mixed-language label; switched to the proper t("enabled")/t("disabled").

Adds tests/test_multi_sdn.py — the first committed coverage for the feature: the
API auth/validation surface (list/validate/create-4xx/404/anon-401/viewer-403 on
every write incl. scan) + the _merge_status_write merge/lock behaviour. 18 pass.
Verified: py_compile, frontend build, adversarial re-review. Live 2-cluster EVPN
E2E still owed (no real EVPN clusters). Refs #612.
2026-08-01 10:04:58 +02:00
mkellermann97
8fac530dd7 updates(#630): carry a per-schedule reboot timeout into scheduled rolling updates
Scheduled rolling updates hardcoded the node reboot/online timeout at 600s, so
the Advanced-options value a user set only took effect for a MANUAL run. Now the
reboot timeout is a first-class schedule field end to end:

- schedules.py: the runner reads + clamps it (60s..2h, same bounds as the manual
  path) instead of the fixed 600; the live scheduler (check_scheduled_updates)
  forwards it into the run config — without that one line the stored value never
  reached the runner and the whole thing was a no-op.
- schema + migration: reboot_timeout column on all three update_schedules
  CREATE TABLE sites plus a PRAGMA-guarded ALTER so existing DBs get backfilled;
  both row->dict paths return it (KeyError-safe on pre-migration rows); INSERT
  and the create endpoint carry it through.
- frontend: reboot/online timeout selector in the schedule modal (disabled when
  the schedule does not reboot), wired through save/load; reuses the existing
  rebootTimeout i18n keys.

Verified: py_compile, frontend build, and a save->DB->load round-trip
(reboot_timeout=2400 survives). Refs #630.
2026-07-29 15:09:57 +02:00
mkellermann97
6a9f852a0a security: fix post-0.9.15 pentest findings (BMC cred exfil, ESXi pw-on-argv, BMC SSRF oracle, Ceph SSH storm)
- nodes.py: BMC test endpoint no longer pairs the stored BMC password with a
  caller-CHOSEN host — a masked/blank pw only falls back to the stored secret
  when the host under test IS the stored host; any other host must carry its
  own full password. Stops credential exfil via a preemptive Basic-auth GET to
  an attacker server.
- redfish.py: _validate_host re-rejects loopback / unspecified / link-local /
  metadata even though allow_private=True is needed for private BMC LANs, so the
  stored credential cannot be aimed back at the PegaProx host. Rejection reasons
  collapsed to a generic string (no resolved-IP / port oracle). Docstring fixed.
- v2p.py + xhm.py: feed the ESXi password over SSHPASS env (sshpass -e) instead
  of -p <pw> on argv, both for the local subprocess ssh/scp and the scp run on
  the PVE node (stdin-piped) — the secret was readable in /proc/<pid>/cmdline
  for the duration of a migration. (owes an ESXi/XCP-ng E2E retest)
- ceph.py: mirror views ran one SSH connect per image/pool; batch over a single
  connection (+ one verbose pool status) so a low-priv cluster.view user cannot
  spam N SSH sessions at a node per view load.
- multi_sdn.py: purge-delete now removes a members zone/controller only when no
  other span on that cluster still shares them, so a full delete cannot collapse
  a co-tenant cross-cluster EVPN span.
2026-07-25 19:18:40 +02:00
mkellermann97
46b1b5b636 fix(sdn): CodeAnt — atomic create-rollback + subnet-input robustness (#612)
Two minor CodeAnt findings on the cross-cluster EVPN code:
- atomic create-rollback tore down vnet/zone/controller wholesale on every
  member, which could delete a pre-existing controller/zone a co-tenant span on
  that cluster reused idempotently. Now tears down only what _apply_on_cluster
  reported it freshly created (result[created]) — same shared-object-safe teardown
  the P3 member-add path already uses.
- a malformed subnet entry in the request body (e.g. a bare number) hit .get() on
  a non-dict and 500ed; now returns a clean 400.

Verified: malformed-subnet -> 400 not 500; all five #612 in-process suites green.
2026-07-25 16:05:25 +02:00
mkellermann97
8030b16eca feat(sdn): cross-cluster EVPN vNets — Phase 3 (membership + drift alerts) (#612)
Two lifecycle extensions on top of Phase 1+2, per Nico's pick.

Membership expand/shrink — you can now grow or shrink a span without delete+
recreate:
- POST /api/multi-sdn/vnets/<id>/members {cluster_id}: collision + reachability +
  SDN-installed pre-flight on the NEW cluster (must share the span's ASN/VNI),
  gated on access to ALL members INCLUDING the new one, builds the same EVPN
  controller/zone/vnet on it, then appends it.
- DELETE .../<id>/members/<cid>[?purge=1]: drops a cluster from the span (keeps
  >=1 member). purge tears the vNet down on that cluster VNET-ONLY — it leaves the
  zone/controller, which another span on that cluster may share.

Drift → alerts — the Phase-2 scanner now emits an operator alert (the existing
push/inbox notification fan-out) on a healthy→drift transition. Edge-triggered
off the prior scan's status, so a vnet that stays drifted alerts ONCE, not every
6h. Fires in detect-only mode too (an alert is a notification, not a cluster
write).

Frontend: per-member remove buttons (shown only when >1 member; − = leave the
vNet, trash = purge) + an add-cluster picker (non-members only) in the expanded
view. i18n all 7 langs.

Adversarial review (3 dimensions x verify) → 3 findings fixed:
- Add-build failure now tears down EXACTLY the objects it freshly created
  (_apply_on_cluster reports ; a mid-build fail at the zone/controller
  step no longer orphans them) — while still never touching a pre-existing shared
  zone/controller.
- The member-list read-modify-write is now serialized under a per-vid lock with a
  re-read inside, so two concurrent admins adding/removing on the same span can't
  lose an update; the remove path also re-checks the >=1-member invariant inside
  the lock.
- The drift alert emits one notification PER drifted member (own cluster_id
  anchor) so every affected tenant's subscribers are woken, not just the first.

Verified in-process (no real same-ASN EVPN clusters — honest limit): 50 tests
green across all five #612 suites incl the new member add/remove, vnet-only purge,
targeted teardown, per-member alert, and edge-triggered no-spam; app boots + 11
routes + scanner. Real 2-cluster EVPN E2E still owed.
2026-07-25 15:07:57 +02:00
mkellermann97
abb4d40160 feat(sdn): cross-cluster EVPN vNets — Phase 2 (edit + drift) (#612)
Builds on Phase 1's create+read layer with the lifecycle + drift half PDM lacks.

Edit fan-out — PUT /api/multi-sdn/vnets/<id>: change alias + add/remove subnets,
fanned out to every member (PVE vnet-alias PUT, subnet POST/DELETE with the
<zone>-<cidr> id URL-encoded) + one apply per cluster. Structural fields
(name/zone/vni/asn/controller) are immutable — a change there rebuilds the span,
so it's rejected with a delete-and-recreate hint. Gated sdn.manage+admin.settings,
per-member check_cluster_access, partial-failure status like create.

Drift detect + reconcile:
- Background scanner (6h, daemon thread started in api/__init__, mirrors
  drift.py) reads live per-member SDN state for each aggregate vnet and persists
  the drift status. DETECT-ONLY by default — it never writes to a cluster unless
  the new global setting multi_sdn_drift_reconcile is opted in.
- Manual POST .../<id>/reconcile (re-assert desired + fix alias drift) and
  POST .../<id>/scan (read-only drift refresh, node.view) for on-demand use.
- When auto-reconcile is opted in, it fixes ONLY the drifted members (never
  reloads an in-sync member), debounces 'missing' (needs two consecutive
  non-healthy passes so a transient partial read can't resurrect intentionally-
  removed SDN), and writes a log_audit trail for the unattended mutation.

Frontend (MultiClusterEvpnView): edit modal (alias/subnets), per-cluster drift
detail, Scan / Reconcile / Edit actions, and an admin-only auto-reconcile toggle.
i18n all 7 langs.

Adversarial review (4 dimensions x verify, 2 workflows) → 7 findings fixed:
_apply_on_cluster now skips the cluster-wide apply when nothing changed (kills
the over-broad blast radius: one drifted member no longer reloads the whole
span, and reapply/reconcile of in-sync members is a no-op); reapply treats
in_sync as done; auto-reconcile is per-member + missing-debounced + audited; the
toggle is hidden from non-admins; edit subnet fold-back de-dupes; and the
pre-existing DUPLICATE it: translation block (silently dropping 317 Italian
keys) is merged into one — Italian is whole again.

Verified in-process (no real same-ASN EVPN clusters — honest limit): 44 tests
green across edit/reconcile/scan/scanner/gate/debounce/no-op-apply/subnet-dedup;
app boots + 9 routes + scanner thread; build + translations parse; auth/injection
review dimensions came back clean. Phase-2 auto-reconcile ships default-off.
2026-07-25 14:33:53 +02:00
mkellermann97
e764df644c feat(sdn): cross-cluster EVPN vNets — Phase 1 (create + read) (#612)
PVE has no cross-cluster SDN primitive: each cluster's /etc/pve/sdn is local. To
make one logical EVPN vNet span several clusters that share a BGP ASN, PegaProx
must create the same EVPN controller/zone/vnet on every member and apply each.
This adds the orchestration layer + the authoritative record PDM lacks. Phase 1
is create + read; edit/alias fan-out + a drift-detect scanner are Phase 2.

New api/multi_sdn.py composes the existing per-cluster SDN passthrough
(api/datacenter.py → PVE /cluster/sdn/*) across N members:
- POST /api/multi-sdn/vnets — validate → per-member collision + reachability
  pre-flight → bounded concurrent fan-out (controller→zone→vnet→subnets→apply,
  idempotent skip-if-exists) → authoritative record. Atomic by default: any
  member failing rolls back ALL members (best-effort, ignores does-not-exist) so
  no half-built L2 span is left behind.
- POST .../validate — dry pre-flight (reachability + collisions), no writes.
- POST .../<id>/apply — idempotent retry of not-yet-applied members.
- GET .../vnets, GET .../<id>?refresh=1 — list/detail + live per-cluster status.
- DELETE .../<id>[?purge=1] — forget the record (default) or also tear the SDN
  objects down on every member.

Collision pre-flight rejects VNI reuse, ASN mismatch, and zone/vnet redefinition,
while treating a same-definition object as idempotent. Writes gated on
sdn.manage + admin.settings (disruptive cluster-wide apply = blast radius);
reads on node.view. Every route gates check_cluster_access on ALL member
clusters (a caller who can't reach a member can't read/mutate it via the
aggregate); deny → 404 on read to avoid confirming existence.

New multi_cluster_vnets table (db.py) is the authoritative record (JSON columns
per the site_recovery_plans convention). Frontend: a new Cloud-level
'Multi-Cluster EVPN' view (sidebar route, ≥2 clusters) — list with member chips +
per-cluster status badges, a create wizard with validate/preview, re-apply and
delete/purge; i18n all 7 languages.

We orchestrate the SDN config objects only — the physical BGP-EVPN underlay
(inter-cluster peering) is the operator's network and is assumed to already peer.

Verified in-process (no real same-ASN EVPN clusters in the lab): 38 tests green —
validation, collision (idempotent/conflict/VNI-clash/ASN-mismatch/501), apply
ordering + bodies + idempotency + offline/failure, canonical-subnet idempotency,
rollback order, rollup, live-status; app boots + 6 routes register + DB table
created; frontend build + translations parse. Adversarial review (4 dimensions,
verified) fixed: atomic rollback now covers the failed member too, subnet
idempotency uses canonical-network equality, topology sidebar state reset.

Real end-to-end validation needs 2+ clusters sharing an ASN with a live EVPN
underlay (lab has one) — an honest limitation, same as #546.
2026-07-25 14:00:24 +02:00
mkellermann97
ce4c83ba3b feat(v2p): opt-in 'wait for confirmation before switchover' (#562)
Adds an optional pre-cutover hold to the ESXi→Proxmox migration: when enabled
in the wizard, the migration stages the disks and then parks in a new
'awaiting_confirmation' phase — with the source VM still running — and waits for
the operator to commit (or cancel) the switchover. This lets the operator
schedule the few seconds of downtime themselves (drain monitoring, flip DNS,
then commit) instead of it landing whenever the copy happens to finish, per
ajoergensen's request.

Backend (core/v2p.py): V2PMigrationTask.await_cutover_confirmation() gate,
inserted before every set_phase('cutover') seam; it's a no-op when the flag is
off (so existing migrations are byte-identical) and only holds while the source
is still running — on modes that already stopped the source (offline / sshfs-
boot / VM-was-off) it logs-and-commits rather than extending downtime. Confirm/
cancel are flags the migration thread polls; cancel (or a generous timeout,
default 24h) raises V2PCutoverCancelled → clean 'cancelled' status with the
source left running and the staging mount/snapshot torn down. The confirmation
wait is never billed as downtime (downtime clock reset on confirm; the primary
live modes already measure from cut_t0 after the gate).

API (api/vmware.py): POST .../migrations/<mid>/confirm-cutover and
.../cancel-cutover (409 unless parked at the gate); to_dict() exposes
wait_for_confirmation + awaiting_confirmation; wait_for_confirmation /
confirmation_timeout accepted on the migrate body.

Frontend (dashboard.js): wizard checkbox + a commit/cancel action on migrations
holding at the gate (full Migrations list + compact per-VM view), new phase/
status styling; i18n for all 7 languages.

In-process verified: gate state machine (park→confirm→proceed, cancel + timeout
→ V2PCutoverCancelled, no-op when off, source-stopped skip, downtime reset),
app boot + both routes registered, frontend build + translations parse.
2026-07-25 13:30:27 +02:00
mkellermann97
6c20c11571 fix(lb): expose proxlb_tags_enabled in GET /clusters so the toggle sticks (#628)
The ProxLB VM Tags toggle saved fine (proxlb_tags_enabled is in
ALLOWED_CONFIG_FIELDS) but the GET /clusters response never returned it, so on
the next refresh the UI reloaded the old value and the toggle snapped back. Add
it to both cluster response builders. Cosmetic — no behaviour change.
2026-07-25 11:23:37 +02:00
MrMasterbay
1cea490c5a security(ssh): harden host-key verification (CodeAnt follow-up)
Five findings from a CodeAnt pass over the SSH host-key work:

- fail CLOSED in verify_transport_host_key: if get_remote_server_key()
  raises we now reject instead of silently proceeding to auth unverified.
  Every caller already wraps the call in its connect try/except, so the
  raise is handled exactly like a changed key.
- port-aware pinning: a non-standard SSH port is keyed as [host]:port in
  known_hosts (port 22 stays the bare host). Previously a non-22 host was
  treated as unknown on every connect and never actually pinned. ssh_pool
  and xhm now pass their real port.
- IPv6 host parsing in remove_host_keys: split(':')[0] truncated IPv6
  addresses at the first colon; a dedicated token parser handles bare and
  bracketed IPv6, bracketed IPv4:port, and comma host-lists.
- standalone WS server: serialize known_hosts writes with a lock so
  concurrent TOFU saves can't corrupt the trust file.
- standalone WS server: warn loudly instead of silently falling back to a
  cwd-relative known_hosts path when PEGAPROX_SSH_KNOWN_HOSTS is unset.

Unit-tested (fail-closed / port pin+reject / IPv6 removal) and live E2E'd
against real ESXi (pin, accept, downgrade-guard + same-type reject-on-change).
.ssh_ws_server.py re-synced from the embedded string in vms.py.
2026-07-21 08:26:33 +02:00
MrMasterbay
aadd53a2e1 fix(ssh): drop pinned host keys when a cluster is removed
Follow-up to the TOFU host-key work: removing a cluster only deleted its DB rows,
leaving the pinned SSH host keys in known_hosts. Re-adding the SAME running node
worked (key still matches), but re-adding a node that had been REINSTALLED in the
meantime presented a new host key and tripped reject-on-change — the reconnect
failed until someone hand-edited known_hosts.

delete_cluster now collects the cluster host + node IPs (manager host, the
node-IP cache, and get_nodes() where reachable) and calls the new
ssh_security.remove_host_keys() before stopping the manager, so a later re-add
re-pins the current key via TOFU regardless of a reinstall.

Verified end-to-end against a live PVE node: add -> pin, simulate reinstall
(stale key blocks connect), delete -> pin removed, re-add -> fresh TOFU connects;
and the baseline remove+re-add of a live (non-reinstalled) node stays green.
2026-07-21 06:56:40 +02:00
MrMasterbay
ee0ac186af security(ssh): real TOFU host-key verification for all paramiko SSH paths
Aikido flagged 10 CRITICAL 'disabled SSH host key verification' findings: every
paramiko connection used AutoAddPolicy/WarningPolicy, which accept an unknown host
key silently — an attacker between the hub and a node (PVE/PBS/ESXi/XCP-ng/storage)
could impersonate it and capture credentials and commands. Worse, the historical
'TOFU' never actually worked: it wrote known_hosts to pegaprox/config/... which does
not exist at runtime, so no key was ever persisted and a changed key was never
rejected — effectively blind-trust on every connection.

New central module pegaprox/utils/ssh_security.py implements trust-on-first-use with
reject-on-change and an opt-in strict mode (PEGAPROX_SSH_STRICT_HOST_KEYS):
- apply_host_key_policy(): loads known_hosts + a custom verifying MissingHostKeyPolicy
  for client.connect paths.
- verify_transport_host_key(): keyboard-interactive auth runs over a manually-built
  paramiko.Transport, which BYPASSES the SSHClient policy entirely — so these paths
  had no verification at all. Now the server key is checked (before any credential is
  sent) on every Transport path; a known host that presents a changed key OR an
  unpinned key type is rejected (blocks key-change + keytype-downgrade MitM).
- known_hosts now resolves to the REAL runtime config dir (<cwd>/config), shared by
  every path including the system-ssh sshpass fallback (StrictHostKeyChecking
  accept-new, upgraded to yes under strict mode) and the standalone terminal WS
  subprocess (inline policy, path handed in via env).

Migrated every paramiko callsite: utils/ssh.py (3 clients + 2 KI transports + sshpass
fallback), utils/ssh_pool.py, utils/vnc_tunnel.py (client; the local forward's
server-side Transport is intentionally left as-is), core/xhm.py, core/pbs.py,
core/xcpng.py, core/manager.py, api/storage.py, api/vms.py (6 client paths, each now
persisting; the remove-node cleanup deliberately does not pin a departing node) and
the embedded ssh-ws server_script (inline TOFU).

Strictly E2E-tested against a live PVE node (192.168.1.2): TOFU first-connect pins
the key; second connect verifies; a changed key (same keytype) is rejected; an
unpinned-keytype downgrade is rejected; strict mode rejects unknown / allows known;
recovery after strict-off; the app boots with the new code; the standalone WS script
starts clean (no NameError). Adversarial-reviewed.

Scope: this covers the paramiko paths (the Aikido criticals). A separate follow-up
hardens the subprocess ssh/scp fleet (StrictHostKeyChecking=no / UserKnownHostsFile=
/dev/null in v2p + HA-fence + sshfs paths), which shares the same known_hosts file.
2026-07-21 01:06:22 +02:00
MrMasterbay
1e33f09866 fix(overview): all-clusters overview + cluster-detail now agree on CPU & guest counts
The all-clusters overview and the cluster-detail summary cards disagreed:
- CPU: overview used core-weighted cluster CPU (Σ cpu×cores / Σ cores, from
  datacenter/status); detail used a plain per-node average -> e.g. 6% vs 7.6%.
- Guest counts: overview total = running+stopped from the polled guests summary,
  which drops paused/suspended guests and could transiently render an impossible
  '16 / 14' (total < running); the detail counted all qemu+lxc -> '16 / 30'.

Fixes:
- backend datacenter/status: add explicit per-type guest totals
  (guests.vms.total / guests.containers.total = every guest of that type, any
  status incl. templates) so the overview denominator equals the detail's
  allVms.length and can never be < running. (proxmox + xcpng paths)
- overview (AllClustersOverview + the cluster-group variant): use the explicit
  total when present, fall back to running+stopped for older backends.
- detail summary cards: read the core-weighted CPU from datacenter/status (same
  source as the overview), falling back to the per-node average until it loads.

Verified live vs the Testi cluster: overview 16/30 == detail 16/30, and both CPU
values now derive from a single source. Rebuilt web/index.html.
2026-07-18 09:45:17 +02:00
MrMasterbay
fc92dd5b47 repl: wire the ZFS branch into the incremental engine (#174)
The incremental engine was RBD-only; the primitive already had zfs_replicate_
dataset (send / send -i / recv) but nothing dispatched to it. Generalise
_execute_replication_incremental: eligibility now accepts a VM whose disks are
all one incremental-capable type — 'rbd' OR 'zfspool' — and requires the target
storage to be that same type (can't diff rbd -> zfs). Per disk it dispatches to
rbd_replicate_disk (Ceph pool/image) or zfs_replicate_dataset (<pool>/<vol>
dataset), and prunes the snapshot chain with the matching helper. Adds
zfs_prune_snapshots + _xcincr_zfs_pool.

Live-verified end-to-end against a file-backed ZFS pool on two nodes (two
managers, source pve1 -> relay -> replica pve2, target storage zfspool): seed
via `zfs send | zfs recv -F` created + adopted the replica dataset (zvol md5
src==replica), then `zfs send -i base | zfs recv` shipped only the delta and
fast-forwarded the replica (zvol md5 src==replica). RBD path unchanged and
still green. Non-eligible / mixed-storage VMs still fall back to the full
clone+migrate path.
2026-07-17 15:53:33 +02:00
MrMasterbay
b349bb04a3 repl: wire incremental RBD replication into the engine + UI (#174 phase 2)
Builds on the tested transfer primitive (core/incremental_repl) to make
cross-cluster replication ship only the snapshot delta for RBD VMs instead
of full-cloning + remote-migrating the whole disk every cycle (#174 aderumier).

Backend (api/vms.py):
- new _execute_replication_incremental(job): eligibility-gate (every disk rbd
  on source AND target storage rbd), guest-snapshot the source, replicate each
  disk's delta via the byte-relay primitive, build/tag the replica VM on the
  first run (adopts the seeded rbd images), advance + prune the snapshot chain,
  update last_snapshot. A stale replica is torn down first behind the same
  tag-safety gate the full path uses (#413), so a mis-set target VMID never
  nukes a bystander. Returns False for non-eligible VMs so _execute_replication
  transparently falls back to the proven full clone+migrate flow.
- opt-in `mode` on the job; schema adds mode + last_snapshot columns (db.py).

Frontend (datacenter.js + translations.js): transfer-mode selector
(Full / Incremental) in the replication dialog + i18n.

Live-verified end-to-end against a real Ceph RBD cluster (source pve1 -> relay
-> replica on pve2): seed created + tagged the replica VM with the adopted disk
(md5 src==replica), then an incremental run after a 50 MB change fast-forwarded
the replica in-place (md5 src==replica) and pruned the old base snapshot; the
tag-safety gate correctly refused an untagged same-VMID VM. Full cross-cluster
(two separate clusters) + the ZFS branch still want a 2-cluster / ZFS lab.
2026-07-17 15:28:18 +02:00
MrMasterbay
969e94294f updates: opt-in Ceph-health gate for rolling updates (#403 part 2)
proxforge asked that a rolling update not keep pulling nodes when Ceph
isn't healthy — on an HCI cluster that can drop data below min_size and
take it offline. Part 1 (the "ceph installed but not deployed -> unknown"
misdetection) already shipped in v0.9.13.3; this is part 2.

New per-run setting `ceph_health_gate`:
- off       -> current behaviour (warn after the wait, then continue)
- degraded  -> HOLD only on genuine data-at-risk: HEALTH_ERR, unknown
              (deployed-but-unreachable), OSDs down, or degraded/undersized/
              incomplete/inactive PGs. Benign HEALTH_WARN (clock skew, a
              leftover flag, backfill/remapped) is tolerated.
- strict    -> HOLD on anything that isn't HEALTH_OK.

When the gate holds, the rolling update pauses with paused_reason
'ceph_unhealthy' + a reason/message, reusing the existing pause/resume
mechanism — the operator restores Ceph, then Continue or Cancel. No effect
on clusters without Ceph (get_ceph_health_summary returns None -> skipped).

- api/settings.py: _ceph_gate_unsafe() helper + parse/validate/store the
  setting + the hold branch in the rolling-update loop.
- security.js: gate selector in the rolling-update settings + start-body
  field; the generic paused-state UI already renders the hold + buttons.
- 6 i18n keys across all languages.

Gate decision unit-tested across HEALTH_OK/ERR/unknown/WARN(noout,
clock-skew,degraded,osd-down,backfill) x off/degraded/strict.
2026-07-17 13:33:56 +02:00
MrMasterbay
9f755196c9 ceph: fix OSD/CephFS/mgr endpoints + build the Ceph create-UI
Backend (api/ceph.py, core/manager.py):
- OSD create/destroy use a fresh password-based root@pam ticket when the
  cluster is API-token-authed (PVE hard-gates OSD ops to root@pam, never
  a token). New manager.create_privileged_session() guards the inline-token
  and non-root@pam cases with a clear message instead of a failed login.
- add POST/DELETE /ceph/mgr/<id> (was GET-only, so a manager could never
  be added or removed from the UI).
- CephFS create forwards to /ceph/fs/<name> (it POSTed the collection URL,
  which PVE has no handler for -> every create returned 501). destroy now
  plumbs remove-storages / remove-pools.
- mirror snapshot-schedule list returns [] on a non-mirrored pool (was a
  500), while still surfacing genuine 501/503 SSH/tooling errors.
- request.get_json(silent=True) across the blueprint so an empty POST body
  no longer raises a raw werkzeug 400/500.

Frontend (datacenter.js, translations.js):
- OSD-create modal with a node selector + disk picker (GET .../disks;
  in-use / existing-OSD disks are shown but disabled).
- Managers section in the Monitors tab (add/remove) + fix the MGR count
  (empty list was truthy -> always reported 1).
- Create-MDS and Create-CephFS modals in the CephFS tab.
- pool and CephFS delete buttons now send confirm_name (+ remove_storages;
  CephFS asks separately before removing the data/metadata pools) - the
  old empty DELETE always failed the backend confirmation gate with 400.
- 18 new i18n keys across all languages.

Verified end-to-end against a live 3-node Ceph cluster (HEALTH_OK).
2026-07-17 13:23:08 +02:00
MrMasterbay
591b6d2770 feat(hardware): out-of-band Redfish monitoring (#609 phase 3)
The credentialed, out-of-band counterpart to in-band ipmitool — reads BMC health
over the management network via the DMTF Redfish API, as a fallback when in-band
is unavailable. A SEPARATE, sharper opt-in with an enforced 5-second delay.

* core/redfish.py — read-only Redfish reader. Pure parsers (thermal/power/system/
  SEL) normalize to the SAME shape as read_node_bmc_inband, so the panel + cache +
  rollup + alerting work unchanged. SSRF-guarded, GET-only, redirects refused.
* core/db.py — node_bmc_endpoints table; BMC password stored encrypted (aes256:),
  masked (********) on every API response, never wiped by a masked re-save.
* core/bmc.py — REDFISH_CONSENT (v1, require_delay_seconds=5) in all 7 languages.
* api/nodes.py — redfish-consent GET/POST + per-node BMC endpoint GET/POST/DELETE/
  test (admin.settings, consent-gated, SSRF-validated host). The per-node read +
  cluster rollup now gate on (in-band OR redfish); the 5-min collector falls back
  to Redfish for nodes with a configured endpoint (own ~1h backoff).
* Frontend — Redfish sub-panel in the node Hardware tab: enable via a 5-second-
  delayed warning modal, then a BMC endpoint form (host/user/password/verify_ssl +
  Save/Test/Remove). i18n across all 7 languages.

Security review (28 agents, 17 positive confirmations) — fixed the 4 confirmed:
* CRITICAL SSRF — a hostile BMC's @odata.id could userinfo-splice the credentialed
  GET onto an internal host (https://<base>@169.254.169.254/...), exfiltrating the
  stored BMC creds. Every follow-on hop is now urljoin'd against the validated base
  and its scheme/host/port asserted equal; @/scheme/protocol-relative refs rejected.
* HIGH — config-restore could flip the new redfish consent without the audited ack;
  the restore-strip now protects BOTH consent keys.
* MEDIUM — malicious BMC could OOM the reader: bodies now streamed + capped at 4 MiB.
* MEDIUM — the cluster rollup route 403'd a Redfish-only deployment; now gated on
  (in-band OR redfish).
Residual (noted, not fixed): DNS-rebind TOCTOU on the base host (IP-pinning not done
autonomously per the codebase's standing SNI-pinning decision).

Tests: +33 (Redfish parsers, SSRF origin-pinning + body-cap regressions, consent +
BMC endpoint CRUD/SSRF/masking, rollup dual-gate). Full suite 347 passed.
2026-07-16 20:53:26 +02:00
MrMasterbay
3aac5211c5 feat(hardware): cluster degraded-hardware rollup + alerting (#609 phase 2)
Cluster-wide in-band BMC health, surfaced + alertable — mirrors the proven #601
temperature pipeline (5-min collector populates a per-manager cache; the 60s alert
loop only READS it, never SSHes).

Backend:
* manager.py: _node_hw_cache/_lock/_backoff + get_cached_node_hardware() and
  get_cluster_hw_rollup() -> {health, available, checked, counts, degraded[]}.
* metrics.py: _node_hw_summary() SSH probe (~1h backoff for nodes without
  ipmitool/BMC), populated in the 5-min collector via run_per_node(cap 8, 90s),
  GATED on the compliance consent + proxmox-only. Off the hot-path.
* alerts.py: 'hardware_health' alert metric (cluster worst + per-node), ok/warning/
  critical mapped to 0/1/2, auto-severity, cache-only.
* nodes.py: GET /clusters/<cid>/hardware/health rollup (consent-gated, cache-only).
* vms.py: compact hardware rollup injected into datacenter/status (cache-only,
  consent-gated) so the overview badge is free.

Frontend: 'hardware_health' alert metric with a Warning-or-worse / Critical-only
threshold selector; degraded-hardware badge in the corporate sidebar, the warning
banner, and the cloud overview. i18n across all 7 languages.

Adversarial review (12 agents): 3 rejected (bounded/cosmetic), 5 confirmed & fixed:
* MED — hwHealth prop was missing at the 2nd (ungrouped) ClusterSidebarItem call
  site, so the sidebar badge never showed for ungrouped clusters — now passed.
* LOW — the rollup endpoint 500'd on non-proxmox clusters — proxmox/callable guard
  now degrades to an empty 'unknown' rollup.
* LOW — a '<' operator on hardware_health builds a silent no-alert rule — the UI now
  only offers '>' and the create/update API pins hardware_health to '>'.
* NIT — alerts now show OK/WARNING/CRITICAL instead of the raw 0/1/2 code.
* NIT — corrected a stale '>=' code comment.

Tests: +11 (rollup endpoint gates, non-proxmox graceful, operator coercion, manager
rollup/cache unit tests). Full suite 322 passed; frontend build clean.
2026-07-16 00:46:25 +02:00
MrMasterbay
ab75a274cc feat(hardware): version the in-band consent warning in all 7 languages (#609 phase 2)
The compliance warning text is now localized (en/de/fr/es/pt/ko/it) while the
VERSION stays language-spanning: _HW_CONSENT_TEXT holds the same v1 warning in
each language, and hw_consent_warning(lang) injects the shared version +
require_delay_seconds so they can never drift per-language (bump all together).

* GET /api/hardware-monitoring/consent?lang=<code> returns the warning in the
  requested language (falls back to English). HW_CONSENT_WARNING kept as an
  English alias for existing callers/tests.
* The acknowledged LANGUAGE is now recorded alongside the version — set consent
  stores ack_lang and the audit line reads 'acknowledged compliance warning v1
  [de]', so the non-repudiation record captures which exact text the user saw.
* Frontend passes the current UI language to the consent GET (?lang=) and the
  enable POST (lang=), and re-fetches the localized warning on language switch.

Translations adversarially verified by 6 native-level reviewers vs the English
source (control IDs CMMC/NIST 800-171 3.4.6+3.1.5/DISA STIG preserved verbatim in
every language). Applied the confirmed fixes: es 'encendido'->'alimentación' (the
operation-list 'power' means power-control, not power-on), es informar+de grammar,
de closing-quote typography, ko install-vs-may modal + 'override' precision.

Tests: +4 integration cases (per-lang warning, unknown-lang fallback, ack_lang
recorded in audit + status, unknown-lang recorded as en). Full suite 311 passed.
2026-07-16 00:13:04 +02:00
MrMasterbay
0d04b4ee79 feat(hardware): ipmitool install route + node Hardware panel (#609 phase 1 step 3-4)
Step 3 — one-click ipmitool install:
* POST /api/clusters/<cid>/hardware/ipmitool/install (admin.settings) mirrors the
  proven StarWind installer but with a FULLY STATIC script (no user-controlled shell
  input), gated: cluster-access -> proxmox-only -> consent-required (the install
  mutates the node, so it must not run before the warning is acknowledged) -> per-node
  name validation -> run_per_node fan-out (cap 8, idempotent) -> audited.

Step 4 — frontend:
* Self-contained HardwareMonitoringPanel: consent fetch, mandatory compliance-warning
  modal (renders the server-versioned warning text, required ack checkbox, generic
  require_delay_seconds countdown ready for the Redfish phase), ipmitool install button
  when missing, sensor/power/FRU/SEL rollup with health badge.
* Wired as a 'Hardware' tab in BOTH node UIs (corporate CorporateNodeDetailView +
  default NodeModal) for parity. 21 i18n keys across all 7 languages.

Hardening from the adversarial review (3 confirmed findings, all fixed):
* MEDIUM — config-restore could enable hardware monitoring WITHOUT the versioned
  acknowledgement audit record (non-repudiation bypass): restore now drops the
  hardware_monitoring key and preserves the live consent state, so consent only ever
  moves through the audit-logged set_hw_monitoring_consent path (settings.py).
* LOW — a non-int stored/posted ack_version raised int() -> 500 on every gate; add
  _hw_int() so the gate fails CLOSED (disabled) instead of crashing.
* NIT — a crafted {"nodes":[123]}/[{...}] blew up _reject_bad_node's regex / set()
  -> 500; validate the raw list + reject non-strings centrally -> clean 400.

Tests: +9 integration cases (install gates admin/consent/proxmox/bad-node + the 3
hardening regressions). Frontend build clean (Babel). Full suite 307 passed.
2026-07-15 23:59:42 +02:00
MrMasterbay
75165738b2 feat(hardware): in-band BMC API route + audited consent gate (#609 phase 1)
Step 1-2 of the #609 in-band hardware-monitoring phase:

* GET  /api/clusters/<cid>/nodes/<node>/hardware  (node.view) — serves the
  parsed in-band sensor/power/FRU/SEL rollup from pegaprox/core/bmc.py, but
  refuses with CONSENT_REQUIRED (403) until the feature is enabled.
* GET  /api/hardware-monitoring/consent  (node.view) — status + the versioned
  compliance warning the UI must present.
* POST /api/hardware-monitoring/consent  (admin.settings) — enable/disable.
  Enabling requires acknowledge=true AND ack_version == the current warning
  version (stale-UI opt-in protection); the acknowledgement is persisted to
  the audit log with who/when/which-version, so opt-in is non-repudiable.

Warning text + version live in bmc.py (single source) so the audit record can
reference an exact version; bumping the version re-prompts.

Tests: 11 full-stack integration cases (consent gate 403->200, admin-only
enable, version-mismatch 400, cross-tenant 403, bad-node 400, audit-record
assertion, disable round-trip). Suite 298 passed.
2026-07-15 23:26:54 +02:00
MrMasterbay
6940ebb3db feat(monitoring): temperature history + temperature alert metric (#601)
Live per-host temperature already shipped via the lm-sensors Sensors panel.
This adds the two pieces #601 asked for on top:

History / chart over time:
- metrics.py collector now records each online node hottest lm-sensors temp
  into the persisted 5-min metrics_history snapshot. SSH fan-out uses the
  SSH-aware run_per_node (max_concurrent=8/cluster) so we never open 30+
  simultaneous sessions; a 1h per-node backoff skips hosts without lm-sensors,
  and the per-manager temp caches are pruned to known nodes.
- new GET /clusters/<id>/nodes/<node>/temperature-history (off-hub, cached).
- node Sensors panel renders a temperature-over-time sparkline + min/max/avg.

Alerting:
- new "temperature" alert metric for node (this node) and cluster (hottest
  node) targets, read from a manager temp-cache the 60s alert loop never SSHes.
  Alerts render in °C (not %); auto-severity uses absolute thresholds
  (>=85 crit, >=75 warn). Alert form gains the metric + a dynamic °C/% unit.
- i18n for all 7 languages.
2026-07-13 23:58:56 +02:00
MrMasterbay
f5b970a94b security(idor): scope vmware/xhm migration list+detail to the caller's clusters
Final re-scan residual: /api/vmware/migrations and /api/xhm/migrations (list + detail) returned
every tenant's migration tasks to any vm.migrate holder. Now filtered/gated so a caller only sees
a migration touching a cluster they can reach (source or target); detail returns 404 otherwise. 277 tests.
2026-07-13 22:16:51 +02:00
MrMasterbay
b4a72a7395 security: re-scan tail — CSRF/http-splitting/CSP, 2 SSRF, SIEM secret masking, power/drift/schedules IDOR
Final batch of the CodeAnt re-scan (all adversarially verified):
- app.py CSRF: the check ran only for JSON/form bodies (if sensitive), so a cross-site
  enctype=text/plain form POST (a browser 'simple request') skipped it — now enforced for every
  state-changing non-exempt /api/*.
- app.py http-response-splitting (x2 redirect handlers): the untrusted request Host was reflected
  into the Location header; now stripped/charset-rejected before use (a configured domain always wins).
- app.py CSP: dropped 'unsafe-eval' from script-src (Babel is pre-compiled, never runs in-browser).
- SSRF: plugins/notifications _send_apprise (prefix blocklist missed decimal/IPv6/metadata) and
  nodes._safe_repo_url (root-run bash curl) now go through the url_security guard.
- siem._row_to_target masks secret settings keys (token/password/api_key/secret/authorization)
  so a siem.view holder can't read the raw credential back out.
- IDOR: power rate routes (get/upsert/delete, __default__ skipped), drift.acknowledge_event
  (gate on the event's cluster), schedules.get_schedules (was fail-open on empty clusters field ->
  now get_user_clusters).

277 passing. Residual (LOW, follow-up): vmware/xhm migration-list per-task cluster filter.
2026-07-13 22:10:56 +02:00
MrMasterbay
76ca96252f security(authz): re-scan batch — VMware/PBS/cluster tenant gates + tenant-branch priv-esc
CodeAnt re-scan hunt (adversarially verified) surfaced a large second wave of missing tenant
gates + a priv-esc the first static_files fix missed:

- static_files.set_user_perms TENANT branch: a non-global-admin could set role='admin' (stored
  unvalidated) or admin.* perms in their OWN tenant_permissions, which resolve to the target's
  EFFECTIVE GLOBAL perms (has_permission runs with no tenant_id) => tenant->global priv-esc.
  Now rejected for non-global-admins.
- new helpers.check_vmware_access (mirrors check_pbs_access) + applied to 21 vmware.py read/
  action routes (get_vmware_vms/detail/hosts/datastores/networks/clusters/drs/ha/... ) that had
  only a role perm and never scoped to the server's linked_clusters => cross-tenant ESXi BOLA.
- 14 more PBS routes gated with check_pbs_access (syslog/rrd/network/dns/time/health/forecast/
  subscription/traffic-control/notifications) — the first sweep missed this route set.
- 13 cluster-scoped routes across pbs/users/clusters/settings/nodes now enforce check_cluster_access
  (run_backup_job_now, get_security_audit, rotate_cluster_api_token, reconfigure/export/update
  cluster config, rolling-update cancel/resume/clear, get_node_dns) — were role-perm only.

Regression tests: VMware cross-tenant route => 403; tenant-admin admin-role/admin.* amplification
=> 403 (both vectors). 279 passing. Remaining re-scan items (power/drift/schedules IDOR, app.py
http-splitting/CSRF/CSP, 2 SSRF, SIEM secret masking) in follow-up commits.
2026-07-13 22:02:46 +02:00
MrMasterbay
cf8b6f16e2 security(authz): fix 3 CodeAnt re-scan criticals (global-perm priv-esc, push fail-open, VMware BOLA)
Second CodeAnt exploitation scan (on the current tip) surfaced 3 critical auth findings, each
verified real and fixed:

- static_files.set_user_perms: the GLOBAL-permissions branch (no tenant_id) only required
  admin.users — which a tenant-scoped admin can hold — so a tenant admin could set a user's
  GLOBAL permissions and grant themselves/anyone global-admin-equivalent perms (priv-esc). The
  per-tenant branch already gated non-global-admins; the global branch now does too (requires
  session role == ROLE_ADMIN).
- push._alert_handler: the cross-tenant scope I added last cycle FAILED OPEN when a subscriber's
  user record was missing/deleted (still in push_subscriptions) — get_user_clusters({}) => None
  => treated as all-cluster admin => received every tenant's alerts. Now fails CLOSED on a
  missing record (and on lookup error). Good catch by the re-scan on my own fix.
- rbac.user_can_access_vmware_vm: the role-permission fallback had NO tenant isolation (the
  Proxmox user_can_access_vm has the equivalent guard) — any vmware.vm.* holder could reach
  every VMware server's VMs cross-tenant. Now gated by the server's linked_clusters
  (admin/unlinked open, else require get_user_clusters overlap), mirroring check_pbs_access.

Regression tests: tenant-admin global-perm PUT => 403 (global admin => 200); ghost/deleted
subscriber gets no cluster alert; VMware cross-tenant => denied, admin/unlinked => allowed.
274 passing. (The ~40 locked re-scan findings are under a parallel adversarial hunt.)
2026-07-13 21:41:37 +02:00