AuthProvider.checkSession and LoginScreen both run on mount. On the OIDC callback
page the session doesn't exist yet, so checkSession's /auth/check returns 401 and
it called logout() — which raced the in-flight callback POST and aborted it
(Firefox nginx-499 / 'Network error during OIDC callback'; Chrome won the timing).
checkSession now stands down when the URL carries the OAuth code+state (the same
signal LoginScreen keys off), leaving the callback handler to own the transition;
it reloads on success and this runs cleanly against the new session.
The in-app updater and update.sh are a git-tree + pip file update that only fits
the source/deploy.sh layout. On an apt/dpkg install the correct path is apt
upgrade (deps are system python3-* packages; a pip run diverges from dpkg and
can't lift the dpkg-owned crypto libs -> fail-closed TLS on restart); on Docker
a file update is discarded at the next image pull.
- settings.py: _detect_install_method() (docker via /.dockerenv + cgroup, apt via
dpkg -S, else source); perform_pegaprox_update refuses on apt/docker with the
right guidance + copy-paste command (409, allow_managed override); check-update
reports install_method / in_app_update_supported / managed_update_{hint,command}.
- update.sh: same detection + guard before touching anything (--force override).
- settings_modal.js: the Install button becomes a package-manager / image hint
with a copyable command on apt/docker; performUpdate handles the 409.
- tests: 5 guard/reporting tests.
Bump PEGAPROX_VERSION + PEGAPROX_BUILD (constants.py, constants.js) and
version.json (version / build / release_date) to 1.0.1 / 2026.08.09, and add
the 1.0.1 changelog entry.
version.json update_files audited for completeness: all shipped application
files (pegaprox/, web/ output, plugins/, static/, images/, misc/, docs/,
examples/, root) are listed; nothing new since 1.0 needs adding — the only
files added on this branch are a CI workflow and tests, both intentionally
not part of update_files.
Low-severity items from the Testing-branch audit:
- power/cost list_rates: scope by the token's floored effective_role, not the
owner's account role, so an admin-owned but scoped token can't enumerate every
cluster's rates (mirrors the upsert admin check).
- i18n relative-time: 'timeAgo' was always appended ('5s vor' in German). Split
into locale-ordered timeAgoSec/timeAgoMin templates for all 8 locales
(de/fr/es/pt lead, en/zh/ko/it trail).
- release-images CI: arm64 builds under qemu emulation; a failed arm64 leg no
longer blocks the amd64 release (continue-on-error on the arm64 matrix legs;
publish already attaches whatever artifacts exist).
- tests: BMC in-band read key->agent->password fallback order (#609) now has
direct coverage (3 tests).
(OIDC cross-source collision finding intentionally left alone — OIDC is hands-off.)
Follow-ups from the full Testing-branch audit:
- add_disk (LXC): a caller-supplied but already-occupied mpN — or 'rootfs' —
was written straight through, REPLACING an existing mountpoint and orphaning
its volume. Now coerce an occupied / rootfs / QEMU-shaped id to the next FREE
mpN, and reject a mount path containing ',' or '=' (mountpoint-option
injection). (+3 regression tests)
- hardening UI: the verbose-mode select-unapplied handler treated the per-control
{status,evidence} object as truthy, so it never pre-selected in verbose mode.
Made it object-aware, matching checkHardening.
- storage: the ISO/template 'From URL' buttons gated on storage.download, but the
download-url endpoint enforces storage.upload — aligned the UI gate.
- i18n(zh): dropped 10 backfill keys that duplicated ones already later in the zh
map (last-wins made them dead; topTalkers text never took effect).
Two Simplified-Chinese strings dropped 'PegaProx' when translated:
clusterManagement (the header, e.g. shown as 'Proxmox VE 集群管理 1.0') and
loginSubtitle. Audited all 55 en strings that mention PegaProx — only these
two had lost it; both now read 'PegaProx — …'.
- sshd_hardening (#433): stop writing 'AllowTcpForwarding no'. PegaProx's VNC
console tunnels to the node over an SSH direct-tcpip channel (utils/vnc_tunnel.py),
which that directive blocks — and the control never documented it. The sed cleanup
still strips a stale directive so a node hardened by an older build heals back to
the sshd default (forwarding on) on re-apply.
- Node hardening UI: an already-applied control rendered a static check icon, so it
could not be ticked — leaving 'Rollback Selected' with nothing to select. Applied
controls now render a real (green-ringed) checkbox; the 'Applied' title badge still
marks their state. (+ appliedTickToRollback DE/EN)
These keys existed in en (incl. the new mpN mount-path fields) but not in
zh, so Chinese users got the English fallback via t(). Adds the add-disk
mount-path strings, guest-detail / monitoring labels (address, services,
filesystems, interfaces, logged-in users, clock skew) and the top-talkers
table + hint.
Resolves two real regressions from #670 that affected ALL languages, not
just Simplified Chinese:
- hoist the can() permission helper to component scope — it was defined
only inside the tab-filter callback, so can('vm.snapshot') / can('vm.backup')
elsewhere threw a ReferenceError and blanked the Snapshot Schedule view.
- drop stray JSX tokens ()))} / )}) that rendered as literal text on the
topology diagram and the report page.
- localize hardcoded live/polling, network, kernel, score, socket labels
and add the Simplified Chinese values (t() falls back to English for the
other locales).
index.html rebuilt on top of the current Testing bundle so it stays
consistent with the mpN add-disk change already on Testing.
- add_disk: containers were sent scsi0/scsiN keys, which PVE rejects
("property is not defined in schema"). Coerce any non-mp id onto the
next free mpN and always pass a container-side mount path (mp=).
- AddDiskModal: for CTs default to the next mp slot, drop the QEMU
bus/format/iothread/ssd controls, add a Mount-Path field (+ DE/EN i18n).
- unlock_vm: plain delete=lock failed for clusters wired up with a root
API token ("Only root may use this option"). Send skiplock when
root@pam and fall back to qm/pct unlock over SSH when the API cannot.
- tests: mpN coercion + slot picking + default mount path + QEMU
untouched; skiplock / token->SSH fallback / not-locked / both-fail.
typeTextToVM() sent each character keysym straight to sendKey(), but qemu's VNC
keysym path does not assert Shift for the symbol row, so every shifted symbol landed
on its base key unshifted (! -> 1, @ -> 2, _ -> -, { -> [, | -> backslash, : -> ;),
making console login impossible for any password containing symbols. Wrap those
characters in a Shift_L press/release around the base-key keysym — the same thing the
browser does when you type them by hand (which already worked). Uppercase is untouched
since qemu shifts A-Z on its own. US-layout map, matching the paste path we already had.
The native <details> "Advanced Settings" accordion had a static ChevronDown and no
`group` class, so the arrow never flipped when the panel was expanded — it always
pointed the same way regardless of open/closed. Add `group` + group-open:rotate-180
so it points down when open, matching the other collapsibles.
Follow-up to 425b2c6 — drop the leftover `tasks.length > 0` condition so the
"show task bar" button stays available whenever the bar is hidden, not only
when a task happens to be around. Kills the empty-state dead end too.
Adopts si458s approach from #649 (he reported #648 and sent the fix).
Co-authored-by: Simon Smith <simonsmith5521@gmail.com>
A `false &&` debug guard slipped into 1.0 and hard-disabled the button that
brings the task bar back. Once you closed the bar with the X inside it there
was no way to get it back short of a reload. Drop the guard so the restore
toggle shows again while the bar is hidden and there are tasks to show.
- SPICE was only in the sidebar right-click menu, so it was easy to miss.
Add it next to the web-console action in the standard + corporate VM
detail views, the resource table (all three layouts) and the Cloud skin.
- Unify the SPICE icon on ExternalLink (it opens an external viewer) so it
reads distinct from the noVNC Monitor icon.
- Wire onOpenSpice through ResourceTable + VmDetailPanel + CorporateVmDetailView
and expose openSpice on the Cloud action bundle.
- QEMU-only + running-gated everywhere (LXC / stopped VMs have no SPICE port).
- Add spiceConsole / spiceConsoleHint / spiceDownloaded / spiceUnavailable
to all 7 locales.
Adds a SPICE option next to the VNC console for QEMU VMs — the same model PVEs
own web UI uses: a downloadable virt-viewer .vv connection file that opens in
remote-viewer, giving full SPICE (audio / USB redirection / multi-monitor) that
an in-browser client cannot.
- core/manager.py get_spice_ticket() gains a `proxy` arg (defaults to the cluster
API host) so remote-viewer tunnels the SPICE stream through the PVE host´s
pveproxy (:3128) — works behind a single public IP / when the node is not
directly reachable.
- api/vms.py: new GET .../<vmid>/spice endpoint, same authz as the VNC console
(vm.console + per-VM ACL). Calls spiceproxy and assembles the [virt-viewer] .vv
(multi-line CA single-line-escaped), served as an application/x-virt-viewer
attachment. Returns 409 with a clear hint when the VM has no SPICE display.
- dashboard.js: a "SPICE" action (qemu-only, running) that downloads the .vv.
Verified: backend py_compile + route registration + .vv-format unit test, and a
LIVE end-to-end run against PVE 9.2.3 — spun up a throwaway qxl VM, confirmed
spiceproxy returned a real SPICE endpoint (tls-port + password + CA, proxy via
the PVE host), our .vv came out well-formed, then destroyed the VM. Frontend
build clean. i18n uses English fallbacks for now (SPICE label is language-neutral).
Follow-up audit for the #644 class — a UI control hidden behind a blanket
isAdmin check while the backend endpoint it calls enforces a granular
non-admin perm, so a non-admin who legitimately holds that perm never sees
the control. Found 13 more (all confirmed against the code) and fixed them;
each control now checks the perm its endpoint requires, admin still implies all:
- storage.js (7): per-file delete + its column header -> storage.delete;
backup-SLA save -> cluster.config; the Storage-Balancing create/delete
cluster, add/remove storage and execute-migration controls -> storage.config
(reuses the hasPerm helper #644 already added).
- security.js (3): the Update Manager Schedule (backup.schedule), Start Rolling
Update and Update-this-Node (node.update) controls; UpdateManagerSection only
destructured isAdmin, so added user + a hasPerm helper.
- dashboard.js (3): the Snapshot-Policies New/Run/Edit/Delete controls
(vm.snapshot, via a canSnapshot prop threaded from the parent`s can()) and the
PBS backup Restore button (vm.backup, via the existing can() helper).
Clean areas (no mismatch): VM lifecycle, node ops, HA/Site-Recovery,
network/firewall/SDN, Ceph/pools/replication — already perm-gated. Build clean.
Backs out the 🇦🇹->🇩🇪 change from 6a3481a. German is spoken in Austria too, so the
AT flag on the `de` entry is a deliberate choice, not a bug. Added an inline note so
it does not get "corrected" again. (The #644 storage-perms fix in that commit stays.)
#644: the datastore action bar (Templates / From URL / Upload / Rescan) was gated
on a blanket `isAdmin`, but the endpoints behind those buttons enforce the
fine-grained storage.download / storage.upload / storage.config perms — so a
non-admin who legitimately holds e.g. storage.upload got no button to click. Each
button now checks its own perm (admin still implies all), so the UI matches what
the backend will actually authorise. Reported by @kglowinska.
#645: the German (de) language entry showed the Austrian flag; swapped 🇦🇹 -> 🇩🇪.
Reported by @ripperrd.
A close review of the #612 feature (backend + frontend + an in-process API run)
came back sound — auth matrix, per-member BOLA, input validation and SSRF surface
are all correct — but turned up a handful of low-severity gaps. Fixed:
Backend (api/multi_sdn.py):
- edit / reconcile / scan did a read-modify-write of per_cluster_status WITHOUT
the per-vid lock and REPLACED it from a pre-fan-out snapshot, so a concurrent
add/remove-member could lose a member entry (and edit could revert a concurrent
membership change in desired_state). New _merge_status_write() takes the per-vid
lock, re-reads inside it, MERGES the fresh results into the stored map, drops
departed members, and preserves the current member_clusters. Self-healing before,
consistent now.
- scan was gated on node.view but it overwrites the shared drift snapshot everyone
sees; gated it like every other span writer (sdn.manage + admin.settings) + the
member-access check.
- create takes a name-scoped lock + re-checks for an existing same-name record
before INSERT, so two admins racing the same name can no longer end up with two
aggregate records for one span (name IS the SDN vnet id; the fan-out is idempotent).
Frontend (vm_modals.js + dashboard.js):
- the whole EVPN feature + every mutation control was shown to node.view-only
users who then 403 on everything; added a canManage prop (isAdmin || sdn.manage
&& admin.settings) that hides Create / Edit / Re-apply / Reconcile / Scan /
Forget / Purge / add- and remove-member, leaving the read-only list. mcevpnReadOnly
note in all 7 locales.
- the validate preview rendered a member whose pre-flight SDN read FAILED as green
"ok" while the banner said "issues found"; added a p.error red branch.
- the auto-reconcile toggle reused the preposition t("on")/t("off") (no "off" key
existed) → mixed-language label; switched to the proper t("enabled")/t("disabled").
Adds tests/test_multi_sdn.py — the first committed coverage for the feature: the
API auth/validation surface (list/validate/create-4xx/404/anon-401/viewer-403 on
every write incl. scan) + the _merge_status_write merge/lock behaviour. 18 pass.
Verified: py_compile, frontend build, adversarial re-review. Live 2-cluster EVPN
E2E still owed (no real EVPN clusters). Refs #612.
- constants.py + web/src/constants.js: PEGAPROX_VERSION "Beta 0.9.15" -> "1.0", build 2026.08.01 (drops the Beta prefix)
- version.json: 1.0 manifest + changelog; 0.9.15 moved to prev_changelog
- web/index.html rebuilt from src
Caps the post-0.9.15 line for the 1.0 milestone. FF Testing -> main + tag is the
release step.
Scheduled rolling updates hardcoded the node reboot/online timeout at 600s, so
the Advanced-options value a user set only took effect for a MANUAL run. Now the
reboot timeout is a first-class schedule field end to end:
- schedules.py: the runner reads + clamps it (60s..2h, same bounds as the manual
path) instead of the fixed 600; the live scheduler (check_scheduled_updates)
forwards it into the run config — without that one line the stored value never
reached the runner and the whole thing was a no-op.
- schema + migration: reboot_timeout column on all three update_schedules
CREATE TABLE sites plus a PRAGMA-guarded ALTER so existing DBs get backfilled;
both row->dict paths return it (KeyError-safe on pre-migration rows); INSERT
and the create endpoint carry it through.
- frontend: reboot/online timeout selector in the schedule modal (disabled when
the schedule does not reboot), wired through save/load; reuses the existing
rebootTimeout i18n keys.
Verified: py_compile, frontend build, and a save->DB->load round-trip
(reboot_timeout=2400 survives). Refs #630.
Two lifecycle extensions on top of Phase 1+2, per Nico's pick.
Membership expand/shrink — you can now grow or shrink a span without delete+
recreate:
- POST /api/multi-sdn/vnets/<id>/members {cluster_id}: collision + reachability +
SDN-installed pre-flight on the NEW cluster (must share the span's ASN/VNI),
gated on access to ALL members INCLUDING the new one, builds the same EVPN
controller/zone/vnet on it, then appends it.
- DELETE .../<id>/members/<cid>[?purge=1]: drops a cluster from the span (keeps
>=1 member). purge tears the vNet down on that cluster VNET-ONLY — it leaves the
zone/controller, which another span on that cluster may share.
Drift → alerts — the Phase-2 scanner now emits an operator alert (the existing
push/inbox notification fan-out) on a healthy→drift transition. Edge-triggered
off the prior scan's status, so a vnet that stays drifted alerts ONCE, not every
6h. Fires in detect-only mode too (an alert is a notification, not a cluster
write).
Frontend: per-member remove buttons (shown only when >1 member; − = leave the
vNet, trash = purge) + an add-cluster picker (non-members only) in the expanded
view. i18n all 7 langs.
Adversarial review (3 dimensions x verify) → 3 findings fixed:
- Add-build failure now tears down EXACTLY the objects it freshly created
(_apply_on_cluster reports ; a mid-build fail at the zone/controller
step no longer orphans them) — while still never touching a pre-existing shared
zone/controller.
- The member-list read-modify-write is now serialized under a per-vid lock with a
re-read inside, so two concurrent admins adding/removing on the same span can't
lose an update; the remove path also re-checks the >=1-member invariant inside
the lock.
- The drift alert emits one notification PER drifted member (own cluster_id
anchor) so every affected tenant's subscribers are woken, not just the first.
Verified in-process (no real same-ASN EVPN clusters — honest limit): 50 tests
green across all five #612 suites incl the new member add/remove, vnet-only purge,
targeted teardown, per-member alert, and edge-triggered no-spam; app boots + 11
routes + scanner. Real 2-cluster EVPN E2E still owed.
Builds on Phase 1's create+read layer with the lifecycle + drift half PDM lacks.
Edit fan-out — PUT /api/multi-sdn/vnets/<id>: change alias + add/remove subnets,
fanned out to every member (PVE vnet-alias PUT, subnet POST/DELETE with the
<zone>-<cidr> id URL-encoded) + one apply per cluster. Structural fields
(name/zone/vni/asn/controller) are immutable — a change there rebuilds the span,
so it's rejected with a delete-and-recreate hint. Gated sdn.manage+admin.settings,
per-member check_cluster_access, partial-failure status like create.
Drift detect + reconcile:
- Background scanner (6h, daemon thread started in api/__init__, mirrors
drift.py) reads live per-member SDN state for each aggregate vnet and persists
the drift status. DETECT-ONLY by default — it never writes to a cluster unless
the new global setting multi_sdn_drift_reconcile is opted in.
- Manual POST .../<id>/reconcile (re-assert desired + fix alias drift) and
POST .../<id>/scan (read-only drift refresh, node.view) for on-demand use.
- When auto-reconcile is opted in, it fixes ONLY the drifted members (never
reloads an in-sync member), debounces 'missing' (needs two consecutive
non-healthy passes so a transient partial read can't resurrect intentionally-
removed SDN), and writes a log_audit trail for the unattended mutation.
Frontend (MultiClusterEvpnView): edit modal (alias/subnets), per-cluster drift
detail, Scan / Reconcile / Edit actions, and an admin-only auto-reconcile toggle.
i18n all 7 langs.
Adversarial review (4 dimensions x verify, 2 workflows) → 7 findings fixed:
_apply_on_cluster now skips the cluster-wide apply when nothing changed (kills
the over-broad blast radius: one drifted member no longer reloads the whole
span, and reapply/reconcile of in-sync members is a no-op); reapply treats
in_sync as done; auto-reconcile is per-member + missing-debounced + audited; the
toggle is hidden from non-admins; edit subnet fold-back de-dupes; and the
pre-existing DUPLICATE it: translation block (silently dropping 317 Italian
keys) is merged into one — Italian is whole again.
Verified in-process (no real same-ASN EVPN clusters — honest limit): 44 tests
green across edit/reconcile/scan/scanner/gate/debounce/no-op-apply/subnet-dedup;
app boots + 9 routes + scanner thread; build + translations parse; auth/injection
review dimensions came back clean. Phase-2 auto-reconcile ships default-off.
PVE has no cross-cluster SDN primitive: each cluster's /etc/pve/sdn is local. To
make one logical EVPN vNet span several clusters that share a BGP ASN, PegaProx
must create the same EVPN controller/zone/vnet on every member and apply each.
This adds the orchestration layer + the authoritative record PDM lacks. Phase 1
is create + read; edit/alias fan-out + a drift-detect scanner are Phase 2.
New api/multi_sdn.py composes the existing per-cluster SDN passthrough
(api/datacenter.py → PVE /cluster/sdn/*) across N members:
- POST /api/multi-sdn/vnets — validate → per-member collision + reachability
pre-flight → bounded concurrent fan-out (controller→zone→vnet→subnets→apply,
idempotent skip-if-exists) → authoritative record. Atomic by default: any
member failing rolls back ALL members (best-effort, ignores does-not-exist) so
no half-built L2 span is left behind.
- POST .../validate — dry pre-flight (reachability + collisions), no writes.
- POST .../<id>/apply — idempotent retry of not-yet-applied members.
- GET .../vnets, GET .../<id>?refresh=1 — list/detail + live per-cluster status.
- DELETE .../<id>[?purge=1] — forget the record (default) or also tear the SDN
objects down on every member.
Collision pre-flight rejects VNI reuse, ASN mismatch, and zone/vnet redefinition,
while treating a same-definition object as idempotent. Writes gated on
sdn.manage + admin.settings (disruptive cluster-wide apply = blast radius);
reads on node.view. Every route gates check_cluster_access on ALL member
clusters (a caller who can't reach a member can't read/mutate it via the
aggregate); deny → 404 on read to avoid confirming existence.
New multi_cluster_vnets table (db.py) is the authoritative record (JSON columns
per the site_recovery_plans convention). Frontend: a new Cloud-level
'Multi-Cluster EVPN' view (sidebar route, ≥2 clusters) — list with member chips +
per-cluster status badges, a create wizard with validate/preview, re-apply and
delete/purge; i18n all 7 languages.
We orchestrate the SDN config objects only — the physical BGP-EVPN underlay
(inter-cluster peering) is the operator's network and is assumed to already peer.
Verified in-process (no real same-ASN EVPN clusters in the lab): 38 tests green —
validation, collision (idempotent/conflict/VNI-clash/ASN-mismatch/501), apply
ordering + bodies + idempotency + offline/failure, canonical-subnet idempotency,
rollback order, rollup, live-status; app boots + 6 routes register + DB table
created; frontend build + translations parse. Adversarial review (4 dimensions,
verified) fixed: atomic rollback now covers the failed member too, subnet
idempotency uses canonical-network equality, topology sidebar state reset.
Real end-to-end validation needs 2+ clusters sharing an ASN with a live EVPN
underlay (lab has one) — an honest limitation, same as #546.
Adds an optional pre-cutover hold to the ESXi→Proxmox migration: when enabled
in the wizard, the migration stages the disks and then parks in a new
'awaiting_confirmation' phase — with the source VM still running — and waits for
the operator to commit (or cancel) the switchover. This lets the operator
schedule the few seconds of downtime themselves (drain monitoring, flip DNS,
then commit) instead of it landing whenever the copy happens to finish, per
ajoergensen's request.
Backend (core/v2p.py): V2PMigrationTask.await_cutover_confirmation() gate,
inserted before every set_phase('cutover') seam; it's a no-op when the flag is
off (so existing migrations are byte-identical) and only holds while the source
is still running — on modes that already stopped the source (offline / sshfs-
boot / VM-was-off) it logs-and-commits rather than extending downtime. Confirm/
cancel are flags the migration thread polls; cancel (or a generous timeout,
default 24h) raises V2PCutoverCancelled → clean 'cancelled' status with the
source left running and the staging mount/snapshot torn down. The confirmation
wait is never billed as downtime (downtime clock reset on confirm; the primary
live modes already measure from cut_t0 after the gate).
API (api/vmware.py): POST .../migrations/<mid>/confirm-cutover and
.../cancel-cutover (409 unless parked at the gate); to_dict() exposes
wait_for_confirmation + awaiting_confirmation; wait_for_confirmation /
confirmation_timeout accepted on the migrate body.
Frontend (dashboard.js): wizard checkbox + a commit/cancel action on migrations
holding at the gate (full Migrations list + compact per-VM view), new phase/
status styling; i18n for all 7 languages.
In-process verified: gate state machine (park→confirm→proceed, cancel + timeout
→ V2PCutoverCancelled, no-op when off, source-stopped skip, downtime reset),
app boot + both routes registered, frontend build + translations parse.
The Top Resources tables in the group and all-clusters overviews (both the
Corporate and modern skins) had static, non-sortable headers and the CPU/RAM
columns didn't line up with their headers. Add a dedicated guest-sort state
(default CPU-desc, so it still opens on the busiest guests), a shared
sortTopGuests comparator + nextGuestSort header state-machine, click-to-sort
on name/cluster/node/CPU/RAM/status with a direction indicator, and
right-align the numeric CPU/RAM cells in the plain-percentage table.
CorporateNodeDetailView fetched /sensors but never rendered it — the Sensors
panel + temperature chart only existed in the default NodeModal's 'system' tab,
so on the Corporate layout temps were invisible even with lm-sensors installed
(reported by jostrasser). Add a dedicated 'sensors' load (sensors +
temperature-history), trigger it on the Hardware tab, and render the same panel
there. Modern/default path untouched. Rebuilt bundle.
The disk section rendered in raw PVE config order, which is effectively random and
looked messy. Order it deterministically: data disks first (by bus then index —
scsi0, scsi1, virtio0, sata0, ide0...), then CD/DVD drives, then firmware volumes
(EFI / TPM) last; LXC keeps rootfs before mp0..n.
Display-only — sorts a copy in the render, no backend or API change. Rebuilt
web/index.html.
Refs #610
An empty CD/DVD drive is stored as `ide2: none,media=cdrom` — no storage volume,
so no colon, so the config parser (which required ':' in the value) dropped it and
the drive was invisible in the Hardware view. Users couldn't see the existing drive
and reached for add/remove-drive instead of simply mounting/ejecting media — exactly
the confusion #610 describes. This matches Proxmox's own GUI, which always shows the
drive.
- backend (manager.py): also parse `none,media=cdrom` entries into the disks list
even without a colon, tagged media=cdrom, storage='', volume='none'.
- frontend (vm_config.js): the row now renders (existing isCdrom path) with a
'No media' badge, the ISO button labelled 'Mount ISO' when empty (vs 'Change ISO'
when media is present), plus a one-click Eject on mounted drives.
- i18n: new `noMedia` key in all 8 language blocks.
Backend mount/eject (PUT .../cdrom, set_cdrom) already supported this — only the
parser hid empty drives. Verified: unit-assert of the parse branch + live E2E (eject
ide2 on a stopped guest -> empty drive now appears in the parsed config). Rebuilt
web/index.html.
Refs #610
When the session became invalid (server restart, idle timeout, revoked token),
authFetch only console.warn'd the 401 and every poll failed silently — the UI
kept showing stale data with no hint the user had been logged out.
Now any 401 (except /auth/* and /sse) raises a session-expired flag that renders
a clear modal overlay (portalled to body so it sits above everything) with a
'Log in' button that runs logout() -> back to the login screen. Idempotent, so a
burst of concurrent 401s only shows the prompt once. Reuses existing i18n
(sessionExpired + login) — no new keys.
Verified live via Playwright: after forcing API calls to 401, the overlay appears
on the next poll. Rebuilt web/index.html.
The all-clusters overview and the cluster-detail summary cards disagreed:
- CPU: overview used core-weighted cluster CPU (Σ cpu×cores / Σ cores, from
datacenter/status); detail used a plain per-node average -> e.g. 6% vs 7.6%.
- Guest counts: overview total = running+stopped from the polled guests summary,
which drops paused/suspended guests and could transiently render an impossible
'16 / 14' (total < running); the detail counted all qemu+lxc -> '16 / 30'.
Fixes:
- backend datacenter/status: add explicit per-type guest totals
(guests.vms.total / guests.containers.total = every guest of that type, any
status incl. templates) so the overview denominator equals the detail's
allVms.length and can never be < running. (proxmox + xcpng paths)
- overview (AllClustersOverview + the cluster-group variant): use the explicit
total when present, fall back to running+stopped for older backends.
- detail summary cards: read the core-weighted CPU from datacenter/status (same
source as the overview), falling back to the per-node average until it loads.
Verified live vs the Testi cluster: overview 16/30 == detail 16/30, and both CPU
values now derive from a single source. Rebuilt web/index.html.
Automation > Alerts only offered enable/disable + delete; changing a rule
meant delete-and-recreate. Add an Edit action that opens the create modal
pre-filled with the alert's current config and saves via the existing
PUT /api/clusters/<id>/alerts/<id> endpoint — no delete/recreate.
Frontend-only; the backend PUT already supports every field:
- Edit button on each alert row (next to delete)
- openEditAlert() pre-fills name/target/metric/operator/threshold/channels/
severity/escalation; modal title + submit button switch to edit mode
- submit branches: editingAlert -> PUT (keeps current enabled), else POST
- form remounts by key so defaultValues re-apply; New Alert resets to create
- new i18n key alertUpdated in all 8 language blocks (editAlert already existed)
Verified live (create -> PUT-edit -> readback): every field updated, enabled
preserved. Rebuilt web/index.html.
Refs #618
Builds on the tested transfer primitive (core/incremental_repl) to make
cross-cluster replication ship only the snapshot delta for RBD VMs instead
of full-cloning + remote-migrating the whole disk every cycle (#174 aderumier).
Backend (api/vms.py):
- new _execute_replication_incremental(job): eligibility-gate (every disk rbd
on source AND target storage rbd), guest-snapshot the source, replicate each
disk's delta via the byte-relay primitive, build/tag the replica VM on the
first run (adopts the seeded rbd images), advance + prune the snapshot chain,
update last_snapshot. A stale replica is torn down first behind the same
tag-safety gate the full path uses (#413), so a mis-set target VMID never
nukes a bystander. Returns False for non-eligible VMs so _execute_replication
transparently falls back to the proven full clone+migrate flow.
- opt-in `mode` on the job; schema adds mode + last_snapshot columns (db.py).
Frontend (datacenter.js + translations.js): transfer-mode selector
(Full / Incremental) in the replication dialog + i18n.
Live-verified end-to-end against a real Ceph RBD cluster (source pve1 -> relay
-> replica on pve2): seed created + tagged the replica VM with the adopted disk
(md5 src==replica), then an incremental run after a 50 MB change fast-forwarded
the replica in-place (md5 src==replica) and pruned the old base snapshot; the
tag-safety gate correctly refused an untagged same-VMID VM. Full cross-cluster
(two separate clusters) + the ZFS branch still want a 2-cluster / ZFS lab.
proxforge asked that a rolling update not keep pulling nodes when Ceph
isn't healthy — on an HCI cluster that can drop data below min_size and
take it offline. Part 1 (the "ceph installed but not deployed -> unknown"
misdetection) already shipped in v0.9.13.3; this is part 2.
New per-run setting `ceph_health_gate`:
- off -> current behaviour (warn after the wait, then continue)
- degraded -> HOLD only on genuine data-at-risk: HEALTH_ERR, unknown
(deployed-but-unreachable), OSDs down, or degraded/undersized/
incomplete/inactive PGs. Benign HEALTH_WARN (clock skew, a
leftover flag, backfill/remapped) is tolerated.
- strict -> HOLD on anything that isn't HEALTH_OK.
When the gate holds, the rolling update pauses with paused_reason
'ceph_unhealthy' + a reason/message, reusing the existing pause/resume
mechanism — the operator restores Ceph, then Continue or Cancel. No effect
on clusters without Ceph (get_ceph_health_summary returns None -> skipped).
- api/settings.py: _ceph_gate_unsafe() helper + parse/validate/store the
setting + the hold branch in the rolling-update loop.
- security.js: gate selector in the rolling-update settings + start-body
field; the generic paused-state UI already renders the hold + buttons.
- 6 i18n keys across all languages.
Gate decision unit-tested across HEALTH_OK/ERR/unknown/WARN(noout,
clock-skew,degraded,osd-down,backfill) x off/degraded/strict.
Backend (api/ceph.py, core/manager.py):
- OSD create/destroy use a fresh password-based root@pam ticket when the
cluster is API-token-authed (PVE hard-gates OSD ops to root@pam, never
a token). New manager.create_privileged_session() guards the inline-token
and non-root@pam cases with a clear message instead of a failed login.
- add POST/DELETE /ceph/mgr/<id> (was GET-only, so a manager could never
be added or removed from the UI).
- CephFS create forwards to /ceph/fs/<name> (it POSTed the collection URL,
which PVE has no handler for -> every create returned 501). destroy now
plumbs remove-storages / remove-pools.
- mirror snapshot-schedule list returns [] on a non-mirrored pool (was a
500), while still surfacing genuine 501/503 SSH/tooling errors.
- request.get_json(silent=True) across the blueprint so an empty POST body
no longer raises a raw werkzeug 400/500.
Frontend (datacenter.js, translations.js):
- OSD-create modal with a node selector + disk picker (GET .../disks;
in-use / existing-OSD disks are shown but disabled).
- Managers section in the Monitors tab (add/remove) + fix the MGR count
(empty list was truthy -> always reported 1).
- Create-MDS and Create-CephFS modals in the CephFS tab.
- pool and CephFS delete buttons now send confirm_name (+ remove_storages;
CephFS asks separately before removing the data/metadata pools) - the
old empty DELETE always failed the backend confirmation gate with 400.
- 18 new i18n keys across all languages.
Verified end-to-end against a live 3-node Ceph cluster (HEALTH_OK).
The credentialed, out-of-band counterpart to in-band ipmitool — reads BMC health
over the management network via the DMTF Redfish API, as a fallback when in-band
is unavailable. A SEPARATE, sharper opt-in with an enforced 5-second delay.
* core/redfish.py — read-only Redfish reader. Pure parsers (thermal/power/system/
SEL) normalize to the SAME shape as read_node_bmc_inband, so the panel + cache +
rollup + alerting work unchanged. SSRF-guarded, GET-only, redirects refused.
* core/db.py — node_bmc_endpoints table; BMC password stored encrypted (aes256:),
masked (********) on every API response, never wiped by a masked re-save.
* core/bmc.py — REDFISH_CONSENT (v1, require_delay_seconds=5) in all 7 languages.
* api/nodes.py — redfish-consent GET/POST + per-node BMC endpoint GET/POST/DELETE/
test (admin.settings, consent-gated, SSRF-validated host). The per-node read +
cluster rollup now gate on (in-band OR redfish); the 5-min collector falls back
to Redfish for nodes with a configured endpoint (own ~1h backoff).
* Frontend — Redfish sub-panel in the node Hardware tab: enable via a 5-second-
delayed warning modal, then a BMC endpoint form (host/user/password/verify_ssl +
Save/Test/Remove). i18n across all 7 languages.
Security review (28 agents, 17 positive confirmations) — fixed the 4 confirmed:
* CRITICAL SSRF — a hostile BMC's @odata.id could userinfo-splice the credentialed
GET onto an internal host (https://<base>@169.254.169.254/...), exfiltrating the
stored BMC creds. Every follow-on hop is now urljoin'd against the validated base
and its scheme/host/port asserted equal; @/scheme/protocol-relative refs rejected.
* HIGH — config-restore could flip the new redfish consent without the audited ack;
the restore-strip now protects BOTH consent keys.
* MEDIUM — malicious BMC could OOM the reader: bodies now streamed + capped at 4 MiB.
* MEDIUM — the cluster rollup route 403'd a Redfish-only deployment; now gated on
(in-band OR redfish).
Residual (noted, not fixed): DNS-rebind TOCTOU on the base host (IP-pinning not done
autonomously per the codebase's standing SNI-pinning decision).
Tests: +33 (Redfish parsers, SSRF origin-pinning + body-cap regressions, consent +
BMC endpoint CRUD/SSRF/masking, rollup dual-gate). Full suite 347 passed.
Cluster-wide in-band BMC health, surfaced + alertable — mirrors the proven #601
temperature pipeline (5-min collector populates a per-manager cache; the 60s alert
loop only READS it, never SSHes).
Backend:
* manager.py: _node_hw_cache/_lock/_backoff + get_cached_node_hardware() and
get_cluster_hw_rollup() -> {health, available, checked, counts, degraded[]}.
* metrics.py: _node_hw_summary() SSH probe (~1h backoff for nodes without
ipmitool/BMC), populated in the 5-min collector via run_per_node(cap 8, 90s),
GATED on the compliance consent + proxmox-only. Off the hot-path.
* alerts.py: 'hardware_health' alert metric (cluster worst + per-node), ok/warning/
critical mapped to 0/1/2, auto-severity, cache-only.
* nodes.py: GET /clusters/<cid>/hardware/health rollup (consent-gated, cache-only).
* vms.py: compact hardware rollup injected into datacenter/status (cache-only,
consent-gated) so the overview badge is free.
Frontend: 'hardware_health' alert metric with a Warning-or-worse / Critical-only
threshold selector; degraded-hardware badge in the corporate sidebar, the warning
banner, and the cloud overview. i18n across all 7 languages.
Adversarial review (12 agents): 3 rejected (bounded/cosmetic), 5 confirmed & fixed:
* MED — hwHealth prop was missing at the 2nd (ungrouped) ClusterSidebarItem call
site, so the sidebar badge never showed for ungrouped clusters — now passed.
* LOW — the rollup endpoint 500'd on non-proxmox clusters — proxmox/callable guard
now degrades to an empty 'unknown' rollup.
* LOW — a '<' operator on hardware_health builds a silent no-alert rule — the UI now
only offers '>' and the create/update API pins hardware_health to '>'.
* NIT — alerts now show OK/WARNING/CRITICAL instead of the raw 0/1/2 code.
* NIT — corrected a stale '>=' code comment.
Tests: +11 (rollup endpoint gates, non-proxmox graceful, operator coercion, manager
rollup/cache unit tests). Full suite 322 passed; frontend build clean.
The compliance warning text is now localized (en/de/fr/es/pt/ko/it) while the
VERSION stays language-spanning: _HW_CONSENT_TEXT holds the same v1 warning in
each language, and hw_consent_warning(lang) injects the shared version +
require_delay_seconds so they can never drift per-language (bump all together).
* GET /api/hardware-monitoring/consent?lang=<code> returns the warning in the
requested language (falls back to English). HW_CONSENT_WARNING kept as an
English alias for existing callers/tests.
* The acknowledged LANGUAGE is now recorded alongside the version — set consent
stores ack_lang and the audit line reads 'acknowledged compliance warning v1
[de]', so the non-repudiation record captures which exact text the user saw.
* Frontend passes the current UI language to the consent GET (?lang=) and the
enable POST (lang=), and re-fetches the localized warning on language switch.
Translations adversarially verified by 6 native-level reviewers vs the English
source (control IDs CMMC/NIST 800-171 3.4.6+3.1.5/DISA STIG preserved verbatim in
every language). Applied the confirmed fixes: es 'encendido'->'alimentación' (the
operation-list 'power' means power-control, not power-on), es informar+de grammar,
de closing-quote typography, ko install-vs-may modal + 'override' precision.
Tests: +4 integration cases (per-lang warning, unknown-lang fallback, ack_lang
recorded in audit + status, unknown-lang recorded as en). Full suite 311 passed.
Step 3 — one-click ipmitool install:
* POST /api/clusters/<cid>/hardware/ipmitool/install (admin.settings) mirrors the
proven StarWind installer but with a FULLY STATIC script (no user-controlled shell
input), gated: cluster-access -> proxmox-only -> consent-required (the install
mutates the node, so it must not run before the warning is acknowledged) -> per-node
name validation -> run_per_node fan-out (cap 8, idempotent) -> audited.
Step 4 — frontend:
* Self-contained HardwareMonitoringPanel: consent fetch, mandatory compliance-warning
modal (renders the server-versioned warning text, required ack checkbox, generic
require_delay_seconds countdown ready for the Redfish phase), ipmitool install button
when missing, sensor/power/FRU/SEL rollup with health badge.
* Wired as a 'Hardware' tab in BOTH node UIs (corporate CorporateNodeDetailView +
default NodeModal) for parity. 21 i18n keys across all 7 languages.
Hardening from the adversarial review (3 confirmed findings, all fixed):
* MEDIUM — config-restore could enable hardware monitoring WITHOUT the versioned
acknowledgement audit record (non-repudiation bypass): restore now drops the
hardware_monitoring key and preserves the live consent state, so consent only ever
moves through the audit-logged set_hw_monitoring_consent path (settings.py).
* LOW — a non-int stored/posted ack_version raised int() -> 500 on every gate; add
_hw_int() so the gate fails CLOSED (disabled) instead of crashing.
* NIT — a crafted {"nodes":[123]}/[{...}] blew up _reject_bad_node's regex / set()
-> 500; validate the raw list + reject non-strings centrally -> clean 400.
Tests: +9 integration cases (install gates admin/consent/proxmox/bad-node + the 3
hardening regressions). Frontend build clean (Babel). Full suite 307 passed.
Bumps PEGAPROX_VERSION -> Beta 0.9.15 (constants.py + web/src/constants.js),
the README badge, and version.json (build 2026.07.14 + changelog). Bundles all
Testing commits since v0.9.14.1: the CodeAnt/pentest security-hardening sweep
(BOLA/IDOR tenant gates, SSRF/RCE/CSRF guards, CWE-117, dependency CVE bumps,
AGPL §7(b) attribution), full Cloud-layout feature parity, temperature history
+ alerting (#601), sidebar/scale perf, the integration+authz test harness, and
community PR #617 (snapshot-name validation, @MatrixNeoKozak). Welcomes
TechniData AG Limited as a Silver sponsor.
Live per-host temperature already shipped via the lm-sensors Sensors panel.
This adds the two pieces #601 asked for on top:
History / chart over time:
- metrics.py collector now records each online node hottest lm-sensors temp
into the persisted 5-min metrics_history snapshot. SSH fan-out uses the
SSH-aware run_per_node (max_concurrent=8/cluster) so we never open 30+
simultaneous sessions; a 1h per-node backoff skips hosts without lm-sensors,
and the per-manager temp caches are pruned to known nodes.
- new GET /clusters/<id>/nodes/<node>/temperature-history (off-hub, cached).
- node Sensors panel renders a temperature-over-time sparkline + min/max/avg.
Alerting:
- new "temperature" alert metric for node (this node) and cluster (hottest
node) targets, read from a manager temp-cache the 60s alert loop never SSHes.
Alerts render in °C (not %); auto-severity uses absolute thresholds
(>=85 crit, >=75 warn). Alert form gains the metric + a dynamic °C/% unit.
- i18n for all 7 languages.
The corporate/modern sidebar rendered every VM in the tree and re-filtered on
each keystroke — janky at 100 nodes / 1000+ VMs. Two changes, no behaviour lost:
- Debounce (250ms) the term used for FILTERING, name highlighting and the
cross-cluster 'N found' match counter, so typing no longer re-filters and
re-counts across all clusters/VMs on every keystroke. The <input>, clear-X
and its visibility stay bound to the immediate value for instant echo; empty
clears synchronously (no lag when wiping the box).
- content-visibility:auto on the tree rows = native windowing: the browser
skips layout + paint for rows scrolled out of the sidebar's scrollport while
keeping them in the DOM, so keyboard nav (querySelectorAll+focus), tooltips,
selection and the #vmid ::after all keep working. Moved the one-shot expand
fade from each row to the tree container so it no longer replays per row on
scroll re-entry (and runs once, not ~1000x). Degrades gracefully where
content-visibility is unsupported.
Adversarially reviewed across 4 dimensions (debounce correctness, CSS
interaction, completeness, regression); coverage unchanged, every VM stays
searchable.
- Corporate sidebar 'show VM IDs' now renders via a CSS ::after on each guest
row (gated on a body flag set from the pref) instead of an extra <span> per
VM, so the pref costs nothing at 1000+ VMs — no per-row re-render, no added
DOM node. data-vmid is only set on named rows (nameless rows already show the
id in their fallback label).
- Cloud (Preview): block an unconstrained allow-all ACCEPT firewall rule (with
confirm) and guard the firewall/SDN submits when no cluster is selected.
Follow-up to 1eae2e3 — the toggle was rendering English fallbacks only. Add
showVmidSidebar/Desc/On/Off to all language blocks (de/en/fr/es/pt/ko/it ×2).
Verified LIVE end-to-end after restarting the dev server onto the new backend:
login + /api/user/preferences return sidebar_show_vmid, the PUT toggle persists,
and the corporate sidebar renders '#<vmid>' next to each guest (confirmed in the
DOM, e.g. 'FiveM-Server#122').