751 Commits

Author SHA1 Message Date
MrMasterbay
bd938b9f17 fix(auth): don't log out during an OIDC callback exchange (#681)
AuthProvider.checkSession and LoginScreen both run on mount. On the OIDC callback
page the session doesn't exist yet, so checkSession's /auth/check returns 401 and
it called logout() — which raced the in-flight callback POST and aborted it
(Firefox nginx-499 / 'Network error during OIDC callback'; Chrome won the timing).
checkSession now stands down when the URL carries the OAuth code+state (the same
signal LoginScreen keys off), leaving the callback handler to own the transition;
it reloads on success and this runs cleanly against the new session.
2026-08-11 19:08:16 +02:00
kglowinska
dd03f7dc0e
fix(oidc): apply custom roles from group mappings instead of dropping them (#682)
* oidc: apply custom roles from group mappings instead of dropping them

A group mapping onto a custom (e.g. tenant-scoped) role never took effect. The
mapping loop ranked roles through a hard-coded {viewer: 0, user: 1, admin: 2}
table read with .get(role, 0), so any custom role tied with viewer and the
strict `>` comparison never fired — while the log line still reported
"Custom group mapping matched", leaving no trace of why the user ended up on
oidc_default_role. Custom roles are offered in the config UI (the Default role
dropdown lists every non-builtin role) and accepted by the settings validator,
so the intent was clearly that they work.

Rank custom roles between user and admin: an explicit mapping onto a custom
role now beats the coarse viewer/user defaults, but can never demote someone
matched by the admin group. A role that came from oidc_default_role is also
tracked separately, so a custom default no longer swallows every mapping.

Adds tests/test_oidc_group_mapping.py — the four custom-role cases fail before
this change, the four no-regression cases (admin precedence, builtin ladder,
unmatched default) pass both before and after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(oidc): don't let a matched mapping demote a configured admin default

Review feedback on #682: role_from_default let any matching mapping replace the
configured default, including when that default is admin — so an install using
oidc_default_role='admin' would be downgraded to viewer/user/a custom role on the
first matching group. The effect is a loss of admin access rather than an
escalation, but it changes long-standing behaviour and can leave an install
without access to admin-only workflows.

Keep the rule that the configured default stands for "no group matched" and may
therefore be replaced by a group that did match, but exempt admin from it. Adds
the regression test, which fails without the exemption.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(oidc): warn when several custom-role group mappings match

Review feedback on #682: custom roles all share one precedence level, so when a
user matches more than one custom-role mapping the first entry in the list wins
and the rest tie out — the outcome depends on mapping order.

Keeping first-match-wins: the mapping list is an ordered list in the UI, so
reading it top-down is the least surprising rule, and it is deterministic for a
given config. Defining a real precedence *between* custom roles would mean
ranking their permission sets against each other, which is a product decision
(permission count is a poor proxy for privilege, and tenant-scoped roles are not
comparable at all) rather than something to settle inside a bug fix.

What was worth fixing is the silence: several matches now log a warning naming
the roles, the one applied, and how to change the outcome.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 19:05:04 +02:00
MrMasterbay
18d03f0476 feat(update): install-method-aware self-update (apt / docker / source)
The in-app updater and update.sh are a git-tree + pip file update that only fits
the source/deploy.sh layout. On an apt/dpkg install the correct path is apt
upgrade (deps are system python3-* packages; a pip run diverges from dpkg and
can't lift the dpkg-owned crypto libs -> fail-closed TLS on restart); on Docker
a file update is discarded at the next image pull.

- settings.py: _detect_install_method() (docker via /.dockerenv + cgroup, apt via
  dpkg -S, else source); perform_pegaprox_update refuses on apt/docker with the
  right guidance + copy-paste command (409, allow_managed override); check-update
  reports install_method / in_app_update_supported / managed_update_{hint,command}.
- update.sh: same detection + guard before touching anything (--force override).
- settings_modal.js: the Install button becomes a package-manager / image hint
  with a copyable command on apt/docker; performUpdate handles the 409.
- tests: 5 guard/reporting tests.
2026-08-11 19:02:44 +02:00
mkellermann97
ae7a7275d7 harden self-update: don't restart into a failed dependency upgrade
Both the in-app updater (api/settings.py) and update.sh now run a crypto/TLS
preflight in the target interpreter (the venv) before bouncing the service. Our
TLS startup is fail-closed (#633), so if a release's dependency bump didn't land
a loadable cryptography/pyOpenSSL pair, a restart takes the service down with no
way back up. On preflight failure we keep the current process running and tell
the operator to install deps + restart manually.

- api/settings.py: gate the restart thread on the preflight; the response now
  carries deps_ok / restarting=false + a clear message instead of a green check.
- update.sh: same preflight before the systemctl restart; exit with the fix
  steps instead of bouncing into a service that won't come back.
2026-08-11 08:20:08 +02:00
MrMasterbay
b57f454736 security(ssrf): pin webhook target to resolved IP + gate custom-template image_url at wget sink
- webhooks._guard_url now returns (ok, reason, url_to_use); public http(s) targets are
  IP-pinned via resolve_and_pin_url so a DNS rebind between the guard check and the POST
  can't swing the request to an internal host. allow_private opt-in skips the pin so an
  internal ntfy/Gotify keeps working. Both send_to_channel / _post_ntfy callers use the
  pinned url. (Aikido #469089218)
- templates_lib deploy: re-run the SSRF guard on tpl['image_url'] right before the node
  wgets it. add_custom_template validates on entry, but built-in catalog and DB-stored
  URLs reached the sink unchecked. (Aikido #469089270)
- tests: webhook-guard 3-tuple contract + metadata/loopback block + allow_private no-pin.
2026-08-10 08:53:20 +02:00
MrMasterbay
d26f044ccd security(authz): tighten backup-job target gate + guard unpooled-VM int()
- _authz_backup_targets: a non-admin backup.schedule holder now needs an
  explicit include-list of VMs they own. all=1 / pool / exclude-mode /
  selMode=all|exclude|pool and an *empty* selection are admin-only — PVE
  treats several of those as "every VM", and the load->edit->save round-trip
  of an admin job re-hits the gate on PUT so a scoped user can't retune it.
- get_vms_without_pool: a single non-numeric vmid no longer int()-throws and
  500s the whole unpooled-VM listing; skip the malformed row instead.
- 2 regression tests (empty/exclude selection + PUT all=1 denied).
2026-08-10 08:50:15 +02:00
MrMasterbay
4a702cd742 security: fix web-push IPv6 SSRF bypass + persist-credentials on CI checkouts (Aikido)
- push.py _is_internal_or_metadata_host (#469089273): stopped splitting the host on
  ':' (which mangled every IPv6 literal, '::1' -> '', bypassing the block) and unwrap
  IPv4-mapped IPv6 (::ffff:a.b.c.d) so metadata/private checks apply. +6 unit tests.
- docker.yml + release-images.yml (#348463387/#348463388): persist-credentials:false on
  the actions/checkout steps that never push.
2026-08-10 08:36:27 +02:00
MrMasterbay
ebc876c3e7 security(authz): per-VM authorization on 5 BOLA routes (Aikido)
The cluster-reach check (check_cluster_access) grants access via the pool/VM-ACL
fallback and expects a downstream per-VM gate — these 5 routes lacked it:
- clusters.py cancel_task (#469089252): parse the task's VMID from the UPID and
  require user_can_access_vm(vm.stop) — no cross-pool/tenant task cancellation.
- storage.py create/update backup job (#469089226): authorize every submitted
  VMID (vm.backup); cluster-wide (all=1) / pool jobs are admin-only.
- static_files.py vms-without-pool (#469089182): filter to accessible VMs.
- search.py cluster tags (#469089237): count only accessible VMs' tags.
- xhm.py migration detail/list (#469089253): require source-VM access, matching
  the plan/start gate.
Admins pass user_can_access_vm unchanged. +5 regression tests.
2026-08-10 08:31:47 +02:00
MrMasterbay
3e2879a681 security(authz): object-level gates on VMware + PBS write/delete/diagnose (Aikido BOLA)
The read routes already call check_vmware_access / check_pbs_access, but the
mutating ones were guarded only by @require_auth(perms=[...]):
- vmware.py update/delete/diagnose (#469089245/#207/#204): a vmware.config holder
  could take over, delete, or probe another tenant's ESXi server by id.
- pbs.py update/delete (#469089284/#244) and auto-storage (#469089213): a
  pbs.config holder could overwrite/delete another tenant's PBS record, or push
  its stored credentials onto an arbitrary cluster's storage config.
Each now calls check_* first (admins + unlinked servers unaffected); auto-storage
also check_cluster_access on every target cluster. +6 regression tests.
2026-08-10 08:18:08 +02:00
MrMasterbay
7ba2a54c5f release: PegaProx 1.0.1
Bump PEGAPROX_VERSION + PEGAPROX_BUILD (constants.py, constants.js) and
version.json (version / build / release_date) to 1.0.1 / 2026.08.09, and add
the 1.0.1 changelog entry.

version.json update_files audited for completeness: all shipped application
files (pegaprox/, web/ output, plugins/, static/, images/, misc/, docs/,
examples/, root) are listed; nothing new since 1.0 needs adding — the only
files added on this branch are a CI workflow and tests, both intentionally
not part of update_files.
v1.0.1
2026-08-09 21:30:30 +02:00
MrMasterbay
a07324be82 fix: audit follow-ups — token-scoped rate reads, i18n relative-time, arm64 CI gate, BMC tests
Low-severity items from the Testing-branch audit:
- power/cost list_rates: scope by the token's floored effective_role, not the
  owner's account role, so an admin-owned but scoped token can't enumerate every
  cluster's rates (mirrors the upsert admin check).
- i18n relative-time: 'timeAgo' was always appended ('5s vor' in German). Split
  into locale-ordered timeAgoSec/timeAgoMin templates for all 8 locales
  (de/fr/es/pt lead, en/zh/ko/it trail).
- release-images CI: arm64 builds under qemu emulation; a failed arm64 leg no
  longer blocks the amd64 release (continue-on-error on the arm64 matrix legs;
  publish already attaches whatever artifacts exist).
- tests: BMC in-band read key->agent->password fallback order (#609) now has
  direct coverage (3 tests).

(OIDC cross-source collision finding intentionally left alone — OIDC is hands-off.)
2026-08-09 21:14:51 +02:00
MrMasterbay
13fefedb31 fix: audit remediation — LXC add-disk slot safety + 3 UI/i18n fixes
Follow-ups from the full Testing-branch audit:
- add_disk (LXC): a caller-supplied but already-occupied mpN — or 'rootfs' —
  was written straight through, REPLACING an existing mountpoint and orphaning
  its volume. Now coerce an occupied / rootfs / QEMU-shaped id to the next FREE
  mpN, and reject a mount path containing ',' or '=' (mountpoint-option
  injection). (+3 regression tests)
- hardening UI: the verbose-mode select-unapplied handler treated the per-control
  {status,evidence} object as truthy, so it never pre-selected in verbose mode.
  Made it object-aware, matching checkHardening.
- storage: the ISO/template 'From URL' buttons gated on storage.download, but the
  download-url endpoint enforces storage.upload — aligned the UI gate.
- i18n(zh): dropped 10 backfill keys that duplicated ones already later in the zh
  map (last-wins made them dead; topTalkers text never took effect).
2026-08-09 21:01:51 +02:00
MrMasterbay
6515d0dca5 i18n(zh): restore the PegaProx brand name in the app title/subtitle
Two Simplified-Chinese strings dropped 'PegaProx' when translated:
clusterManagement (the header, e.g. shown as 'Proxmox VE 集群管理 1.0') and
loginSubtitle. Audited all 55 en strings that mention PegaProx — only these
two had lost it; both now read 'PegaProx — …'.
2026-08-09 20:37:07 +02:00
MrMasterbay
f335057d0a security(authz): scope cost-rate reads too (Aikido IDOR follow-up)
Same gap as the power-rate fix: GET /api/cost/rates listed every cluster's
cost rates to any authed user, and GET /api/cost/rates/<id> had no tenant gate.
Scope the list via get_user_clusters (keep __default__) and add the
check_cluster_access gate to the single-cluster read, mirroring api/power.py.
+ 3 regression tests.
2026-08-09 20:13:48 +02:00
MrMasterbay
82f7cfe0c8 fix(hardening): keep VNC-tunnel forwarding + make applied CIS controls rollback-selectable
- sshd_hardening (#433): stop writing 'AllowTcpForwarding no'. PegaProx's VNC
  console tunnels to the node over an SSH direct-tcpip channel (utils/vnc_tunnel.py),
  which that directive blocks — and the control never documented it. The sed cleanup
  still strips a stale directive so a node hardened by an older build heals back to
  the sshd default (forwarding on) on re-apply.
- Node hardening UI: an already-applied control rendered a static check icon, so it
  could not be ticked — leaving 'Rollback Selected' with nothing to select. Applied
  controls now render a real (green-ringed) checkbox; the 'Applied' title badge still
  marks their state. (+ appliedTickToRollback DE/EN)
2026-08-09 20:08:31 +02:00
MrMasterbay
965ec2ce6e security(authz): close 2 Aikido code-audit gaps (portal reboot perm + power-rates IDOR)
- client_portal _vm_power: reboot was gated on vm.start, so a start-only
  portal user could reboot (stop+start) a guest. Now gates on vm.restart,
  matching the main VM API and the scheduler. (Aikido #469089251)
- api/power.py list_rates (GET /api/power/rates): returned every cluster's
  power/cost rates to any authed user. Now scoped via get_user_clusters
  (admins / default tenant see all), always keeping the shared __default__
  fallback row. (Aikido #469089241)
- tests: tests/test_aikido_batch3.py (5) — scoped vs admin list + the
  reboot/start/shutdown permission mapping.
2026-08-09 20:00:23 +02:00
MrMasterbay
d6da9c1e60 i18n(zh): backfill 13 Simplified Chinese strings missing from the zh map
These keys existed in en (incl. the new mpN mount-path fields) but not in
zh, so Chinese users got the English fallback via t(). Adds the add-disk
mount-path strings, guest-detail / monitoring labels (address, services,
filesystems, interfaces, logged-in users, clock skew) and the top-talkers
table + hint.
2026-08-09 19:53:38 +02:00
MrMasterbay
efa1972897 Merge #677: fix critical Chinese-audit issues (i18n, @ranydb)
Resolves two real regressions from #670 that affected ALL languages, not
just Simplified Chinese:

- hoist the can() permission helper to component scope — it was defined
  only inside the tab-filter callback, so can('vm.snapshot') / can('vm.backup')
  elsewhere threw a ReferenceError and blanked the Snapshot Schedule view.
- drop stray JSX tokens ()))} / )}) that rendered as literal text on the
  topology diagram and the report page.
- localize hardcoded live/polling, network, kernel, score, socket labels
  and add the Simplified Chinese values (t() falls back to English for the
  other locales).

index.html rebuilt on top of the current Testing bundle so it stays
consistent with the mpN add-disk change already on Testing.
2026-08-09 19:47:01 +02:00
MrMasterbay
50d77140f5 fix(vm): container disks use mpN mountpoints + reliable unlock fallback
- add_disk: containers were sent scsi0/scsiN keys, which PVE rejects
  ("property is not defined in schema"). Coerce any non-mp id onto the
  next free mpN and always pass a container-side mount path (mp=).
- AddDiskModal: for CTs default to the next mp slot, drop the QEMU
  bus/format/iothread/ssd controls, add a Mount-Path field (+ DE/EN i18n).
- unlock_vm: plain delete=lock failed for clusters wired up with a root
  API token ("Only root may use this option"). Send skiplock when
  root@pam and fall back to qm/pct unlock over SSH when the API cannot.
- tests: mpN coercion + slot picking + default mount path + QEMU
  untouched; skiplock / token->SSH fallback / not-locked / both-fail.
2026-08-09 18:14:23 +02:00
MrMasterbay
fdb893d03f Merge main into Testing: reconcile the PR-guard commit so main stays FF-able 2026-08-09 15:00:17 +02:00
MrMasterbay
f34c8d3b8a ci: auto-retarget external PRs opened against main onto Testing
Contributors open PRs against the default branch (main), but we develop on Testing
and FF main only at release. This pull_request_target workflow retargets external
PRs opened against main onto Testing + leaves a friendly note. Maintainers
(OWNER/MEMBER/COLLABORATOR) and a base:main label are exempted. It never checks out
PR code, so the write token is not exposed to untrusted input.

NOTE: GitHub reads pull_request_target workflows from the PR base branch, so this
must also live on main to fire for main-targeted PRs (reaches main at the next FF).
2026-08-09 14:59:13 +02:00
MrMasterbay
286432d88e ci: auto-retarget external PRs opened against main onto Testing
Contributors open PRs against the default branch (main), but we develop on Testing
and FF main only at release. This pull_request_target workflow retargets external
PRs opened against main onto Testing + leaves a friendly note. Maintainers
(OWNER/MEMBER/COLLABORATOR) and a base:main label are exempted. It never checks out
PR code, so the write token is not exposed to untrusted input.

NOTE: GitHub reads pull_request_target workflows from the PR base branch, so this
must also live on main to fire for main-targeted PRs (reaches main at the next FF).
2026-08-09 14:34:08 +02:00
MrMasterbay
978f0a9701 Merge main into Testing: OIDC username-collision fix (#651, @rmalchow)
Brings rmalchow s #651 (keep the full preferred_username so bob@corp.com and
bob@partner.com no longer collapse onto one account) onto Testing. It had been
merged to main by mistake (a silent gh retarget failure), so merge main back in
to land it where our work actually ships and keep Testing a descendant of main.
2026-08-09 14:23:13 +02:00
Nico Schmidt
48877f41db
Merge pull request #651 from rmalchow/fix/486-oidc-username-truncation
oidc(#486): keep the full preferred_username instead of truncating at…
2026-08-09 14:18:04 +02:00
Nico Schmidt
05075d6dd9
Merge pull request #676 from ranydb/fix/language-switcher-translation-context
fix(i18n): provide translator to language switcher
2026-08-09 14:07:20 +02:00
yanjunpu
b94340bfef fix(i18n): address critical Chinese audit issues 2026-08-09 17:31:07 +08:00
gyptazy
d7f46bc395 feature: Build ARM64 / aarch64 hw arch artifacts (#674)
Landed on Testing from gyptazy s PR #674 (fast-forwards to main at the next
release, per our Testing-first workflow). Verified on the contributor s own fork
run: LXC + VM artifacts for BOTH amd64 and arm64 built, boot-tested and published.

Follow-up (non-blocking): the arm64 matrix legs share the publish success gate, so
a future arm64 failure would also hold the amd64 release — decouple later.
2026-08-09 10:20:52 +02:00
mkellermann97
c995284124 ci(docker): harden the Testing docker workflow (daily security scan)
Two non-blocking findings from the Aikido + CodeAnt daily scan on the new
docker-testing.yml:
- checkout persisted GITHUB_TOKEN into .git/config (Aikido, med). This job never
  pushes git, so set persist-credentials: false — a compromised downstream action
  can no longer lift the token from the checkout.
- workflow_dispatch could be fired from any branch and publish it as
  pegaprox-testing:latest (CodeAnt). Guard the job on refs/heads/Testing so a
  dispatch from elsewhere is a harmless no-op.
2026-08-09 08:46:02 +02:00
yanjunpu
cebadfd045 fix(i18n): provide translator to language switcher 2026-08-09 13:36:33 +08:00
Nico Schmidt
57e20e16e9
Merge pull request #670 from ranydb/feat/simplified-chinese-i18n
feat(i18n): add Simplified Chinese support
2026-08-08 21:42:52 +02:00
yanjunpu
2b4d327d74 feat(i18n): add Simplified Chinese support 2026-08-09 03:16:52 +08:00
MrMasterbay
df09fe22a6 ci(docker): build a dev image on every Testing push
New workflow, kept separate from the release-tag-only docker.yml so it never
touches the release build. On every push to Testing it builds the Dockerfile and
publishes to its own package ghcr.io/pegaprox/pegaprox-testing (:latest + a
:sha-<short> immutable tag). amd64-only + no-cache (clean build each time so the
apt security upgrade re-runs), concurrency cancels superseded builds.

The pegaprox-testing GHCR package is created private on the first run — flip it to
public once in the package settings.
2026-08-08 19:42:40 +02:00
MrMasterbay
f24c5a5d65 fix(test): model an empty ssh_key on the fake BMC manager (#609 CI)
The in-band BMC read now tries key -> agent -> password SSH (26b313d). The fake
manager is a MagicMock, so mgr.config.ssh_key was a truthy auto-attribute and the
read went down the (unstubbed) key branch instead of the agent path the hardware
tests stub + assert -> test_read_after_consent_returns_hardware_200 got available=False.
Real PegaProxConfig.ssh_key is a string (empty when unset); set it to '' on the fake
so it models a node with no key configured. Full suite 401 green.
2026-08-08 02:00:45 +02:00
MrMasterbay
26b313d94b fix(hardware): fall back to key/password SSH for the in-band BMC read (#609)
read_node_bmc_inband only ever tried _ssh_run_command_output (agent/known-key auth).
On a node reachable purely by password — the common root@pam setup with no SSH key
deployed — that returns nothing, so the in-band IPMI enable failed with
"no response from node (SSH unavailable?)" even though the stored cluster password
would have worked. Every other node-SSH path (HA, maintenance) already does the
key -> agent -> password fallback; this one did not.

Add the same fallback chain, reusing the existing helpers. INBAND_PROBE_CMD always
echoes a marker on a live shell, so an empty/None result reliably means the auth
method did not connect and it is safe to fall through. The password helper uses
sshpass -e via the SSHPASS env var (not argv), so no credential lands on the
command line.

Verified live against a root@pam cluster with no SSH key: the old agent-only path
returned None (-> "SSH unavailable"); with the fallback the password auth connects
and the read now correctly reports "ipmitool is not installed" on the nodes.
Reported by si458 on #609.
2026-08-08 01:27:00 +02:00
MrMasterbay
63ce5de69d fix(v2p): refresh a stale ESXi session before the migration plan/start
On a long-lived server the ESXi REST/CIS session expires, and unlike every other
vmware route (get_vmware_vms, single-VM GET, snapshots, performance, ... all call
mgr.ensure_connected() first) the migration-plan and migrate-start handlers went
straight to get_vm(). get_vm() then hit /api/vcenter/vm/<id> on the dead session and
returned {error: VM not found}, so the wizard reported the VM missing and the
migration never started — i.e. "VMware migration stopped working" after the box had
been up a while, even though the VM list (which does reconnect) still showed it.

Add mgr.ensure_connected() at the top of get_vmware_migration_plan and
start_vmware_migration, mirroring the other routes. Reproduced + confirmed live
against a real ESXi 8.0.3 host: plan 400d on the stale session, then returned the
full plan once the session was refreshed; a full 90GB VM migrate then ran end-to-end.
2026-08-07 18:20:12 +02:00
MrMasterbay
77bfe71d98 security: enforce authz/validation gaps from the Aikido Testing-branch pentest (batch 2/2)
Second half of the adversarially-verified findings (batch 1 = 35d4078). Each fixed,
code-reviewed and covered by tests/test_aikido_batch2.py (17 new, full suite 401 green).

- power: only a global admin (effective_role) may overwrite the shared __default__
  power-rate row — a cluster.config holder edits only its own cluster
- metrics exporter: /api/metrics requires an admin-role token, not any valid token
  (it emits cluster-wide, cross-tenant infra gauges)
- portal: build_authz_user in _vm_power so a token's effective_role is honoured;
  invalidate the user's other sessions on portal password change
- cluster-groups: treat a global (tenant_id NULL) group as admin-only for the
  delete + balance-now writes, matching the earlier update fix
- vm-tags: reject a non-numeric vmid before the global DELETE+rewrite, and roll back
  save_vm_tags on error so a mid-loop failure can't persist a partial table wipe
- datacenter/multipath: allowlist the path_selector policy, and reject non-member
  nodes before SSH (no more `_get_node_ip(node) or node` fallback to a raw hostname)
- storage: pin http download-url fetches to the validated IP (DNS-rebind); https is
  left as the hostname since the node's TLS cert check already defeats a rebind
- multi-sdn: advertise an in-flight span's zone/controller so a concurrent purge
  can't tear down infra a create is still building (TOCTOU)
- ws-token validate: enforce node.shell for the standalone node SSH shell path
  (shell=node) — the VM termproxy path is unaffected
- SSE: scope vmware_vms / vmware_vm_detail to the server's linked_clusters instead
  of broadcasting guest_info/performance to every client; scope the portal audit
  task feed to the cluster it happened on (portal writers now set cluster=)
- LDAP: authoritative re-sync — rebuild LDAP-sourced perms/tenant_permissions from
  the current group mapping instead of only unioning them in, so group removal revokes
2026-08-07 17:31:57 +02:00
MrMasterbay
35d40781bd security: enforce authz gaps from the Aikido Testing-branch pentest (batch 1/2)
Adversarially verified all 45 AI-pentest findings against the current code; this batch
fixes the 10 HIGH + 2 MED confirmed REAL (17 were by-design/known admin-wide debt, 2
already fixed, 1 false positive). Full in-process test suite is green — the one case that
failed was a test that encoded the very BMC credential-exfil this now guards; corrected it
and added the negative case.

- rbac.py: user_can_access_vmware_vm honours effective_role (token-scoped), like its Proxmox
  twin, so an admin-owned viewer-scoped API token no longer gets the full-admin VMware bypass.
- nodes.py / vmware.py: reject a masked password paired with a CHANGED host - the stored BMC /
  VMware secret can no longer be shipped to a caller-chosen host (credential exfiltration).
- schedules.py: create/update enforce the same per-VM ACL as the live action (build_authz_user
  + user_can_access_vm), not just cluster reachability.
- pbs.py: restore (overwrite/test) enforces per-VM authorization on the destination VMID;
  backup-diff now requires pbs.datastore.view instead of plain pbs.view.
- users.py: create/update cap delegated permissions and custom roles to what the caller holds
  (a tenant admin.users delegate can no longer mint accounts with admin.* perms via the
  permissions[] list or the custom-role tier fallback); update_user rejects a foreign-tenant
  custom role; update_tenant is tenant-scoped like get_tenant_quota.
- groups.py: a global (tenant_id NULL) cluster group is admin-only for writes.
2026-08-07 15:31:13 +02:00
MrMasterbay
96a43f22a6 fix(migrations): route the ESXi + cross-hypervisor migration lists to their handlers (#654)
The @bp.route + @require_auth decorators for GET /api/vmware/migrations and
/api/xhm/migrations were attached to the _migration_reachable / _xhm_reachable IDOR
helpers instead of the list handlers — a copy-paste slip when those NS-Jul-2026 helpers
were inserted between the decorators and the intended functions. Flask therefore called
the helper with no args (TypeError: missing 't') and returned 500, so ESXi/XCP-ng
migration failed immediately and list_vmware_migrations / xhm_list were never routed.

Move the decorators back onto the real list handlers; the reachability/IDOR filter is
unchanged (still applied inside the handlers). Verified via url_map: /api/vmware/migrations
now maps to vmware.list_vmware_migrations and /api/xhm/migrations to xhm.xhm_list.

Confirmed live in #654 (miketate1985) and flagged by the CodeAnt full-repo sweep
(vmware.py:1130, xhm.py:194).
2026-08-06 21:30:06 +02:00
mkellermann97
54a8f3ca9d fix(console): hold Shift when pasting shifted symbols into the VNC console (#653)
typeTextToVM() sent each character keysym straight to sendKey(), but qemu's VNC
keysym path does not assert Shift for the symbol row, so every shifted symbol landed
on its base key unshifted (! -> 1, @ -> 2, _ -> -, { -> [, | -> backslash, : -> ;),
making console login impossible for any password containing symbols. Wrap those
characters in a Shift_L press/release around the base-key keysym — the same thing the
browser does when you type them by hand (which already worked). Uppercase is untouched
since qemu shifts A-Z on its own. US-layout map, matching the paste path we already had.
2026-08-06 15:21:12 +02:00
mkellermann97
bdac211d56 fix(ui): rotate the chevron on the Advanced Settings panel when open (#652)
The native <details> "Advanced Settings" accordion had a static ChevronDown and no
`group` class, so the arrow never flipped when the panel was expanded — it always
pointed the same way regardless of open/closed. Add `group` + group-open:rotate-180
so it points down when open, matching the other collapsibles.
2026-08-06 09:41:52 +02:00
MrMasterbay
2ca4804872 fix(update): scheduled rolling updates evacuate local disks by default (#630)
Unattended scheduled updates now default allow_local_disks=True instead of only
following the cluster balance_local_disks flag — nobody is watching a 3am run, so
pausing it because one local-disk VM cannot live-migrate is the worst outcome. On
shared-storage clusters this is a no-op; a per-schedule override is honoured if set.

Also surface allow_local_disks + reboot_timeout in the manual run log summary so a
paused evacuation can be diagnosed from the log (was: only evacuation_timeout shown).
2026-08-06 08:22:16 +02:00
MrMasterbay
8c22bc7084 deps: trim requirements.txt comments
The per-pin CVE-history had grown into an essay. Stripped it — that context lives
in git blame/log anyway. Kept only short guardrail notes where a blind bump actually
breaks the app: the cryptography/pyopenssl/fido2 interlock and the SQLCipher platform
fallback. All 32 pins + version specifiers unchanged; dry-run still resolves clean.
2026-08-06 07:19:45 +02:00
MrMasterbay
2d4c2a29bf deps: raise cryptography floor to 50 (pin the vulnerable 49 out)
Now that pyopenssl 26.4.0 permits it, bump the cryptography floor 49 -> 50 so no
lockfile or older resolution can fall back onto cryptography 49 (CVE-2026-69247).
It stays unreachable in our code (zero PKCS#7, no RSA decryption — only Fernet/
AESGCM + RSA signing for ACME), but there is no reason to keep allowing it.

Effective range with pyopenssl's <51 cap is 50.x. Verified: the two pins co-resolve,
pip check clean, and the full crypto surface + self-signed cert gen run on 50.0.0.
2026-08-06 07:16:28 +02:00
MrMasterbay
bf21a12b33 deps: bump pyopenssl floor to 26.4.0 so cryptography can move to 50 (#650)
pyopenssl 26.3.0 capped cryptography <50, which is what actually held us on 49.
26.4.0 lifts that cap to <51. Our cryptography pin (>=49.0.0,<52) already permitted
50, so only the pyopenssl floor moves — the resolver now lands on cryptography 50.0.0
+ pyopenssl 26.4.0. Clears the Aikido cryptography-50 finding.

Done as a clean floor bump rather than the #650 autofix PR, which appended a duplicate
`cryptography==50.0.0` that conflicted with the pin above (and pinned exactly 50).

Verified on 50/26.4.0: `from OpenSSL import crypto`, _generate_self_signed, the ACME
RSA-sign path and the full Fernet/AESGCM/x509 surface all import + run clean, pip check OK.
2026-08-06 07:08:50 +02:00
ruben
9bb78ebb7a oidc(#486): bind the legacy-username fallback to the same subject
The back-compat lookup added in the previous commit adopted an existing
account whenever its key matched the local part of the derived username. That
re-opened the very collision the commit set out to close, just one step later:
bob@partner.com logging in found the 'bob' row that bob@corp.com had
provisioned pre-change and inherited its role and tenant.

A legacy row is now only reused when it is provably the same identity -- its
auth_source is an OIDC-family one and its stored oidc_sub equals the incoming
sub claim. Local and LDAP rows are never adopted. A sub is only unique within
an issuer and no issuer is recorded on the user row, so installs that repoint
at a different IdP still have to rekey by hand; likewise a pre-change account
that never completed an OIDC login has no oidc_sub and now needs a manual
rekey rather than being silently adopted.

'+' joins the allowed char set. sanitize_username() permits it, and dropping
it merged identities the same way dropping '@' did -- a+b@example.com derived
onto a real ab@example.com account.

oidc_derive_username() takes the already-loaded users table. Both callers in
the login path had one in hand and were loading it a second time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:36:45 +02:00
ruben
3f00e828ec oidc(#486): keep the full preferred_username instead of truncating at '@'
Username derivation split preferred_username at '@', so bob@corp.com and
bob@partner.com both provisioned onto a single 'bob' account and whoever
logged in second inherited the first one's role and tenant.

Keep the whole claim. '@' also has to join the allowed char set, or dropping
the split just yields 'bobexample.com'. Nothing downstream constrains it:
sanitize_username() already permits '@', LDAP provisioning already produces
such names, username is an unconstrained TEXT PRIMARY KEY, and no username
reaches a filesystem path, shell command or interpolated SQL.

Existing accounts keep their current key via a back-compat lookup, so no
install loses permissions and nobody is locked out under auto_create_users:
false. Installs that already hold a truncated 'bob' therefore keep the
collision -- closing that needs a rekey migration, left out deliberately.

The derivation was duplicated in the auto_create_users pre-check, with a
comment recording that a previous divergence caused 403s. Both call sites now
share one helper.

Adds the first test coverage for OIDC username derivation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:09:14 +02:00
mkellermann97
bfb10272aa fix(maintenance): honour local-disk evacuation on the maintenance-mode paths (#629)
The anti-affinity fix threaded --with-local-disks through the balancer, but the
manual maintenance-mode toggle (api/vms.py) and the scheduled-update path
(api/schedules.py) still called enter_maintenance_mode() without allow_local_disks,
so local-disk VMs bailed out of a maintenance evacuation even with "Balance VMs
with Local Disks" enabled. Pass the same cluster balance_local_disks flag through
both call sites (the rolling-update path in settings.py already did). Off by
default, so no behaviour change unless the setting is on.
2026-08-05 09:42:35 +02:00
MrMasterbay
df4633d828 fix(ui): keep the task-bar restore toggle up the whole time the bar is hidden (#648, #649)
Follow-up to 425b2c6 — drop the leftover `tasks.length > 0` condition so the
"show task bar" button stays available whenever the bar is hidden, not only
when a task happens to be around. Kills the empty-state dead end too.

Adopts si458s approach from #649 (he reported #648 and sent the fix).

Co-authored-by: Simon Smith <simonsmith5521@gmail.com>
2026-08-04 15:39:50 +02:00
mkellermann97
425b2c6302 fix(ui): restore the task-bar toggle button (#648)
A `false &&` debug guard slipped into 1.0 and hard-disabled the button that
brings the task bar back. Once you closed the bar with the X inside it there
was no way to get it back short of a reload. Drop the guard so the restore
toggle shows again while the bar is hidden and there are tasks to show.
2026-08-04 15:31:04 +02:00
mkellermann97
c798fd987c feat(console): surface SPICE in every VM console entry point + i18n
- SPICE was only in the sidebar right-click menu, so it was easy to miss.
  Add it next to the web-console action in the standard + corporate VM
  detail views, the resource table (all three layouts) and the Cloud skin.
- Unify the SPICE icon on ExternalLink (it opens an external viewer) so it
  reads distinct from the noVNC Monitor icon.
- Wire onOpenSpice through ResourceTable + VmDetailPanel + CorporateVmDetailView
  and expose openSpice on the Cloud action bundle.
- QEMU-only + running-gated everywhere (LXC / stopped VMs have no SPICE port).
- Add spiceConsole / spiceConsoleHint / spiceDownloaded / spiceUnavailable
  to all 7 locales.
2026-08-04 15:05:44 +02:00