Second half of the adversarially-verified findings (batch 1 = 35d4078). Each fixed,
code-reviewed and covered by tests/test_aikido_batch2.py (17 new, full suite 401 green).
- power: only a global admin (effective_role) may overwrite the shared __default__
power-rate row — a cluster.config holder edits only its own cluster
- metrics exporter: /api/metrics requires an admin-role token, not any valid token
(it emits cluster-wide, cross-tenant infra gauges)
- portal: build_authz_user in _vm_power so a token's effective_role is honoured;
invalidate the user's other sessions on portal password change
- cluster-groups: treat a global (tenant_id NULL) group as admin-only for the
delete + balance-now writes, matching the earlier update fix
- vm-tags: reject a non-numeric vmid before the global DELETE+rewrite, and roll back
save_vm_tags on error so a mid-loop failure can't persist a partial table wipe
- datacenter/multipath: allowlist the path_selector policy, and reject non-member
nodes before SSH (no more `_get_node_ip(node) or node` fallback to a raw hostname)
- storage: pin http download-url fetches to the validated IP (DNS-rebind); https is
left as the hostname since the node's TLS cert check already defeats a rebind
- multi-sdn: advertise an in-flight span's zone/controller so a concurrent purge
can't tear down infra a create is still building (TOCTOU)
- ws-token validate: enforce node.shell for the standalone node SSH shell path
(shell=node) — the VM termproxy path is unaffected
- SSE: scope vmware_vms / vmware_vm_detail to the server's linked_clusters instead
of broadcasting guest_info/performance to every client; scope the portal audit
task feed to the cluster it happened on (portal writers now set cluster=)
- LDAP: authoritative re-sync — rebuild LDAP-sourced perms/tenant_permissions from
the current group mapping instead of only unioning them in, so group removal revokes
Adversarially verified all 45 AI-pentest findings against the current code; this batch
fixes the 10 HIGH + 2 MED confirmed REAL (17 were by-design/known admin-wide debt, 2
already fixed, 1 false positive). Full in-process test suite is green — the one case that
failed was a test that encoded the very BMC credential-exfil this now guards; corrected it
and added the negative case.
- rbac.py: user_can_access_vmware_vm honours effective_role (token-scoped), like its Proxmox
twin, so an admin-owned viewer-scoped API token no longer gets the full-admin VMware bypass.
- nodes.py / vmware.py: reject a masked password paired with a CHANGED host - the stored BMC /
VMware secret can no longer be shipped to a caller-chosen host (credential exfiltration).
- schedules.py: create/update enforce the same per-VM ACL as the live action (build_authz_user
+ user_can_access_vm), not just cluster reachability.
- pbs.py: restore (overwrite/test) enforces per-VM authorization on the destination VMID;
backup-diff now requires pbs.datastore.view instead of plain pbs.view.
- users.py: create/update cap delegated permissions and custom roles to what the caller holds
(a tenant admin.users delegate can no longer mint accounts with admin.* perms via the
permissions[] list or the custom-role tier fallback); update_user rejects a foreign-tenant
custom role; update_tenant is tenant-scoped like get_tenant_quota.
- groups.py: a global (tenant_id NULL) cluster group is admin-only for writes.
Three critical/high auth-bypass findings from a CodeAnt AI-exploitation scan, each
independently verified as REAL (adversarially cross-checked + reproduced) and fixed:
- auth.py / users.py — off-boarding bypass: require_auth did get_user() or {}, so a
DELETED user resolved to {} -> {}.get('enabled', True)==True passed the disabled
check and the role fell back to the stale session role; get_user_clusters({})==None
= all-cluster read. delete_user's inline session purge was also DEAD (operated on a
stale active_sessions binding after load_sessions rebind) and never revoked pgx_ API
tokens. Fix: require_auth fails CLOSED when the record is gone (ACCOUNT_DELETED,
covers session AND token auth); delete_user uses invalidate_all_user_sessions() +
revokes api_tokens.
- push.py — cross-tenant alert leak: _alert_handler wrote every cluster alert (cluster/
node/VM names + live metric %) into EVERY push subscriber's inbox and woke them, with
zero tenant scoping. Fix: scope recipients to subscribers who can reach
alert_data['cluster_id'] via get_user_clusters (admin/default-tenant still get all;
cluster-less system alerts go to all; fail closed on lookup error).
- groups.py — cross-tenant cluster hijack: assign_cluster_to_group / rename_cluster only
gated the source cluster with check_cluster_access, whose ACL/pool fallbacks pass on
mere VM-level reach; a non-admin with admin.groups + one foreign VM-ACL grant could
move that cluster into their own group and gain cluster-wide tenant membership. Fix:
require real tenant ownership (get_user_clusters(include_pools=False)) of the source
cluster before re-grouping/renaming.
Regression tests (integration harness): test_integration_auth_deleted / _groups / _push
drive each through the real app. 260 passing (was 249).
api/groups.py read user.get('tenant') everywhere, but db.get_user()/
get_all_users() map the 'tenant' column to the dict key 'tenant_id' — so it
was ALWAYS None. Effect: get_cluster_groups + get_cluster_group_status fell
into the admin branch for every tenant user (cross-tenant group + metrics
leak), and trigger_xclb_balance_now had no ownership check at all — a
tenant-scoped cluster.config user could kick off REAL cross-cluster VM
migrations on another tenant's group (BOLA).
- new _user_tenant(user) helper: reads tenant_id; admin + implicit 'default'
tenant stay unscoped (None) so single-tenant installs are unaffected
- route all groups.py tenant lookups through it
- add group-ownership gate to balance-now (before the spawn) and lb-history
Verified: tenant-A user sees only tenant-A + global groups (not tenant-B);
balance-now/lb-history on a foreign tenant's group -> 403; own group passes
the gate; admin sees all.
Closes the 10 MEDIUM CWE-117 "Improper Output Neutralization for Logs"
findings from the 2026-06-04 Aikido scan. Eight were genuine user-input-
to-logger paths (URL params, request body, login-time username); two
(site_recovery, vmware-mgr-name) are admin-controlled config strings
but flagged anyway for consistency — applied the same fix shape so a
future grep for `logging.*{` stays clean and the next scan doesn't
re-litigate them.
Re-used the existing `sanitize_log_message()` helper from
pegaprox/utils/sanitization.py (already running on every audit log
write via utils/audit.py) — no new helper needed, no new dependency.
Aliased to `_sl` in each affected file to keep the f-strings compact.
Helper strips CR / LF / U+2028 / U+2029, leaves TAB and the rest of
the payload alone (legitimate in some action strings).
Callsites fixed (file:line, after the import-block shift):
* api/realtime.py:70 — WebSocket client connected: {username}
* api/realtime.py:183 — [WS-TOKEN] user '{data['user']}' no access
* api/static_files.py:177 — add_pool_member: cluster={cluster_id}
* api/static_files.py:186 — Request data: vmid={vmid}
* api/static_files.py:193 — Cluster {cluster_id} not found
* api/static_files.py:210 — Adding VM {vmid} to pool {pool_id} on {host}
* api/site_recovery.py:228 — Force-deleting plan '{plan['name']}'
* api/vmware.py:129 — [VMware:{mgr.name}] Skipped auto-connect
* api/webauthn.py:333 — [WebAuthn] auth_complete failed for {username}
* api/groups.py:439 — Error getting group status for {group_id}
Verified: all 6 modules import clean. sanitize_log_message roundtrip
test on a literal `evil\ninjected: secret_admin_action` returns
`evil injected: secret_admin_action` — newline collapsed to space,
content preserved. Plain str/grep over the 6 files confirms 10/10
callsites carry the _sl wrap.
E2E-shaped test against the running dev server requires a server
restart to pick up the imports — function-level smoke is the
load-bearing check here, the changes are pure str-wrap on the
log message path.