rfc-0012: Isolation Backend interface - #2048
Conversation
|
All contributors have signed the DCO ✍️ ✅ |
|
I have read the DCO document and I hereby sign the DCO. |
Tighten RFC 0012 for review without changing its structure or tone: - Remove the undefined `BoundaryIntent` from `attach`, and state that `claim` is the only transition that binds sandbox identity, policy, agent, and resources; an attached boundary holds no claimed or untrusted workload. - Add the runtime interface definitions (`BoundaryExec`/`ExecSession`, `BoundaryPortForward`, `EventSource`) and their semantics: owned distinct stdio, PTY resize, placement-neutral exit and signals, stable repeated `wait`, validated loopback targets, and single-consumer events. - Make boundary confirmation mandatory (`MUST`) and lifetime-long, and state retry, termination, and cleanup semantics (only attach is auto-retryable; a lost `start_agent` returns the existing process; descendants are owned; `wait_terminated` and an idempotent `shutdown`). - Require trusted admission to supply the expected backend id, contract version, policy digest, and capabilities that a backend cannot lower; add the no-silent-weakening policy invariant; mark `VerifiedBoundaryDescriptor` privately constructed and distinguish envelope from contract version; make `ResourceBinding` driver-issued, bound, and non-wideable. - Make the rollout executable and forbid a silent default backend. topology-matrix: rename "sidecar proxy" to sidecar-assisted (the proxy stays with the supervisor), frame sidecar and node as composite backends, move the single-pod outer sandbox to the shared-kernel cell, and note a node enforcer alone is not `restricted`-compatible. codebase-grounding: re-pin anchors to the RFC's parent ba21bb3 (driver.rs and proxy.rs line numbers refreshed) and add a permalink base; correct the event wording so a denial carries `Evidence` plus request-specific L7 data. Set state: review and link PR NVIDIA#2048. Signed-off-by: Jordan Ganoff <jordan.ganoff@docker.com>
Tighten RFC 0012 for review without changing its structure or tone: - Remove the undefined `BoundaryIntent` from `attach`, and state that `claim` is the only transition that binds sandbox identity, policy, agent, and resources; an attached boundary holds no claimed or untrusted workload. - Add the runtime interface definitions (`BoundaryExec`/`ExecSession`, `BoundaryPortForward`, `EventSource`) and their semantics: owned distinct stdio, PTY resize, placement-neutral exit and signals, stable repeated `wait`, validated loopback targets, and single-consumer events. - Make boundary confirmation mandatory (`MUST`) and lifetime-long, and state retry, termination, and cleanup semantics (only attach is auto-retryable; a lost `start_agent` returns the existing process; descendants are owned; `wait_terminated` and an idempotent `shutdown`). - Require trusted admission to supply the expected backend id, contract version, policy digest, and capabilities that a backend cannot lower; add the no-silent-weakening policy invariant; mark `VerifiedBoundaryDescriptor` privately constructed and distinguish envelope from contract version; make `ResourceBinding` driver-issued, bound, and non-wideable. - Make the rollout executable and forbid a silent default backend. topology-matrix: rename "sidecar proxy" to sidecar-assisted (the proxy stays with the supervisor), frame sidecar and node as composite backends, move the single-pod outer sandbox to the shared-kernel cell, and note a node enforcer alone is not `restricted`-compatible. codebase-grounding: re-pin anchors to the RFC's parent ba21bb3 (driver.rs and proxy.rs line numbers refreshed) and add a permalink base; correct the event wording so a denial carries `Evidence` plus request-specific L7 data. Set state: review and link PR NVIDIA#2048. Signed-off-by: Jordan Ganoff <jordan.ganoff@docker.com>
efaeff7 to
c31cbda
Compare
bb9cefb to
b0b6a44
Compare
|
I'd be interested to see if we can support MXC using this interface, #2071. |
|
@maxamillion @derekwaynecarr @mrunalp thoughts? |
e19656f to
0924382
Compare
|
@TaylorMutch thanks for the review and discussion last week. Based on our discussion, I've made the following changes:
I've also updated my POC implementation for the current in-pod supervisor topology to align with these updates. Please let me know what you all think! |
|
This pull request has had no activity for 14 days and is now marked stale. It may be closed in 7 days if there is no further activity. |
Tighten RFC 0012 for review without changing its structure or tone: - Remove the undefined `BoundaryIntent` from `attach`, and state that `claim` is the only transition that binds sandbox identity, policy, agent, and resources; an attached boundary holds no claimed or untrusted workload. - Add the runtime interface definitions (`BoundaryExec`/`ExecSession`, `BoundaryPortForward`, `EventSource`) and their semantics: owned distinct stdio, PTY resize, placement-neutral exit and signals, stable repeated `wait`, validated loopback targets, and single-consumer events. - Make boundary confirmation mandatory (`MUST`) and lifetime-long, and state retry, termination, and cleanup semantics (only attach is auto-retryable; a lost `start_agent` returns the existing process; descendants are owned; `wait_terminated` and an idempotent `shutdown`). - Require trusted admission to supply the expected backend id, contract version, policy digest, and capabilities that a backend cannot lower; add the no-silent-weakening policy invariant; mark `VerifiedBoundaryDescriptor` privately constructed and distinguish envelope from contract version; make `ResourceBinding` driver-issued, bound, and non-wideable. - Make the rollout executable and forbid a silent default backend. topology-matrix: rename "sidecar proxy" to sidecar-assisted (the proxy stays with the supervisor), frame sidecar and node as composite backends, move the single-pod outer sandbox to the shared-kernel cell, and note a node enforcer alone is not `restricted`-compatible. codebase-grounding: re-pin anchors to the RFC's parent ba21bb3 (driver.rs and proxy.rs line numbers refreshed) and add a permalink base; correct the event wording so a denial carries `Evidence` plus request-specific L7 data. Set state: review and link PR NVIDIA#2048. Signed-off-by: Jordan Ganoff <jordan.ganoff@docker.com>
Introduces RFC 0012, the runtime-selectable Isolation Backend contract: a pluggable component that establishes and enforces an agent's isolation boundary across network, filesystem, syscall, and identity, while the supervisor stays the policy authority (proxy, policy, audit) and the agent's only egress. Includes the supporting topology matrix and codebase-grounding notes. Signed-off-by: Jordan Ganoff <jordan.ganoff@docker.com>
Tighten RFC 0012 for review without changing its structure or tone: - Remove the undefined `BoundaryIntent` from `attach`, and state that `claim` is the only transition that binds sandbox identity, policy, agent, and resources; an attached boundary holds no claimed or untrusted workload. - Add the runtime interface definitions (`BoundaryExec`/`ExecSession`, `BoundaryPortForward`, `EventSource`) and their semantics: owned distinct stdio, PTY resize, placement-neutral exit and signals, stable repeated `wait`, validated loopback targets, and single-consumer events. - Make boundary confirmation mandatory (`MUST`) and lifetime-long, and state retry, termination, and cleanup semantics (only attach is auto-retryable; a lost `start_agent` returns the existing process; descendants are owned; `wait_terminated` and an idempotent `shutdown`). - Require trusted admission to supply the expected backend id, contract version, policy digest, and capabilities that a backend cannot lower; add the no-silent-weakening policy invariant; mark `VerifiedBoundaryDescriptor` privately constructed and distinguish envelope from contract version; make `ResourceBinding` driver-issued, bound, and non-wideable. - Make the rollout executable and forbid a silent default backend. topology-matrix: rename "sidecar proxy" to sidecar-assisted (the proxy stays with the supervisor), frame sidecar and node as composite backends, move the single-pod outer sandbox to the shared-kernel cell, and note a node enforcer alone is not `restricted`-compatible. codebase-grounding: re-pin anchors to the RFC's parent ba21bb3 (driver.rs and proxy.rs line numbers refreshed) and add a permalink base; correct the event wording so a denial carries `Evidence` plus request-specific L7 data. Set state: review and link PR NVIDIA#2048. Signed-off-by: Jordan Ganoff <jordan.ganoff@docker.com>
Contract completion: add MediationIngress/MediatedConnection so the proxy receives connections placement-neutrally; complete the runtime types (BoundaryExitStatus/BoundarySignal, BoundaryTerminal::resize, IdentitySource Send+Sync) and a machine-readable BackendErrorKind; name the AdmittedBoundary- Requirements admission block retained by VerifiedBoundaryDescriptor; show claim_id/generation in ClaimContext; require a canonical bounded descriptor encoding; define capability comparison. Mark MediationIngress, the admission block, and claim generation as contract additions the in-pod POC does not yet implement. Security semantics: make confirmation mandatory and lifetime-long, but correct the enforcement model: Landlock/seccomp are monotonic while netns/nftables are mutable, so require the mutation authority fenced from the workload, atomic default-deny updates, and fail-closed on lease loss rather than calling the boundary irreversible. "Confirms effective enforcement", not "reads rules back". Bind Attested identity to sandbox and claim id/generation. Tighten the execution-domain invariant (admission-visible named capability + audit, whole descendant trees). Distinguish structural pre-Running ordering from dynamic post-termination rejection; note adoption must remove legacy bypass paths. Lifecycle: drop the "establishes before supervisor boot" contradiction; add shutdown to the supervisor sequence; rewrite the in-pod migration appendix (policy-free attach, netns/rules at claim/bind, agent().wait(), wait_terminated, idempotent shutdown). The RFC is implemented only after the in-pod backend passes its release gates. Topology: narrow node-enforcer claims to "network-setup capabilities off the pod"; nuance NetworkPolicy (coarse L3/L4 defense in depth, not a proxy replacement); drop the proxy-policy-reprograms-boundary claim; label the containment table a kernel-compromise ceiling. Grounding: re-pin to the RFC's parent a5161d0 (process.rs and proxy.rs line numbers refreshed; others unchanged and reverified). Signed-off-by: Jordan Ganoff <jordan.ganoff@docker.com>
This removes a lot of detail to keep the RFC succinct.
The RFC should describe the end state, and not intermediate milestones or POCs.
- Introduced topology nomenclature to align with how we've been talking about this in other conversations - Simplified the explanation in the proposal section - Clarified binary identity - Simplified the initial contract to require the minimal interface necessary to satisfy all known topologies - Confirmed this will work with the proposed mxc (RFC 0013) proposal
- Repin codebase-grounding.md to 8eacb47 (sidecar supervisor topology, NVIDIA#2076); update capabilities line numbers (1534→2538, 1540→2544), init container line numbers (191→423, 993→1506, 1185→2113), and remove the stale "no native sidecars today" claim. Add openshell-network-init and openshell-supervisor-network sidecar entries and expand the rg pattern. - Add Implementation column to topology-matrix.md; mark Co-located/in-pod and Same-pod composite as implemented (original topology and NVIDIA#2076 respectively); remaining patterns noted as proposed. Signed-off-by: Jordan Ganoff <jordan.ganoff@docker.com>
All line numbers and function names verified against the post-sidecar state of the codebase (commit 8eacb47, NVIDIA#2076). Changes: - process.rs: ProcessHandle::spawn 440→527, netns param 446→535; drop_privileges call sites 603/700→710/812, enforcement 613/705→721/818-819; enter_netns_and_sandbox now documented at ssh.rs:1245 - CLONE_NEWNET call sites: process.rs:589→695, ssh.rs:619/1186→653/1262, supervisor_session.rs:610→735, netns/mod.rs:363→342 (226 unchanged) - CLONE_NEWNS: was one unshare at :393; now unshare at :449 and a new setns at :480 added for sidecar mount-namespace entry - nft fail-open: line 264→265, return Ok(()) range 272-277→277; note that the sidecar path (netns/mod.rs:477) requires nft and returns an error if absent, fixing the invariant bug for the sidecar topology - nft_ruleset.rs: policy accept 41→53; accept rules 43-49→56-92; reject rules now at 106+ - VM driver MASQUERADE: runtime.rs:417/436→418/437 - Agent command: main.rs:331→601; sleep infinity driver.rs:1886→2937, clarify it is set via SANDBOX_COMMAND env var - OPA evaluation: proxy.rs:1611 / evaluate_opa_tcp renamed to authorize_egress_intent at proxy.rs:1955; NetworkInput built at :2032 - openshell.proto: clarify no lifecycle Attach; note AttachSandboxProvider (provider record attachment, not isolation lifecycle) - README.md appendix: update pinned commit reference a5161d0→8eacb477 Signed-off-by: Jordan Ganoff <jordan.ganoff@docker.com>
0924382 to
3f5bc31
Compare
russellb
left a comment
There was a problem hiding this comment.
Reviewed this against the in-progress cni-sidecar supervisor topology (#2606), which moves nftables installation from an in-pod NET_ADMIN init container to a node CNI DaemonSet. It lines up well with this RFC — it's a concrete instance of the "Delegated backend components" row, and the direction (privilege out of the agent container, one supervisor operating the boundary) is exactly what we built toward.
Most of my notes are inline. The substantive one is invariant 6: as written it assumes the backend can detect standing-enforcement loss at runtime, which a node-delegated backend structurally can't do from inside the pod. I think the contract needs a small amount of give for delegated backends there and in the Ready/confirmation wording; everything else is either a clean fit or a mechanical reshape on our side.
| |---|---|---|---|---| | ||
| | **Co-located/in-pod** | With the workload | Backend runs in the supervisor process; network mediation is co-located | Trusted components share the workload's host, guest, or application kernel, depending on the runtime | Implemented (original topology) | | ||
| | **Same-pod composite** | With the workload | Backend runs with the supervisor; network mediation may run in a sidecar | Components share the workload's kernel | Implemented (#2076) | | ||
| | **Delegated backend components** | With the workload | A node or remote helper establishes some controls; network mediation may be co-located or delegated | Depends on which trusted components remain with the workload | Proposed | |
There was a problem hiding this comment.
Concrete grounding for this row: the cni-sidecar topology (#2606) implements exactly this shape — a node CNI DaemonSet establishes standing enforcement (egress-redirect nftables in the pod netns at CNI ADD), while network mediation stays co-located as the in-pod network sidecar. Might be worth citing it here the way the two rows above cite #2076, since it's the first real instance of "a node helper establishes some controls" and it surfaces the confirmation/loss-detection questions I've flagged on the README.
| 3. The backend and network mediation authorize an operation only when the complete effective policy permits it; there is no silent weakening. | ||
| 4. Agent activation, `exec`, and forwarding occur only through the active backend, and every workload process remains in the compute driver's provisioned execution environment. | ||
| 5. Shared infrastructure preserves strict per-boundary lifecycle, policy, identity, enforcement, and cleanup isolation. | ||
| 6. Loss of required enforcement ends `Running` and terminates all workload processes. Network-mediation unavailability denies outbound connections and never enables direct egress. |
There was a problem hiding this comment.
This is the one place cni-sidecar can't satisfy as written. The standing enforcement is node-installed nft/iptables state, and reading it back to detect loss requires NET_ADMIN — which this topology deliberately withholds from the pod. We tried an in-pod self-verify in #2606 and had to drop it: nft list/iptables -L return EPERM in the unprivileged sidecar, which fail-closed the check and crash-looped every healthy pod. (Your own codebase-grounding.md — the "reading it back does not prove 'only the proxy can egress'" row — makes a related point that reading rules is weak evidence anyway.)
For a node-delegated backend, detection and termination realistically live out of pod, on the node helper's reconcile cadence. Could the invariant explicitly allow, for delegated backends: (a) loss-detection that is periodic with a documented bound rather than in-band/instantaneous, and (b) termination effected by a control-plane or node actor rather than the in-pod supervisor? As written, the RFC's flagship "delegated backend" example can't claim conformance.
| The states have normative meanings: | ||
|
|
||
| - **Bound:** the topology descriptor and trusted sandbox context are bound to the same resource, and the network-mediation ingress is available. No untrusted workload code is running. | ||
| - **Ready:** the supervisor has initialized network mediation, the backend has confirmed standing enforcement, and the backend is prepared to ensure the admitted launch-time controls are in force before untrusted execution. |
There was a problem hiding this comment.
Same root issue as invariant 6, at confirmation time. cni-sidecar confirms standing enforcement via an out-of-pod gate (the CNI labels its node openshell.ai/cni-ready, the gateway pins a required nodeAffinity, and a wait-cni-coverage init container blocks the pod until the node acks coverage) — the in-pod component cannot self-confirm. Line 187 ("how a backend confirms its Ready conditions is implementation-specific") plus the helper-coordination language mostly covers this in spirit, but the normative "the backend has confirmed" reads in-band. A sentence explicitly blessing a provisioning-time / out-of-pod confirmation signal would remove the ambiguity.
|
|
||
| 1. Workload egress is denied except through network mediation for the boundary's lifetime. | ||
| 2. No untrusted instruction executes until every admitted control applicable to that process is in force. | ||
| 3. The backend and network mediation authorize an operation only when the complete effective policy permits it; there is no silent weakening. |
There was a problem hiding this comment.
Worth acknowledging cluster-scoped gating for node backends. cni-sidecar's node component is shared infrastructure whose "is this pod enforced at all" decision depends on cluster-scoped namespace registration — pods in unregistered namespaces are passed through without enforcement (a deliberate blast-radius guard). The per-netns rules keep invariant 5's per-boundary isolation intact, but a strict reading of "no silent weakening" collides with that pass-through unless an admitted cni-sidecar sandbox is guaranteed to be in a registered namespace before it runs. We hold that coupling (registration marker + wait-cni-coverage), so we stay conformant — but the contract should note that a delegated node backend may gate enforcement on cluster-scoped registration, with conformance resting on the admission↔registration coupling.
|
|
||
| - **Topology:** the placement and arrangement provisioned by the compute driver. | ||
| - **Isolation boundary:** the effective network, filesystem, syscall, and binary-identity enforcement around the workload. | ||
| - **Supervisor:** the trusted role that operates an active boundary through the Isolation Backend contract, from `attach` until termination. Each active boundary has exactly one logical supervisor. The in-sandbox supervisor process hosts this role today; another topology may host it in another trusted component. |
There was a problem hiding this comment.
Minor modeling note for delegated topologies: cni-sidecar runs two in-pod processes — the network sidecar (network mediation + policy authority + gateway session) and the process-supervisor (network-only; owns start_agent/exec), coupled over a control socket so they live and die together. That maps fine onto the RFC's NetworkMediationIngress vs RunningBoundary split as one logical supervisor composed of two processes, but the singular "one logical supervisor" phrasing invites the question. A clause confirming the logical supervisor may span multiple coupled processes would help.
Note
This RFC is a draft and open for feedback.
Summary
Adds RFC 0012 for the Isolation Backend, a proposed pluggable component that establishes and enforces an agent's isolation boundary.
OpenShell runs untrusted agent code inside an isolation boundary: the network, filesystem, syscall, and identity constraints that decide what the agent can reach and what can reach it. Building that boundary takes privilege, and today that privilege lives inside the agent's own container, beside the code it is meant to confine. That blocks restricted and multi-tenant clusters, whose Pod Security Standards reject the capability set the in-pod setup needs.
This RFC proposes making the boundary a pluggable component, the Isolation Backend, that separates the privileged work that builds the boundary from the supervisor that operates it as the policy authority (proxy, policy, audit) and the agent's only egress. The supervisor drives any backend through one runtime contract, so the privileged setup can move out of the agent's container (a sidecar, a separate pod, a microVM, a node component, and eventually outside the agent's kernel) without changing how the supervisor operates it. Each placement is a new backend, not a new supervisor.
Feedback especially welcome on: the lifecycle state boundaries and what each one guarantees, the provenance-based identity and attestation model, and the resource-binding split between the backend and the compute driver.
Related to #1737.
Checklist