Docker brings Sandbox Kit specification to the CNCF and packages AI agent permissions into OCI images
The Sandbox Kit Specification, now at version 3, turns what an agent like Claude Code or Codex can access into a common, versionable, auditable OCI artifact. Docker brought the proposal to the CNCF after testing it with AWS, Snyk, Datadog, and others.
The problem: agent permissions that never become an artifact
Agents like Claude Code and Codex install packages, call APIs, and use credentials on behalf of whoever triggers them. The problem, according to Docker, is that the grants that make this useful (bind mounts, broad tokens, open firewall rules) tend to live in shell history, dashboards, and the memory of whoever configured them, not in a reviewable artifact.
Docker calls this exactly the kind of fragmentation that OCI was created to avoid. Without a common standard, every agent runtime vendor would invent its own way of declaring what the agent can or can't do, repeating the pre-OCI scenario of incompatible image formats between Docker, rkt, and other runtimes.
The answer is the Sandbox Kit Specification, licensed under Apache 2.0, which Docker announced it is bringing to the CNCF during WeAreDevelopers, on September 24. The idea is simple to state: package the agent, its tools, and a typed list of the hosts, credentials, and volumes it requests inside a common OCI image, the same kind that already runs in any existing registry, scanner, or CI pipeline.
How the v3 specification works
In version 3, a Kit stopped being its own artifact type. There's no custom media type and no satellite file: the manifest carries a single declaration, vnd.docker.sandbox.kit.descriptor. This means a Kit is built with docker buildx build, pulled with docker pull, can be scanned and signed by tools that already exist, and can even be used as a base in a FROM.
Pinning the Kit's digest pins, at the same time, both the content and the permissions associated with it. This solves a practical supply chain problem: you can't silently change what an agent can access without also changing the hash of the image that distributes it.
The permissions themselves are typed, versioned declarations, such as com.docker.sandbox/network-policy@2 for network rules and com.docker.sandbox/credential@1 for credentials. Versioning by capability makes it possible to evolve the semantics of one permission category without breaking the others.
The GitHub CLI example: deny wins
In the example published alongside the specification, a Kit allows access to api.github.com but explicitly denies the DELETE verb on /repos/**. The resolution rule is straightforward: when one permission grants and another denies the same scope, the denial wins.
Credentials can be proxy-managed: a runtime that implements the specification injects the real token only into requests to the domains named in the Kit, while inside the sandbox there's only a sentinel value. In practice, the agent never sees the real credential, which reduces the risk of it leaking into logs, prompts, or the output of a poorly thought-out command.
An important detail: a Kit only requests permissions; the host decides. Without a runtime that implements the specification, the annotation stays inert, and if a mandatory permission can't be satisfied, the agent's launch is refused instead of running with partial access.
Mixins: composing permissions in layers
A launch combines a workload Kit, which provides the root filesystem, with any number of layered mixins. The order in which mixins are applied isn't the order of flags on the command line: it follows a dependency graph based on provides and requires.
Resolution fails in two cases: when a requires isn't met, or when two Kits declare the same provides. Overlapping declarations try to reconcile, with network rules merged by union; declarations incompatible with each other produce an error instead of an ambiguous permission.
Every descriptor also reduces to a normalized set of grants. That opens the door for a runtime that locks updates to record that set and block any new version of the Kit that expands it, including a version that removes a denial rule that used to exist.
Who has already tested it and what the CNCF said
Docker says it built the specification together with AWS, Box, Datadog, Dynatrace, JFrog, NanoClaw, OpenClaw, Palo Alto Networks, and Snyk, among other partners. The company draws a direct parallel with the donation of the image format and runc that resulted in the creation of the OCI itself, years ago.
CNCF's CTO, Chris Aniszczyk, commented on the announcement:
Standards are what let an ecosystem move fast without fragmenting, and few companies understand that better than Docker. By delivering Sandbox Kits as standard OCI images, Docker is giving the industry an open, repeatable way to package an AI agent, its tools, and its guardrails as one artifact.
Chris Aniszczyk, CTO of the CNCF
The published material doesn't say whether the specification has already been formally accepted into any CNCF program, nor what maturity level (sandbox, incubating, graduated) it would hold. Until the governance decision, Docker itself continues to maintain the project.
What changes for teams running AI agents in production
For teams in Brazil that already use Docker in build and deploy pipelines, the practical change is that reviewing an AI agent's permissions now fits into the same flow that already exists for container images: docker pull, vulnerability scanner, signing, and cluster admission policy. There's no need to adopt a new, separate governance tool, just to learn how to read a new descriptor.
This matters especially for those exposing agents like Claude Code or Codex to corporate repositories and credentials. Today, knowing exactly what an agent can access usually depends on opening configuration scattered across multiple places; with a Kit, that surface becomes a single artifact, versioned and with a fixed digest, that can be audited before approving a deploy.
The central caveat, acknowledged by Docker's own material, is that the specification only really works with a runtime that implements it, and today the only conformant runtime is Docker Sandboxes, which runs agents in microVMs with their own kernel. In other words, portability between different runtimes, the core promise of any open standard, hasn't been demonstrated in practice yet, because there's no second competing runtime to test it against.
How to start testing
Anyone who wants to experiment can find the project in the docker/sandbox-kit-spec repository, with runnable examples via the sbx CLI, such as sbx run ./hello --kit ./gh. Existing registry, scanner, and signing tools work with Kits without modification, because the artifact is still a common OCI image.
What actually requires learning is the descriptor's grammar and the semantics specific to each capability, such as the differences between declaring a network policy and declaring a proxy-managed credential. Docker says it keeps receiving feedback about Kit or runtime responsibilities that the specification still can't express, which signals that v3 isn't the final version of the design.
Translated from the Brazilian Portuguese original · Read the original
Flaw in ChatGPT app for Mac allowed theft of users' sensitive data
Researchers at the Objective-See Foundation found a loophole