Grant nothing by default. That is the short answer for any team that now has to decide what to allow on machines where Docker Agent is already installed. The longer answer is a mapping exercise: every capability Docker’s docker-agent README declares is, read from the operator’s side, a permission request. “Built-in tools + any MCP server (local, remote, or Docker-based)” is a request for an unbounded tool grant. “Push agents to any OCI registry, pull and run them anywhere” is a request to onboard executable artifacts through your registry. The provider list, OpenAI, Anthropic, Gemini, AWS Bedrock, Mistral, xAI, Docker Model Runner, defines your credential-exposure surface. Treat the README as a stack of requests, decide each one against a boundary you already enforce, and the decision stops being about whether agents are interesting and becomes about which gates they pass.
Two facts set the stakes. First, per the same README, the plugin ships pre-installed in Docker Desktop 4.63 and later, so docker agent is likely sitting on updated developer machines right now; the question is not whether to evaluate it but whether your policies reach it before someone’s experiment does. Second, every product fact in this article is vendor self-report: the README documents no independent audit of the permission model, no benchmark of the product, no failure modes, no rate limits. Vendor documentation also has an obvious commercial incentive to make onboarding sound frictionless. So every capability below carries an implicit tag: unverified until tested in your estate.
What each declaration actually requests
Docker Agent, per its README, is a Docker CLI plugin that defines agents in declarative YAML, attaches tools to them, and runs them, with multi-agent delegation between defined agents. Read as a permission manifest, the YAML is the unit of review: it names the tools, the model provider, and (by extension) the credentials the run will need. That is genuinely convenient for governance, because it gives you a static artifact to inspect before anything executes. It is also only as good as your discipline about reading it, because Docker’s README does not document any authorization layer the plugin enforces on top of what the YAML declares.
Here is the mapping this article works through:
| Declared capability (README) | What it actually requests | Boundary to apply | Evidence status |
|---|---|---|---|
| Built-in tools + any MCP server (local, remote, Docker-based) | Unbounded tool execution surface | Per-tool allowlist enumerated from the YAML; dry-run first | Vendor claim; MCP tool risk documented independently |
| Push/pull agents as OCI artifacts | Onboarding executables via registry | Existing image signing, scanning, registry allowlists | Vendor claim; image risk documented independently |
| Provider-agnostic models incl. Docker Model Runner | API keys or local-model access in agent context | Scoped, short-lived keys; audit each action | Vendor claim; no independent key-handling audit |
| Multi-agent delegation | One agent invoking others, compounding tool reach | Review the delegation graph, not just the entry agent | Vendor claim only |
| Pre-installed in Docker Desktop 4.63+ | Presence on developer machines by default | MDM/policy decision, not a vetting signal | Vendor claim only |
The rest of the article walks the rows that need argument.
Gate 1: Enumerate the MCP toolset from the YAML before granting any server
The phrase “any MCP server” deserves the most scrutiny, because the MCP ecosystem is large enough that wholesale grants are unmanageable on their face. MCPEvol-Bench catalogs 123 MCP servers exposing 1,272 tools across nine functional categories, with Software Development the largest, at 19.5% of the corpus. Nobody needs 1,272 tools. The practical move is to allowlist by category and then by individual tool: a documentation agent gets Media & Documentation tools, a coding agent gets a defined subset of Software Development tools, and nothing else is reachable.
There is a second reason for tight scoping that has nothing to do with malice. The MCP-Universe benchmark evaluates LLMs on realistic tool use across 11 MCP servers and 231 tasks in six domains, and its premise is that real MCP tool use is genuinely hard for current models. A model that fumbles tool selection in a benchmark will fumble it in your estate; the difference is that in your estate the tool it mis-calls might have write access. Fewer granted tools means a smaller blast radius for ordinary model error, not just for adversarial input. Note that these benchmark numbers characterize LLMs and the MCP ecosystem generally; Docker’s README reports nothing about docker-agent’s own tool-use reliability.
Mechanically, MCP connects models to tools over JSON-RPC 2.0 on STDIO and SSE transports, per the MCP-Universe paper. “Local, remote, or Docker-based” in the README therefore describes three different trust relationships, not three flavors of the same one: a local stdio server is code executing on the host with the agent’s privileges, a remote server is an external party receiving whatever the agent sends it, and a Docker-based server is at least nominally container-isolated. Each needs its own review question: what does this code do, where does this data go, what is actually isolated.
Gate 2: OCI agent images meet your existing supply-chain controls
The OCI packaging story is the part that should make platform teams relax slightly, and then only slightly. Because agents push to and pull from standard registries, per the README, agent onboarding routes through infrastructure you already govern: signing, vulnerability scanning, registry allowlists. If your estate blocks unsigned images today, an unsigned agent image should fail the same gate tomorrow with no new policy work. This is a real advantage of the packaging choice, and it is fair to say so.
It is not a safety guarantee. A vulnerability analysis of 2,500 Docker Hub images found that “the most severe vulnerabilities originate from two of the most popular scripting languages, JavaScript and Python”. Those languages are hard to avoid in this ecosystem: the MCP servers MCPEvol-Bench curates ship as NPM/TypeScript packages, so a granted toolchain is likely to carry them. OCI packaging shifts the review burden onto your gates; it does not reduce the risk inside the artifact. So the gate checklist is the one you already run: verify the signature against a publisher you have decided to trust, scan the layers, confirm the registry is on your allowlist, and pin a digest rather than a mutable tag. The new work is not the mechanism; it is the decision about which agent publishers earn trust, and on that question Docker’s documentation is silent.
Gate 3: Provider keys, from OpenAI to Docker Model Runner
Every entry on the provider list is a credential that will sit within reach of an agent’s execution context. The least-privilege research agenda is directly relevant here: a paper on task-conditioned least privilege for terminal and MCP agents proposes that “each action is audited before execution and again from observed effects along six dimensions of risk.” Translated into key policy: issue scoped, revocable, short-lived credentials per agent rather than a shared team key; log what the key was used for; and review whether the observed calls match the task the agent was granted for. Docker Model Runner changes the credential question (local inference, no external API key) but not the auditing one: the model’s outputs still drive tool calls that need review.
The harder problem is what flows the other direction. The ANX protocol proposal, whose comparative chapter surveys existing agent interaction methods, states the uncomfortable finding plainly: “All surveyed methods either expose sensitive data to LLM context or lack application‑level, unbypassable human confirmation.” Assume docker-agent is not the exception; Docker’s documentation does not show it is. Whatever the agent can read, files, environment variables, tool outputs, should be treated as potentially visible to the model provider on the next API call. That assumption drives concrete choices: no production secrets in the working directory, no broadly-scoped cloud credentials in the environment, and human approval gates you control at the application level rather than ones you hope the framework provides. This is one place where an exception would change the recommendation: if docker-agent’s own documentation, once independently reviewed, demonstrates unbypassable confirmation for destructive actions, the approval-gate requirement relaxes. Docker’s README neither confirms nor denies it.
Runtime placement is a latency decision, not a security one
Teams often treat “run the MCP server locally” as an isolation posture. The performance evidence does not support reading anything security-related into placement. An analysis of MCP server architecture patterns found that “transport overhead is dominated by network RTT, not by the protocol layer,” with in-host transport (stdio, loopback streamable-http) adding well under a millisecond. So choose local versus remote versus Docker-based servers on latency, data-residency, and operational grounds. Isolation comes from the boundary you draw, container policy, network egress rules, credential scope, not from the transport. A local stdio server running with your user’s full privileges is not safer than a remote one; it is faster, and it can read your SSH keys.
Guardrails worth borrowing
The academic work does not mention docker-agent, but three of its patterns import cleanly as grant conditions:
- Signed manifests. Verifiable manifest signing work for MCP pipelines proposes treating “each MCP tool-use manifest as a first-class security object whose canonical representation must be policy-validated, freshness-checked, digitally signed, verified before execution, and linked to tamper-evident audit evidence.” The Docker Agent YAML is a natural candidate for exactly this treatment: store it in version control, require review on changes, sign approved versions, and alert on drift.
- Pre-execution audit with re-audit from effects. The least-privilege framework audits before execution and again from what the action actually did. In practice this means logging tool calls with their arguments and periodically reconciling logs against granted scopes.
- Dry-run-only first exposure. A safety-bounded MCP gateway pattern from medical AI exposes only “selected MCP tools” in dry-run configurations, with “a policy and audit layer” separating “proposal validation from any device-side execution.” If a medical device gateway can stage exposure this way, a developer tool rollout can. Docker’s README does not document a native dry-run mode; that is a verification task. The pattern is worth imposing externally (for example, pointing the agent at read-only or stub MCP servers first) if the product lacks it.
The unverified ledger
Honesty about what this guide cannot establish, collected in one place so it does not dilute the argument above:
- Docker socket behavior: the README says nothing about whether, or how, docker-agent mounts or uses the Docker socket. Withhold socket and host access until you have observed what it requests in your own testing. Silence in the README is not proof of absence in either direction.
- Rate limits, cost behavior, failure modes: Docker’s documentation covers none of these for the product.
- The Desktop 4.63+ pre-install, the provider list, and the toolset claims are README self-report from a vendor with an interest in adoption, and likely to drift within a release cycle.
- The independent papers characterize MCP, LLMs, and container images generally. None of them audited docker-agent.
The checklist
For a team with image-signing and registry policy already in place, the pre-flight sequence resolves to this:
- Read the YAML. Enumerate every declared MCP server and built-in tool. Deny anything not on a per-tool allowlist; grant by category where the MCPEvol-Bench taxonomy helps you draw the line.
- Gate the image. Signature verification, layer scanning, registry allowlist, pinned digest. The Docker Hub vulnerability analysis is your reminder that JavaScript and Python images, the runtimes behind MCPEvol-Bench’s NPM-packaged MCP servers, concentrate severe vulnerabilities.
- Scope the keys. Short-lived, per-agent, per-provider credentials; Docker Model Runner where local inference removes the external exposure; audit per the pre-execution and observed-effects pattern.
- Dry-run first. Selected tools only, validation separated from execution, per the gateway pattern.
- Assume the confirmation gap. Until independently disproven, operate as though sensitive data can reach model context and no unbypassable approval exists, per the ANX paper.
- Withhold socket and host access until docker-agent’s actual behavior is verified in your estate.
My judgment: none of this is a reason to ban the tool, and the OCI packaging genuinely lowers adoption cost by reusing gates you already trust. I would pilot it on work that tolerates inspection and restart, behind the checklist above. What I would not do is treat “ships with Docker Desktop” as a vetting signal, because the only entity that has vetted docker-agent so far, as far as the evidence shows, is Docker.
Frequently Asked Questions
Is Docker Agent pre-installed in Docker Desktop?
per the same README, the plugin ships pre-installed in Docker Desktop 4.63 and later, so docker agent is likely sitting on updated developer machines right now; the question is not whether to evaluate it but whether your policies reach it before someone’s experiment does.
Does Docker Agent have a native dry-run mode?
Docker’s README does not document a native dry-run mode; that is a verification task. The pattern is worth imposing externally (for example, pointing the agent at read-only or stub MCP servers first) if the product lacks it.
Does the Docker Agent README document Docker socket usage?
Docker socket behavior: the README says nothing about whether, or how, docker-agent mounts or uses the Docker socket. Withhold socket and host access until you have observed what it requests in your own testing. Silence in the README is not proof of absence in either direction.

Join the discussion
Share a useful perspective or ask a question about this article.