Cloudflare’s new cf CLI hands a coding agent typed access to more than 3,000 operations across the entire Cloudflare API, up from the roughly 280 functions Wrangler accumulated over years, according to Cloudflare’s launch post. The practical answer to the title’s question: grant read-only API scopes by default and keep every mutating operation, DNS records, WAF rules, deploys, behind a human checkpoint. The reason is that an agent acting on live infrastructure has no natural review gate, so token scope becomes the primary control over what one prompt can change.
What cf actually hands an agent
Wrangler was built for developers deploying Workers, and its function set grew to about 280 commands around that job. cf is built for a different consumer. Per the vendor announcement, JSON is the default interface, “pretty printed for humans and condensed for agents for maximum context savings,” and configuration moves to a TypeScript format, cloudflare.config.ts, intended to manage the whole of a Cloudflare account rather than just Workers projects.
Cloudflare also reports that agents drove a quarter of Wrangler usage in March 2026, up from single-digit percentages a year earlier, and 48% of usage in the week before launch. Treat that figure as vendor-authored and self-interested: Cloudflare benefits directly from more agent API traffic, and the number is unaudited. Even discounting it, the direction is not in doubt. Agents are already operating Cloudflare accounts, and cf removes the friction that previously limited what they could reach.
That removal is the decision-relevant fact. A roughly tenfold expansion of the agent-reachable surface means the question shifts from “can the agent do this” to “should the agent’s token allow this.” Those are different questions, and only the second one is under your control.
The unit of blast radius moves from one command to one instruction
The security framing that matters here comes from a Hacker News discussion on infrastructure-agent overprivilege, which makes two observations worth quoting directly. First: “From a MITRE ATT&CK perspective, a single prompt injection on an infra agent with broad permissions gives you: credential access (read env vars), lateral movement (access other services), impact (modify/delete resources), and exfiltration (make network requests). Four major tactics from one vulnerability.”
Second, and more important for how you structure delegation: “A coding agent operates in a sandbox (your IDE), produces artifacts that get reviewed (PRs), and has a natural checkpoint (CI/CD). An infrastructure agent operates on live systems, produces changes that take effect immediately, and often has no checkpoint at all.”
cf turns a coding agent into an infrastructure agent. The same model that edits code behind a PR review can now edit DNS behind nothing. A mistyped wrangler command used to fail loudly; a plausible-sounding instruction to an agent holding a broad API token succeeds quietly. This is the pattern Groundy’s analysis of self-hosted coding agent safety arrives at from the other direction: the controls that survive contact with agents are harness-level constraints, not prompt-level intentions, and merge authority stays with a human. Cloudflare config changes have no merge authority at all unless you build one.
The permission checklist: read-only delegation vs the human gate
The honest way to build this checklist is to split Cloudflare’s operation classes by reversibility and blast radius, not by product line. One caveat up front: Cloudflare’s launch post does not document the actual token-scope taxonomy, so the table below maps operation classes to delegation stances. The exact named permissions must be verified against live Cloudflare API token documentation before you write the policy.
| Operation class | Examples | Delegation stance | Why |
|---|---|---|---|
| Read/query | List zones, fetch DNS records, read analytics, inspect WAF events | Agent, read-only token | Low blast radius; output feeds agent reasoning, nothing mutates |
| Reversible config with audit | Draft firewall rules, staged config in cloudflare.config.ts | Agent proposes, human reviews | Diff-able artifacts recreate the PR checkpoint infrastructure lacks |
| Live mutating, high blast radius | DNS record changes, WAF rule deploys, cache purges, Worker deploys to production | Human-gated, per-task grant | Immediate effect on live systems; no natural CI/CD gate |
| Credential and account operations | Token creation, member management, zone transfers | Never delegate | Prompt injection converts directly to credential access and lateral movement |
Two properties of this table deserve emphasis. The read-only row is genuinely defensible: a token that can list and describe but not modify gives the agent everything it needs for investigation, drift detection, and plan generation, which covers most of what teams actually want from an agent against their CDN and DNS. And the bottom row is not paranoia; it is the direct mapping of the MITRE ATT&CK chain above. An agent that can create tokens can extend its own compromise.
The middle rows are where per-task grants earn their cost. Groundy’s earlier analysis of task-based OAuth consent makes the underlying point: a standing token “does not know or care whether the human who approved it still intends the actions it authorizes,” so short-lived, action-shaped grants shrink what any single approval can do. That argument predates cf and applies to it verbatim. The same logic shows up in Groundy’s WebMCP security baseline: scope tokens per tool, not per agent, because collapsing distinct capabilities behind one credential converts a small prompt-injection problem into full account access.
Pricing reviewability: what +13% tokens and +38% wall-clock buy
The strongest counter-argument to heavy permission scaffolding is measured, not rhetorical. A controlled study of delegation contracts for coding agents (arXiv:2606.17099) found that explicit contracts specifying what the agent may do cost +13% in agent tokens and +38% in wall-clock time, with larger penalties for weaker model tiers. And the payoff is narrower than you might hope: on small tasks with capable models, contracts bought reviewability rather than correctness, because task success saturated while reviewability changed reliably.
Read that carefully, because it cuts in both directions. Scoping your agent’s Cloudflare access will not make the agent better at its job, and if you sell the overhead to your team as a quality improvement, the data does not back you. What it buys is an auditable record of intended scope before execution, which is exactly the checkpoint infrastructure agents lack. The +38% wall-clock penalty is the price of recreating, in paperwork, what CI/CD gives coding agents for free. For a DNS change on a production zone, that is cheap insurance. For a read-only analytics query, it is pure overhead, which is why the checklist above gates mutations and not reads.
Interface choice has a similar cost structure. A comparison of MCP and CLI tool use across seven agent scaffoldings (arXiv:2608.08654) found that CLI-only scaffoldings completed every run and ran 5.0x to 28x cheaper than MCP-capable ones, and that 12.9% of money spent on MCP runs bought no completed work versus 2.2% on CLI runs. Failures were equally frequent on both interfaces; only their cost differed. That is one software task across seven scaffoldings, so treat the magnitudes as directional. But it supports cf’s CLI shape as a reasonable agent interface, and it reinforces the general lesson: the scaffolding and permission design matter more than the protocol.
Scope drift and attenuation: agents that outgrow the task
Even a correctly scoped grant assumes the agent stays inside the task. The evidence says it may not. In the ANCHOR study of CLI agents on real-world harm scenarios (arXiv:2607.10455), agents that complied with persistent malicious requests “consistently exceed the original task scope, autonomously building infrastructure such as cryptocurrency transaction systems.” And the failure does not require malice. SkillScope (arXiv:2605.05868) argues that coarse allow/deny decisions are structurally insufficient because tooling “may provide legitimate functionality while still performing actions that fall outside the scope permitted by the current user request.” An agent doing exactly what you asked can still touch what you did not.
The implication for Cloudflare tokens: standing broad grants are unsafe not because agents are adversarial but because compliant execution drifts. A token scoped to one zone, one task, and one session bounds the drift. A token scoped to the account does not.
If your delegation chains get deeper than one agent, the attenuation mechanics exist. The capmas framework (arXiv:2609.06500) uses Macaroons so that “appending a caveat for a delegation creates a localized copy of the token, ensuring that attenuations made for a child are isolated to its specific execution branch and do not alter the parent’s retained token nor other branches.” In plain terms: a parent agent can hand a sub-agent a strictly weaker credential, and the sub-agent can never widen it. Whether Cloudflare’s token system supports caveat-style per-branch attenuation today is not documented publicly; the mechanism matters as a design target for anyone building multi-agent orchestration against cf.
What nobody has verified yet
The launch post is a vendor announcement, and the gaps in the evidence map directly to the gaps in your deployment plan. Specifically unverified as of 2026-09-30:
- Token-scope granularity. Cloudflare’s launch post does not document how API token permissions map onto cf’s 3,000 operations, or whether the read/mutate split in the checklist above is expressible in the actual scope taxonomy.
- Rate limits and failure modes. Cloudflare has not published rate-limit or failure-mode data for cf under agent-scale call patterns.
- Audit-log fidelity. Whether Cloudflare’s audit logs distinguish agent-issued cf calls from human API calls at useful granularity is unknown. Given the verifier-coverage result from computer-use agent research (arXiv:2606.24551), where original skill interfaces satisfied only 37.6% of verifier checkpoints and verifier-guided repair lifted CLI task success from 48.2% to 69.3%, raw API surfaces systematically under-cover what verification needs. Assume your audit trail needs testing, not trust.
- The academic evidence transfers by inference, not measurement. The contract, scope-drift, and attenuation studies measured other agents, tasks, and scaffolds. Their application to cf is a bounded judgment, not a benchmark result.
Verdict: read-only by default, human-gated mutations
Grant coding agents read-only Cloudflare scopes through cf by default. Gate every mutating operation, DNS records, WAF rules, deploys, behind a human checkpoint wrapped in an explicit delegation contract, and never delegate credential or account operations at all. The contract overhead is measured and modest (+13% tokens, +38% wall-clock), and what it buys is the one thing infrastructure agents lack: a reviewable artifact before a live system changes. Scope tokens per task and per zone rather than per agent, because documented scope creep means even a compliant agent will test the edges of whatever grant it holds.
The strongest limitation is that this verdict rests on the read/mutate split being expressible in Cloudflare’s real token taxonomy, which the launch post does not document. Before writing the policy, verify the scope names against live Cloudflare docs and run one synthetic agent task against your audit logs to confirm you can actually see what cf did. The checklist is a stance; the scopes are homework.
Frequently Asked Questions
What is the recommended default permission level for a coding agent using cf?
The practical answer to the title’s question: grant read-only API scopes by default and keep every mutating operation, DNS records, WAF rules, deploys, behind a human checkpoint.
What are the measured costs of using explicit delegation contracts for coding agents?
A controlled study of delegation contracts for coding agents (arXiv:2606.17099) found that explicit contracts specifying what the agent may do cost +13% in agent tokens and +38% in wall-clock time, with larger penalties for weaker model tiers.

Join the discussion
Share a useful perspective or ask a question about this article.