groundy
Developer Tools

Debugging Vercel From the Terminal: When traces search Beats the Dashboard

Vercel's CLI traces search command offers fast terminal triage for recent incidents, but its one-hour window and 100-span cap require careful scope verification.

Published 4 references
A bubble-speckled green resin dinosaur with a skeptical eye lifts a yellow link with its tweezer-like snout above an ivory dish of green links, casting hard shadows on a warm ivory background.
On this page13 sections

Vercel added vercel traces search to the CLI in its September 29, 2026 changelog: a terminal-side query against trace spans that returns up to 100 spans from the last hour by default, with filters for status, deployment and trace ID, a subset of KQL for ad-hoc queries, and --json output aimed at scripts and coding agents. For incidents less than an hour old, this is now the fastest first probe. Everything else, the dashboard still owns.

What the command actually does, and what it doesn’t promise

The headline behavior is deliberate breadth without prerequisites, as Vercel’s changelog puts it: “By default, vercel traces search returns up to 100 spans from the last hour, newest first, so you can search without first finding a request or trace ID.” That matters operationally. The old terminal path against Vercel started with an identifier you had to find somewhere else, usually the dashboard. This command inverts the order: you sweep first, then identify.

On top of the default sweep sit dedicated flags. The changelog lists filters for “environment, service, span name, status, deployment ID, request ID, trace ID, or span ID,” plus --since and --until to narrow the time range. For everything those flags don’t cover, there is --query, which accepts “Vercel’s supported subset of Kibana Query Langugage (KQL),” including duration and span and resource attributes. The changelog’s own example targets slowness directly:

vercel traces search --query 'span.duration > 2000'

That returns spans slower than two seconds, which is the single most useful shape of query in incident triage.

The output contract completes the picture. “Add --json to output one JSON object per span, giving scripts and agents structured data to analyze.” One object per span is pipe-friendly: jq can group, count, or filter it without parsing tables, and an agent can consume it without screenshots. This is the same pattern Vercel applied when it moved feature-flag segments into the terminal with scriptable --json output, and it signals a consistent direction: the CLI is being built for programmatic consumers, and humans at a keyboard get the same structure.

Two things the changelog does not promise are worth stating plainly, because both are easy to assume. It does not say how far back traces are retained, and it does not enumerate the KQL subset. Both gaps become decision-relevant below.

A reproducible terminal procedure

The flags compose into a short triage sequence that works for any fresh incident, without touching the dashboard until you have a reason to.

Step 1: error-span sweep. Query the last hour for failed spans across the affected service or environment:

vercel traces search --status error --json

Narrow by service or environment if the project is noisy; the changelog doesn’t print those flag spellings, so run vercel traces search --help for the exact names. The --json output lets you count spans by name or deployment ID in one pipe, which answers the first triage question, is this one bad deploy or a broad regression, faster than any chart.

Step 2: slow-span query. If errors are absent but latency is the complaint:

vercel traces search --query 'span.duration > 2000' --since 1h --json

Adjust the threshold to the endpoint’s baseline. One check before you script this: the changelog says only “one JSON object per span” and never enumerates the fields, so run a single --json call and confirm the span objects actually carry duration values and trace IDs. The handoff to the next step depends on both.

Step 3: full-trace pull. Once a trace ID is in hand from the verified output of steps 1 or 2, pull every span in that request:

vercel traces search --trace-id <id> --json

This is the moment the terminal genuinely beats the dashboard: no clicking through a waterfall UI, no screenshot into a chat thread, just a structured artifact you can paste into an issue, diff against a healthy trace, or hand to an agent.

Step 4: escalate deliberately. If the sweep comes back empty and the incident is real, the explanation is almost always scope, not absence: the one-hour window, the 100-span cap, or an incident that predates the window entirely, where the CLI’s reach is undocumented and the dashboard or an external trace store is the reliable surface. That is the subject of the next section.

The silent defaults: a short window and a hard cap

The changelog frames the defaults as convenience, and they are. They are also a truncation mechanism, and the changelog does not spell out the failure mode.

An incident that started ninety minutes ago produces nothing in a default sweep, and whether the terminal can reach further back is undocumented. An incident that old may belong to the dashboard or an external trace store. A deploy that pushed error rates just above baseline in a high-traffic service can push the relevant spans past the 100-span cap within minutes; newest-first ordering then guarantees you see the tail of the incident, not its onset. Neither condition produces a warning in the output as described. The command returns results and stops, and an empty or partial result set looks identical to “nothing is wrong.”

This is consequence analysis from the stated defaults, not a documented bug, and the distinction matters. The changelog says the command returns “up to 100 spans from the last hour.” What happens to span 101, or to a span from minute 61, is inference: the changelog describes --since and --until only as a way to narrow the time range, and it never says whether the default hour can be widened or how far back the underlying trace store goes. The changelog is likewise silent on retention windows, plan gating, and sampling behavior. Before you trust an empty sweep as an all-clear, verify your plan’s actual retention in Vercel’s documentation. Groundy has hit this exact wall before: when Runtime Logs gained CDN cache fields, the retention window turned out to be the real ceiling on the feature’s usefulness, not the telemetry itself.

The practical rule: treat the default sweep as a probe for the last hour, never as evidence about anything older.

Why the short window exists

It is tempting to read the one-hour default as stinginess. Independent evidence suggests it is closer to cost alignment. A study of log retention economics in small-scale cloud deployments found that cutting retention from 90 days to 14 days reduced log storage costs by up to 78 percent while preserving more than 97 percent of operationally useful logs. The same study notes that retention policies are “frequently configured far beyond operational requirements,” often defaulting to 90 days in small and early-stage deployments.

The authors tested 7, 14, 30, and 90-day windows on synthetic log datasets designed to mimic real volume and access patterns, so the numbers are directional rather than a measurement of any production tracing product. But the study’s own conclusion transfers: longer retention windows provide diminishing operational returns while disproportionately increasing storage cost and query overhead. A CLI default that optimizes hard for the fresh-incident case is consistent with a vendor that prices that tradeoff the same way.

That reframing changes the debugging implication. The bottleneck for older incidents is not this command; it is your retention and sampling configuration. If your team investigates week-old incidents regularly, the fix is a longer retention tier or an external trace store, not a better terminal flag.

Terminal or dashboard: the decision table

Decision axisTerminal (vercel traces search)Dashboard
Time scope1 hour / 100 spans by default; --since/--until adjust the window, though the changelog describes them only as narrowing, and any upper bound is unverifiedRetention-dependent; not specified in the changelog
FiltersDedicated filters for environment, service, span name, status, deployment/request/trace/span IDPoint-and-click filtering; breadth not measured here
Ad-hoc queries--query with a supported, undefined subset of KQL (duration, span/resource attributes confirmed)Vendor UI; expressiveness not compared against --query
Output contractOne JSON object per span via --json; pipes into jq, scripts, agentsVisual charts and waterfalls; export capability not assessed here
Best-fit taskFresh incidents, precise queries, automation, agent consumptionOlder incidents, aggregate trends, visual waterfall analysis, sharing with non-terminal stakeholders

The honest summary of that table: the terminal wins on precision and automation for recent incidents, and the dashboard wins on everything requiring history, aggregation, or an audience. This is a division of labor, not a replacement. Vercel’s own observability trajectory points the same way; when redirect and rewrite telemetry landed in the dashboard, the value was externalizing a routing mental model for whoever is on call. The question is only which surface you query first.

Wiring observability into an agent’s toolset

The changelog’s phrase “giving scripts and agents structured data to analyze” is an explicit invitation, and it forces a permission decision most teams have not made deliberately.

If a coding agent can run vercel traces search, it holds whatever scope the CLI token behind it carries. The safe pattern is a dedicated read-only observability token: the agent can sweep error spans, run duration queries, and pull traces, but cannot deploy, mutate environment variables, or touch production configuration. Read-only telemetry access lowers the cost of agent-assisted debugging substantially, because the agent can do the repetitive sweep-and-group work while the human reads the shortlist. Vercel’s CLI has broken CI scripts before with a quiet breaking change under a UX-framed changelog entry, which is another argument for constraining what automation can invoke rather than assuming benign defaults.

Here is the caveat the changelog does not supply: it says nothing about scopes, tokens, or permissions. The read-only recommendation is design judgment, not documented capability. Whether Vercel’s token model actually supports a scope narrow enough for this pattern is a question for the docs, and it should be the first thing you verify before handing an agent the keys.

There is also a quieter benefit to --json as the agent interface. Research on CLI agents documents a selective-observation bottleneck: the agent starts with a partial view of a large workspace and must find the task-relevant evidence in it. One paper proposes σ-Reveal, a mechanism that selects token-budgeted workspace context before the agent’s first action. The setting differs from querying a trace store, but the lesson transfers: context selection is a first-class problem for terminal agents, and one-object-per-span JSON applies the same principle at the source.

Does the KQL subset cover your queries?

This is the loose thread that decides whether the command survives contact with production. The changelog says --query accepts Vercel’s supported subset of KQL and names duration plus span and resource attributes as supported fields. It never enumerates the subset.

That omission is not academic. Real incident queries get complicated fast: duration thresholds combined with status, attribute matches on specific HTTP routes, negations to exclude health checks. Which of those parse, which silently degrade, and which error out depends entirely on a field list the changelog does not publish. The correct posture is to test your three or four most common production queries against --query before building any runbook or agent prompt around them, and to treat the dedicated flags, which are fully specified, as the dependable layer.

The counter-evidence: agents may prefer the GUI anyway

Everything above assumes the terminal is the right surface for agents. A matched benchmark complicates that assumption. In a 440-task desktop benchmark spanning 18 applications, GUI and skill-mediated CLI agents received identical user goals, initial states, and executable final-state verifiers. The strongest GUI agent (GPT-5.4) achieved a 59.1% full pass rate; the strongest CLI agent (Codex GPT-5.5) reached 48.2% under the original skill layer.

Those are desktop tasks, not Vercel tracing, so the transfer is by analogy. But the direction is consistent with the context-selection problem above: giving an agent a terminal does not automatically make it better at finding what matters. For the specific use case here, the resolution is that --json is not really a CLI interface in the benchmark’s sense. It is a structured API that happens to be invoked from a shell, which sidesteps most of what makes raw terminal output hard for agents. The benchmark result is best read as a warning against handing an agent unstructured terminal scrollback, not against structured observability access.

For humans and supervised agents running precise queries, the terminal-first case stands on its own merits: one command replaces a dashboard session, and the output is already an artifact.

The verdict, and what to verify first

Use vercel traces search --json as the first probe for any Vercel incident less than an hour old: error-status sweep, span.duration slowness query, then a full-trace pull by trace ID once you have confirmed the JSON objects carry the fields the handoff needs. Keep the dashboard for incidents older than the window, aggregate trends, and anything you need to show another human. If you wire the command into an agent’s toolset, do it behind read-only observability credentials, and treat that scoping as a requirement to verify, not a given.

The limits on all of this are real. Every feature fact in this article traces to Vercel’s own changelog; no independent user reports existed at the time of writing. Retention windows, plan gating, sampling behavior, the --json field list, and the full KQL field list are unverified, and the “older incidents silently disappear” framing is inference from a stated default, not a documented failure. The retention-economics numbers come from synthetic datasets measuring storage cost, not debuggability, and the GUI-versus-CLI benchmark measures desktop tasks, not tracing. Check Vercel’s documentation for retention and the KQL subset before you build a runbook on this command, and treat an empty sweep as a question about scope, not an answer about your system.

Frequently Asked Questions

By default, vercel traces search returns up to 100 spans from the last hour, newest first, so you can search without first finding a request or trace ID.

How do I query for slow spans using the CLI?

vercel traces search —query ‘span.duration > 2000’

Use vercel traces search --json as the first probe for any Vercel incident less than an hour old: error-status sweep, span.duration slowness query, then a full-trace pull by trace ID once you have confirmed the JSON objects carry the fields the handoff needs.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Search trace spans from the Vercel CLIvercel.comAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy