groundy
Agents & Frameworks

Trying StepFun's Step 5 in Claude Code: What to Test Before Switching

Vercel's AI Gateway enables a low-cost trial of StepFun's Step 5 Preview in Claude Code, but vendor-reported benchmarks show it trails incumbents, requiring rigorous local A/B

Published 6 references
A skeptical green resin dinosaur raises one chunky foot above an ivory footprint mold, with an open doorway behind it. A yellow ankle band and hard shadows punctuate the warm ivory scene.
On this page12 sections

Vercel’s AI Gateway now exposes StepFun’s Step 5 Preview as stepfun/step-5-preview, and a single setup command, vercel ai-gateway setup after npm i -g vercel@latest, makes it selectable inside Claude Code, Codex, or Cursor, according to Vercel’s changelog. That makes a trial nearly free to start: the gateway bills at provider pricing with no markup and no platform fee, including on Bring Your Own Key (BYOK) requests, and its routing rules, retries, and failover give you a documented path back to your incumbent model. What it does not make free is the evaluation itself. StepFun’s own benchmark table places Step 5 Preview behind the models most Claude Code users already run, and every capability figure on StepFun’s and Vercel’s pages is vendor-reported. So: trial it, in a bounded and budget-capped way, but treat the gateway setup as the easy part and your own repository testing as the real work.

This article maps each thing worth verifying to the gateway feature that makes the verification cheap, then lays out a one-week protocol. None of it requires trusting StepFun’s numbers, because none of it depends on them.

What actually changed

The announcement is a distribution change, not a capability finding. Vercel’s changelog positions Step 5 Preview as “built for agentic coding, professional knowledge work, and financial analysis,” with text and image input and a 1M-token context window sized for large codebases, document stacks, or screenshots in a single request. Those are Vercel’s and StepFun’s words about the model, and the changelog itself establishes availability and setup only, nothing in it tests tool calling inside Claude Code.

The genuinely useful part of the announcement is operational. Per the same changelog, AI Gateway provides “a unified API for calling models, tracking usage and cost, and configuring retries, failover, and performance optimizations for higher-than-provider uptime,” plus custom reporting, Zero Data Retention support, budgets for API keys, and routing rules. That is the machinery a serious trial needs regardless of which model prompted it.

One disclosure worth keeping in view: Vercel benefits commercially from gateway adoption even while charging no inference platform fee. The no-markup pricing is favorable, but the distribution push is not neutral, so weight Vercel’s positioning accordingly.

Step 5 Preview on paper

According to StepFun’s model page, Step 5 Preview is a sparse Mixture-of-Experts model with 600B total parameters and 27B active per token, supporting a 1M-token context window and vision input. StepFun states it is available now through its products and API, with open weights scheduled for October 15, a week from today, which matters if you want the option of self-hosting later rather than committing to a hosted-only dependency.

The capability numbers deserve their own treatment, because they are the strongest counter-evidence in this whole story. On StepFun’s own DeepSWE v1.1 table, Step 5 Preview (High) scores 67.7, ahead of Kimi K3 (Max) at 67.5 and GLM-5.3 (Max) at 66.9 but behind GPT-6 Astra (Max) at 74.1 and Claude Opus 5 (Max) at 74.0. Read that again: the vendor’s own numbers place its flagship roughly six points behind the two models it most wants to displace on agentic coding. StepFun’s page also reports single-run anecdotes, a peak of 508 TFLOPS on an H100 MLA kernel after roughly 22 hours of iteration versus 493 for Claude Opus 5, a Pokémon Red run sustained past 3,000 turns and 6 million tokens per the page, and a climate study coordinating 950 web fetches in one agent action per StepFun. These are interesting as engineering demonstrations and weak as purchasing evidence; they are vendor-reported, unreplicated, and none of them ran inside your coding agent.

The only capability source cited here is StepFun itself.

Decision axisWhat the evidence saysVendor-reported?
Agentic coding capabilityDeepSWE v1.1: 67.7, trailing GPT-6 Astra (74.1) and Claude Opus 5 (74.0) on StepFun’s own tableYes, entirely
Context and input1M tokens, text and image, sparse MoE 600B/27B activeYes
Trial costProvider pricing, no markup, no platform fee, BYOK supported per Vercel; StepFun’s page does not list Step 5’s per-token pricesPartially
Safety netGateway retries, failover, routing rules documentedNo, this is platform mechanics
GovernanceZero Data Retention support, per-API-key budgets documentedNo, but verify your configuration

Checkpoint 1: tool-call reliability on your repositories, not vendor leaderboards

A 1M-token context window says nothing about whether the model calls tools correctly inside Claude Code. Tool-call reliability is a property of the model, the agent harness, and your repository’s conventions together, and nothing in StepFun’s or Vercel’s announcements tests tool calling inside Claude Code. The vendor leaderboard gap (67.7 versus 74.0 for Claude Opus 5, per StepFun’s own numbers) is the only signal available, and it points toward a capability deficit, not parity.

The useful testing method already exists. The SWE-CC benchmark paper scores coding-agent resolution through the SWE-bench suite after extracting patch contents from model final outputs, which means you can take a set of real issues from your own repositories, run Step 5 against them in Claude Code, and score resolved-versus-unresolved the same way the benchmark does. SWE-CC was demonstrated on other agents, not Step 5, so treat it as a method to copy, not a result to expect.

A hypothetical shape for this checkpoint: pick ten closed issues from a repository your team knows well, run each through Claude Code pointed at stepfun/step-5-preview, and compare pass rates against the same ten issues run with your incumbent model. Same harness, same prompts, different model, that is the comparison the vendor table cannot give you.

Checkpoint 2: cost, where output tokens dominate

“No platform fee” does not mean “cheap.” It means you pay StepFun’s provider pricing directly, and StepFun’s page does not list per-token prices, so you cannot precompute the trial’s cost. What you can precompute is the structure of the bill. The pricing literature surveyed in arXiv:2606.11690 shows commercial providers charge asymmetrically: GPT-5.5 at $5.00/M input versus $30.00/M output, Claude Sonnet 4.6 at $3.00/M versus $15.00/M, Gemini 3.1 Pro at $2.00/M versus $12.00/M, per the same paper. Output tokens run five to six times input prices across all three.

For an agentic coding workload this asymmetry is the whole game. A coding agent reads a lot (repository context, tool results, file contents) and writes a lot too (patches, reasoning, retries). A model with a 1M-token window invites you to fill that window, and if the agent is verbose or loops on a hard problem, output tokens accumulate at the expensive rate. Two practical consequences: meter the trial by token class, not by request count, using the gateway’s usage and cost tracking; and set a per-API-key budget before the first run, which the changelog documents as a built-in feature, so a looping agent exhausts a cap instead of a credit card.

Checkpoint 3: failover as the trial’s safety net

The reason a trial through the gateway is low-risk is that the exit is a routing rule, not a migration. Vercel documents retries, failover, and routing rules as first-class gateway features. Configure failover from stepfun/step-5-preview back to your incumbent model before you start, and two failure modes become boring: the new model is unavailable, and the new model produces garbage on a task class and you want to stop routing that class to it.

There is a subtlety worth testing explicitly: failover handles request-level failures (errors, timeouts), not quality-level failures (a confident but wrong patch). Retries will not save you from a model that fails politely. That is what Checkpoint 1 is for, and why the failover configuration is a complement to evaluation, not a substitute for it.

Checkpoint 4: governance, Zero Data Retention and budgets

If your code cannot leave a retention boundary, this checkpoint gates everything else. The changelog states the gateway includes Zero Data Retention support, but “support” is a capability of the platform, not a guarantee about your configuration or about StepFun’s handling on the provider side. Verify the ZDR setting is actually active for the key and route you use, and confirm whether routing through the gateway to StepFun’s API satisfies your data-handling requirements; neither Vercel’s changelog nor StepFun’s page settles that question. Per-key budgets, covered above, double as a governance control: they limit blast radius if a key leaks or a script runs away.

Checkpoint 5: harness drift and policy compliance

Repointing an agent at a new model is not a configuration-free act. A study of agentic coding tool adoption, Harness Engineering for Agentic AI Coding Tools, found configuration mechanisms are heavily tool-specific: 72.8% of Cursor repositories adopt Rules and 62.3% of Gemini repositories use Settings in the study’s sample, while no other mechanism exceeds 20% adoption across Claude, Copilot, Cursor, or Gemini, per the same study. The implication for a model switch is that the prompt-engineering and configuration investment in your current agent does not automatically transfer, and the same repository conventions may land differently on a model that never saw your rules file during training.

The second half of this checkpoint is policy compliance: does the agent’s output actually follow your repository’s contribution rules, not just pass tests? SWE-CC exists because correct patches routinely violate repository policy. Its authors show compliance can be evaluated after each execution in 120 to 215 ms, per the paper, on a local MacBook Pro (Apple M4 Pro, 24 GB unified memory), cheap enough to run on every trial task, on a laptop, without new infrastructure. During your Step 5 trial, check both axes: did it resolve the issue, and did the patch respect the repo’s conventions. A model that resolves more issues while violating contribution policy is a regression wearing a benchmark win.

A one-week trial protocol

  1. Day 1, setup and guardrails. Run vercel ai-gateway setup, select stepfun/step-5-preview in Claude Code. Configure a per-key budget, enable usage tracking, set a routing-rule failover to your incumbent model, and verify Zero Data Retention settings if your code requires them.
  2. Days 2–3, resolution baseline. Run your ten-issue benchmark set under both models, same harness and prompts. Score resolution SWE-bench-style. Record tool-call failures (malformed calls, wrong arguments, abandoned tasks) separately from wrong answers; they have different causes.
  3. Day 4, compliance pass. Run per-execution policy checks on every patch from both arms, using the SWE-CC method as the template.
  4. Day 5, cost accounting. Pull the gateway’s usage report. Split spend into input and output tokens for each arm, using your provider’s actual rates (StepFun’s page does not list Step 5’s rates; pull them from the gateway report itself). Project what a team-wide month would cost at each model’s observed token mix.
  5. Decision gate. Switch only if Step 5 matches or beats the incumbent on resolution and compliance at a cost you can defend, with ZDR confirmed. Otherwise keep the routing rule pointing at the incumbent and revisit after the open-weights release or the first independent benchmark.

What this evidence cannot tell you

The verdict is: run the bounded trial, defer the team-wide switch. The evidence supports a trial because the gateway makes it reversible and financially bounded by construction. It does not support a switch, because the only capability data on the pages cited here is StepFun’s own, and it shows Step 5 trailing on the vendor’s agentic-coding table.

Be honest about the gaps. Nothing here establishes Step 5’s tool-call reliability, latency, or uptime in Claude Code. StepFun’s page does not list per-token pricing, so cost projections wait on your own metering. The Pokémon Red and climate-study results are single-run vendor anecdotes. And a 2025 survey by MIT’s NANDA initiative, which the linked paper cites as Challapally et al. (2025), found the overwhelming majority of enterprise generative-AI pilots fail to produce measurable impact, locating the failure in a learning gap, deployed systems that neither retain feedback nor adapt to context. A trial that ends in a vibe (“it felt fast”) is that failure mode. A trial that ends in retained, scored runs on your own repositories is the one piece of evidence the vendors cannot give you, and the only one that should move your routing rules permanently.

Frequently Asked Questions

How do I set up Step 5 Preview in Claude Code?

Vercel’s AI Gateway now exposes StepFun’s Step 5 Preview as stepfun/step-5-preview, and a single setup command, vercel ai-gateway setup after npm i -g vercel@latest, makes it selectable inside Claude Code, Codex, or Cursor, according to Vercel’s changelog.

How can I ensure the trial does not exceed my budget?

set a per-API-key budget before the first run, which the changelog documents as a built-in feature, so a looping agent exhausts a cap instead of a credit card.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Vercel's changelogvercel.comAccessed
  2. StepFun's model pagestepfun.comAccessed
  3. SWE-CC benchmark paperarxiv.orgAccessed
  4. arXiv:2606.11690arxiv.orgAccessed
  5. Harness Engineering for Agentic AI Coding Toolsarxiv.orgAccessed
  6. linked paperarxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy