groundy
Models & Research

GLM-5.2 vs Kimi K2.7 Code: Two Open-Weight Bets on Agentic Coding

GLM-5.2 and Kimi K2.7 Code shipped a day apart as open-weight coding models. One leads the open-weights index; the other wins on price and a claimed token-efficiency gain.

Published Updated 16 references
On this page9 sections

Moonshot shipped Kimi K2.7 Code on June 12, 2026. Zhipu opened GLM-5.2 on its Coding Plan the next day. Two Chinese labs, two open-weight models pointed at the same job (autonomous coding agents), released inside the same 24 hours, in the same week a US export-control directive led Anthropic to disable access to Claude Fable 5 and Mythos 5 for every customer worldwide. The timing is not subtle, and launch-week coverage of both releases read them directly against the suspension, as Kilo Code’s head-to-head did when it framed downloadable weights as the hedge against exactly this kind of withdrawal.

The two models look like twins on paper and turn out to be opposites. GLM-5.2 leads with a benchmark sheet, a million-token context window, and the top open-weights score on the one independent index that ranks it. Kimi K2.7 Code leads with a lower price and an efficiency claim, ships a benchmark table it mostly loses, and reads as a coding-and-efficiency release rather than a capability upgrade. These are opposite bets on what an open-weight coding model should compete on. Choosing between them has almost nothing to do with the few benchmark points separating their generations, and almost everything to do with whether you are buying capability or buying a cheaper meter.

Two labs, two bets, one week

GLM-5.2 comes from Zhipu AI (operating internationally as Z.ai), which listed in Hong Kong in January 2026 before a run that carried its market cap to a record above $112 billion in late May. Zhipu calls GLM-5.2 the strongest open-source model on standard coding benchmarks, per the GLM-5 repository. The pitch is capability plus reach: a 744B-parameter mixture-of-experts model with 40B active per token, a 1M-token window, openly released weights, and an Anthropic-compatible endpoint that Planet Tools’ comparison describes as a drop-in for Claude Code, Cline, Kilo Code, OpenClaw, Goose, and Roo, documented harness by harness in our migration checklist. That strongest-open-model label has since moved inside Zhipu’s own lineup: the GLM-5 repository now fronts GLM-5.3 and GLM-5.3-Flash, GLM-5.2’s HuggingFace card carries a “newer version available” pointer to GLM-5.3, and the Coding Plan now sells GLM-5.3, GLM-5.3-Flash, GLM-5.2, and GLM-5-Turbo. Zhipu says GLM-5.3 uses the same base model as GLM-5.2 with every gain coming from post-training, claims a 50% improvement over GLM-5.2 on its in-house Z.ai Code Bench, and reports open-source SOTA on Terminal Bench 3.0 and Agents’ Last Exam, per the GLM-5 repository. Those are vendor claims, not independent measurements, and this comparison sticks to what can be checked independently. GLM-5.2 remains the Zhipu model with an independent record, though buyers should price the newer pair before committing.

Kimi K2.7 Code comes from Moonshot AI, which has shipped five major K2 releases in under a year, from K2 in July 2025 to K2.7 Code in June 2026, per ModemGuides. It is a coding-focused build on K2.6 rather than a new base model, carrying forward the trillion-parameter MoE design with 32B active per token across 384 experts, per Moonshot’s model card. The pitch is economics: openly released weights under a Modified MIT license, an API priced below GLM-5.2 on every axis, and a claim that the model spends about 30% fewer reasoning tokens than its predecessor. What it does not lead with is a capability win. Across the six benchmarks Moonshot ran itself, K2.7 Code trails GPT-5.5 and Claude Opus 4.8 in 11 of 12 head-to-head cells, the subject of our breakdown of the K2.7 Code launch.

The strategic backdrop is the same for both. Chinese-origin models reached 61% of OpenRouter token volume in February 2026, with programming the largest usage category and agentic workflows generating over half of all output tokens on the platform. Usage tracks price more than quality. Open weights route around the export restrictions that bind closed models. When Washington can push Fable 5 off the market worldwide but cannot switch off a file on HuggingFace, the release calendar starts to look like policy. That is the stakes layer under what is otherwise a straightforward question: which of these two should you actually run.

Here is the shape of the decision before the detail.

GLM-5.2 (Z.ai / Zhipu)Kimi K2.7 Code (Moonshot)
ReleasedJune 13, 2026June 12, 2026
Total / active params744B / 40B (GitHub)1T / 32B (model card)
Context window1,000,000 (model card)256K (model card)
Vision inputNo vision encoder on the model card (card)Yes (400M MoonViT) (model card)
LicenseMIT (card)Modified MIT (card)
Intelligence Index (independent)51, leading open weights (Artificial Analysis)no independent Index writeup
API price, in / out per 1M$1.40 / $4.40 (AA)$0.95 / $4.00 (ModemGuides)

The architecture: 744B sparse against a trillion-parameter MoE

Both models are sparse mixtures of experts, and in both cases the headline parameter count is the least useful number on the page. What sets inference cost is the active-parameter slice, not the total.

GLM-5.2 carries 744B total parameters under the repository’s 744B-A40B designation and activates 40B per token, matching GLM-5.1’s size exactly, per the GLM-5 repository and Artificial Analysis. Our self-hosting breakdown works its memory math from a 753B count; either figure lands in the same hardware class, and the active slice, not the total, is what sets compute per token. Kimi K2.7 Code is the larger model on paper at a trillion parameters, but activates only 32B per token across 384 experts (8 routed plus 1 shared) over 61 layers, per its model card. So the trillion-parameter model is the cheaper one to run per token: 32B active against 40B. Raw size is a memory-footprint problem, not a compute-per-token problem, and the two models invert the ranking depending on which you care about.

The decoding stacks differ in ways that matter for long-running agents. GLM-5.2 uses what Zhipu calls IndexShare sparse attention, reusing a single token-selection indexer across every four sparse-attention layers to cut per-token FLOPs by 2.9x at 1M context, plus an improved multi-token-prediction layer that raises speculative-decoding acceptance length by up to 20%, per the GLM-5 repository. IndexShare is what makes the million-token window usable rather than nominal, and our benchmark explainer walks through the mechanism. Kimi K2.7 Code carries the K2 line’s Multi-head Latent Attention across 61 layers, ships native INT4 weights, and adds a 400M-parameter MoonViT vision encoder, per Moonshot’s card.

That vision encoder is a real asymmetry. Kimi K2.7 Code can read a screenshot, a UI mockup, or a rendered error directly, which removes a manual transcription step for tasks like reproducing a design or debugging a layout. GLM-5.2’s own pitch is text: its card stakes the claim on coding and reasoning benchmarks and a 1M window, and the multimodal model in Zhipu’s open line is GLM-5.3-Flash, a 320B-total, 18B-active build on a 30T-token multimodal pre-training corpus with a hybrid sparse-and-linear attention design, per the GLM-5 repository. That is a different model at a different tier. Moonshot’s benchmarks do not isolate vision-grounded coding performance, so the capability is real but unmeasured. If your workflow never touches an image, it is also irrelevant.

Context and speed: a million tokens, and who finishes first

GLM-5.2 exposes a 1,000,000-token input window, a 5x context jump over GLM-5.1’s 200K, and its model card’s evaluation footnotes include runs capped at 131,072 output tokens. Kimi K2.7 Code holds the K2 line’s 256K window, paired with prompt caching that drops repeated-context input to $0.19 per million tokens.

The two approaches answer different questions. A 1M window lets you load an entire repository and its dependency surface in one pass and skip the retrieve-then-reason loop most coding agents run. Kimi’s bet is that you rarely need a million tokens live, and that caching the 256K you do reuse is cheaper than paying attention compute over a window four times larger.

On wall-clock speed the early hands-on evidence inverts what the spec sheets suggest. In AkitaOnRails’s timed Rails 8 build, Kimi K2.7 Code finished a verified end-to-end app in 22 minutes for about $0.30, while GLM-5.2 was the slowest Tier A run at 43 minutes on Z.ai’s coding endpoint, which was throttled to 12 to 55 tokens per second during the test. One run on one throttled endpoint is not a speed benchmark, and Moonshot has announced a “6x High-Speed Mode” as coming soon, which is vendor-stated and unconfirmed. But on checkable evidence, GLM-5.2’s column holds the context advantage, and the finish-time evidence sits with Kimi, with endpoint throttling as the confound.

The benchmark problem, and the one independent number in it

Here is the honest difficulty in any GLM-5.2-versus-Kimi comparison: the two vendors published benchmark tables you cannot merge. GLM-5.2’s card reports public benchmarks (SWE-bench Pro, Terminal-Bench 2.1, GPQA, AIME, HLE, DeepSWE). Kimi K2.7 Code’s card leans on Moonshot’s own suites (Kimi Code Bench v2, Kimi Claw 24/7) plus third-party benchmarks Moonshot ran itself (Program Bench, MLS Bench Lite, MCP Atlas, MCP Mark Verified). Two benchmark names appear on both cards. On MCP-Atlas, GLM-5.2 posts 76.8 on the 500-task public set and K2.7 Code posts 76.0, a near-tie under divergent configurations: Zhipu with a 10-minute task timeout and a Gemini-3.0-Pro judge, Moonshot under the official 100-tool-call budget averaged over three runs. Program Bench, a 200-task suite that asks an agent to recreate a program’s behavior from a compiled binary and its documentation, is the more instructive overlap. GLM-5.2’s card lists 63.7, run in Claude Code 2.1.156 at max reasoning effort over a 400K context; Moonshot’s card lists K2.7 Code at 53.6, run in Kimi Code CLI over a 262K context. That is a ten-point gap in Zhipu’s favor, each side scoring in its own harness. The harness effect is measurable in the reference models: across the two tables GPT-5.5 lands within two points (70.8 vs 69.1) while Opus 4.8 moves eight (71.9 vs 63.8). One shared benchmark lands as a tie, the other as a gap, and neither is a controlled comparison.

The one independent number comes from outside both vendors. The Artificial Analysis Intelligence Index is a third-party composite, and its GLM-5.2 writeup scores the model at 51, the leading open-weights result.

ModelIntelligence IndexNote
GLM-5.251leading open weights (AA)
MiniMax-M344open
DeepSeek V4 Pro44open
Kimi K2.643open, Kimi’s previous generation

The same writeup puts GLM-5.2 ahead of the open field on GDPval-AA v2, its real-world agentic metric, at 1524, effectively level with GPT-5.5 at 1514. K2.7 Code has no score in that writeup; its independent evidence base is practitioner runs, and those point the same direction as Moonshot’s table. Researcher Elliot Arledge ran K2.7 Code against K2.6 and Claude Fable 5 on KernelBench-Hard, a public GPU-kernel benchmark, and published full run logs: K2.7 authored real Triton kernels on five of six problems where K2.6 had leaned on library wrappers, two of those kernels failed on the model’s own bugs, and the MoE kernel result regressed from K2.6’s 0.222 to 0.157. His summary: “more honest but not more capable.” Moonshot has also not submitted K2.7 Code to DeepSWE, the independent coding benchmark practitioners singled out at launch for producing a 70-point spread across models, where K2.6 scored 24%, per the same VentureBeat report. Zhipu’s card, by contrast, carries a vendor-run DeepSWE score of 46.2. Read together, the evidence says K2.7 Code is a coding-and-efficiency release, not a general upgrade. A buyer who treats it as a smarter Kimi is misreading the release.

The vendor coding numbers, taken on their own terms, broadly agree. GLM-5.2 self-reports 62.1% on SWE-bench Pro, up from GLM-5.1’s 58.4% and ahead of GPT-5.5’s 58.6% on the same benchmark, the same GPT-5.5 that beat Kimi K2.7 Code across nearly every cell of Moonshot’s table. Kimi K2.7 Code published no SWE-bench score at launch and its card carries none, so treat any SWE-bench number attached to K2.7 Code as unverified. Kimi’s one published win is MCP Mark Verified at 81.1, ahead of Opus 4.8’s 76.4 but behind GPT-5.5’s 92.9, per Moonshot’s card. MCP Mark scores Model Context Protocol tool-server integration, which is closer to real agent work than single-shot generation, so it is a genuine win in the one place it lands.

BenchmarkGLM-5.2Kimi K2.7 CodeReference: Opus 4.8
SWE-bench Pro62.1 (card)not published69.2 (Zhipu’s card)
Program Bench63.7 (card)53.6 (card)71.9 (Zhipu’s card) / 63.8 (Moonshot’s card)
DeepSWE46.2, vendor-run (card)not submitted (VB)58 (Zhipu’s card)
Terminal-Bench 2.181.0 Terminus-2 / 82.7 best harness (card)not published85.0 Terminus-2 / 78.9 best harness (Zhipu’s card)
MCP Mark Verifiednot published81.1 (card)76.4 (Moonshot’s card)
AIME 202699.2 (card)not published95.7 (Zhipu’s card)
HLE (with tools)54.7 (card)not published57.9 (Zhipu’s card)

Two caveats travel out of that table. The Opus column is worth reading closely, because every figure in it comes from a competitor’s harness, not from Anthropic: Zhipu’s own card lists Opus 4.8 at 85.0 under its Terminus-2 harness and 78.9 in its best-reported-harness row, a six-point spread on one model inside a single vendor’s table, before any cross-vendor disagreement enters. Zhipu’s card illustrates the harness sensitivity on its own model too: GLM-5.2 scores 81.0 under Terminus-2 and 82.7 under Claude Code, with both configurations published. And GLM-5.2’s 99.2% AIME is a near-ceiling math score with the highest contamination exposure in its suite; competition math saturates training crawls, so it measures recall as much as reasoning. Its 40.5% raw HLE (54.7 with tools) is the more honest signal, because that benchmark is built to resist saturation. That row also carries a convention from the card worth quoting exactly: GLM-5.2’s 54.7 is the default text-only subset, while Opus 4.8’s 57.9 is asterisked as a full-set result, so the two cells mix annotation sets.

Price: cheaper tokens against a stronger, pricier model

This is where Kimi’s strategy pays off, and it is the cleanest win either model has.

Both models are metered per token, and Kimi K2.7 Code is cheaper on every axis: $0.95 per million input tokens, $0.19 cached, and $4.00 output against GLM-5.2’s $1.40 input, $0.26 cached, and $4.40 output. Then the efficiency claim compounds it: if K2.7 Code really spends roughly 30% fewer thinking tokens per coding task than K2.6, as Moonshot’s card claims, its effective cost-per-task advantage is wider than the per-token spread alone, though practitioners publicly questioned whether the gain holds outside Moonshot’s own suites. That claim cuts directly at GLM-5.2’s known weakness. Artificial Analysis flags GLM-5.2 as verbose: 43,000 output tokens per Intelligence Index task, up from 26,000 for GLM-5.1 and above Kimi K2.6’s 35,000, the kind of generation bloat that turns a per-token edge into a per-task deficit. AA puts GLM-5.2’s cost per Index task at roughly $0.46 against $0.31 for K2.6.

The subscription tiers are a near-tie at the entry point, and the two have been lined up directly by third-party reviewers since launch. Zhipu’s GLM Coding Plan starts at $18/month for the Lite tier, with higher Pro and Max tiers whose exact prices Zhipu has not published, per The Planet Tools’ breakdown, and the plan page now sells GLM-5.3, GLM-5.3-Flash, GLM-5.2, and GLM-5-Turbo. Moonshot’s Kimi Code membership plans are listed from $19/month as of June 2026. One caution on the subscription path: Zhipu’s own plan FAQ fields questions about usage limits, concurrent connections, and balance deductions appearing after purchase, so an always-on agent is worth running against the plan before you trust the flat fee. And the self-host path zeroes the per-token line for either model, moving the bill to hardware.

The summary is uncomfortable for GLM-5.2 on this axis and only this axis: it is the more expensive model to run, per token and per task, and it is competing on an independent capability lead and a context window, not on price. Kimi K2.7 Code is the budget instrument. For a coding agent that loops hundreds of times per task, that is the number that shows up on the invoice.

Open weights, two licenses, two hardware bills

Both models ship openly, and the licenses are close but not identical. GLM-5.2’s HuggingFace card declares MIT, with no regional limits, and Artificial Analysis lists the same. Kimi K2.7 Code is Modified MIT: standard MIT behavior until a product crosses 100 million monthly active users or US$20 million per month in revenue, at which point it must display “Kimi K2.7 Code” in its interface, per the license text pulled by ModemGuides. For all but the largest products that clause never fires.

Open weights do not make either model cheap to hold. GLM-5.2 is the smaller model but still 744B parameters, which works out to roughly 750GB in FP8 and about 1.5TB in BF16 before KV-cache overhead, per our self-hosting cost breakdown. Kimi K2.7 Code is the larger model at a trillion parameters: about 610GB in its shipped native INT4, roughly 2TB in FP16, and a 340GB floor for the smallest community GGUF (measured on its identical-architecture sibling K2.6), which wants 350GB or more of combined RAM and VRAM for usable speeds, per ModemGuides’ self-hosting breakdown. The MoE memory penalty applies to both: every expert stays resident even though only a fraction fires per token. Neither is a single-card proposition. Kimi wants an 8-way H200-class node (around 640GB of VRAM) for production serving of the native INT4 build, and GLM-5.2’s full-precision footprint needs more than one node. The practical difference is which quantization you are handed: Kimi ships an official INT4 build, while the GLM repository hands you BF16 and FP8 and leaves deeper compression to the community.

One popularity metric sits on the model pages themselves: GLM-5.2’s HuggingFace card showed about 765,000 downloads in the last month as of late September 2026, alongside the banner pointing to GLM-5.3 as the newer version. Downloads measure curiosity rather than quality, but the self-host crowd is still pulling GLM-5.2 three months after launch, with its successor already out.

Agentic coding: the near-tie that splits by use case

Strip away the leaderboards and the question becomes narrower: which one writes and ships better code inside an agent loop. Here the independent evidence is thin, and it lands closer to a tie than the Intelligence Index does.

Integration is a near-wash. GLM-5.2 speaks the Anthropic Messages API, and official integrations shipped for eight harnesses at launch, Claude Code, Cline, OpenCode, Roo Code, Goose, Crush, OpenClaw, and Kilo Code, per our migration checklist; a base-URL and model-name swap usually suffices. Kimi K2.7 Code offers both OpenAI- and Anthropic-compatible endpoints, broader surface, but the OpenAI path carries two real gotchas documented in the community integration guide: multi-turn tool use requires replaying the prior reasoning_content block or the tool loop breaks, and tool_choice accepts only auto or none. Both need adapter awareness despite the compatible-looking API. Reasoning control diverges sharply: GLM-5.2 accepts only High and Max effort presets, per the GLM-5 repository, while Kimi K2.7 Code forces thinking on through preserve_thinking with no instant mode and temperature fixed at 1.0, so a one-line rename pays full reasoning overhead it cannot opt out of. For hard tasks that is a feature; for trivial high-volume edits it is the efficiency pitch arguing with itself.

The week-one hands-on comparisons split predictably:

  • Kilo Code scored the two across planning and building and gave the nod to GLM-5.2 in both phases (planning 9.0 to Kimi’s 8.1, building 15/15 to 14/15), noting GLM-5.2’s plan landed within 0.1 of Claude Fable 5’s 9.1 at roughly a tenth of the list price. The builds were near-identical in behavior: given GLM’s plan, both services returned the same answers for all 200 test user IDs.
  • AkitaOnRails ran a timed Rails 8 build and scored it almost even: GLM 5.2 at 87/100, with the cleanest dependency injection in the field but the slowest Tier A run at 43 minutes, and Kimi K2.7 Code at 86/100, with a verified end-to-end result in 22 minutes for about $0.30, marred by a regression where it drops the system prompt via with_instructions. The same reviewer had clocked GLM 5.1 at 46/100 on the same benchmark months earlier, a generational leap that tracks the Intelligence Index story.
  • Across the cluster of reviews the pattern is consistent: GLM-5.2 is the more careful planner and the stronger one-shot builder; Kimi K2.7 Code is the faster, cheaper finisher on timed runs. Both still cede the hardest tasks to Opus 4.8 and GPT-5.5, which remains AkitaOnRails’s standing advice for serious programming work.

For teams already on Claude Code or Cline, our GLM-5.2 migration checklist covers the harness specifics.

The asterisks both launches share

Both releases rest on vendor-reported tables with no independent replication of the coding numbers. GLM-5.2’s card at least runs public benchmarks and publishes its harness configurations; Kimi’s card leans on in-house suites for its headline gains, and the efficiency claim is self-measured. The Artificial Analysis Index is the one independent anchor, and it tempers the picture rather than settling it.

Each model also carries a specific risk a benchmark sheet hides. GLM-5.2’s predecessor failed publicly in at least one independent test (GLM-5.1 scored 46/100 in AkitaOnRails’s benchmark after inventing API calls that crashed on turn two), and while 5.2 fixed that exact bug, its throttled endpoint in the same test is a reminder that plan quotas and endpoint throttling shape real throughput. The release-day analysis in Zhipu ships GLM-5.2 with zero benchmarks flags the gap between availability and trust, and the self-hosting cost reality is steeper than the MIT label implies. Kimi K2.7 Code’s risk is quieter: forced thinking means its cost advantage is real on hard tasks and can evaporate on trivial ones, the documented with_instructions regression is the kind of integration footgun that only surfaces in production, and Arledge’s kernel failures are a reminder that Moonshot’s headline numbers are Moonshot’s own.

Neither launch hands a procurement team a number it can defer to. The comparison work has moved off the leaderboard and onto your repository.

How to choose

The decision resolves to a few axes, and on most of them the two models answer different questions rather than beat each other.

  • General capability and long context: GLM-5.2. It holds the leading open-weights score on the one independent index and the only million-token window of the two. If you want the strongest open coding model with an independent record and can absorb its higher per-token cost, this is it. Caveat: Zhipu has since shipped GLM-5.3 and claims a 50% coding gain on its in-house bench, per the GLM-5 repository, on vendor-run benchmarks only.
  • Cost per task at volume: Kimi K2.7 Code. Cheaper on every per-token axis and more token-efficient per coding task by its own measurement, it is the budget instrument for agents that loop hundreds of times, provided you run it against real pull requests first.
  • Timed builds against careful planning: Kimi finished the one timed hands-on build faster and cheaper; GLM-5.2 planned better and caught more edge cases in Kilo Code’s two-phase test. Long-horizon autonomy claims for either model rest on vendor suites and single-run reviews.
  • Vision in the coding loop: Kimi K2.7 Code, via MoonViT. GLM-5.2’s card lists no vision encoder; Zhipu’s multimodal option is GLM-5.3-Flash, a different model at a different tier.
  • Reasoning-budget control: GLM-5.2, with explicit effort presets. Kimi forces thinking on.

The broader read is that this is no longer a story about whether Chinese open-weight models are competitive. GLM-5.2 holding the top open-weights slot and trading blows with closed frontier models on developer-relevant tasks, the same week the US restricted those closed models, is the story. For the full field these two sit in, DeepSeek, Qwen, Kimi, ByteDance’s Doubao, Baidu’s ERNIE, and Zhipu’s own GLM line, our map of the Chinese model ecosystem puts them in context, and the task-level coding benchmarks collect the numbers against GPT and Claude.

The answer neither vendor prints on a slide is the same one that closed every comparison this year: run both on your own codebase and measure cost per merged pull request, not benchmark cells. GLM-5.2 gives you the stronger independent record and the bigger bill. Kimi K2.7 Code gives you the cheaper meter and the faster timed finish. Both stopped handing you a leaderboard to hide behind, which is not a relief from evaluation. It is a transfer of the work to you.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. GLM-5 Model Family. zai-org GitHub repositorygithub.comAccessed
  2. GLM-5.2 Model Card. Hugging Face, zai-orghuggingface.coAccessed
  3. GLM-5.2 Guide. Z.ai Documentationdocs.z.aiAccessed
  4. Kimi K2.7 Code Model Card. Hugging Face, moonshotaihuggingface.coAccessed
  5. GLM Coding Plan Pricing. Z.aiz.aiAccessed
  6. GLM-5.2 vs Kimi K2.7 Code: Which Model Wins? Kilo Codeblog.kilo.aiAccessed
  7. Claude Opus 4.8. Anthropicanthropic.comAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy