GLM-5.2 vs DeepSeek vs Kimi: Open Coding Model Compared

GLM-5.2 vs DeepSeek vs Kimi: Open Coding Model Compared
GLM-5.2 vs DeepSeek vs Kimi: Open Coding Model Compared

Picking an open-weight coding model in 2026 means choosing between three very different trade-offs: GLM-5.2, DeepSeek V4 Pro, and Kimi K3 each lead on a different axis, and none of them wins everywhere. Here's how they actually compare, using each vendor's own published benchmarks and official API pricing.

Key Takeaways

  • Kimi K3 tops the Artificial Analysis Intelligence Index among open-weight models (57), ahead of GLM-5.2 (51) and DeepSeek V4 Pro (44) — but its weights aren't public yet and it costs 2-17x more per token than the other two.
  • GLM-5.2 still posts the highest SWE-bench Pro score of the three (62.1%), a real-world software-engineering benchmark, and it's the one available now on Siray at 20% off official pricing.
  • DeepSeek V4 Pro remains the cheapest at $0.435 / $0.87 per million input/output tokens.
Open Models Parameters
Open Models Parameters

Three Models, Three Different Bets

GLM-5.2 (Zhipu AI / Z.ai, June 2026) is a 744B-parameter MoE model (40B active) tuned for coding, reasoning, and agentic tool-use, with a 1M-token context window and an MIT license.

DeepSeek V4 Pro (DeepSeek, April 2026) is a larger 1.6T-parameter MoE model (49B active), also 1M-token context and MIT-licensed, built with a stronger emphasis on raw competitive-programming performance.

Kimi K3 (Moonshot AI, July 2026) is the newest and largest — a 2.8T-parameter MoE model, the largest open-weight release to date, with a new Kimi Delta Attention + Attention Residuals architecture, native vision, and "always-on" max-reasoning. Its context window now matches the other two at 1M tokens (1,048,576), a big jump from predecessor Kimi K2.6's 256K cap. The catch: K3's full weights aren't public yet — it's API-only under a "Modified MIT" license, with open-weight release scheduled for late July 2026.

Artificial Analysis Intelligence Index: Ranking the Three

Artificial Analysis runs an independent Intelligence Index across reasoning, coding, and agentic-work evaluations — useful here because it scores all three on the same third-party yardstick rather than each company's self-reported numbers. On the latest revision (v4.1, each model's top reasoning-effort setting):

  • Kimi K3: 57 — highest open-weight score on the index, #4 across all models including closed frontier systems.
  • GLM-5.2: 51
  • DeepSeek V4 Pro: 44

The lead isn't free: Artificial Analysis notes K3 runs slower and more token-hungry per task than the other two, on top of its weights not yet being downloadable.

SWE-Bench Pro: Where It's Still Published

SWE-bench Pro grades a model on fixing actual GitHub issues rather than isolated coding puzzles. Kimi K3 hasn't published an official score on it — Moonshot's release materials lead with SWE-bench Verified (76.8%) and Terminal-Bench 2.1 (88.3) instead, different benchmarks not directly comparable to the Pro scores below. Among the two that do publish SWE-bench Pro, GLM-5.2 leads at 62.1% (still the top open-weight score on this benchmark overall) against DeepSeek V4 Pro's 55.4%.

Pricing and Context Window Compared


AA Intelligence Index (v4.1)
SWE-bench Pro
Context Window
License
Official Pricing (input / output per 1M tokens)
Available on Siray
GLM-5.2
51
62.1%
1M tokens
MIT
$1.40 / $4.40 (Siray: $1.12 / $3.52, 20% off)
✅ Yes
Kimi K3
57
not published
1M tokens
Modified MIT (weights pending)
$3.00 / $15.00
Reference only
DeepSeek V4 Pro
44
55.4%
1M tokens
MIT
$0.435 / $0.87
Reference only
Cost per task source: artificial analysis
Cost per task source: artificial analysis

K3's jump to 1M tokens erases context window as a differentiator — all three now match, unlike the K2.6 generation that capped at 256K. Price doesn't match, though: K3's output-token rate runs roughly 17x DeepSeek V4 Pro's and 4x GLM-5.2's, a real cost against which its index lead has to be weighed, on top of its weights not being self-hostable yet.

Why GLM-5.2 on Siray

GLM-5.2's Intelligence Index score (51) sits close behind Kimi K3's 57 at a fraction of the price, and it still posts the highest SWE-bench Pro score of the three. It's also the only one of these three you can call today through Siray — at 20% off Zhipu/Z.ai's official pricing, through the same API key used for Siray's other 300+ models, no separate Zhipu/Z.ai signup required. (For a closer look at GLM-5.2's own specs and pricing, see GLM-5.2 on Siray: Open-Source Coding Model at 20% Off.)DeepSeek V4 Pro and Kimi K3 appear here only as public benchmark and pricing references — neither is currently available through Siray. No single open-weight coding model wins on every axis in 2026: pick by index score, benchmark-specific performance, price, or whether you need self-hostable weights today.

FAQ

Can I call DeepSeek V4 Pro or Kimi K3 through Siray?

Not currently. This comparison uses their official public pricing and benchmark scores as reference points; only GLM-5.2 is available through Siray today.

Which model tops the Artificial Analysis Intelligence Index?

Kimi K3, at 57 (v4.1, top reasoning-effort setting), ahead of GLM-5.2 (51) and DeepSeek V4 Pro (44). K3's full weights aren't publicly downloadable yet, and its official API pricing runs several times higher than the other two.

Which model scores highest on SWE-bench Pro?

GLM-5.2, at 62.1%, ahead of DeepSeek V4 Pro (55.4%). Kimi K3 hasn't published an official SWE-bench Pro score.

Is GLM-5.2 cheaper on Siray than going direct to Zhipu/Z.ai?

Yes — Siray prices GLM-5.2 at 20% off Zhipu/Z.ai's official API rate ($1.12 / $3.52 per million input/output tokens vs. $1.40 / $4.40 official), billed through the same key as Siray's other models.

Try GLM-5.2 on Siray →