GLM-5.2 vs DeepSeek vs Kimi: Open Coding Model Compared
Picking an open-weight coding model in 2026 means choosing between three very different trade-offs: GLM-5.2, DeepSeek V4 Pro, and Kimi K3 each lead on a different axis, and none of them wins everywhere. Here's how they actually compare, using each vendor's own published benchmarks and official API pricing.
Key Takeaways
- Kimi K3 tops the Artificial Analysis Intelligence Index among open-weight models (57), ahead of GLM-5.2 (51) and DeepSeek V4 Pro (44) — but its weights aren't public yet and it costs 2-17x more per token than the other two.
- GLM-5.2 still posts the highest SWE-bench Pro score of the three (62.1%), a real-world software-engineering benchmark, and it's the one available now on Siray at 20% off official pricing.
- DeepSeek V4 Pro remains the cheapest at $0.435 / $0.87 per million input/output tokens.

Three Models, Three Different Bets
GLM-5.2 (Zhipu AI / Z.ai, June 2026) is a 744B-parameter MoE model (40B active) tuned for coding, reasoning, and agentic tool-use, with a 1M-token context window and an MIT license.
DeepSeek V4 Pro (DeepSeek, April 2026) is a larger 1.6T-parameter MoE model (49B active), also 1M-token context and MIT-licensed, built with a stronger emphasis on raw competitive-programming performance.
Kimi K3 (Moonshot AI, July 2026) is the newest and largest — a 2.8T-parameter MoE model, the largest open-weight release to date, with a new Kimi Delta Attention + Attention Residuals architecture, native vision, and "always-on" max-reasoning. Its context window now matches the other two at 1M tokens (1,048,576), a big jump from predecessor Kimi K2.6's 256K cap. The catch: K3's full weights aren't public yet — it's API-only under a "Modified MIT" license, with open-weight release scheduled for late July 2026.
Artificial Analysis Intelligence Index: Ranking the Three
Artificial Analysis runs an independent Intelligence Index across reasoning, coding, and agentic-work evaluations — useful here because it scores all three on the same third-party yardstick rather than each company's self-reported numbers. On the latest revision (v4.1, each model's top reasoning-effort setting):
- Kimi K3: 57 — highest open-weight score on the index, #4 across all models including closed frontier systems.
- GLM-5.2: 51
- DeepSeek V4 Pro: 44
The lead isn't free: Artificial Analysis notes K3 runs slower and more token-hungry per task than the other two, on top of its weights not yet being downloadable.
SWE-Bench Pro: Where It's Still Published
SWE-bench Pro grades a model on fixing actual GitHub issues rather than isolated coding puzzles. Kimi K3 hasn't published an official score on it — Moonshot's release materials lead with SWE-bench Verified (76.8%) and Terminal-Bench 2.1 (88.3) instead, different benchmarks not directly comparable to the Pro scores below. Among the two that do publish SWE-bench Pro, GLM-5.2 leads at 62.1% (still the top open-weight score on this benchmark overall) against DeepSeek V4 Pro's 55.4%.
Pricing and Context Window Compared
AA Intelligence Index (v4.1) | SWE-bench Pro | Context Window | License | Official Pricing (input / output per 1M tokens) | Available on Siray | |
GLM-5.2 | 51 | 62.1% | 1M tokens | MIT | $1.40 / $4.40 (Siray: $1.12 / $3.52, 20% off) | ✅ Yes |
Kimi K3 | 57 | not published | 1M tokens | Modified MIT (weights pending) | $3.00 / $15.00 | Reference only |
DeepSeek V4 Pro | 44 | 55.4% | 1M tokens | MIT | $0.435 / $0.87 | Reference only |

K3's jump to 1M tokens erases context window as a differentiator — all three now match, unlike the K2.6 generation that capped at 256K. Price doesn't match, though: K3's output-token rate runs roughly 17x DeepSeek V4 Pro's and 4x GLM-5.2's, a real cost against which its index lead has to be weighed, on top of its weights not being self-hostable yet.
Why GLM-5.2 on Siray
GLM-5.2's Intelligence Index score (51) sits close behind Kimi K3's 57 at a fraction of the price, and it still posts the highest SWE-bench Pro score of the three. It's also the only one of these three you can call today through Siray — at 20% off Zhipu/Z.ai's official pricing, through the same API key used for Siray's other 300+ models, no separate Zhipu/Z.ai signup required. (For a closer look at GLM-5.2's own specs and pricing, see GLM-5.2 on Siray: Open-Source Coding Model at 20% Off.)DeepSeek V4 Pro and Kimi K3 appear here only as public benchmark and pricing references — neither is currently available through Siray. No single open-weight coding model wins on every axis in 2026: pick by index score, benchmark-specific performance, price, or whether you need self-hostable weights today.
FAQ
Can I call DeepSeek V4 Pro or Kimi K3 through Siray?
Not currently. This comparison uses their official public pricing and benchmark scores as reference points; only GLM-5.2 is available through Siray today.
Which model tops the Artificial Analysis Intelligence Index?
Kimi K3, at 57 (v4.1, top reasoning-effort setting), ahead of GLM-5.2 (51) and DeepSeek V4 Pro (44). K3's full weights aren't publicly downloadable yet, and its official API pricing runs several times higher than the other two.
Which model scores highest on SWE-bench Pro?
GLM-5.2, at 62.1%, ahead of DeepSeek V4 Pro (55.4%). Kimi K3 hasn't published an official SWE-bench Pro score.
Is GLM-5.2 cheaper on Siray than going direct to Zhipu/Z.ai?
Yes — Siray prices GLM-5.2 at 20% off Zhipu/Z.ai's official API rate ($1.12 / $3.52 per million input/output tokens vs. $1.40 / $4.40 official), billed through the same key as Siray's other models.