Kimi K3.7 Scores 66.7% on SWE-bench with Open Weights. Do the Math on That
Back to All Posts

Kimi K3.7 Scores 66.7% on SWE-bench with Open Weights. Do the Math on That

Moonshot released Kimi K3.7 on July 9 with a 512K context window, full open weights and a 66.7% score on SWE-bench Verified. It costs $0.60 per million input tokens on the hosted API.

Sit with that combination for a moment, because each element individually is unremarkable and together they are not.

The Comparison That Stings

Claude Opus 4.6, the model at the top of our leaderboard in April, scored 58.4% on SWE-bench Verified at $15 per million input, closed weights. Kimi K3.7 beats it by eight points at four percent of the price, with weights you can download.

Three months. That is the gap between a closed frontier model and an open one that beats it. The compression of that cycle is the single most important trend in the industry right now and it gets far less attention than any individual model launch.

What You Get

  • 512K context window. Larger than most, smaller than the 1M and 2M tiers, and in practice the useful range is broader than either because the degradation curve is gentler.
  • Open weights. Downloadable, self-hostable, fine-tunable. Large, but within reach of a serious multi-GPU node.
  • Strong agentic tool use. This is where K3.7 is genuinely differentiated. Long tool chains hold together well, which is not true of most models at this price.
  • A Turbo tier at $1.20 per million for latency-sensitive work, and a distilled Mini at $0.15 that runs on a single high-memory GPU.

What You Do Not Get

Multimodal is weak at 74.2 on MMMU. Writing quality is noticeably behind the frontier, which matters if the output is customer-facing prose rather than code. English instruction-following is good but not flawless, and the failure mode is subtle: it follows the letter of a complex instruction while missing the intent more often than the top-tier models do.

And the hosted API runs in China, which is a hard stop for a lot of buyers. The open weights are the answer to that, but self-hosting a model this size is a real commitment. Estimate your hardware requirement with the RAM calculator before you plan around it.

Who Should Actually Switch

If you are running high-volume agentic coding workloads and paying frontier prices for them, you should at minimum benchmark K3.7 against your own tasks this week. The potential saving is large enough that an afternoon of evaluation has an obvious expected value.

If your agentic work is low volume, the switching cost probably exceeds the saving. If it is customer-facing and quality-sensitive, stay where you are.

The Strategic Read

Moonshot is not trying to make money on inference at $0.60 per million. Nobody is. This is ecosystem capture, the same play Meta is running with Llama, executed with a model that happens to be better at the specific thing developers care most about.

The consequence for buyers is straightforward and good: the price of competent agentic coding is falling much faster than the price of frontier reasoning. Budget accordingly, and route accordingly. See the full spec on the Kimi K3.7 page.

Try Our Token Calculator

Want to optimize your LLM tokens? Try our free Token Calculator tool to accurately measure token counts for various models.

Go to Token Calculator