DeepSeek V4.5 Doubled the Context and Refused to Raise the Price
DeepSeek shipped V4.5 in early June. The context window doubled to 256K. Tool calling improved substantially. Multilingual output quality went from adequate to good. The price did not move: $0.45 per million input tokens, $1.80 per million output.
That last sentence is the whole story.
The Strategy Nobody Else Can Copy
Every other provider prices to a margin. DeepSeek prices to a floor and lets everyone else decide whether to follow. So far, mostly, they have not been able to. The gap between DeepSeek V4.5 and the nearest Western equivalent on cost per useful token is still roughly an order of magnitude.
Put it concretely. A workload that runs 500 million input tokens and 100 million output tokens a month costs about $405 on V4.5. The same workload on a mid-tier hosted frontier model runs somewhere between $2,000 and $4,000 depending on which one you pick. On a premium reasoning tier it can clear $12,000. Model those numbers for your own usage on the annual cost calculator.
Is It Good Enough?
For a large and growing set of tasks, yes. V4.5 posts 88.7 on MMLU, 90.8 on HumanEval and 58.6 on SWE-bench Verified. That is not frontier. It is roughly where the frontier was nine months ago, at a twentieth of the price.
The honest framing is that model choice has stopped being a single decision. You are not picking one model, you are building a routing layer. Cheap model for the eighty percent of requests that are easy, expensive model for the twenty percent that are not, and a classifier deciding which is which. DeepSeek V4.5 is an extremely strong candidate for the cheap side of that split.
R2 Is the Other Half
DeepSeek R2, the reasoning sibling from March, still exposes full chain of thought and still scores 96.4 on MATH and 82.6 on GPQA. Those are numbers that were exclusive to premium closed reasoning endpoints a year ago. R2 charges $0.40 per million input.
Between V4.5 for general work and R2 for hard reasoning, DeepSeek now anchors the bottom of the price curve in both directions. Compare them directly on the comparison tool.
The Caveats, Stated Plainly
- Data residency. The hosted API runs in China. For a lot of buyers that is the end of the conversation. The weights are open, so self-hosting is available, but self-hosting a model this size is a real infrastructure commitment.
- Multimodal is weak. 72.8 on MMMU is well behind everything else at this tier. If your workload involves images, look elsewhere.
- Long-context degradation. The 256K window is real but the useful middle is smaller than the number suggests. Same caveat as everyone else, just worth restating.
The Thing to Actually Take Away
The price floor moving down is more consequential than the ceiling moving up. A new frontier model being 3% better on a benchmark changes what is possible for a handful of teams. A competent model getting 20x cheaper changes what is economically viable for everyone.
Full details on the DeepSeek V4.5 page.