Open Weights vs Closed Models: In 2026 the Price Advantage Has Vanished

For three years the logic was simple: closed models are better, downloadable models are cheaper. In July 2026 that sentence no longer holds. The strongest open model on the market costs more than a closed model that beats it on the intelligence index, and the cheapest open model costs one fiftieth of the first.

Let us compare three models that are open, or announced as such: Kimi K3 (Moonshot AI), GLM-5.2 (Zhipu, trading as Z.ai) and DeepSeek V4.

The numbers

MetricKimi K3GLM-5.2DeepSeek V4
Architecture2.8T MoE744B MoE, 40B activeMoE (Pro-Max variant)
Licencenot yet published, weights due 27 JulyMITMIT
Context1.05M tokens1M tokens1M tokens on the Flash variant
Price (input / output per 1M tokens)$3.00 / $15.00around $1.10 / $4.10$0.14 / $0.28 (V4-Flash)
AA Intelligence Index5751not published for V4 Pro-Max
SWE-bench Verified93.4% (Vals AI)not reported80.6% (vendor, Pro-Max)
Terminal-Bench88.3% on 2.1 (KimiCode harness)81.0% on 2.0not reported
GPQA93.5%91.2%not reported
Humanity’s Last Exam56.0%54.7%not reported

For reference, the closed models at the top of Artificial Analysis’s Intelligence Index v4.1: Claude Opus 5 at 61 ($5 / $25), Claude Fable 5 at 60, GPT-5.6 Sol at 59.

Three things that table is saying

The technical gap has almost closed. Kimi K3 sits at 57 on the Intelligence Index against Claude Opus 5’s 61. Four points. In 2024 the lag between downloadable models and the frontier was measured in twelve to eighteen months. Today it is measured in weeks and a handful of index points.

At the top, the price advantage is gone. Kimi K3 costs $3 and $15 per million tokens. Claude Opus 5, which beats it on nearly everything, costs $5 and $25. That is less than two to one, where the historical expectation for an open model was an order of magnitude. Moonshot priced K3 as a frontier product rather than a budget alternative, which is internally consistent: 2.8 trillion parameters are not cheap to serve. Artificial Analysis measures K3 at roughly 34 tokens per second on Moonshot’s API, with around seven seconds to first token.

The value now lives in the middle tier. DeepSeek V4-Flash at $0.14 and $0.28 with a million token context is in the same table as Kimi K3 and costs a little over one percent as much on input. GLM-5.2 at around $1.10 and $4.10 delivers an index of 51 under an unrestricted MIT licence. Among closed models, Artificial Analysis puts a Grok 4.5 Intelligence Index task at about $0.31, five times cheaper than Claude Sonnet 5.

What “open” actually means now

Three models, three different meanings of the same word.

GLM-5.2 is open in the full sense. MIT licence, downloadable weights, no usage restrictions, fine tuning and inspection permitted. At 40 billion active parameters it is also realistically servable in house on ordinary infrastructure.

Kimi K3 is open in the announced sense. Weights are expected on 27 July 2026 and the licence has not been published. Even once they land, 2.8 trillion parameters require multi node GPU clusters: within reach of teams that already run that kind of infrastructure, not of a team running an experiment on a single pod.

DeepSeek V4 is open and light. MIT, self hostable, with a near free Flash variant. It gives up ground on headline benchmarks and wins everywhere volume matters more than the last percentage point.

The architecture that is emerging

The pattern consolidating in real deployments is not “pick a model.” It is routing: a cheap open model handles somewhere between 72% and 96% of traffic, a closed frontier model catches the hard cases, and a gateway sits in front of both. Analyses of production workloads report that this configuration beats either model on its own, on quality and on cost.

The macro context explains why. BenchLM’s Token Price Index reads 12 for July 2026 against a base of 100 in March 2023, meaning frontier token prices are 88% below that base, with a median blended $4.50 per million tokens across thirteen flagship models. GPT-4 level capability, which cost roughly $30 per million tokens in 2023, is now available for under a dollar.

In practice

If your workload is high volume and low difficulty, the economically rational choice is no longer the strongest open model. It is the smallest open model that clears your quality bar, with a frontier fallback for the rest.

If you need a downloadable model for data sovereignty, compliance or fine tuning reasons, GLM-5.2 is currently the soundest option: clean licence, manageable size, verified scores.

Kimi K3 is worth watching for one reason, but an important one: if the weights and licence genuinely arrive, it becomes the first downloadable model with independently verified frontier scores. 27 July is the date to watch.

Leave a Reply

Your email address will not be published. Required fields are marked *