← Blog

How Close Are Open-Weight Models to the Frontier, Really? The Answer Depends on How You Measure

Mads Kristiansen

CTO, Liviate

How Close Are Open-Weight Models to the Frontier, Really? The Answer Depends on How You Measure

In June, Jamie Dborin from the inference company Doubleword published a blog post with a striking prediction: A frontier open-source LLM will be released on December 3, 2026. The headline was, of course, meant provocatively — but the analysis behind it hits on a central point in this summer's debate over open vs. closed models.

The prediction comes from a chart circulating on Twitter/X showing the gap between the best open-weight models and the closed frontier models on the Artificial Analysis Intelligence Index — a composite index that tries to assess models' overall capability. If you draw a trend line through the gap and extend it forward in time, it crosses zero around December 3 of this year.

But as Dborin himself points out, tongue firmly in cheek — before you liquidate your pension and move to a remote island — that's not the whole picture.

A single benchmark can fool you

The big problem is that any single index doesn't give a complete picture of a model's abilities. So Doubleword repeated the analysis across 18 different benchmarks from Artificial Analysis — everything from math and coding to agentic tasks and terminal work.

The result is thought-provoking:

Measure Result
Intelligence Index alone Gap closes around Dec. 2026
Average across all 18 benchmarks Nearly flat — just under 5 months behind throughout the period
Coding index From 15 months behind down to just 1–2 months
Most other benchmarks Moderately widening gap over time

In other words: if you only look at one index, you can predict an "open source singularity" by Christmas — but if you look across many measures, open models consistently sit around five months behind, and the gap may even be widening again outside of coding.

The backdrop: where do we actually stand in summer 2026?

Recent weeks have brought a series of notable open-weight releases — Kimi K3 from Moonshot AI (2.8 trillion parameters), GLM-5.2 from Z.ai, and Qwen3.8-Max from Alibaba. It's worth putting them in context:

The frontier gap is genuinely small. According to the State of Open Source AI report (July 2026), the strongest closed model (Claude Opus 5) scores 61 on the Artificial Analysis Intelligence Index v4.1, versus 57 for the strongest open model (Kimi K3). That puts Kimi K3 in fourth place overall — ahead of three of the largest closed labs. The Epoch Capabilities Index shows the same picture: Kimi K3 at 156 versus GPT-5.6 Sol's 162 — a six-point gap, roughly one release cycle.

But it depends on the type of task. On frontend coding, open models are actually leading (Kimi K3 debuted at the top of LMArena's Frontend Code Arena). On agentic terminal work, it's a close race (88.3 versus Sol's 88.8). Closed models, on the other hand, still hold a clear edge in deep reasoning, long-context reliability and professional knowledge tasks.

An independent LessWrong study also points to something important: on private benchmarks, where the data isn't publicly available, open models are still around 8–10 months behind — while the gap is only 4–6 months on public benchmarks. That suggests a fair amount of benchmark overfitting among open models.

Why does it matter?

There are at least three reasons this topic deserves attention:

The measurement problem is real

The story illustrates just how hard it is to measure LLM quality at all. Depending on which index you pick, you get wildly different conclusions about how close open source is to the frontier. That matters concretely for companies choosing a model provider — and for the regulatory debate that leans on these numbers.

The economics have already shifted

Even though the frontier gap technically persists, production usage is already dominated by open-weight models, according to the report: the seven most-used models by token volume on OpenRouter are all open-weight, and the majority of production tokens are now routed through them. The price drop has been dramatic — inference pricing fell roughly 50x over three years. When you can get within a few points of frontier-level performance at a fraction of the price (Kimi K3 costs about a third), the math quickly becomes compelling for most workloads.

The sovereignty angle

For European companies, the trend feeds directly into the debate over AI sovereignty and EU-hosted infrastructure (which we've also covered here on the blog before). If the best models' capabilities can be downloaded as open weights and run locally, or with European hosting providers, instead of through American APIs or cloud services that transfer data to third countries — then that changes the compliance picture significantly under both GDPR and the EU AI Act.

Conclusion

Doubleword's analysis doesn't deliver a definitive answer — if anything, it elegantly demonstrates how easy it is to fool yourself with a single number. The most honest picture right now looks like this:

Open-weight models have genuinely caught up to the frontier in coding and agentic tasks to near-parity — but still trail by several months overall, especially when measured broadly or on private benchmarks that can't be trained against.

Whichever side of the "open source singularity" debate you land on, the story underscores one thing clearly: the model choices of the future are less about "open vs. closed" as a black-and-white question, and more about matching the right model to the right task — with an eye on capability, price and data jurisdiction alike.

Want to talk about what's actually blocking your AI initiatives?

Mads Kristiansen is happy to have a no-obligation chat.

Book a meeting →