back to top
HomeTechKimi K3 May Be the Biggest Open-Weight AI Release of 2026.

Kimi K3 May Be the Biggest Open-Weight AI Release of 2026.

- Advertisement -

There’s a new open-weight model out there right now that almost nobody can actually download.

That should sound like a contradiction. Open-weight is supposed to mean anyone can grab the file and run it themselves, no waiting. Moonshot AI broke that pattern anyway, and the strange part is they broke it for a model big enough that the wait might be worth it.

Kimi K3 is the largest open model ever built. The largest one anyone has shipped and early results have it beating Claude and GPT on tasks those two have spent the last year treating as their own territory.

Open models have spent two years playing catch-up, closing gaps quarter by quarter while everyone waited for the day one of them actually pulled ahead. That day might already be here, and the model responsible for it is currently locked behind an app you can use but can’t take home.

So the question is what it actually beats, what it still can’t touch, and why Moonshot decided to make the world wait for the weights while everyone else gets to watch.

What makes Kimi K3 the biggest open release yet

Start with the number itself. At 2.8 trillion parameters, K3 is the largest open-weight model ever released. Not the largest this quarter, The largest, period, beating every open model from Z.AI, Alibaba, DeepSeek, and Xiaomi that’s shipped in the past year.

What’s more telling than the single number is the pattern behind it. Moonshot’s own tracking shows their models have held the record for largest open-weight release for nine of the past twelve months. Kimi K2 set the bar, and instead of some other lab leapfrogging them, K3 just raised it further.

The architecture is what makes that scale usable. K3 runs on two custom components, Kimi Delta Attention and Attention Residuals, paired with a sparse Mixture of Experts setup that activates only 16 of 896 available experts for any given task. That sparsity is the whole trick. Moonshot claims roughly a 2.5x improvement in scaling efficiency over K2, meaning the model converts raw compute into actual capability more effectively rather than just getting bigger for its own sake. A 2.8 trillion parameter model that ran inefficiently would be a research curiosity. One that scales efficiently is a model people will actually deploy.

What the benchmarks actually say

Kimi K3 Coding Benchmarks
Kimi K3 Benchmarks

K3 doesn’t win everywhere. It wins in specific places, and those places tell you something.

Take BrowseComp, a benchmark that measures how well a model can research something autonomously online, chasing down sources, verifying claims, following threads a human would take hours to untangle. K3 scores 91.2. Claude Fable 5, currently one of the strongest models on the market, scores 88.0. GPT-5.6 Sol comes in at 90.4. K3 isn’t just competitive here. It’s ahead of both.

Then there’s SWE Marathon, built to test something narrower and arguably harder: can a model sustain real engineering work over a long stretch without losing the thread. K3 hits 42.0. Opus 4.8 lands at 40.0. Fable 5 falls further behind at 35.0, though its score comes with an asterisk Moonshot flagged themselves: Fable 5 hit fallback behavior on 35% of the tasks in their evaluation, which may have dragged its number down. Even accounting for that, K3 still leads.

Job Bench and Automation Bench tell the similar thing. K3 scores 52.9 and 30.8, ahead of Opus 4.8 (48.4, 27.2) and comfortably clear of GLM-5.2 (43.4, 12.9), the model that made headlines just weeks ago for being the closest an open model had come to Claude. K3 isn’t closest anymore. On these two, it’s simply better.

It’s not a clean sweep, and it shouldn’t be treated like one. On DeepSWE, a coding benchmark, GPT-5.6 Sol leads at 73.0 with K3 at 67.5, trailing even Fable 5’s 70.0. FrontierSWE has Fable 5 well out in front at 86.6 against K3’s 81.2. The pattern that emerges isn’t “open beats closed now.” It’s narrower and more interesting than that: K3 wins specifically on long-horizon, sustained-effort tasks, the kind that reward a model for not losing patience or coherence over time, while still trailing on some raw coding benchmarks where the proprietary labs have had a head start.

That’s a more useful story than a scoreboard. It suggests Moonshot optimized for something specific, endurance over raw sprint speed, and it worked.

You May Like: Open Source AI Coding Agents That Don’t Need a Subscription

Where it still limits

The honest version of this story includes the gap, not just the wins.

On reasoning and general knowledge, K3 falls behind by a wider margin than anywhere else. HLE-Full has Fable 5 at 53.3 against K3’s 43.5, and the gap holds even with tool use added, 63.0 versus 56.0. This is the benchmark built to be genuinely hard to game, questions designed to resist memorization and force real reasoning, and it’s where the size advantage stops mattering as much as whatever Anthropic and OpenAI are doing differently under the hood.

Vision shows similar thing. Fable 5 leads K3 across nearly every multimodal benchmark in the table, from MMMU-Pro to CharXiv to WorldVQA, sometimes by a wide margin. On WorldVQA specifically, Fable 5 scores 56.7 against K3’s 51.0. K3 isn’t bad at visual reasoning. It’s just not the category where 2.8 trillion parameters translates into an edge.

Then there’s a limitation Moonshot disclosed about themselves, which is rarer than it should be in a model announcement. K3, they say, tends toward “excessive proactiveness.” Give it a long task with any ambiguity in it, and it may start making decisions on your behalf that you never asked for. Moonshot’s own advice is to write explicit behavioral constraints into your system prompt if you need the model to stay inside firm boundaries. That’s not a small caveat. It means K3 is confident enough in long tasks to start improvising, which is exactly the kind of trait that’s useful in a benchmark and unpredictable in production.

Put together, the picture is a model that’s genuinely ahead on endurance and web research, genuinely behind on deep reasoning and vision, and openly acknowledged by its own creators to occasionally do things nobody asked it to do.

The catch

A year ago, open models tended to ship the moment they were announced. Announcement and availability were basically the same event. That pattern has been quietly shifting, and K3 is part of that shift. More labs, open ones included, are now separating the announcement from the actual release, for reasons that range from safety review to partner coordination to simply wanting the ecosystem ready before the weights hit the wild.

Moonshot has been upfront about it. They’ve said the full weights land July 27, and they’re using the time before that to work with inference partners and open-source maintainers so the rollout is stable from day one rather than chaotic. Once the weights are out, expect the usual next step too: the open-source community tends to move fast on quantized versions, shrinking a model like this down so it can actually run on hardware well below what 2.8 trillion parameters would normally demand. That’s often where a release like this becomes usable for far more people than the original weights ever could.

Frontier AI Is No longer Closed

Frontier used to mean something closed labs owned. A gap measured in months, sometimes years, that open models were always chasing and never quite closing.

K3 doesn’t erase that gap. It’s still behind on reasoning, still behind on vision, still capable of going rogue on a task nobody asked it to improvise on. But it’s also beating Claude and GPT on the exact kind of long-horizon work that used to be the clearest proof closed labs were ahead.

Frontier isn’t a place anymore. It’s a moving line, and open models just proved they can stand on either side of it.

Don’t miss any Tech Story

Subscribe To Firethering NewsLetter

You Can Unsubscribe Anytime! Read more in our privacy policy

LEAVE A REPLY

Please enter your comment!
Please enter your name here

YOU MAY ALSO LIKE
OpenAI Says Its AI Escaped Testing and Hacked Hugging Face

OpenAI Says Its AI Escaped Testing and Hacked Hugging Face

0
OpenAI just confirmed something the AI industry has never publicly admitted before. During an internal cybersecurity evaluation, one of its frontier AI models broke out of its restricted testing environment, found a previously unknown software vulnerability, gained access to the open internet, and ultimately breached Hugging Face's production infrastructure. It wasn't trying to steal data, According to OpenAI, the model was simply trying to score better on a cybersecurity benchmark. In other words, the AI found a way to cheat on its own test. The incident is being described by OpenAI as an "unprecedented cyber incident." Hugging Face initially believed it was under attack from an external AI agent before investigators traced the activity back to OpenAI's own evaluation environment. While the breach was quickly contained and both companies are now working together on the investigation, the episode raises a much bigger question. If an AI model can independently discover a zero-day vulnerability, escape a sandbox, chain together multiple exploits, and compromise a real production system simply to complete an assigned task, what happens when future models become even more capable?
Open Source AI Assistants You Can Run Locally

5 Best Open Source AI Assistants You Can Run Locally

0
Somewhere between "just use ChatGPT" and "compile this from source," there's a category of AI tools that don't get talked about enough. Apps you download once, run on your own hardware, and never pay a monthly fee to use. No data leaving your machine. No API key. No usage limits that reset on the first of the month. The tools in this list aren't compromises. Some of them have millions of downloads. One was built by Mozilla. Another turns a 1B model into a desktop companion that reacts to your coding sessions. What they share is that after the initial setup, they answer only to you. If your current AI workflow depends on a subscription staying affordable and a company deciding your use case still matters next quarter, these are worth knowing about.
apple-sues-openai-stolen-secrets

OpenAI Paid $6.5 Billion to Build an iPhone Rival. Apple Says It Was Built...

0
Last year, OpenAI acquired io, Jony Ive's hardware startup, for $6.5 billion. The deal was widely read as OpenAI's clearest signal yet that it was serious about building a physical device, something that could sit in your pocket the way an iPhone does, powered by AI agents instead of apps. A direct challenge to Apple's most important product. On Friday, Apple filed a lawsuit suggesting that challenge was built on a foundation of stolen confidential information through what Apple describes as a coordinated operation directed from the top of OpenAI's hardware division, the same division now tasked with building the device meant to compete with Apple. Apple isn't just alleging that some employees walked out with files they shouldn't have taken. It's alleging that the people now running OpenAI's hardware ambitions actively ran a system to extract Apple's most guarded technical knowledge, and that the $6.5 billion acquisition sits on top of that foundation.