back to top
HomeTechKimi K3 May Be the Biggest Open-Weight AI Release of 2026.

Kimi K3 May Be the Biggest Open-Weight AI Release of 2026.

- Advertisement -

There’s a new open-weight model out there right now that almost nobody can actually download.

That should sound like a contradiction. Open-weight is supposed to mean anyone can grab the file and run it themselves, no waiting. Moonshot AI broke that pattern anyway, and the strange part is they broke it for a model big enough that the wait might be worth it.

Kimi K3 is the largest open model ever built. The largest one anyone has shipped and early results have it beating Claude and GPT on tasks those two have spent the last year treating as their own territory.

Open models have spent two years playing catch-up, closing gaps quarter by quarter while everyone waited for the day one of them actually pulled ahead. That day might already be here, and the model responsible for it is currently locked behind an app you can use but can’t take home.

So the question is what it actually beats, what it still can’t touch, and why Moonshot decided to make the world wait for the weights while everyone else gets to watch.

What makes Kimi K3 the biggest open release yet

Start with the number itself. At 2.8 trillion parameters, K3 is the largest open-weight model ever released. Not the largest this quarter, The largest, period, beating every open model from Z.AI, Alibaba, DeepSeek, and Xiaomi that’s shipped in the past year.

What’s more telling than the single number is the pattern behind it. Moonshot’s own tracking shows their models have held the record for largest open-weight release for nine of the past twelve months. Kimi K2 set the bar, and instead of some other lab leapfrogging them, K3 just raised it further.

The architecture is what makes that scale usable. K3 runs on two custom components, Kimi Delta Attention and Attention Residuals, paired with a sparse Mixture of Experts setup that activates only 16 of 896 available experts for any given task. That sparsity is the whole trick. Moonshot claims roughly a 2.5x improvement in scaling efficiency over K2, meaning the model converts raw compute into actual capability more effectively rather than just getting bigger for its own sake. A 2.8 trillion parameter model that ran inefficiently would be a research curiosity. One that scales efficiently is a model people will actually deploy.

What the benchmarks actually say

Kimi K3 Coding Benchmarks
Kimi K3 Benchmarks

K3 doesn’t win everywhere. It wins in specific places, and those places tell you something.

Take BrowseComp, a benchmark that measures how well a model can research something autonomously online, chasing down sources, verifying claims, following threads a human would take hours to untangle. K3 scores 91.2. Claude Fable 5, currently one of the strongest models on the market, scores 88.0. GPT-5.6 Sol comes in at 90.4. K3 isn’t just competitive here. It’s ahead of both.

Then there’s SWE Marathon, built to test something narrower and arguably harder: can a model sustain real engineering work over a long stretch without losing the thread. K3 hits 42.0. Opus 4.8 lands at 40.0. Fable 5 falls further behind at 35.0, though its score comes with an asterisk Moonshot flagged themselves: Fable 5 hit fallback behavior on 35% of the tasks in their evaluation, which may have dragged its number down. Even accounting for that, K3 still leads.

Job Bench and Automation Bench tell the similar thing. K3 scores 52.9 and 30.8, ahead of Opus 4.8 (48.4, 27.2) and comfortably clear of GLM-5.2 (43.4, 12.9), the model that made headlines just weeks ago for being the closest an open model had come to Claude. K3 isn’t closest anymore. On these two, it’s simply better.

It’s not a clean sweep, and it shouldn’t be treated like one. On DeepSWE, a coding benchmark, GPT-5.6 Sol leads at 73.0 with K3 at 67.5, trailing even Fable 5’s 70.0. FrontierSWE has Fable 5 well out in front at 86.6 against K3’s 81.2. The pattern that emerges isn’t “open beats closed now.” It’s narrower and more interesting than that: K3 wins specifically on long-horizon, sustained-effort tasks, the kind that reward a model for not losing patience or coherence over time, while still trailing on some raw coding benchmarks where the proprietary labs have had a head start.

That’s a more useful story than a scoreboard. It suggests Moonshot optimized for something specific, endurance over raw sprint speed, and it worked.

You May Like: Open Source AI Coding Agents That Don’t Need a Subscription

Where it still limits

The honest version of this story includes the gap, not just the wins.

On reasoning and general knowledge, K3 falls behind by a wider margin than anywhere else. HLE-Full has Fable 5 at 53.3 against K3’s 43.5, and the gap holds even with tool use added, 63.0 versus 56.0. This is the benchmark built to be genuinely hard to game, questions designed to resist memorization and force real reasoning, and it’s where the size advantage stops mattering as much as whatever Anthropic and OpenAI are doing differently under the hood.

Vision shows similar thing. Fable 5 leads K3 across nearly every multimodal benchmark in the table, from MMMU-Pro to CharXiv to WorldVQA, sometimes by a wide margin. On WorldVQA specifically, Fable 5 scores 56.7 against K3’s 51.0. K3 isn’t bad at visual reasoning. It’s just not the category where 2.8 trillion parameters translates into an edge.

Then there’s a limitation Moonshot disclosed about themselves, which is rarer than it should be in a model announcement. K3, they say, tends toward “excessive proactiveness.” Give it a long task with any ambiguity in it, and it may start making decisions on your behalf that you never asked for. Moonshot’s own advice is to write explicit behavioral constraints into your system prompt if you need the model to stay inside firm boundaries. That’s not a small caveat. It means K3 is confident enough in long tasks to start improvising, which is exactly the kind of trait that’s useful in a benchmark and unpredictable in production.

Put together, the picture is a model that’s genuinely ahead on endurance and web research, genuinely behind on deep reasoning and vision, and openly acknowledged by its own creators to occasionally do things nobody asked it to do.

The catch

A year ago, open models tended to ship the moment they were announced. Announcement and availability were basically the same event. That pattern has been quietly shifting, and K3 is part of that shift. More labs, open ones included, are now separating the announcement from the actual release, for reasons that range from safety review to partner coordination to simply wanting the ecosystem ready before the weights hit the wild.

Moonshot has been upfront about it. They’ve said the full weights land July 27, and they’re using the time before that to work with inference partners and open-source maintainers so the rollout is stable from day one rather than chaotic. Once the weights are out, expect the usual next step too: the open-source community tends to move fast on quantized versions, shrinking a model like this down so it can actually run on hardware well below what 2.8 trillion parameters would normally demand. That’s often where a release like this becomes usable for far more people than the original weights ever could.

Frontier AI Is No longer Closed

Frontier used to mean something closed labs owned. A gap measured in months, sometimes years, that open models were always chasing and never quite closing.

K3 doesn’t erase that gap. It’s still behind on reasoning, still behind on vision, still capable of going rogue on a task nobody asked it to improvise on. But it’s also beating Claude and GPT on the exact kind of long-horizon work that used to be the clearest proof closed labs were ahead.

Frontier isn’t a place anymore. It’s a moving line, and open models just proved they can stand on either side of it.

Don’t miss any Tech Story

Subscribe To Firethering NewsLetter

You Can Unsubscribe Anytime! Read more in our privacy policy

LEAVE A REPLY

Please enter your comment!
Please enter your name here

YOU MAY ALSO LIKE

The Biggest AI Companies Are All Building Their Own Chips. That’s Not a Coincidence.

0
Anthropic confirmed this week it's hiring a custom silicon team to design chips for running Claude. The announcement was quiet a job listing, a spokesperson confirmation, no big launch event. Easy to file under "interesting but expected" and move on. But zoom out for a second. OpenAI shipped its first custom inference chip in June. Google has been running models on its own TPUs for years. Meta has designed and deployed its own silicon. Mistral is reportedly exploring the same path. And now Anthropic. Five of the most important AI labs in the world, all arriving at the same decision, within roughly the same window. None of them are copying each other. All of them looked at the same competitive landscape and reached the same conclusion independently. That kind of convergence doesn't happen by accident. It happens when an entire industry agrees that the thing everyone assumed was someone else's problem is actually the problem and that whoever solves it first has an advantage that's very hard to close later.
AI Was Supposed to Stop Cheating. Instead, 58,000 Students Must Retake Their Exams

AI Was Supposed to Stop Cheating. Instead, 58,000 Students Must Retake Their Exams.

0
UNAM runs the largest university in Mexico. Every year, hundreds of thousands of students take an entrance exam that determines whether they get in. This year, for the first time, the whole thing went remote. They deployed a lockdown browser, AI webcam monitoring, and one human supervisor per 150 applicants. The kind of setup that sounds serious on paper. Then the scores came in. Students hitting 100 or above jumped from 3.5 percent in previous years to 16.3 percent this year. At the very top end, scores of 110 or higher went from 0.9 percent to 5.5 percent. Not a small shift. Not noise. A roughly fivefold increase in top scores, in one year, under one new format. An expert commission investigated. Their conclusion: administer the entire exam again, in person, to around 58,000 people. The rector apologized to students who hadn't cheated. They now have to prepare for and sit another exam anyway.
Claude Chats Ended Up on Google Search. Here's How It Happened

Claude Chats Ended Up on Google Search. Here’s How It Happened.

0
A single line typed into Google was all it took. Type "site:claude.ai/share" into the search bar, and over the weekend, it surfaced a long list of conversations people had shared through Claude, Anthropic's AI chatbot. Not conversations they'd shared with the world on purpose. Conversations they'd shared with one person, or thought they had.Some of what turned up reads like exactly the kind of thing you'd never want indexed anywhere. Medical records. Children's names and phone numbers. Internal company documents marked for employees only. This wasn't a hack, no one broke into anything. It was a feature working exactly as built, surfacing exactly what people had typed into it, in ways most of them almost certainly never intended.