Kimi K3 Just Landed — And It's Closing the Gap on the World's Best Models

Moonshot AI's Kimi just got a serious upgrade. Kimi K3, the company's new flagship model, dropped this week as a 2.8-trillion-parameter system — making it the largest openly released model the world has seen so far.
The story is where Kimi K3 lands when you put it next to the models everyone already benchmarks against: Claude and GPT.
The Numbers, Quickly
Before we get to the "who's winning" question, here's what Kimi K3 actually is:
- 2.8 trillion parameters — built on a new hybrid linear-attention architecture (Kimi Delta Attention) paired with something called Attention Residuals
- A 1-million-token context window, letting it hold enormous amounts of information — entire codebases, long documents, sprawling conversations — without needing to compress or summarize along the way
- Native multimodal support, so it can reason over images as well as text
- Available now in the Kimi app, desktop client, and command line, split into two variants — a general-purpose "Max" flagship and a "Swarm Max" version built for large-scale parallel search and batch processing
This isn't Moonshot's first swing at scale, either. Kimi K3 is reportedly the ninth time in the last twelve months that a Kimi model has pushed the ceiling on open-model size. This is a team that has made "biggest open model" something of a habit.
So — Does It Beat Claude and GPT?
Here's where it gets interesting, and where the honest answer is: almost, but not quite.
On independent benchmarks measuring real-world knowledge work — things like GDPval-AA v2 (which scores AI performance across 44 professions and 9 industries) — Kimi K3 posted a strong result, edging out Anthropic's Claude Opus 4.8 Max. That's a genuinely impressive result for an open model going up against a leading closed one.
But zoom out one tier further, and the picture is different. Across the board, K3's overall standing places it just behind two models: Claude Fable 5 and GPT-5.6 Sol — currently considered the frontier of the field. On a separate benchmark for long-horizon agentic knowledge work (AA-Briefcase), Kimi K3 landed in second place overall, again trailing Fable 5 but ahead of GPT-5.6 Sol's own numbers on that particular test.
The pattern that emerges is this: Kimi K3 has closed most of the distance to the absolute frontier, but the frontier itself hasn't stood still. It's not that Kimi is chasing last year's best models — it's chasing this year's, and getting remarkably close.
Why This Matters More Than the Parameter Count
It's tempting to treat "trillions of parameters" as the whole story, but raw size and real capability aren't the same thing — a bigger model doesn't automatically mean a smarter one, and Moonshot's own efficiency work (their newer attention architecture) has been arguably more important to Kimi K3's results than sheer scale.
What actually matters here is the trend line. A year ago, the idea of an open-weight model built in China trading blows with the best closed models from Anthropic and OpenAI on real knowledge-work benchmarks would have sounded aspirational. Now it's a fair description of where Kimi K3 sits: not the leader, but no longer a distant follower either.
The Open Question
That leaves an interesting question hanging over the rest of the AI field: is this gap between open and closed frontier models shrinking because open models are catching up — or because the frontier's growth rate is slowing down enough for others to close in? Kimi K3 doesn't answer that question on its own, but it's a very useful data point.
Either way, it's worth watching where the next model — Kimi's or otherwise — lands relative to this one.