Better than Fable 5 and GPT-5.6? Chinese unveil new Kimi V3 model, but independent measurements are still missing Home News Chinese Moonshot AI unveiled Kimi K3 – a model with 2.8 trillion parameters and a context of up to 1 million tokens According to the manufacturer, it crushes Chinese competitor GLM 5.2, but against Western leaders (Fable 5, GPT-5.6, Opus 4.8), it's more of a draw All figures come from Moonshot, some even from an internal benchmark – independent tests are still missing Sdílejte: Jakub Kárník Published: 23. 7. 2026 12:30 Advertisement The label “beats the best Western models” is attached to almost every other announcement from China in 2026. Now it’s Moonshot AI‘s turn with its new flagship model Kimi K3. On paper, it’s a real powerhouse: 2.8 trillion parameters, Mixture-of-Experts architecture, and a context window of up to 1 million tokens. The ambitions are clear – to compete with the best that OpenAI, Anthropic, and domestic rivals offer today. A trillion parameters and memory for an entire repository Benchmarks to be taken with a grain of salt The real battlefield is called code A trillion parameters and memory for an entire repository The biggest draw is precisely the million-token context window. A programmer can load an entire repository, several scientific papers, or a stack of legal documents into a single conversation without having to aggressively shorten the text. Moonshot, meanwhile, primarily targets Kimi K3 at developers – it’s not meant to function as a code autocompleter, but as a partner who understands large projects, refactors across files, and can debug errors. In addition, it offers image processing, so-called agent swarms, and integration into its own Kimi Code tool. However, the full million tokens will only be enjoyed by subscribers of higher tiers; the cheaper level stops at 256 thousand. Benchmarks to be taken with a grain of salt And now for the main thing, what all the hype is about. Against Chinese rival GLM 5.2, Kimi K3 is truly in a different league – in the DeepSWE test, it scored 67.5 points against 46.2, and in SWE Marathon, even 42 against a mere 13. There’s no debate there. However, the headline “beats Claude Fable” is already a marketing exaggeration. Against the Western elite, Kimi K3 only sometimes wins: in Program Bench (77.8) and SWE Marathon (42.0), it is indeed first, but in DeepSWE, both Fable 5 and GPT-5.6 Sol outperform it, and in Moonshot’s own Kimi Code Bench 2.0, it finished with 72.9 points behind Fable 5 (76.9). It’s no surprise that a manufacturer ensures a good score in its own test – that’s precisely why internal benchmarks should be viewed with a critical eye until independent measurements arrive. The real battlefield is called code Regardless of how the numbers turn out after re-measurement, the direction is clear. Kimi Code is joining the same league as Claude Code from Anthropic or Codex from OpenAI, and the battle for programmers is becoming the main arena of the entire industry. For developers, this is good news – competition drives prices down and capabilities up. But let’s save the rest of the enthusiastic superlatives until Kimi K3 passes through the hands of someone other than its creators. Do you trust benchmarks from manufacturers themselves, or do you wait for independent tests? Source: Data Science in Your Pocket, Moonshot AI About the author Jakub Kárník Jakub is known for his endless curiosity and passion for the latest technologies. His love for mobile phones started with an iPhone 3G, but nowadays… More about the author Sdílejte: AI Anthropic Čína Claude OpenAI