All posts

Kimi K3 vs Claude Fable 5 vs GPT 5.6 Sol: What the Numbers Say, and What They Miss

Bill XuBill Xu
Published on Jul 21, 2026 18min read

Kimi K3 vs Claude Fable 5 vs GPT 5.6 Sol: What the Numbers Say, and What They Miss

Three frontier models landed within five weeks of each other. Claude Fable 5 went GA on June 9. GPT 5.6 Sol followed on July 9. Kimi K3 arrived on July 16 as the largest open weight model ever released. Composite leaderboards suggest they are nearly interchangeable. The benchmark detail says otherwise. This post lays out the verified numbers first, then what those differences mean when a model has to carry real application work from research to deployment.

The Three Contenders

Kimi K3 is Moonshot AI's flagship: a mixture of experts model with 2.8 trillion total parameters, 16 of 896 experts active per token, native image and video input, and a 1 million token context window. It is open weight, with full weights scheduled for July 27 under a modified MIT license, per The Decoder.

Claude Fable 5 is Anthropic's most capable widely released model. It holds the top spot on most composite rankings and commands the highest price of the three.

GPT 5.6 Sol is OpenAI's current flagship, generally available since July 9. It sits between the other two on both price and most capability measures.

The One Fair Comparison

Only one benchmark measures all three models with a single methodology run by one party: the Artificial Analysis Intelligence Index. Version 4.1 aggregates nine evaluations, including GDPval, Terminal Bench 2.1, SciCode, Humanity's Last Exam, and GPQA Diamond. The scores: Claude Fable 5 at 60, GPT 5.6 Sol at 59, Kimi K3 at 57, with Claude Opus 4.8 just behind at 56. That is a three point spread on a scale where the gap between frontier and mid tier models runs fifteen points or more. On general intelligence, these three are effectively peers.

aa_kimik3_index-scaled.jpg

The prices are not peers. Fable 5 charges $10 per million input tokens and $50 per million output. Sol charges $5 and $30. K3 charges $3 and $15, with cached input at $0.30. On identical workloads, Fable costs more than three times what K3 does, for a capability gap most tasks will never surface.

44cfa6cf4a19aeffa60687cb44facbddd23a0d45-1024x1024.webp

Per task cost tells a subtler story. K3 uses more tokens than its rivals to finish equivalent work, so its real cost per task lands closer to Sol's than the sticker price implies. Cheap tokens and cheap tasks are not the same thing.

Coding: A Genuine Three Way Split

The coding benchmarks do not crown a single winner. They split by the shape of the work.

GPT 5.6 Sol leads DeepSWE at 73.0 and Terminal Bench 2.1 at 88.8, with K3 essentially tied on the terminal test at 88.3. Fable 5 owns FrontierSWE at 86.6, a full five points ahead of K3 and fifteen ahead of Sol. K3 takes Program Bench at 77.8 and, more notably, SWE Marathon at 42.0, a test of long sustained coding sessions where both GPT 5.5 and GLM 5.2 collapse to the low teens.

1d9chlgn6rtp4tqfnnmjg.webp

Read the note at the bottom of that chart before treating any single number as decisive: Fable 5 results include potential fallbacks, and Sol results include potential cyberguards. Scaffolding differences move these scores by points, not decimals.

Agentic Work: Fable Leads, K3 Is Closer Than Expected

On broad agentic evaluations, Fable 5 holds the top of the board. It leads GDPval v2 Elo at 1760 against Sol's 1748 and K3's 1668, and it takes JobBench at 57.4 and AA Briefcase Elo at 1583. When a task requires sustained judgment across many steps, Fable is still the ceiling.

The surprise is where K3 lands. It wins Automation Bench, SpreadsheetBench 2, and BrowseComp at 91.2, and it places second on nearly everything else, ahead of Opus 4.8 and GPT 5.5 across the board. On visual agent tasks, Fable leads CharXiv at 93.5 with K3 close behind at 91.3. For an open weight model priced at a fraction of its rivals, second place on frontier agentic benchmarks is the actual headline.

1d9chlbnf2ena6205244g.webp

The Caveats That Most Comparisons Skip

The published coding scores were collected on different agent scaffolds, so they are not clean model to model comparisons. Several charts above come from Moonshot's own release materials, and vendors choose which benchmarks to publish. Elo based rankings shift as new battles accumulate. And no benchmark measures the qualities that decide daily use: instruction adherence over long sessions, recovery from bad states, and consistency across reruns. Treat the numbers as a map of tendencies, not a verdict.

What These Differences Mean Inside a Multi Model Workflow

The practical question is not which model wins every benchmark. It is where each model fits within a complete build process.

Kimi K3 is most relevant when cost, deployment control, and open model access shape the decision. Claude Fable 5 is better aligned with tasks that require careful interpretation, consistent writing, and sustained work across large amounts of context. GPT 5.6 Sol is the stronger general choice when a workflow moves repeatedly between planning, code, interface decisions, debugging, and tool use.

This is why model selection should happen at the task level rather than the platform level. Research, product planning, frontend implementation, copy, and quality review do not necessarily benefit from the same model. A multi model system can assign each stage according to its requirements, then compare outputs before committing to a direction.

That approach matters for Atoms. Race Mode is designed around parallel execution: the same brief can be approached through different model behaviors, while the user evaluates the resulting structure, reasoning, and implementation. The value is not access to more model names. It is the ability to make model differences visible within the work itself.

The Bottom Line

Fable 5 buys the highest ceiling on judgment heavy, long horizon work, at the highest price. Sol buys consistency: rarely first, never far from it, with the strongest terminal and repository level coding results. K3 buys frontier adjacent capability at roughly a third of the cost, with open weights and a 1 million token context as the differentiators no closed model can match.

Three point differences on a composite index will not decide your build. Task fit, cost structure, and deployment constraints will.