Ad
Skip to content

Moonshot's Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex math

Moonshot's AI model Kimi K3 is getting a lot of attention in the Western AI community. The big question is how close it actually gets to the best Western models. Two new data points paint a mixed picture. In the Code Arena: Frontend benchmark, which ranks models based on human preference ratings, Kimi K3 scores 1,679, beating Claude Fable 5 (1,631), GPT-5.6 Sol (1,618), and every other tested model by a wide margin. It's the first time a Chinese model has claimed the top spot on this benchmark.

The picture looks different for hard math. According to data from Epoch AI, Kimi K3 hits only about 39 percent accuracy on FrontierMath Tier 4, the benchmark's hardest expert-level math tasks. Models from OpenAI and Anthropic score close to 90 percent there in some cases.

Kimi K3 falls well behind top Western models on complex math tasks. | Image: Epoch AI
Ad

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Source: via X | Frontier Math