Kimi K3 is a very good model with excellent benchmarks. Assuming its weights are released as planned it will become, purely in terms of raw capability, the strongest open model.

Do not get carried away. Do not judge Kimi K3 only its relative strengths. In aggregate it is several months behind the closed model frontier, at least four and my median guess is six, with the post-training closer and the pre-training farther out. This is less months than before, but the months are denser now.

It is somewhat distilled. It likely outperforms on benchmarks relative to practical performance. All its benchmarks are scored at maximum effort, typically a lot more tokens than are used in similar tests by Fable or Sol. Performance looks jagged. Kimi will be excellent at some things, less so at other things.

We will know more over the coming weeks. For now access is spotty and not that many people have actually had the chance to try Kimi K3, so I have larger error bars than usual around its capabilities. Alas, time waits for no one, so we press on.

It is the largest open model so far at 2.8T, on [...]

---

Outline:

(03:07) DeepSeek Moments: Here We Go Again

(05:47) We Had a Moment (Reprise from June 2025)

(10:03) The Story Since Then

(16:19) The Kimi K3 Announcement, Pitch and Basic Facts

(19:34) On Modern Benchmaxxing

(21:16) Other People's Benchmarks

(26:15) Benchmarks Are Not The Real World

(27:17) Technical Safeguards? What Are Those?

(30:53) Things Kimi Can Do

(32:06) Things Kimi Cannot Do

(33:40) Things It Is Not Easy To Get Kimi To Do

(37:02) Open Weight Models Are Unsafe And Nothing Can Fix This

(40:34) Dean Ball Attempts To Be Constructive

(58:24) Trump Administration Considering Executive Order Banning Chinese Open Models Within the United States

(01:01:53) OpenAI Employees Are Relatively Bullish On This One

(01:03:30) Kimi K3 Is Relatively Strongest At Typical Agentic Coding, Front End Work and 3D

(01:06:06) Reactions

(01:10:14) Who Are You?

(01:12:09) How Did They Do It?

(01:15:00) Conclusion

---

First published:
July 20th, 2026

Source:
https://www.lesswrong.com/posts/t7oZyAFej8FZrfbtY/on-kimi-k3-its-capabilities-and-related-discontents

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Scatter plot titledLine graph titledBar graph comparing AI model performance percentages, with Kimi K3 highlighted.Bar graph showingBar graph comparing AI coding benchmarks titledBar graphs comparing AI models acrossBar graph titledBar graph titledA ranking table of AI models by ECI score.Bar graph titledBar graph showingBar chart titledBar graph titledTriangular heatmap matrix comparing AI language models' writing style similarity in bits.Text about keeping humor light and finalizing writing.Bar graph showing

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Podden och tillhörande omslagsbild på den här sidan tillhör zvi. Innehållet i podden är skapat av zvi och inte av, eller tillsammans med, Poddtoppen.