Model launches

Open-model landscape: Kimi K3, GLM 5.2, Qwen and the distillation debate

Hosts Nathan Lambert and Florian Brand review the recent surge in Chinese open models after Kimi K3’s release, compare them to GLM 5.2 and Qwen, and discuss infrastructure, data and talent factors that may explain rapid progress.

Open-model landscape: Kimi K3, GLM 5.2, Qwen and the distillation debate

In a recent Interconnects podcast, Nathan Lambert and Florian Brand reviewed the rapid developments in the open‑weight model ecosystem after the Kimi K3 release. They compared Kimi K3 with GLM 5.2 and Qwen, surveyed other Chinese and Western providers, discussed why Chinese labs appear to be advancing quickly, and debated the practical impact of distillation (SFT) versus large‑scale reinforcement learning (RL). They also considered cybersecurity implications and made short‑term predictions about which labs will remain competitive.

Takeaways on Kimi K3 and GLM 5.2

  • Kimi K3’s debut has generated a lot of activity. Its API has experienced heavy load and errors, and access has been uneven across users and subscription tiers. The model weights were not yet broadly available at the time of the discussion, and the hosts expected intense post‑training activity once weights are released.
  • GLM 5.2 continues to play an important role: many developers use it via fast internal endpoints and value its speed and adequacy for many grunt tasks.
  • In practice: Kimi K3 showed surprising strengths in some tasks such as identifying community trends and produced simpler, more readable code in many coding tasks; however, on niche or edge cases Codex‑level refinement can still outperform it.

Why might Chinese models be advancing quickly?

The hosts outlined several plausible, not mutually exclusive explanations:

  • Cost and capital efficiency: Chinese labs might be able to convert capital to compute, data and talent more cheaply, lowering the effective cost of model iterations.
  • Focus and team structure: many research teams they observed appeared tightly focused on a single model goal rather than multiple side projects.
  • Growing domestic chip supply and workaround of export limits: they discussed increased use of domestic chips and implied that workarounds for export restrictions have increased available compute, especially for inference workloads.
  • Data and environments: there are early signs of more purchases of external data and environment tooling in China, though the scale and effect are hard to quantify from public signals.

Chinese providers and their positions

The podcast covered Kimi, Zhipu / GLM, Qwen (Alibaba), DeepSeek, MiniMax, Ling, Meituan, LongCat and others. Highlights:

  • Qwen excels at smaller models and developer adoption, helping Alibaba capture cloud developer mindshare, but its largest models have not always ranked as highly in absolute performance.
  • DeepSeek iterates very rapidly on preview endpoints, uploading new checkpoints frequently.
  • MiniMax’s licensing and openness have shifted over time; its future licensing choices will matter.
  • Many Chinese models provide substantial internal value for companies, but they don’t always immediately produce a community developer breakout.

The U.S. and global open‑model ecosystem

Hosts also reviewed players outside China: Thinking Machines (Inkling models), Nemotron, Poolside, Reflection, Nvidia and Google’s Gemma family. Key notes:

  • Several U.S. teams are focusing on models that are easy to fine‑tune for domain tasks, which can drive adoption even if they’re not frontier benchmark leaders.
  • Shipping weights and enabling efficient inference remain major infrastructural hurdles for rapid adoption.

Frontier vs. near‑frontier and security implications

  • The hosts distinguished between near‑frontier capabilities (sufficient for many coding and productivity tasks) and true frontier capabilities (new scientific discoveries, drug design, novel proofs). They expect the latter to remain concentrated in the very frontier labs for longer.
  • Security concern: if defenders (enterprises, security teams) are legally or practically prevented from using certain high‑quality open models, defenders could fall behind attackers who can access those models. This disparity could exacerbate cyber risk.

Distillation (SFT) vs RL — how much does a teacher model help?

  • The discussion drilled into whether extracting reasoning traces and completions from top closed models (distillation/SFT) materially accelerates catch‑up. The hosts’ view:
    • Distillation and SFT do have effects (e.g., shaping manner, behavior and some capabilities), but on their own they rarely yield frontier‑level performance.
    • Large RL runs (millions of agent rollouts) are where major capability improvements happen; these require fast and cost‑effective judge models and substantial infrastructure.
    • They challenged claims that distillation’s role has risen so much that it should be specifically legislated against; available research and practice do not yet support a clear, sweeping conclusion.

Cybersecurity and access inequality

If policy or legal pressure effectively prevents some organizations from using the best open‑weight models, defenders could be forced to rely on weaker models for defensive analysis. The hosts argued this would create structural disadvantages in cyber defense.

Predictions and leaderboard thinking

  • Near term, Kimi and Zhipu/GLM looked like front‑of‑pack contenders; DeepSeek and Qwen were viewed as close competitors, with MiniMax a potential surprise later in the year.
  • It’s possible a U.S. company (Nvidia, Thinking Machines, Reflection or similar) could enter the close‑competitor cluster by year‑end if a smaller, highly useful model breaks through, but that outcome is uncertain.
  • They did not expect an explosion in open models of 5–10+ trillion parameters within the short term; instead, fine‑tunability, infrastructure, and iteration cadence are central.

Conclusions

The open‑weight landscape is evolving fast, driven by rapid iteration cycles and growing infrastructure in multiple regions. The decisive factors for competitive positioning are nuanced licensing choices, weight availability, inference tooling, data and environment purchases, and the ability to fine‑tune effectively for domain tasks. Distillation has a role, but the hosts emphasized that massive RL runs and engineering scale typically determine the largest performance leaps.

Footnote

The hosts also mentioned Nathan Lambert’s upcoming book on sharing post‑training knowledge; the book is shipping soon and was noted at the start of the episode.