Model launches

Moonshot AI launches Kimi K3, a 2.8‑trillion‑parameter open‑source language model

Beijing-based Moonshot AI has released Kimi K3, a 2.8‑trillion‑parameter large language model that the company says is now the largest open‑source model in the world.

Moonshot AI launches Kimi K3, a 2.8‑trillion‑parameter open‑source language model

Beijing‑based Moonshot AI, backed by Alibaba, on Thursday introduced Kimi K3, a large language model with 2.8 trillion parameters that the company says is now the world’s largest open‑source AI model. The launch, timed ahead of the 2026 World Artificial Intelligence Conference in Shanghai, represents a significant step for the open‑source AI movement and a notable comeback for Moonshot after market setbacks over the past 18 months.

Access and planned weight release

Researchers who reviewed the company’s technical documentation say Moonshot plans to publish the model’s full trained weights on July 27. In the meantime, Kimi K3 can be tried at kimi.com; sign‑up requires a Google account or phone number and does not require a credit card.

Technical features and innovations

Kimi K3 is described as a frontier‑class LLM with 2.8 trillion total parameters — roughly 75% larger than DeepSeek’s V4 Pro at about 1.6 trillion parameters in the company’s timeline. The model supports a 1‑million‑token context window, includes native visual understanding, and offers an always‑on reasoning mode the company calls “thinking mode.”

Moonshot attributes K3’s architecture to two in‑house innovations: Kimi Delta Attention, a hybrid linear attention mechanism, and Attention Residuals, a drop‑in replacement for residual connections that the company says yields consistent scaling gains. Both techniques were previously published by Moonshot on GitHub as open research.

API compatibility, pricing and promotion

Kimi K3’s API is compatible with the OpenAI SDK to ease integration for developers already building on OpenAI or Anthropic toolchains. Pricing is set at $3 per million input tokens and $15 per million output tokens, with cached input tokens reduced to $0.30 per million. A promotional top‑up rebate running through August 12 offers up to 30% back in vouchers for API credits of $1,000 or more.

Benchmarks: competitive with top proprietary models

Public leaderboard data and a private evaluation by analytics firm Artificial Analysis show Kimi K3 performing at or near the top of major benchmarks.

  • On GDPval‑AA v2, which measures real‑world tasks across 44 occupations and 9 industries, Kimi K3 scored 1,687, placing third behind Claude Fable 5 Max (1,815) and GPT‑5.6 Sol Max (1,747.8), and ahead of Claude Opus 4.8 (1,600).
  • On AA‑Briefcase, a private agentic benchmark from Artificial Analysis that tests long‑horizon knowledge work, K3 scored 1,527, earning second place ahead of GPT‑5.6 Sol Max (1,495) and behind Fable 5 Max (1,587).
  • K3 achieved 91.2/100 on BrowseComp, a benchmark for long‑horizon, high‑difficulty information seeking — a state‑of‑the‑art result according to the company.

Moonshot says these results were produced in a single‑agent setup using the 1‑million‑token context window without context compression or additional context management techniques, suggesting that long raw context plus strong retrieval can outperform complex multi‑agent workarounds in some tasks.

Demonstrations: 48‑hour autonomous chip design and astrophysics task

Beyond benchmarks, Moonshot showcased proof‑of‑concept demonstrations intended to highlight K3’s agentic capabilities. In a documented demo, Kimi K3 was tasked with designing a physical chip to run a nanoscale version of itself. Over 48 hours of continuous autonomous agent operation, K3 completed the chip’s full construction pipeline — from architectural design through optimization and verification — using open‑source electronic design automation tools.

The result was a functional chip design of 4 square millimeters that achieved timing convergence at 100 MHz and, in simulation, decoded over 8,700 tokens per second. Moonshot notes this is not a production chip but a demonstration of long‑range autonomous agent capability: sustaining coherent, multi‑step technical work over 48 hours, including reading documentation, making design decisions, running verification loops and iterating on failures.

In another example, the company reported that K3 reproduced the universal I‑Love‑Q relation — a task that typically takes a senior researcher one to two weeks — in roughly two hours by reading and cross‑validating more than 20 papers and implementing a complete numerical pipeline.

Moonshot AI’s fall and recovery

Kimi K3 must be seen in the context of Moonshot AI’s recent trajectory. Founded in 2023 by Yang Zhilin, a Tsinghua University graduate who previously did research at Google and Meta, Moonshot rose rapidly. The Kimi platform gained traction for long‑text analysis and AI search in 2024. By early 2026 the company had raised roughly $1.5 billion and saw its valuation rise from $2.5 billion to $4.3 billion, with reports that it sought a new round at $5 billion.

The landscape shifted after DeepSeek released a low‑cost R1 model in January 2025, which disrupted the Chinese AI market and pushed Moonshot down in monthly active user rankings. Moonshot pivoted toward open‑source releases — Kimi K2 in July 2025 and K2.5 in January 2026 — as part of a comeback strategy that culminates in K3.

Training a model of this scale required major compute resources and months of preparation, suggesting the architectural and infrastructure choices behind K3 were settled well before public release.

Geopolitical implications and enterprise impact

Publishing the full weights on July 27 is a strategic decision. Moonshot’s timeline positions K3 far above competitors such as DeepSeek (1.6T), Xiaomi (1.02T) and Alibaba (397B). By releasing the largest open‑source model, Moonshot aims to attract the global open‑source AI developer community.

This fits a broader pattern among Chinese AI firms that open‑source models to showcase capability and expand developer communities, an approach Reuters has said helps China counter U.S. efforts to limit Beijing’s tech progress. Several Chinese players — DeepSeek, Alibaba, Tencent and Baidu — have released open models, but none at this parameter count.

For enterprise technology leaders, a 2.8‑trillion‑parameter open‑source model that achieves near‑frontier performance provides a base for fine‑tuning, self‑hosting and building proprietary layers without exclusive API dependence on OpenAI or Anthropic. The trade‑off is infrastructure: inference at this scale demands substantial GPU resources. Moonshot’s Mooncake project, which won Best Paper at FAST 2025, introduced KV‑cache‑centric disaggregated serving designed to make extreme‑scale inference more practical.

Kimi Code and the product lineup

Moonshot also emphasizes developer tools. Kimi Code, its open‑source coding agent competing with Anthropic’s Claude Code and Google’s Gemini CLI, received major updates (versions 0.25.0 and 0.26.0) on the day K3 launched, adding subagent tooling, background task management and security fixes.

Kimi Code’s CLI has over 3,100 stars on GitHub and integrates with VSCode, Cursor and Zed. The coder subagent now supports background tasks, todo lists, plan mode, skill invocation and nested agents, enabling multi‑layered autonomous workflows in software engineering.

Moonshot’s model tiering comprises three offerings: K3 as the flagship ($3/$15 per million tokens for input/output), K2.7 Code as a specialized coding model ($0.95/$4), and K2.6 as a general‑purpose option ($0.95/$4). All three support context windows of 256,000 tokens or more; K3 provides the full 1‑million‑token window. Context caching is automatic and requires no cache ID, TTL or extra parameter.

What Kimi K3 means for the future

Kimi K3’s release forces a rethink of some enterprise AI assumptions. The frontier performance gap between open‑source and proprietary models appears to have narrowed considerably. If independent evaluations confirm K3’s scores — and once the weights are public on July 27 — closed‑source providers will face increased pressure to justify premium pricing solely on capability.

The development also signals a shift in the locus of AI innovation: China’s ecosystem has produced a model competitive with top systems from firms that had direct access to advanced Nvidia hardware. Architectural innovations like hybrid linear attention suggest algorithmic efficiency may be as important as raw compute.

Finally, K3’s demonstrated agentic abilities — chip design, compressed multi‑week research, long‑horizon information seeking — point toward a future where models autonomously execute complex, multi‑day projects rather than only answering questions. For enterprises, the value proposition may move from productivity augmentation toward autonomous technical workforces.

China’s state news agency Xinhua framed the release as a national milestone, quoting Liu Tieyan, dean of the Zhongguancun Academy in Beijing, who said the wave of Chinese open‑source models has moved from isolated breakthroughs to collective advancement, offering new solutions and paths for global AI development.

Two years ago Moonshot AI was a scrappy startup; 18 months ago it was a cautionary tale. Today it is the maker of the world’s largest open‑source AI model — one that, given 48 hours and an internet connection, can design a chip to run itself. The frontier remains a race, and that field just got more crowded.