Moonshot AI Launches Kimi K3 For Advanced Reasoning, Coding, And Knowledge Work

Chinese artificial intelligence startup Moonshot AI has launched Kimi K3, a 2.8 trillion-parameter model designed for advanced reasoning, coding, and knowledge work. The company describes it as the world’s largest open-weight AI model.

The release places Moonshot in more direct competition with U.S. AI developers including OpenAI and Anthropic. Early evaluations indicate that Chinese companies are narrowing the performance gap with leading proprietary systems while offering developers greater control over deployment.

Open-weight models allow developers to download, operate and customize the underlying model parameters. This differs from closed systems such as ChatGPT and Claude, which are primarily accessed through hosted products and application programming interfaces.

Kimi K3 includes a one-million-token context window, allowing it to process large volumes of text, images and code within a single session. The model is designed to support longer assignments that require maintaining context, using tools and completing multiple steps with limited human supervision.

Moonshot said Kimi K3 remains behind Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol in some overall comparisons. However, the company reported stronger performance than several other leading systems across selected coding, agentic and hardware optimization tests.

Independent evaluator Arena ranked Kimi K3 first in its Frontend Code benchmark, which uses blind comparisons to assess AI-generated web interfaces. Kimi K3 received 1,679 points and finished ahead of Anthropic’s Fable 5 in that evaluation.

Vals AI placed Kimi K3 second overall behind Fable 5 and ahead of GPT-5.6 Sol. Artificial Analysis also found that its performance was comparable with OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.8 across several complex, multistep tasks.

The model uses a mixture-of-experts architecture with 896 available experts, of which 16 are activated for each token. This design is intended to provide the capacity of a very large model without using every parameter for every request.

Moonshot also introduced architectural improvements called Kimi Delta Attention and Attention Residuals. The company said these technologies improve scaling efficiency and help information move more effectively through the model during lengthy tasks.

Kimi K3’s application programming interface pricing starts at $0.30 per million input tokens when cached information is reused. Uncached input costs $3 per million tokens, while generated output is priced at $15 per million tokens.

Although Kimi K3 will be available as an open-weight model, operating it privately could remain impractical for many companies because of its size. Running the full system may require substantial computing hardware, high-speed networking and specialized technical expertise.

Moonshot plans to release Kimi K3’s complete model weights on July 27. The release will allow independent developers and researchers to evaluate the company’s performance claims more extensively.

Until the full weights become available, some technical and benchmark claims remain dependent on Moonshot’s disclosures and tests conducted through hosted access. Independent evaluation will be important in determining how reliably the model performs across real-world workloads.

The launch comes as Chinese AI developers including Moonshot, Z.ai and MiniMax accelerate the release of increasingly capable models. Their progress is challenging assumptions that China’s AI industry remains significantly behind the leading U.S. laboratories.

Moonshot is backed by major Chinese technology companies including Alibaba and Tencent. The company has also reportedly explored raising $2 billion at a valuation of approximately $30 billion ahead of a potential Hong Kong listing.

Kimi K3 could appeal to businesses and governments seeking advanced AI capabilities without becoming fully dependent on a U.S. model provider. Its long-term impact will depend on independent testing, deployment costs and whether developers can translate its technical scale into reliable performance across commercial applications.