Moonshot Ai

Moonshot AI unveils Kimi K3

Beijing startup releases a 2.8T, 1M‑token model and promises open weights later in July

Beijing startup releases a 2.8T, 1M‑token model and promises open weights later in July

Beijing‑based Moonshot AI announced Kimi K3 on July 16, 2026, describing a 2.8‑trillion‑parameter, multimodal model with a 1‑million‑token context window and a pledge to publish full weights later in July.

Moonshot says the model is built on two new architectural ideas — Kimi Delta Attention and Attention Residuals — and scales as a Mixture‑of‑Experts (MoE) system that selectively activates a small slice of experts per token. Those design choices are intended to make million‑token contexts practical while keeping inference costs competitive.

At launch Kimi K3 is available through Moonshot’s Kimi services and API, and the company set a date — July 27, 2026 — for the public release of the model weights so developers can download and run the system themselves. Moonshot also published pricing tiers for API usage and said it will ship a technical report alongside the weights.

Moonshot’s own benchmark tables show K3 matching or approaching the performance of leading closed‑source frontier models on many agentic coding and reasoning tasks, and the company highlighted examples such as kernel optimization, compiler construction and multi‑file engineering sessions. Those vendor results are central to the company’s launch narrative.

Independent trackers and early community tests posted within 24 hours also put K3 at or near the top on several coding and web‑development benchmarks, with Arena.ai and some composite indexes ranking it above or alongside models such as Anthropic’s Fable and OpenAI’s GPT‑5.x in specific tests. Analysts and platform maintainers described the initial results as striking but partial.

Not all third‑party counters have finalized composite rankings, and some researchers cautioned that many launch screenshots cherry‑pick tasks or different harnesses, making apples‑to‑apples comparisons difficult until the weights and full evaluation harnesses are public. Community replication and lower‑level checks are expected to accelerate after the promised weights drop.

Moonshot’s open‑weights pledge has already changed conversations about developer access and procurement. Industry observers say easier access to large, customizable models could shift routine engineering, summarization and internal agent workloads toward cheaper self‑hosted or hybrid deployments. That tectonic shift is what many in Silicon Valley call the open‑weight insurgency.

The K3 announcement also lands in a charged geopolitical moment. U.S. export controls and debate over AI supply chains mean China’s domestic pipeline for chips and inference infrastructure is a closely watched variable, and Moonshot’s ties to large local cloud and hardware partners were noted in coverage of the model’s debut. The timing coincided with the World AI Conference in Shanghai.

Practical self‑hosting remains nontrivial: Moonshot recommends supernode deployments with dozens of accelerators for full‑scale inference and says it is contributing implementation patches to major inference projects to ease large‑context serving. Community teams will likely first produce quantized, smaller footprints before many organizations attempt full on‑premises runs.

The U.S. industry has reacted in several ways — some investors and startups framed K3 as proof that open models can erode the premium tier, while others said the release will spur more investments in inference optimizations, tooling and proprietary safety controls. Policy watchers also flagged that broader access changes the enforcement and audit calculus for sensitive capabilities.

What to watch next: the July 27 weights release, independent benchmarkers’ full replication reports, community‑built quantizations and the first months of real‑world deployments. Those milestones will determine whether K3’s early promise holds across diverse workloads beyond the curated launch examples.

Kimi K3’s debut is a clear signal that open‑weights labs are pushing harder into frontier territory, compressing the lead once held by closed, expensive systems. For now, Moonshot’s claims and the early community scores merit attention — and careful verification once the promised weights and technical report are public.