The most efficient thinker in the Kimi family. Coding and agent capabilities closing in on flagship K3 — with adjustable thinking effort and a 1M-token context window on every membership tier, rolled out with zero configuration.
Announced in the official Kimi Code changelog on September 11, 2026. Every claim below traces to that release note — no speculation, no rumor.
Officially described as "performance close to K3" — Moonshot's flagship. Coding and agent capabilities are improved across the board over K2.7 Code, with gains concentrated in long-horizon, multi-step engineering work.
Significantly more efficient thinking than K2.7 Code — itself a model that cut reasoning-token usage 30% versus K2.6. Less overthinking, more doing: fewer tokens burned per solved task.
The same three thinking levels as K3 — low / high / max — with max as the default. Turn thinking off entirely and requests are still served, now by K2.8 in instant mode.
A context window of up to one million tokens — available across all membership tiers. The flagship K3 reserves its 1M window for higher tiers; K2.8 Preview democratizes it.
The model ID is unchanged — still kimi-for-coding. Every client, IDE plugin, CI pipeline, and third-party tool keeps working exactly as before, now on a stronger model.
With thinking switched off, requests to the K3 series and K2.8 Preview are both served by K2.8 Preview (no thinking). One model now powers both fast answers and deep deliberation across the lineup.
K2.8 Preview introduces K3-grade thinking control to the mainstream lineup. Pick a level to see how the trade-off shifts.
The out-of-the-box setting. K2.8 reasons at full depth before acting — the mode behind its near-K3 results on long-horizon coding and agentic tasks.
Whole monorepos, marathon sessions, and massive document sets — in a single uninterrupted conversation. On K3, 1M requires a higher membership tier. On K2.8 Preview, it's everywhere.
* K3 unlocks 1M only on higher tiers (Allegretto and above); K2.8 Preview ships 1M on every tier. ≈ figures are order-of-magnitude approximations for scale, not official measurements.
Three generations of Kimi coding models, side by side — from the K2.7 Code specialization to the 2.8-trillion-parameter K3 flagship.
| Dimension | K2.7 Code | K2.8 Preview | K3 (flagship) |
|---|---|---|---|
| RELEASE | June 2026 · GA | Sep 11, 2026 · rolling out | 2026 · open-sourced |
| POSITIONING | Specialist coding model | Near-K3 coding for everyone | Flagship · 2.8T params |
| THINKING | Thinking on / off | low / high / max (default max) | low / high / max |
| CONTEXT | Long-context | 1M — all tiers | 1M · higher tiers only |
| ARCHITECTURE | — | Not disclosed for preview | KDA hybrid linear attention + Attention Residuals |
| VISION INPUT | — | — | Native visual understanding |
| ACCESS | Kimi for Coding | All Kimi Code users, zero config | Moderato membership+ |
The open-weight generation: topped SWE-Bench Pro at 58.6% in its launch comparison and introduced the Thinking / Instant split that would define the family.
+10.4% Program-Bench, +11.4% MCP Mark Verified, +76.2% SWE Marathon over K2.6 — while reasoning-token usage dropped 30%. The efficiency story starts here.
Fully rolled out in Kimi Code. Near-K3 coding performance, K3-grade thinking controls, 1M context on every tier — under the unchanged model ID kimi-for-coding.
The 2.8-trillion-parameter ceiling: KDA hybrid linear attention, native visual understanding, and benchmark wins over Opus 4.8 and GPT-5.5. K2.8 Preview closes most of the gap at every tier.
The preview label is honest. As of September 15, 2026, these details remain undisclosed by official sources — treat anything claiming them with suspicion.
Moonshot has not published K2.8 Preview's parameter count or architectural details. The only official architecture disclosure in the family is K3's: 2.8 trillion parameters on KDA hybrid linear attention with Attention Residuals. Any site claiming exact K2.8 specs is extrapolating.
The changelog describes K2.8's gains qualitatively — "performance close to K3," "improved across the board," "significantly more efficient thinking" — without SWE-Bench, Program-Bench, or MCP Mark figures. Hard numbers are expected with the stable release or a technical report.
Because the model ID didn't change, existing kimi-for-coding billing carries over — but Moonshot hasn't published K2.8-specific pricing or quota announcements. Current plan limits apply as before.
Preview implies refinement in flight: behavior may shift as Moonshot tunes the model toward a stable K2.8. The changelog notes no breaking changes — same ID, same interfaces — but pinning expectations to the changelog rather than anecdotal benchmarks is the safe move.
The official changelog covers Kimi Code. Secondary reports (Pandaily, Sept 14, 2026) say the same model also reached Kimi Work users — marked as third-party coverage here, not confirmed by the changelog we sourced.