S
pending
Kimi K3: Moonshot's 2.8T Open Model That Designed Its Own Chip in 48 Hours
Grounded / Real
Inflated / Uruttu
Original Content
Kimi K3 just released
🔹 2.8 Trillion Parameters, 1 Million Context, Natively Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
This a Fable 5, GPT 5.6-sol level model
Some interesting underrated gems in Kimi-K3 release:
> an early K3 wrote the majority of the kernels in the late development stages
> it built a triton-class compiler from scratch, MiniTriton, that delivers performance on par with or better than Triton and torch.compile
> then it designed a chip, by a model, for a model, in one 48-hour autonomous run
On many tasks this model performs equivalent to Fable,Opus 4.8 and GPT-5.6 at half or even lesser cost
Excellent at front end code generation.
The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens.
This is a crazy release from Moonshot.ai
🔹 2.8 Trillion Parameters, 1 Million Context, Natively Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
This a Fable 5, GPT 5.6-sol level model
Some interesting underrated gems in Kimi-K3 release:
> an early K3 wrote the majority of the kernels in the late development stages
> it built a triton-class compiler from scratch, MiniTriton, that delivers performance on par with or better than Triton and torch.compile
> then it designed a chip, by a model, for a model, in one 48-hour autonomous run
On many tasks this model performs equivalent to Fable,Opus 4.8 and GPT-5.6 at half or even lesser cost
Excellent at front end code generation.
The first-party API is priced at $3.00/$15.00 per 1M input/output tokens, with cached input discounted 90% to $0.30 per 1M tokens.
This is a crazy release from Moonshot.ai
Validated Content
All benchmark and technical claims match Moonshot's own announcement exactly, word-for-word in several cases. This is a highly accurate post.
Confirmed — every specific claim matches Moonshot's official announcement
2.8 Trillion Parameters, 1M Context, Natively Multimodal — confirmed exactly. Kimi K3 launched July 16, 2026 as a 2.8T MoE model (104B active params) with a 1,048,576-token context window and native text/image/video understanding.
Kimi Delta Attention → up to 6.3x faster decoding in million-token contexts — confirmed word-for-word from Moonshot's own launch post.
Attention Residuals → ~25% higher training efficiency at <2% additional cost — confirmed word-for-word from Moonshot's own launch post.
"Built for long-horizon agentic coding and self-evolving workflows" — matches Moonshot's own framing exactly.
Live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API; open weights by July 27, 2026 — confirmed. Weights were released July 26-27, 2026, matching Moonshot's stated timeline.
"Fable 5, GPT 5.6-Sol level model" — reasonably accurate characterization. Independent coverage (VentureBeat) describes Fable 5 and GPT-5.6 Sol as K3's closest competitors, with K3 leading on some benchmarks and trailing on most others — "level" is a fair, if slightly generous, summary.
"An early K3 wrote the majority of the kernels in late development stages" — confirmed. Moonshot's own blog states an early version "handled the majority of the team's kernel optimization works."
MiniTriton, a from-scratch Triton-class compiler matching or beating Triton/torch.compile — confirmed exactly, directly from Moonshot's tech blog and independent coverage.
Designed a chip, by a model, for a model, in one 48-hour autonomous run — confirmed. K3 autonomously designed, optimized, verified, and simulated a functional chip (4mm², Nangate 45nm, ~8,721 tokens/sec) in a single 48-hour run.
Excellent at frontend code generation — confirmed; K3 ranks #1 of 99 models on WebDev Arena, the first open model ever to top that leaderboard.
Pricing: $3.00/$15.00 per 1M input/output tokens, cached input at $0.30 (90% discount) — confirmed exactly.
Minor caveat worth noting
"Performs equivalent to Fable, Opus 4.8, and GPT-5.6 at half or lesser cost, on many tasks" — this is a reasonable summary but slightly favorable framing. Independent coverage (VentureBeat) notes K3 led in 4 of 8 real-world task benchmarks but finished second to Fable 5 in most others, and Moonshot's own launch post reportedly acknowledges it still trails the strongest proprietary systems overall. "Equivalent on many tasks" is defensible, but not universally true.
All of the flashy autonomous-work claims (chip design, MiniTriton, kernel optimization) come from Moonshot's own launch materials — independent, third-party verification of these specific demos hadn't happened at time of writing, per multiple outlets (e.g., Yahoo Finance/Investing.com explicitly flag this).
Bottom line: This is an unusually accurate summary — nearly every number and claim is a direct, verbatim match to Moonshot's own announcement and independently reported coverage. The only softness is that the flashiest claims (autonomous chip design, MiniTriton) are self-reported by Moonshot and not yet independently reproduced, and the "equivalent to Fable/Opus/GPT-5.6" framing is on the generous end of how it's actually being described elsewhere.
Confirmed — every specific claim matches Moonshot's official announcement
2.8 Trillion Parameters, 1M Context, Natively Multimodal — confirmed exactly. Kimi K3 launched July 16, 2026 as a 2.8T MoE model (104B active params) with a 1,048,576-token context window and native text/image/video understanding.
Kimi Delta Attention → up to 6.3x faster decoding in million-token contexts — confirmed word-for-word from Moonshot's own launch post.
Attention Residuals → ~25% higher training efficiency at <2% additional cost — confirmed word-for-word from Moonshot's own launch post.
"Built for long-horizon agentic coding and self-evolving workflows" — matches Moonshot's own framing exactly.
Live on Kimi.com, Kimi Work, Kimi Code, and the Kimi API; open weights by July 27, 2026 — confirmed. Weights were released July 26-27, 2026, matching Moonshot's stated timeline.
"Fable 5, GPT 5.6-Sol level model" — reasonably accurate characterization. Independent coverage (VentureBeat) describes Fable 5 and GPT-5.6 Sol as K3's closest competitors, with K3 leading on some benchmarks and trailing on most others — "level" is a fair, if slightly generous, summary.
"An early K3 wrote the majority of the kernels in late development stages" — confirmed. Moonshot's own blog states an early version "handled the majority of the team's kernel optimization works."
MiniTriton, a from-scratch Triton-class compiler matching or beating Triton/torch.compile — confirmed exactly, directly from Moonshot's tech blog and independent coverage.
Designed a chip, by a model, for a model, in one 48-hour autonomous run — confirmed. K3 autonomously designed, optimized, verified, and simulated a functional chip (4mm², Nangate 45nm, ~8,721 tokens/sec) in a single 48-hour run.
Excellent at frontend code generation — confirmed; K3 ranks #1 of 99 models on WebDev Arena, the first open model ever to top that leaderboard.
Pricing: $3.00/$15.00 per 1M input/output tokens, cached input at $0.30 (90% discount) — confirmed exactly.
Minor caveat worth noting
"Performs equivalent to Fable, Opus 4.8, and GPT-5.6 at half or lesser cost, on many tasks" — this is a reasonable summary but slightly favorable framing. Independent coverage (VentureBeat) notes K3 led in 4 of 8 real-world task benchmarks but finished second to Fable 5 in most others, and Moonshot's own launch post reportedly acknowledges it still trails the strongest proprietary systems overall. "Equivalent on many tasks" is defensible, but not universally true.
All of the flashy autonomous-work claims (chip design, MiniTriton, kernel optimization) come from Moonshot's own launch materials — independent, third-party verification of these specific demos hadn't happened at time of writing, per multiple outlets (e.g., Yahoo Finance/Investing.com explicitly flag this).
Bottom line: This is an unusually accurate summary — nearly every number and claim is a direct, verbatim match to Moonshot's own announcement and independently reported coverage. The only softness is that the flashiest claims (autonomous chip design, MiniTriton) are self-reported by Moonshot and not yet independently reproduced, and the "equivalent to Fable/Opus/GPT-5.6" framing is on the generous end of how it's actually being described elsewhere.