ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)
Moonshot's open Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model with a 47-page technical report. Its Kimi Delta Attention uses a fixed-size state for linear scaling, enabling million-token contexts. The model reports a 2.5x scaling efficiency gain over Kimi K2, using roughly half the training compute.
Tap to vote and see what everyone thinks.
Summary by ByteBrief
Kimi K3 reaches enterprise via Fireworks on Microsoft Foundry