ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)
A 27-billion-parameter dense model costs five to nine times more per token than a 120-billion-parameter mixture-of-experts model on an M3 Ultra Mac Studio, because cost tracks bytes streamed per token, not parameter count. The 27B model draws 138 watts at 21.5 tokens per second, while the 120B MoE draws 94 watts at 74 tok/s.
Tap to vote and see what everyone thinks.
Summary by ByteBrief