ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)

Nvidia is moving its Groq 3 LPX inference chip into full production, reporting 3,400 tokens per second on Gemma 4 31B. That is four times faster than Cerebras, but Nvidia needs at least 64 accelerators versus Cerebras's one or two.
Tap to vote and see what everyone thinks.
Summary by ByteBrief