ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)

AMD is acquiring Canadian startup Taalas, which hard-codes model weights directly into inference chips, locking each chip to a single model. A demo chip hit over 16,000 tokens per second per user running Llama 3.1-8B. Google is reportedly working on a similar approach for Gemini.
Tap to vote and see what everyone thinks.
Summary by ByteBrief