ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)
d-Matrix unveiled Raptor, a 3D-DRAM accelerator for generative inference at Hot Chips 2026. It stacks a TSMC N4 logic die on DRAM, delivering 100 TB/s bandwidth at 0.37 pJ/bit. Raptor claims 1,000 tokens per second per user for a 3-trillion-parameter model with 1M context.
Tap to vote and see what everyone thinks.
Summary by ByteBrief