ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)
Princeton researcher Yifan Zhang has published a technical report proposing the Recurrent Looped Transformer (RLT). The architecture carries the decoder's final hidden state and its layerwise sliding-window attention cache into the next token, across both prompt and response, with no reset at the boundary. This design aims to fix 96 blocks per token with unbounded temporal depth.
Tap to vote and see what everyone thinks.
Summary by ByteBrief