ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)

Tencent's WeChat AI team detailed WeLM-HD4-80B and WeLM-HD4-617B models using Hidden Decoding, which expands tokens into internal computation streams without enlarging the Transformer backbone. The 617B model activates 23 billion parameters and outperformed autoregressive baselines across nine benchmarks, with training costs at 4.4 times baseline.
Tap to vote and see what everyone thinks.
Summary by ByteBrief