ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)
Z.ai released GLM-5.3-Flash, a 320B-parameter multimodal MoE with 18B active parameters, and Alibaba's Qwen team shipped Qwen3.8-Flash-Next, a 125B model with 6B active parameters. Both independently use a 3:1 linear-to-full attention hybrid, 2048-token context caps, four gated residual branches, and the Muon optimizer.
Tap to vote and see what everyone thinks.
Summary by ByteBrief