ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)

GLM-4.7-Flash hits 88.5 tokens per second at 4K context on a single RTX 3090 but drops to 16.8 at 128K. The model beats Qwen3.5:27b at short context but loses at long context. Two RTX 3090s in tensor parallel reach 498 tokens per second batched.
Tap to vote and see what everyone thinks.
Summary by ByteBrief