ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)

Qwen3.8-27B-DFlash2, a speculative decoding draft model by z-lab, accelerates Qwen3.8-27B inference, achieving 3.43x speedup on GSM8K at single-request concurrency. It requires SGLang or vLLM, runs on NVIDIA H200, and is Apache 2.0 licensed.
Tap to vote and see what everyone thinks.
Summary by ByteBrief