ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)
HarnessDev, a framework from ByteDance Seed and collaborators, evaluates the runnable harness code an LLM writes, not the answer it produces. On Terminal-Bench 2.1, GPT-5 solves 35.2% of tasks in Terminus 2 but 49.6% in Codex CLI with identical weights, showing harness design is critical.
Tap to vote and see what everyone thinks.
Summary by ByteBrief