ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)
BenchMIRT, a new method from Ai2, audits LLM benchmarks at the individual prompt level using multidimensional Item Response Theory. Trained on 100 LLMs across 16 benchmarks and 34K questions, it independently recovered safety and general reasoning dimensions, revealing that BBQ and WMDP align more with reasoning than safety.
Tap to vote and see what everyone thinks.
Summary by ByteBrief