ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)
LLM predictions can replace human outcomes in A/B tests only under strict assumptions, not by design. On the Upworthy dataset, raw gpt-4o-mini predictions recovered just 39% of the human treatment effect. Machine learning calibration recovered it, but assumptions remain unverifiable for new treatments.
Tap to vote and see what everyone thinks.
Summary by ByteBrief