ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)
Researchers present Dust, the first zeroth-order method competitive with backprop for pretraining transformer language models. Dust perturbs activations independently at every token, allowing one forward pass to evaluate a virtual population in parallel. From 1M tokens up, Dust is on the order of 10^3 to 10^4 times more efficient than a transformer implementation of EGGROLL, a state-of-the-art evolution strategy method.
Tap to vote and see what everyone thinks.
Summary by ByteBrief