ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)

Mixture-of-experts (MoE) models allow powerful LLMs to run on consumer GPUs with limited VRAM by offloading expert weights to system RAM. An RTX 3080 Ti with 12GB VRAM runs the 35B-parameter Qwen3.6-35B-A3B model at 25 tokens/second using llama.cpp's --n-cpu-moe flag.
Tap to vote and see what everyone thinks.
Summary by ByteBrief