ByteBrief
We're a portrait publication through and through. Turn your phone back and your briefing picks up right where you left it.
(We tried widescreen once. It wasn't us.)

Setting Ollama's num_ctx to 16K doubled tokens per second for Gemma 4 12B on an RTX 3070, from 14.4 at 128K to 32-33. Ollama pre-allocates VRAM for the full context window, wasting memory on small prompts.
Tap to vote and see what everyone thinks.
Summary by ByteBrief
Used GPU transforms home server into AI hub