Tweet by destraynor

April 9, 2026

Our AI group published a novel finding from pre-training research In short: you can cut KV-cache memory in half by sharing most of the attention structure across heads + keeping small per-head differences—without hurting model quality or speed I'll explain this as best I can https://t.co/BRtsEIQZSy

Author
destraynor
Date
April 9, 2026