Tweet by destraynor
April 9, 2026
Our AI group published a novel finding from pre-training research In short: you can cut KV-cache memory in half by sharing most of the attention structure across heads + keeping small per-head differences—without hurting model quality or speed I'll explain this as best I can https://t.co/BRtsEIQZSy
- Author
- destraynor
- Date
- April 9, 2026
- Canonical URL
- /tweets/destraynor-2042305356705984935-925f44