Tweet by citrini
June 16, 2026
AI workloads are creating new memory objects that are too large for HBM but too valuable for cold storage. It’s simply to expensive to use DRAM for everything if it’s possible to use NAND for some of it. We’ve gotten three announcements focused on flash from three major players in the past couple weeks: Nvidia is using flash to store reusable KV cache via CMX. This targets long-context, multi-turn and agentic inference by extending GPU memory with a shared KV-cache tier, allowing flash to become context memory for long-running agents. AMD is using flash to store some system memory until it is predicted to be hot enough to go to DRAM. They’re buying MEXT as a software layer to accomplish this. Apple is using flash to store some model weights until needed. Apple stores inactive experts in flash on the device side (AFM 3 Core Advanced, published a week ago). If you can get the system to know what data is likely to be needed next, flash goes from simple storage to cheaper memory.
- Author
- citrini
- Date
- June 16, 2026
- Canonical URL
- /tweets/citrini-2066712653297299797-2ae29b