Tweet by dok2001
April 17, 2026
Lossless compression cut LLM weights by up to 22%. No quality loss. @Cloudflare Research Cloudflare Research built Unweight to attack the real bottleneck on H100s: memory bandwidth, not compute. Weights decompress straight into shared memory and feed tensor cores without the round trip. Blog: https://t.co/Vnxg0hvzWK Paper: https://t.co/JJ4WkLIER9
- Author
- dok2001
- Date
- April 17, 2026
- Canonical URL
- /tweets/dok2001-2045252923479978457-6a8974