Tweet by dok2001

April 17, 2026

Lossless compression cut LLM weights by up to 22%. No quality loss. @Cloudflare Research Cloudflare Research built Unweight to attack the real bottleneck on H100s: memory bandwidth, not compute. Weights decompress straight into shared memory and feed tensor cores without the round trip. Blog: https://t.co/Vnxg0hvzWK Paper: https://t.co/JJ4WkLIER9

Author
dok2001
Date
April 17, 2026