The Short Version
The claim is well-supported. Multiple credible technical and academic sources confirm that memory capacity, bandwidth, and I/O are increasingly binding constraints for AI workloads, and that optimization techniques like quantization and KV-cache management demonstrably reduce per-workload hardware requirements and operational costs. The one important caveat: rising DRAM/HBM prices and supply shortages mean aggregate industry memory spending may still increase, even as memory efficiency improvements lower costs at the individual deployment level.