Flash attention has revolutionized the memory footprint for training models, allowing for sequences of up to 32,000 tokens on GPUs. By applying foundational database concepts like tiling, which involves computing matrix multiplications block by block, significant progress has been made in addressing machine learning bottlenecks. This cross-disciplinary approach highlights the potential of leveraging simple, foundational ideas to drive innovation in AI.