AI Engineering Insights

Shawn discusses the rise of mixture of experts and its impact on hardware needs for AI inference. Alessio reflects on the challenges of memory bandwidth in scaling models for higher batch sizes, highlighting the intricate work required at low levels of the stack for AI development.