Shubho emphasizes the challenges of scaling complex AI architectures like wavenets and attentional models. He highlights the absence of an efficient software middle layer that can seamlessly transition experiments from a single GPU to multiple GPUs without performance loss. The conversation also touches on the iterative nature of model training, where tweaking and retraining are crucial for achieving scalable solutions.