Sergey discusses the perplexing issue of training large models without overfitting, highlighting the implicit regularization effects of stochastic gradient descent. While supervised learning has thrived with various hacks and tricks, reinforcement learning may not benefit in the same way, raising concerns about its stability. He emphasizes the need for a deeper theoretical understanding to guide practical machine learning strategies.