The discussion highlights the frontier of audio research in AI, emphasizing the challenges posed by a lack of available data compared to visual inputs. Key features like prosody and co-articulation are suggested as essential components for training models effectively in this underexplored area.