Learning Safeguards

Prem discusses the challenges of online learning, particularly the risks posed by adversarial attacks that can skew model training. He emphasizes the importance of developing classifiers to detect offensive content and explores innovative approaches like "unlikelihood learning" to teach systems what to avoid. By leveraging crowdsourced data, the aim is to enhance the robustness of machine learning systems against potential exploits.