Tackling Toxic Content
Iz and Pradeep discuss strategies to prevent models from generating toxic content, including filtering out a significant portion of toxic data during pre-training and using control tokens during inference to guide the model towards non-toxic outputs. The conversation highlights the challenges in aligning models to avoid toxic responses and the need for continued research in this area.In this clip
From this podcast

NLP Highlights
141 - Building an open source LM, with Iz Beltagy and Dirk Groeneveld
Related Questions