Reward Model Challenges

Misha discusses the complexities of training reward models for chatbots, highlighting how a poorly designed reward system can lead to models that avoid answering questions altogether. Sonya adds that while pre-training techniques are well-established, post-training remains an evolving field, emphasizing the need for clarity in understanding the roles of both stages in AI development.