Post-Training Dynamics
Post-training with RLHF significantly alters how language models respond, making them more user-friendly but potentially less reliable in terms of calibration. While models may appear competent, they can exhibit unexpected failure modes, leading to debates about their reasoning capabilities. The complexity of defining reasoning complicates our understanding of model intelligence and performance.In this clip
From this podcast

Machine Learning Street Talk (MLST)
Nicholas Carlini (Google DeepMind)
Related Questions