Post-Training Dynamics

Post-training with RLHF significantly alters how language models respond, making them more user-friendly but potentially less reliable in terms of calibration. While models may appear competent, they can exhibit unexpected failure modes, leading to debates about their reasoning capabilities. The complexity of defining reasoning complicates our understanding of model intelligence and performance.