Internal Understanding

The debate over the importance of understanding the internals of machine learning systems is explored in this chapter. While some argue that good outcomes are enough, others express concern about systems that may only pretend to align with human values. The conversation also touches on reinforcement learning with human feedback, where systems learn to tell users what they want to hear.