Nathan discusses the need for a new evaluation tool to assess models more effectively. He highlights the efficiency of reward models over generative language models in scoring text, pointing towards a potential shift in evaluation methods that could revolutionize the field.