Testing LLMs' Limits

Peter discusses his approach to evaluating large language models, particularly in the medical field. He emphasizes the importance of models being able to disagree and engage in meaningful collaboration, noting that many current models still fall short. With advancements like memory features, there's hope for future improvements, but the journey toward true AGI remains ongoing.