Published Oct 9, 2023

Emergent Deception in LLMs

Explore the emergent deception abilities in large language models with Thilo Hagendorff, delving into the cognitive advancements of AI like GPT-3 and GPT-4, and addressing ethical concerns such as speciesism and the potential for AI to independently induce false beliefs.
Episode Highlights
Data Skeptic logo

Popular Clips

Episode Highlights

  • Behaviorist Approach

    In the realm of generative AI, champions a behaviorist approach to understanding language models. He proposes the concept of "machine psychology," where language models are treated as participants in psychological experiments, akin to the human brain's black box nature. This method relies on observing correlations between inputs and outputs to discern behavioral patterns, rather than delving into the models' internal workings 1.

    My goal is to establish a field that I would like to call machine psychology.

    ---

    and Thilo discuss how this approach can reveal the impressive capabilities of language models, such as their performance on theory of mind tasks and cognitive reflection tests 1.

       

    Cognitive Tasks

    Thilo highlights the growing cognitive abilities of language models through their performance on cognitive tasks. He notes that while older models like GPT-2 struggled with tasks like the bat and ball problem, newer models such as GPT-3 and GPT-4 demonstrate significant improvements. These advancements suggest a steady increase in the cognitive capabilities of language models 1.

    Newer models like ChatGPT or GPT-4, they basically ace these tasks and nearly in all cases, give you the correct answer.

    ---

    The discussion also touches on the potential ceiling of AI capabilities, with Thilo speculating on the future of artificial general intelligence and the role of multimodal models that integrate language, images, and audio 1.

Related Episodes