Understanding Deception Circuits
The discussion delves into the complexities of how models, particularly larger ones, develop circuits that can identify and replicate behaviors like deception. Sholto highlights the intricate interplay of different model heads, which can either reinforce or suppress certain behaviors. The conversation raises critical questions about the reliability of labels used to define deceptive outputs and the potential for misunderstanding these complex behaviors.In this clip
From this podcast

Dwarkesh Podcast
Sholto Douglas & Trenton Bricken - How to Build & Understand GPT-7's Mind
Related Questions
What are induction heads in relation to large language models (LLMs) as discussed in the episode Sholto Douglas & Trenton Bricken - How to Build & Understand GPT-7's Mind and the clip Understanding Deception Circuits
What are induction heads in relation to large language models (LLMs) as discussed in the episode Sholto Douglas & Trenton Bricken - How to Build & Understand GPT-7's Mind and the clip Reasoning Circuits Explored?
What are induction heads in relation to large language models (LLMs)?