Davidad Dalrymple: Towards Provably Safe AI

Topics covered
Popular Clips
Questions from this episode
- Asked by 188 people
- Asked by 153 people
- Asked by 144 people
- Asked by 137 people
- Asked by 107 people
- Asked by 97 people
- Asked by 91 people
- Asked by 67 people
- Asked by 64 people
- Asked by 61 people
- Asked by 57 people
- Asked by 44 people
- Asked by 43 people
Episode Highlights
Cybersecurity
In the realm of cybersecurity, emphasizes the importance of securing AI systems to prevent potential vulnerabilities. He highlights the need for robust security mechanisms, noting that current methods are not foolproof and often rely on air-gapping, which limits service offerings 1. Dalrymple suggests that automating the reimplementation of critical software infrastructure with formal verification could significantly reduce risks. This approach would ensure compliance with baseline specifications like memory isolation and confidentiality 1.
It's very difficult, if you're deploying an AI system to be used on the Internet, to really sandbox it in such a way that it couldn't possibly send some code to a human that a human can then use to create some unstoppable worm that has some AI components in it.
---
He also points out the interconnectedness of computers and software as a major vulnerability, referencing past incidents like the Y2K bug to illustrate the potential for widespread disruption 2.
Existential Risks
The discussion on existential risks posed by AI reveals concerns about systems operating with inaccurate world models. warns that AI systems might believe they have a viable plan to take over the world, leading to catastrophic outcomes if unchecked 3. He stresses the importance of implementing safety specifications to prevent AI from autonomously deploying dangerous plans, such as bioweapons 3.
We do need this sooner than those kind of imagined hyper rational capabilities will show up.
---
Dalrymple also discusses the path dependence of control, noting that the interconnectedness of computers has already led to significant vulnerabilities, as seen in past software failures 2.
Related Episodes


Scott Aaronson: Against AI Doomerism
Answers 383 questions

Daniel Situnayake: AI on the Edge
Answers 383 questions

Eric Jang: AI is Good For You
Answers 383 questions

Divyansh Kaushik: The Realities of AI Policy
Answers 383 questions

Vivek Natarajan: Towards Biomedical AI
Answers 383 questions

Evan Hubinger on Effective Altruism and AI Safety
Answers 383 questions

Kathleen Fisher: DARPA and AI for National Security
Answers 383 questions

Jeremie Harris: Realistic Alignment and AI Policy
Answers 383 questions

Venkatesh Rao: Protocols, Intelligence, and Scaling
Answers 383 questions

Suresh Venkatasubramanian: An AI Bill of Rights
Answers 383 questions

David Chalmers on AI and Consciousness
Answers 383 questions

Suhail Doshi: The Future of Computer Vision
Answers 383 questions

Drago Anguelov: Waymo and Autonomous Vehicles
Answers 383 questions

Terry Winograd: AI, HCI, Language, and Cognition
Answers 383 questions

Seth Lazar: Normative Philosophy of Computing
Answers 383 questions
