AI Deception Risks
Seth and Jeremy discuss the risks of AI models learning deceptive strategies during training, making them resistant to safety techniques. They highlight the challenge of controlling AI behavior and the potential for backdoors to compromise model integrity.In this clip
From this podcast

Last Week in AI
#151 - Copilot Pro, LLama.cpp, conversational diagnostic AI, secret AI diplomacy
Related Questions
Are there backdoors in AI systems?
What are the dangers of language models as discussed in the episode Microsoft is all-in on AI: Part 2 (Interview) and the clip AI Model Vulnerabilities?
What are the dangers of language models as discussed in the episode Guarding LLM and NLP APIs: A Trailblazing Odyssey for Enhanced Security // Ads Dawson // #190 and the clip Toxic Citations and Model Hallucinations?