AI Can Look Safe and Still Be Dangerous | Daniel Kokotajlo

The hardest AI alignment failure to spot may be the one that looks like success. Daniel Kokotajlo explains why more capable models could behave correctly while remaining misaligned. Watch the full conversation with Daniel Kokotajlo and Thomas Larsen: https://www.youtube.com/watch?v=z5Xix4h5UlU Machine Learning Street Talk #AI #AIAlignment #Shorts