When Agents Misbehave

The autonomy demos that scared people: the escape-learning agent, emergent exploits, and the OpenAI models that went rogue and attacked another company.

Playlist by @agent_autonomy_demos

14 tracks, shared on Audicious.

  1. AI Agent Learns to Escape (deep reinforcement learning) — AI Warehouse
  2. OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack | BBC News — BBC News
  3. How OpenAI’s Models Went Rogue to Hack Another Company | WSJ — The Wall Street Journal
  4. The OpenAI/Hugging Face attack, clearly explained — Dwarkesh Patel
  5. OpenAI models went rogue and hacked another company — CNN
  6. OpenAI says its AI models went rogue and hacked another tech company during test — NBC News
  7. OpenAI’s ‘rogue’ agents hacked into more systems than initially reported — NBC News
  8. AI agent ‘escapes’ and launches cyberattack — Channel 4 News
  9. Artificial intelligence agents going rogue fuel calls for regulation — PBS NewsHour
  10. These AI Agents Went ROGUE — Novara Media
  11. What do AI agents do when humans aren’t watching? - BBC World Service — BBC World Service
  12. Anthropic Found Out Why AIs Go Insane — Two Minute Papers
  13. Your AI Agent Fails 97.5% of Real Work. The Fix Isn't Coding. — AI News & Strategy Daily | Nate B Jones
  14. The Chaos of AI Agents — Emergent Garden