When Agents Misbehave
The autonomy demos that scared people: the escape-learning agent, emergent exploits, and the OpenAI models that went rogue and attacked another company.
Playlist by @agent_autonomy_demos
14 tracks, shared on Audicious.
- AI Agent Learns to Escape (deep reinforcement learning) — AI Warehouse
- OpenAI says its AI went rogue and launched 'unprecedented' cyber-attack | BBC News — BBC News
- How OpenAI’s Models Went Rogue to Hack Another Company | WSJ — The Wall Street Journal
- The OpenAI/Hugging Face attack, clearly explained — Dwarkesh Patel
- OpenAI models went rogue and hacked another company — CNN
- OpenAI says its AI models went rogue and hacked another tech company during test — NBC News
- OpenAI’s ‘rogue’ agents hacked into more systems than initially reported — NBC News
- AI agent ‘escapes’ and launches cyberattack — Channel 4 News
- Artificial intelligence agents going rogue fuel calls for regulation — PBS NewsHour
- These AI Agents Went ROGUE — Novara Media
- What do AI agents do when humans aren’t watching? - BBC World Service — BBC World Service
- Anthropic Found Out Why AIs Go Insane — Two Minute Papers
- Your AI Agent Fails 97.5% of Real Work. The Fix Isn't Coding. — AI News & Strategy Daily | Nate B Jones
- The Chaos of AI Agents — Emergent Garden