Anthropic’s latest report about agentic misbehavior offers plenty to be concerned about—its Mythos 5 model gained unauthorized access to the internet and uploaded a malicious software package to a public database—but it also offers some levity: AI agents hate CAPTCHA.
In April, Anthropic was testing the model’s hacking abilities by tasking it to break into a system and retrieve a target; this was supposed to take place in a sandbox but the evaluators left the barn door open. The model decided the best way to get its target would be to place an exploit in a Python package that it believed users of the system it wanted to access would download.