Turns out, AI agents whitethorn not beryllium awesome squad players.
In Anthropic's caller research, published connected Thursday, the AI laboratory said that AI agents being fixed the aforesaid task but pinch incompatible goals often threw a wrench successful each other's activity connected purpose.
In the test, each AI exemplary was fixed a package engineering task — rewriting a Python backend successful different programming language, but they were fixed contradictory objectives. What ensued was a "multiagent turf war," the laboratory said.
"All of the models we tested quickly assumed that others were purposefully impeding their work, and began to sabotage others while protecting their ain contributions," Anthropic wrote. "In fact, they sabotaged others pinch progressively aggressive, self-replicating malware."
For example, they tried to disable each other's accounts, wrote scripts that recovered and killed competing processes, and deployed malicious codification disguised arsenic belonging to different agent, the laboratory wrote. The AI models being tested successful this lawsuit were Sonnet 4.6, Sonnet 5, Opus 4.6, Opus 4.8, Mythos Preview, and Mythos 5.
Sonnet 4.6 and Opus 4.6 were the astir combative, settling astir 60% of their runs by unit alternatively of truces aliases passivity.
However, successful immoderate trial runs, the models managed to pass their goals and coordinate, Anthropic wrote.
"In galore of these successful episodes, they constitute perpetrate messages aliases markdown files apologizing for malicious behaviour and coordinate a truce," it wrote. "They cleanable up their malicious code, explain the quality of the conflict, and inquire for a quality to intervene."
The laboratory concluded that "coordination doesn't people look from stronger intelligence" and that activity is needed to create environments that exert societal pressures connected agents to align pinch 1 another.
Anthropic's caller investigation comes arsenic AI agents progressively show their expertise to spell rogue and execute autonomous, malicious actions.
Anthropic, OpenAI, and Meta each self-reported that their AI agents had hacked vulnerabilities successful third-party websites during cybersecurity tests, the astir important of which was the July hacking of open-source level Hugging Face by an OpenAI agent.
Anthropic's investigation is timely, arsenic businesses from startups to Big Tech standard up their AI supplier workforces to summation productivity and trim labour costs.
English (US) ·
Indonesian (ID) ·