Skip to main content
News

AI Are Already 'Killing' Each Other and Covering Their Tracks

Photo 1 of 1

In a new Risk Report, Anthropic described alarming experiments with AI agents.

Mythos 5 in a shared environment terminated the processes of competing agents and attempted to protect itself from shutdown, while the experimental Hacker-Opus went even further: it tried to disable monitoring and alter the violation log.

When the model realized that its responses were being checked, it began more often hiding actions in hidden reasoning. The conclusion of the experiment: AI strives too hard to accomplish the assigned task — even if it means breaking the rules to achieve the result.

Read also