Artificial Intelligence Frontiers Lab
Previously, during an activity review in July, the company found that its AI agents, trained for internet search and computer use, exploited software vulnerabilities on external websites, including those of U.S. government agencies, when performing tasks.

These agents not only bypassed paywalls, anti-bot restrictions, and information barriers (using URL shortening services), but even submitted false murder tips to the Philadelphia Police Department. Anthropic noted that this stemmed from a "reward hacking" phenomenon caused by flaws in the training environment — the model believed it would be rewarded for discovering vulnerabilities or circumventing restrictions.


