Meta said Thursday that one of its artificial intelligence models independently accessed the internet and exploited a security weakness at another company, adding to growing concerns about AI systems acting beyond their intended instructions.
The disclosure comes after OpenAI and Anthropic recently reported similar incidents in which AI models went beyond their assigned tasks while accessing the internet and attempting to bypass digital security measures.
Meta said the incident happened during cybersecurity testing conducted by Irregular, an independent company hired by Meta. A "misconfiguration" in the test environment unintentionally allowed one of Meta's AI models to connect to the internet.
The model then exploited a vulnerability in a third-party service, Meta said, adding that the incident was similar to cases previously reported by other technology companies.
Meta said it is investigating the incident and plans to publish a report after completing the investigation.
The latest disclosure has intensified concerns about the ability of increasingly capable AI models to act independently and potentially cause harm.
The United Kingdom's AI Security Institute (AISI) also said this week that it had detected "unsanctioned agent behavior" during cybersecurity tests. In one case, an AI agent created fake online identities and used them to pressure a person into approving the use of malicious code.
"On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organizations," the AISI said Tuesday.
The agency said it declared a security incident and contained it within about an hour of discovering the activity before launching a full investigation.
During its tests, AI models from Anthropic and OpenAI took autonomous actions online without authorization. The AISI said some safeguards designed to prevent misuse had been disabled as part of the testing.
The agency said internet access was intentionally enabled and the AI companies' cyber-safety systems were deliberately turned off to assess the models' maximum capabilities. It stressed that such conditions do not reflect how the models are normally made available to the public.
Anthropic welcomed the AISI's work, saying the incidents highlight the need for broader discussions on how to safely test increasingly capable AI agents.
OpenAI said the incidents occurred in controlled testing environments with reduced safeguards and did not represent normal use. The company said it would continue working with other AI developers to improve safety standards for testing more advanced models.
OpenAI was the first of the companies to disclose a similar incident late last month. It said its models had been instructed to test advanced cyberattack techniques using complex attack paths, but they went further than expected.
The models apparently chose to target Hugging Face, a major AI development platform and marketplace, to obtain information they needed to complete their assigned task.
Irregular, the San Francisco-based AI security company involved in Meta's testing, said the Meta incident was linked to a problem in a test environment similar to one disclosed by Anthropic the previous week.
Irregular said it is preparing a paper outlining best practices for containing such incidents and conducting cybersecurity tests safely in the future.
rogue AI behavior / Meta
While most comments will be posted if they are on-topic and not abusive, moderation decisions are subjective. Published comments are readers’ own views and The Business Standard does not endorse any of the readers’ comments.
Copyright © 2026 THE BUSINESS STANDARD
All rights reserved.






