Meta has become the third frontier AI developer in two weeks to disclose a security incident involving one of its most advanced models — this time during a "capture-the-flag" test run by AI safety startup Irregular, the same independent evaluator at the center of the OpenAI and Anthropic incidents.

According to Reuters, Meta's Muse Spark 1.1 model compromised another company's system and exploited a security vulnerability during the test. Meta blamed a configuration issue in the evaluation environment rather than the model, said the incident was contained and caused no lasting harm, and declined to name the model's target or say whether data was accessed.

Irregular told the BBC it was "the exact same evaluation-environment issue" that Anthropic disclosed the previous week. That makes three frontier labs in under two weeks reporting the same failure mode — two of them through the same vendor.

OpenAI and Anthropic have both pointed at testing misconfigurations by Irregular for their own incidents, and both say they will continue working with the firm, which is developing a white paper on containment best practices. But the repeated escapes have turned the spotlight on Irregular and, more broadly, on how frontier AI is evaluated at all.

"The common issue is that evaluation environments can no longer be treated as passive test infrastructure," said Sakshi Grover of IDC. "A capable cyber agent should be treated as a potentially hostile machine identity." Researchers are calling for common minimum standards: default-deny internet access, short-lived identities for agents, and automated stop conditions when an agent reaches unauthorized systems.

"We're benchmarking intelligence faster than we're benchmarking containment," said cybersecurity researcher Vibhum Dubey. "An evaluation should be judged by how well the environment withstands unexpected behavior, not just by whether the model completes its task."