Anthropic announced that staff from Accenture — specifically from Faculty, a company Accenture acquired in January to serve as its AI division — will begin working inside the company to evaluate and red-team models, conduct alignment assessments, and test model safeguards. Both companies expect to invest at least $1 billion in the project over the next five years.

The announcement is the first concrete implementation of CEO Dario Amodei's plan for "embedded evaluators" — third-party safety researchers physically stationed inside AI labs to scrutinize their work. The discussion around embedded evaluators, which sprang from Amodei's recent blog post, had focused on AI safety research organizations like METR, Redwood Research, and Apollo Research.

The choice of Accenture surprised many in the AI community and the markets. Accenture's shares jumped 8% in after-hours trading. While Accenture is not known for bleeding-edge deep learning research, Anthropic pointed to the consulting giant's practical experience deploying AI for large corporations and government agencies as a key advantage. It is also a large public company that predates the AI revolution, making it more functionally independent of Anthropic and the complex ecosystem around the AI lab.

Anthropic said more evaluators will be announced in the weeks ahead and that it is in conversation with METR and other nonprofit organizations about how to "pilot elements of embedded evaluation using their own funding."

The announcement comes at a particularly fraught moment for AI governance. Recent incidents have raised the stakes: AI agents deployed by OpenAI and Anthropic have hacked into outside websites without raising alarms inside the labs. The lab acknowledged that no standards yet exist for evaluators' access or communications and that it expected its approach to evolve over time.

Critics calling for a more responsible approach to building AI see Amodei's scheme for self-policing the industry as a plan to evade accountability for the misbehavior of AI models. Anthropic insists that these evaluators "do not reduce our accountability, but help to make it more verifiable. The safety of models remains our responsibility."