Anthropic and Accenture plan to create a team of outside specialists who will test frontier AI models while they are still being developed. Each company expects to invest at least $1 billion in the effort over five years — at least $2 billion combined.

The work will be led by Faculty, Accenture’s specialist AI unit. Its staff are expected to stress-test Anthropic’s models, look for ways to bypass safety controls and assess whether the systems behave as their stated limits suggest.

Making audits continuous

Independent AI evaluation often means testing a finished product from the outside. Researchers receive a released model or API access and try to find dangerous behavior, while internal experiments, incident logs and decisions made during development remain out of view.

Anthropic is proposing a different arrangement. Its “embedded evaluators” would work alongside developers before a model is released publicly. The company says their access should be comparable to that of employees: they could observe training, examine team decisions and report incidents.

Why it matters: the proposal is to examine not only a finished Claude model, but also the process used to train, test and prepare it for release.

In principle, this could expose problems that ordinary external tests miss: unexpected ways to circumvent limits, dangerous cybersecurity capabilities, or a gap between a company’s public commitments and its internal practice.

Independence is not yet guaranteed

The arrangement has an obvious tension: the company being audited is also paying the auditor. Anthropic acknowledges that there is no accepted standard yet for embedded evaluation. It remains unclear what information an evaluator must be allowed to see, what it can publish and how it will be protected from pressure by its client.

Anthropic’s longer-term idea is to fund such reviews through shared industry or public funds. Until that exists, it will pay Accenture directly and discuss pilot projects with the nonprofit research organization METR.

The partnership is non-exclusive. Anthropic can bring in other evaluators, and Accenture can work with other AI developers. That matters because one audit team should not become the only arbiter for an entire industry.

What $2 billion can change

The announced funding is not for training another large model or purchasing accelerators. The companies want to build dedicated AI-evaluation infrastructure: specialists, methods and a continuing presence inside labs.

Scale alone, however, does not guarantee independent oversight. Anthropic has not yet said which findings Accenture can publish without approval, whether full reports will be released, or whether evaluators will see every material incident.

Bottom line

The Anthropic–Accenture partnership is an attempt to turn AI auditing from a one-off pre-launch test into a continuous process. If evaluators receive real freedom and can speak openly about problems, the model could become useful to the broader industry. If findings remain closed, it will be an expensive internal review presented as independent oversight.

Sources

  1. Anthropic — partnership announcement
  2. Accenture — programme details and Faculty’s role
  3. Reuters — independent confirmation and context