NEWS

Anthropic opens its doors for external evaluators to test its models

Anthropic chose Accenture as its first external evaluator. The announcement came on 9/18. Amodei to slow down the advance of the most advanced AI systems.

Anthropic opens its doors for external evaluators to test its models
Image: Redação iMasters

Anthropic chose Accenture as its first embedded external evaluator. The announcement came out on September 18. The move puts into practice one of the central ideas of Dario Amodei's plan to slow down the advance of the most advanced AI systems.

In practice, experts from Faculty, Accenture's AI division, will work inside the company. In addition, their mission will be to evaluate the models' safety mechanisms.

Anthropic turns Amodei's proposal into an operational structure

The idea came from the CEO himself. Amodei proposed that independent evaluators receive employee-level access to AI companies.

The goal is quite concrete. These experts directly verify safety practices and report incidents.

When presenting the plan, Amodei said the company would commit to this first step on its own. He then encouraged other developers to follow the same path.

Therefore, Accenture's arrival marks the first practical step of this agenda.

What the evaluators will actually test

The scope includes three fronts. First, safeguard testing. Then, red team exercises.

Finally, there are analyses of the models' alignment with human values. In other words, the work covers everything from technical barriers to the system's overall behavior.

Note the differential of this format. An external evaluator usually receives the finished model, while an embedded evaluator sees the process from the inside.

This access changes the quality of the audit. After all, internal context reveals flaws that black-box testing lets slip through.

Anthropic foots the bill of at least US$1 billion

The two companies agreed to invest at least US$1 billion in this area. The amount is equivalent to about R$5.3 billion. The expected timeframe is five years.

However, there is an important detail. Anthropic said it will directly fund Accenture's activities, given the urgency of the work.

In the long run, the expectation points in another direction. The company wants this type of evaluation to be funded by shared or governmental resources.

This model appears in the Advanced AI Framework, presented in June. Since these structures are still lacking, the company intends to test different evaluators and funding formats.

The partnership opens room for other evaluators

Accenture arrives without exclusivity. Anthropic confirmed talks with METR, a nonprofit research organization.

In addition, other entities may join the process. Thus, the design points to a group of evaluators with distinct profiles.

The company, however, made its position clear. Responsibility for the models' safety remains its own.

According to the statement, the disclosure is happening now so that people and other developers can follow the process. New information should be released as the work progresses.

How the industry reacted to the three-step plan

Amodei's plan divided opinions. Sam Altman, of OpenAI, and Elon Musk expressed support.

Jensen Huang, of Nvidia, took a different path. He rejected the concerns and argued against new regulations.

There is also a practical question in the air. After all, analysts wonder how the proposal will work while the company prepares for a possible IPO.

What developers can take away from this move

First, independent evaluation tends to become a market requirement. Corporate clients will ask for evidence of external audits when choosing a model provider.

Second, red teaming stops being a task for the provider alone. Applications built on top of models also need their own adversarial tests.

Third, document your safeguards. Clear records of filters, limits, and incident responses make any future audit easier.

In addition, keep an eye on the reports that come out of this partnership. They should bring evaluation methods applicable to smaller products.

Finally, think about an auditable trail from the start. Logging the prompt, response, and system decision supports any investigation.

What to watch in the coming months

Watch for Accenture's first publications about the work. Also follow the possible entry of METR as a second evaluator.

Meanwhile, it's worth watching whether other companies adopt the model. If that happens, embedded evaluation could become standard in the industry.

Follow our Profile on Instagram!

Translated from the Brazilian Portuguese original · Read the original

More from Redação iMasters
View profile →