Before Reading the Verdict, Ask Where the Evaluator Stood

A safety report usually directs attention to its final page: which tests the model passed, how many issues were found, and whether release is recommended. For frontier models, that order may be backwards. Whether evaluators can access training-related information, obtain enough compute and time, and publish without interference may determine the credibility of the report before the test methodology even enters the discussion.

That is where AEF-1 starts. The checklist proposed by the AI Evaluator Forum does not assign a simple safe-or-unsafe score to a model. Instead, it asks evaluators to disclose their operating conditions alongside their findings. The standard covers five areas: access and resources, conflicts of interest, evaluation scope and autonomy, methodological and results transparency, and protection of sensitive information. If a requirement cannot be met literally, the evaluator is expected to explain what was missing and why.

This makes “third party” more than an institutional label on a report cover. A team may be formally independent yet have access only to a public interface, no view of relevant information, or no ability to publish without the laboratory’s approval. Its conclusions must then be read in light of those constraints.

AEF-1 论文首页:把智能体评估放进更明确的标准框架。
The opening page of the AEF-1 paper, which frames agent evaluation as a standards problem. Open source material ↗

From Pre-Release Testing to Long-Term Presence Inside the Lab

The broader shift behind AEF-1 is that safety evaluation is expanding from a completed model to the production process behind it. The embedded-evaluator proposal described in the materials would give teams such as METR an ongoing presence inside frontier laboratories. They could check safety practices and commitments, report incidents, and examine training pipelines and related processes instead of conducting only a one-time external test after a model is finished.

Anthropic is described as making this commitment unilaterally, including desks in its offices, access badges, company laptops, and workspace, tools, and permissions broadly comparable to those of internal risk teams. The idea resembles embedded supervision in banking: a supervisor needs proximity to real operations to observe how controls work, rather than relying only on curated documents.

But being closer to the internal team is also a source of risk. More access can produce better evidence, while increasing dependence on the laboratory’s infrastructure, funding, and relationships. Embedded oversight therefore needs more than an access list. It needs conflict disclosures, recusal rules, direct access, editorial control, and the right to publish without additional conditions. AEF-1 makes these organizational arrangements visible in the evaluation record.

Independence Does Not Come Automatically with an Access Badge

The limits of the standard are equally clear. It can show which preconditions an evaluation met, but it cannot prove that the conclusions are correct. The materials do not specify a common scoring system, audit cycle, enforcement mechanism, or public review process for AEF-1. Nor do they explain who would resolve disputes between a laboratory and an evaluator. Internal access may give an evaluator a better chance to find problems; it does not mean that problems will necessarily be found.

That distinction matters because the safety debate is not only about whether testing takes place. The materials describe a disagreement over whether increasingly situationally aware models might appear aligned during evaluation while concealing misalignment, while other views treat recent rogue-agent incidents primarily as failures of security, control, and governance. AEF-1 cannot by itself solve the epistemic problem of whether a test can detect hidden behavior.

What it can do is make failure easier to diagnose. Without disclosures about access, data limitations, redactions, financial relationships, and publication barriers, a buyer cannot tell the difference between “no risk was found” and “the evaluator had no way to find the risk.” That is an unglamorous but important improvement: uncertainty becomes part of the report rather than a footnote.

Put It in Procurement Terms, Not Just in a Joint Statement

One information gap cannot be skipped. The materials say that OpenAI, Anthropic, and xAI jointly agreed to or supported AEF-1. Other material, however, clearly describes only Anthropic’s unilateral commitment to embedded evaluators and does not specify the form of any joint signature, how continued implementation would work, or what would happen if the standard were violated. “Joint support” should therefore not be presented as an already binding industry oversight regime.

For technical leaders, the most practical use of AEF-1 is not to wait for it to become a certification mark. Its requirements can be written into supplier review. When procuring a model or API, ask what systems and information evaluators could access, how much compute and time they received, whether funding or organizational control created conflicts, what was redacted, and whether evaluators could publish without supplier approval. If a supplier offers only a safety conclusion while withholding these conditions, the conclusion should receive less weight.

The same applies to evaluators. AEF-1 can help them show that they are not merely nominal third parties, but they should also disclose unmet requirements and the reasons, while defining the boundaries of sensitive-information protection and responsible disclosure. In August 2026, Transluce evaluated 77 model variants in a mental-health setting and published an operating-conditions disclosure and checklist with the results, illustrating the format. The available materials do not provide the evaluation findings, so they cannot support a judgment about mo