Key takeaways
- On 18 September 2026, Anthropic announced a new arrangement with Accenture for independent evaluators to work inside the AI developer, with access intended to be comparable to an employee’s.
- The development points towards a stronger model of AI assurance. Organisations increasingly need evidence about how important AI systems are tested and governed, rather than relying only on provider claims, certifications or pre-deployment reviews.
- New Zealand professional services firms do not need embedded evaluators of their own, but they should strengthen supplier assurance, independent testing and ongoing review for AI that handles sensitive information or materially affects professional work.
AI assurance is starting to look more like financial audit and cybersecurity testing than traditional software procurement.
On 18 September 2026, Anthropic announced that Accenture, through its specialist AI business Faculty, will conduct independent evaluation of its frontier AI models. The work is expected to include model evaluation, red teaming, alignment assessments and testing of safeguards.
The unusual part is where that evaluation will happen. Rather than assessing a finished product solely from the outside, evaluators are intended to work inside the AI company with access comparable to employees, allowing them to observe models as they are developed and examine the decisions and safeguards surrounding their deployment.
For New Zealand firms, the immediate lesson is not that every AI provider needs an embedded auditor. It is that independent evidence is becoming a more important part of demonstrating that AI controls actually work.
Moving beyond provider assurances
Most organisations adopting generative AI depend heavily on information supplied by the vendor. Procurement reviews may examine security documentation, contractual commitments, privacy terms, certifications and published descriptions of model safeguards. Those remain important, but they have limitations. A provider can describe its governance framework without independently demonstrating how effectively the framework operates in practice.
Anthropic’s approach attempts to narrow that gap. Its stated intention is that evaluators will be able to observe model development, assess whether safety commitments are being followed, identify blind spots and report incidents. Anthropic also makes an important distinction. Independent evaluation does not transfer responsibility away from the AI provider. It provides another mechanism for making the provider’s claims and controls more verifiable.
That principle is equally relevant to organisations buying AI.
Assurance should match the risk
Not every use of AI warrants an independent assessment. An employee using an approved AI tool to brainstorm headings for an internal presentation presents a very different risk from an AI system analysing privileged legal documents, reviewing financial information, recommending client actions or connecting directly to an organisation’s email and document repositories. The higher the potential consequence, the stronger the evidence an organisation should expect before relying on the system.
For professional services firms, this suggests a tiered approach. Lower-risk uses may be adequately managed through approved tools, policies, training and normal technology controls. Higher-risk uses may justify deeper supplier due diligence, documented testing, independent review and ongoing monitoring.
This is consistent with New Zealand’s Responsible AI Guidance for Businesses, which takes a proportionate, risk-based approach and emphasises accountability and systematic risk management throughout the AI lifecycle.
Testing the system you actually use
Independent model testing also raises another important issue. Testing a foundation model is not necessarily the same as testing the AI system deployed inside a firm.
A professional services organisation may combine a commercial model with its own prompts, document repositories, client information, retrieval systems, workflows, permissions and automated actions. Each additional component can alter the risk. An AI system may therefore perform well in the provider’s general evaluations but fail in the specific environment in which the firm uses it. For example, an internal AI assistant might retrieve information accurately but have excessive permissions that allow it to expose documents from another client matter. An agent may produce reliable summaries but be given authority to send messages or alter records without sufficient human approval.
The practical assurance question is therefore broader than “Is the model safe?” It is also “Is our implementation of this model safe for this particular purpose?”
Red teaming is becoming more relevant
Anthropic says the new arrangement will include red teaming, where testers deliberately challenge systems to find weaknesses or behaviours that ordinary testing may miss. That approach has long been used in cybersecurity. Its application to AI can include attempts to bypass safeguards, extract information, manipulate instructions, trigger inappropriate actions or identify situations where a model behaves unexpectedly.
For New Zealand professional firms, formal AI red teaming will not be necessary for every system. But the underlying idea is useful. Before deploying a higher-risk AI workflow, someone should deliberately test how it fails. That might include incorrect source material, ambiguous instructions, conflicting documents, attempts to access restricted information, prompt injection, fabricated citations or requests outside the system’s intended purpose.
Testing only whether AI works under normal conditions can leave important risks undiscovered.
Independence needs to mean something
The Anthropic announcement also highlights an unresolved issue in AI assurance: what counts as independent? Anthropic will initially fund Accenture’s evaluation work directly. Anthropic acknowledges that there are not yet established standards governing what embedded evaluators should be able to access, how findings should be reported or how independent evaluation should be funded. That is an important qualification.
Independent assurance should not be treated as a label. Organisations relying on an assessment should understand who commissioned it, who paid for it, what was tested, what access the evaluator received, whether limitations were imposed and whether significant findings are disclosed.
A familiar audit principle applies: the value of assurance depends partly on the independence, scope and evidence behind it.
Supplier assurance is becoming an AI governance function
Most New Zealand professional services firms will consume AI rather than develop frontier models. Their most important control may therefore be supplier assurance. Traditional technology due diligence often focuses on security, privacy, availability, data location and contractual protections. AI introduces additional questions.
Firms may need to understand how models are evaluated, how harmful or inaccurate behaviour is identified, how material model changes are communicated, how incidents are handled and whether independent testing is undertaken. This does not mean requesting confidential technical information from every provider. It means matching the depth of the review to the importance of the system and the information or decisions entrusted to it.
For high-risk uses, a vendor questionnaire alone may no longer be sufficient evidence.
What this means for your organisation
Classify AI systems by risk. Identify which systems handle sensitive client information, contribute to professional advice, influence important decisions or can take actions in other systems. Apply stronger assurance requirements to these uses.
Ask suppliers for evidence, not just statements. For higher-risk tools, understand what testing has been completed, whether independent evaluation occurs, what the scope was and how significant findings are addressed.
Test your implementation. Do not assume a provider’s model evaluation covers your configuration. Test retrieval, permissions, integrations, prompts, data boundaries and automated actions in the environment in which the AI will actually operate.
Test failure as well as success. Include adversarial and unusual scenarios. Check what happens when the system receives incorrect information, conflicting instructions, malicious content or requests for information outside a user’s permissions.
Record assurance decisions. Keep evidence of what was reviewed, what testing occurred, what limitations were identified and who accepted any residual risk.
Reassess after material changes. A new model, integration, data source or agent capability may alter the original risk assessment. Define which changes trigger another review.
From trust to evidence
The Anthropic and Accenture arrangement is an early model and important details remain unresolved. It should not be treated as an established assurance standard. Its significance is the direction of travel. As AI becomes more capable and more deeply integrated into business systems, organisations will increasingly need ways to verify that controls work rather than simply accepting that they exist.
For New Zealand professional services firms, that does not require creating an elaborate new assurance regime. It means applying a familiar risk principle to AI: the more important the system, the stronger the evidence you should require before trusting it with client information, professional work or the ability to act.
Sources
- Anthropic, Partnering with Accenture on embedded evaluation, 18 September 2026
- New Zealand Ministry of Business, Innovation & Employment, Responsible AI Guidance for Businesses
- NIST, AI Risk Management Framework and AI Resource Center

