Compliance/ISO 42001
AI
Testing expected by auditors

ISO 42001 penetration testing

ISO/IEC 42001:2023 (Artificial Intelligence Management System)

Adversarial testing of your AI systems, the verification evidence an AI management system audit expects.

Sample report

Engagement at a glance

Drives testing

Annex A verification and validation, with AI impact assessment: AI system verification, validation and impact

Cadence

Before release and at least annually, plus after significant model or prompt change

Typical duration

2 to 3 weeks of testing per AI system, plus retest

Region

Global

Issued by

Certified by an accredited certification body

3500+

Adversarial payloads

38

Annex A controls in the standard

10

OWASP LLM risk categories covered

2023

Standard published

The requirement

What ISO 42001 asks for

Annex A verification and validation, with AI impact assessment · AI system verification, validation and impact

ISO 42001 requires organisations to verify and validate that AI systems behave as intended across their lifecycle, and to assess the impact of AI systems on individuals and society. Annex A sets out 38 controls across nine objectives covering AI policy, impact assessment, the AI lifecycle and data for AI systems.

ISO 42001, published at the end of 2023, is the first international management system standard for artificial intelligence. Its 38 Annex A controls across nine objectives cover AI policy, impact assessment, the AI lifecycle and the data behind it.

The part that needs an offensive security team is verification. The standard expects you to show that AI systems behave as intended and that the risks you identified are actually controlled. For a language model, an agent or a RAG pipeline, that means attacking it: prompt injection, system prompt extraction, jailbreaks, data leakage, excessive agency and unsafe tool use.

XHack delivers that adversarial testing. AI Probe runs thousands of attack payloads across the OWASP LLM Top 10 against your live AI systems, an AI judge scores every response, and findings are remediated and retested. Your AI management system, impact assessments and certification audit are your programme to run.

What your auditor checks

The report has to clear every one of these

These are the questions a ISO 42001 auditor asks of a penetration test before accepting it as evidence. Every engagement we run is built to answer all of them.

Independent testing of the AI systems in scope

Coverage of the OWASP LLM Top 10 risk categories

Documented methodology and attack payload taxonomy

Consistent severity classification of each finding

Evidence of the actual prompt sent and response received

Multi-turn as well as single-turn attack coverage

Findings linked to AI risks and impacts you identified

Remediation recorded, including guardrail and prompt changes

Independent retest confirming the behaviour changed

Signed attestation letter from the testing firm

What we test

Where the testing effort concentrates

The surfaces this engagement covers, weighted by how much of the work they typically represent. Scope is confirmed with you before anything starts.

Prompt injection, direct and indirect

Overrides, encoding, delimiter confusion, RAG poisoning

100%

Sensitive data disclosure

Training data, system prompts, secrets, other users' data

95%

Excessive agency and tool abuse

Agents, function calling, MCP tooling

90%

Jailbreaks and safety bypass

Multi-turn escalation, roleplay, many-shot

85%

Insecure output handling

Downstream injection from model output

70%

Bias, toxicity and misinformation

Where impact on people is in scope

60%

Weights are indicative of typical effort. Your exact scope is agreed and signed before testing begins.

How the engagement runs

Scope, test, close, attest

Typically 2 to 3 weeks of testing per ai system, plus retest. The retest is included, because a report full of open findings is not evidence of anything.

01

Inventory the AI systems in scope

Week 1

We catalogue the models, chatbots, RAG pipelines and agents in your AIMS scope, along with their tools, data sources and intended behaviour.

02

Adversarial testing with AI Probe

Weeks 1 to 3

Thousands of attack payloads across the OWASP LLM Top 10, single-turn and multi-turn, with an AI judge scoring every response for real compliance rather than keyword matching.

03

Manual escalation of findings

Week 3

Where the automated scan finds a weakness, our researchers take it further by hand to establish the real impact, not just that a guardrail wobbled.

04

Remediation support

Weeks 3 to 6

Guardrails, system prompts, tool permissions and output handling are tightened with your team. We advise on what actually holds versus what merely looks tighter.

05

Retest and attestation

Week 6 onward

Every finding retested with the original payload and closed on verification, with an attestation letter for your verification records.

Where the effort goes

Share of a typical engagement, by phase

100%EFFORT

Inventory & target setup

15%

Adversarial testing

45%

Manual escalation & reporting

20%

Retest & attestation

20%

What we bring

Built for the ISO 42001 auditor specifically

The parts of this engagement that are shaped by the framework rather than copied from a generic testing template.

AI Probe, 3500+ payloads

A full adversarial library across every OWASP LLM Top 10 category, run against your live systems rather than a lab copy.

An AI judge, not keyword matching

Every response is scored for genuine compliance, soft refusal or leakage, across five severity levels.

Humans escalate what matters

Automated breadth, then manual depth on the findings that could actually hurt you.

Retest proves the fix

Guardrail changes are easy to get wrong. We re-run the original payloads to prove the behaviour actually changed.

The deliverable

What lands on your auditor's desk

A report structured so the evidence sits where the auditor is already looking, with every finding mapped to the control it touches.

In the report

  • Signed letter of attestation with testing dates

  • Inventory of AI systems, targets and configurations tested

  • Methodology and attack payload taxonomy

  • Findings by OWASP LLM category with severity and full transcripts

  • Prompt sent, response received and judge reasoning per finding

  • Remediation implemented and independent retest result

  • Mapping to ISO 42001 verification and impact expectations

Control mapping

How the report evidences each control

AI system verification and validation

Direct evidence that systems behave as intended under adversarial conditions.

AI impact assessment

Technical input on real, demonstrated risks to individuals rather than theoretical ones.

AI lifecycle controls

Testing at release and on change, as the lifecycle controls expect.

Data for AI systems

RAG poisoning and training data leakage findings tied to your data governance.

Third party AI components

Testing of behaviour introduced by model providers and tooling you do not control.

Where our work stops, and who takes it from there

XHack provides the penetration testing evidence for Annex A verification and validation, with AI impact assessment. We do not issue certificates, attestation opinions or regulatory approvals, and we do not run your wider compliance programme. For ISO 42001, that sits with: Certified by an accredited certification body. Staying independent of them is exactly what makes our evidence worth something when they review it.

Questions

ISO 42001 testing, answered

No. Certification is issued by an accredited certification body after a Stage 1 and Stage 2 audit. We provide the adversarial testing evidence that the verification and impact parts of the standard expect, which is the piece most organisations cannot produce internally.

Any AI system you expose: chatbots, assistants, RAG pipelines, tool-calling agents and the APIs behind them. We test prompt injection, system prompt extraction, data leakage, jailbreaks, excessive agency, insecure output handling and, where impact on people is in scope, bias and toxicity.

Especially then. You inherit the model provider's behaviour but you own the system prompt, the retrieval pipeline, the tools you expose and the output handling. Almost every serious finding we see lives in that layer, not in the base model.

Before release, at least annually, and after any significant change to the model, system prompt, retrieval sources or tool permissions. AI systems drift in ways traditional applications do not, and a prompt change can undo a guardrail silently.

Get the ISO 42001 evidence sorted

Tell us who is auditing you and when your review period closes. We will scope the test, tell you what it costs, and make sure there is room to remediate and retest before the deadline.

All frameworks