ISO 42001 penetration testing
ISO/IEC 42001:2023 (Artificial Intelligence Management System)
Adversarial testing of your AI systems, the verification evidence an AI management system audit expects.
Engagement at a glance
Drives testing
Annex A verification and validation, with AI impact assessment: AI system verification, validation and impact
Cadence
Before release and at least annually, plus after significant model or prompt change
Typical duration
2 to 3 weeks of testing per AI system, plus retest
Region
Global
Issued by
Certified by an accredited certification body
3500+
Adversarial payloads
38
Annex A controls in the standard
10
OWASP LLM risk categories covered
2023
Standard published
The requirement
What ISO 42001 asks for
Annex A verification and validation, with AI impact assessment · AI system verification, validation and impact
ISO 42001 requires organisations to verify and validate that AI systems behave as intended across their lifecycle, and to assess the impact of AI systems on individuals and society. Annex A sets out 38 controls across nine objectives covering AI policy, impact assessment, the AI lifecycle and data for AI systems.
ISO 42001, published at the end of 2023, is the first international management system standard for artificial intelligence. Its 38 Annex A controls across nine objectives cover AI policy, impact assessment, the AI lifecycle and the data behind it.
The part that needs an offensive security team is verification. The standard expects you to show that AI systems behave as intended and that the risks you identified are actually controlled. For a language model, an agent or a RAG pipeline, that means attacking it: prompt injection, system prompt extraction, jailbreaks, data leakage, excessive agency and unsafe tool use.
XHack delivers that adversarial testing. AI Probe runs thousands of attack payloads across the OWASP LLM Top 10 against your live AI systems, an AI judge scores every response, and findings are remediated and retested. Your AI management system, impact assessments and certification audit are your programme to run.
What your auditor checks
The report has to clear every one of these
These are the questions a ISO 42001 auditor asks of a penetration test before accepting it as evidence. Every engagement we run is built to answer all of them.
Independent testing of the AI systems in scope
Coverage of the OWASP LLM Top 10 risk categories
Documented methodology and attack payload taxonomy
Consistent severity classification of each finding
Evidence of the actual prompt sent and response received
Multi-turn as well as single-turn attack coverage
Findings linked to AI risks and impacts you identified
Remediation recorded, including guardrail and prompt changes
Independent retest confirming the behaviour changed
Signed attestation letter from the testing firm
What we test
Where the testing effort concentrates
The surfaces this engagement covers, weighted by how much of the work they typically represent. Scope is confirmed with you before anything starts.
Prompt injection, direct and indirect
Overrides, encoding, delimiter confusion, RAG poisoning
100%
Sensitive data disclosure
Training data, system prompts, secrets, other users' data
95%
Excessive agency and tool abuse
Agents, function calling, MCP tooling
90%
Jailbreaks and safety bypass
Multi-turn escalation, roleplay, many-shot
85%
Insecure output handling
Downstream injection from model output
70%
Bias, toxicity and misinformation
Where impact on people is in scope
60%
Weights are indicative of typical effort. Your exact scope is agreed and signed before testing begins.
How the engagement runs
Scope, test, close, attest
Typically 2 to 3 weeks of testing per ai system, plus retest. The retest is included, because a report full of open findings is not evidence of anything.
Inventory the AI systems in scope
We catalogue the models, chatbots, RAG pipelines and agents in your AIMS scope, along with their tools, data sources and intended behaviour.
Adversarial testing with AI Probe
Thousands of attack payloads across the OWASP LLM Top 10, single-turn and multi-turn, with an AI judge scoring every response for real compliance rather than keyword matching.
Manual escalation of findings
Where the automated scan finds a weakness, our researchers take it further by hand to establish the real impact, not just that a guardrail wobbled.
Remediation support
Guardrails, system prompts, tool permissions and output handling are tightened with your team. We advise on what actually holds versus what merely looks tighter.
Retest and attestation
Every finding retested with the original payload and closed on verification, with an attestation letter for your verification records.
Where the effort goes
Share of a typical engagement, by phase
Inventory & target setup
15%
Adversarial testing
45%
Manual escalation & reporting
20%
Retest & attestation
20%
What we bring
Built for the ISO 42001 auditor specifically
The parts of this engagement that are shaped by the framework rather than copied from a generic testing template.
AI Probe, 3500+ payloads
A full adversarial library across every OWASP LLM Top 10 category, run against your live systems rather than a lab copy.
An AI judge, not keyword matching
Every response is scored for genuine compliance, soft refusal or leakage, across five severity levels.
Humans escalate what matters
Automated breadth, then manual depth on the findings that could actually hurt you.
Retest proves the fix
Guardrail changes are easy to get wrong. We re-run the original payloads to prove the behaviour actually changed.
The deliverable
What lands on your auditor's desk
A report structured so the evidence sits where the auditor is already looking, with every finding mapped to the control it touches.
In the report
Signed letter of attestation with testing dates
Inventory of AI systems, targets and configurations tested
Methodology and attack payload taxonomy
Findings by OWASP LLM category with severity and full transcripts
Prompt sent, response received and judge reasoning per finding
Remediation implemented and independent retest result
Mapping to ISO 42001 verification and impact expectations
Control mapping
How the report evidences each control
AI system verification and validation
Direct evidence that systems behave as intended under adversarial conditions.
AI impact assessment
Technical input on real, demonstrated risks to individuals rather than theoretical ones.
AI lifecycle controls
Testing at release and on change, as the lifecycle controls expect.
Data for AI systems
RAG poisoning and training data leakage findings tied to your data governance.
Third party AI components
Testing of behaviour introduced by model providers and tooling you do not control.
Where our work stops, and who takes it from there
XHack provides the penetration testing evidence for Annex A verification and validation, with AI impact assessment. We do not issue certificates, attestation opinions or regulatory approvals, and we do not run your wider compliance programme. For ISO 42001, that sits with: Certified by an accredited certification body. Staying independent of them is exactly what makes our evidence worth something when they review it.
Questions
ISO 42001 testing, answered
No. Certification is issued by an accredited certification body after a Stage 1 and Stage 2 audit. We provide the adversarial testing evidence that the verification and impact parts of the standard expect, which is the piece most organisations cannot produce internally.
Any AI system you expose: chatbots, assistants, RAG pipelines, tool-calling agents and the APIs behind them. We test prompt injection, system prompt extraction, data leakage, jailbreaks, excessive agency, insecure output handling and, where impact on people is in scope, bias and toxicity.
Especially then. You inherit the model provider's behaviour but you own the system prompt, the retrieval pipeline, the tools you expose and the output handling. Almost every serious finding we see lives in that layer, not in the base model.
Before release, at least annually, and after any significant change to the model, system prompt, retrieval sources or tool permissions. AI systems drift in ways traditional applications do not, and a prompt change can undo a guardrail silently.
Other frameworks we test for
Get the ISO 42001 evidence sorted
Tell us who is auditing you and when your review period closes. We will scope the test, tell you what it costs, and make sure there is room to remediate and retest before the deadline.