
Table of contents
16
By Salman Khan, OSCP+, Founder of XHack, SRT (Synack Red Team member)
Read this in 30 seconds: Nobody has banned AI-run penetration tests. But every framework that actually demands a test still expects a named human to sign the report, and the cost of guessing wrong lands on you, not your vendor.
- In July 2026 CREST became the first body to formally accredit AI-enabled pentesting. Its rule is simple: AI assists, a qualified human validates, and AI output never automatically becomes a finding. Ten firms were accredited by September.
- The PCI Council told a QSA in writing that “being able to operate a tool is not the same as being a qualified penetration tester.” That reply was private, not published, and it was about scan tools rather than AI agents. It is still the closest thing to a ruling anyone has.
- FedRAMP names a person. Its guidance requires a “Penetration Test Team Lead” holding a recognised credential. Search its documents for the word “AI” and you get nothing.
- ISO 27001 and SOC 2 never mention penetration testing at all. Neither has said a word about AI. That silence is real latitude, and also real risk.
- Every AI pentest vendor with a compliance product keeps a human in that path. Cobalt says outright that its autonomous product “does not produce compliance attestation reports.”
If you are buying AI pentesting and something in your life has the word “audit” attached to it, this is the question that actually matters. Not whether the agent finds bugs, because AI agents are genuinely good at that now. Whether the report survives contact with the person reviewing it.
I run XHack, so I sell both an AI agent and human-led testing. Weigh that as you like. What follows is what the standards actually say, quoted, including the parts that are inconvenient for anyone selling automation.
The short version is that AI penetration testing compliance is not settled law. It is a gap that everyone is quietly working around, and one accreditation body has just published the first real answer.

For frameworks that actually require a penetration test, you need a named, qualified human to stand behind the report. Today, in 2026, that is not optional in practice.
For frameworks that do not require a test at all, and that is more of them than most people realise, you have real room to move.
The reason is not that anyone banned AI. It is that nobody has approved it either, and an auditor who is unsure defaults to asking who did the work.
For years this question had no answer anywhere. Then CREST, the international accreditation body for pentest firms, published one.
On 28 July 2026 it launched an accreditation specifically for AI-enabled penetration testing. Two parts: a new “Responsible AI Use” domain covering how a provider governs AI, and an Annex B to the existing penetration testing standard.
By 3 September, ten firms had been accredited.
The substance is what matters, and it is refreshingly blunt. AI should support expertise rather than bypass it. Qualified testers validate what the AI produces. And the line that should be on a poster somewhere: AI output should not automatically become a penetration testing finding.
CREST also published numbers that explain why it bothered. 69% of penetration testing providers already use AI. 76% increased their use in the past twelve months. 85% expect clients to demand transparency about it.
So the industry is already there. The paperwork is catching up.
One quote from the announcement is worth keeping, from Chris Oakley at LRQA: regulators and auditors “quickly move from ‘is AI used?’ to ‘how is AI governed?'”
That is the whole shift in one sentence. Nobody serious is asking whether you used AI. They are asking who checked it.
Here is where each framework actually stands on AI penetration testing compliance. The important thing is that “forbidden”, “silent”, and “allowed but auditors still want a human” are three different situations, and they get mashed together constantly.
| Framework | AI-run test alone? | With a human? | Where it really sits |
|---|---|---|---|
| PCI DSS 4.0 | No | Yes, and this is the working model | Council told a QSA privately that tools cannot replace manual skilled testing |
| FedRAMP | No | Yes, as a tool used by the 3PAO team | Guidance names a credentialed “Penetration Test Team Lead” |
| DORA / TIBER-EU | No | Presumably, but no regulator has said so | Requires “proven threat-intelligence and red-team expertise” |
| ISO 27001 | Not forbidden | Recommended by convention | The standard never says “penetration test” at all |
| SOC 2 | Not forbidden | Auditor’s discretion | The Trust Services Criteria never mention pentesting |
| NIS2 | Not forbidden | No guidance either way | Directive does not name testing or testers |
| HIPAA | No test required yet | Not applicable yet | A proposed rule would add annual testing, still pending |
Two patterns jump out.
The strict frameworks got strict by naming people and credentials, not by mentioning AI. FedRAMP and DORA both predate this question entirely. Search their documents for “AI” and you find nothing, because nobody was thinking about it when they were written.
And the relaxed frameworks are relaxed because they never specified anything. ISO 27001 and SOC 2 do not require a penetration test in their own text. Auditors ask for one out of convention, which means the convention is what you are actually negotiating with.
In October 2025 a working QSA named Jeffrey B. Hall got tired of the ambiguity and asked the PCI Council directly. Some of his colleagues were seeing automated tool output submitted as pentest reports and disagreeing about whether they could refuse it.
He published the Council’s written reply in full. The relevant lines:
“Automated tools can support penetration testing but do not replace the need for manual, skilled testing as specified in PCI DSS.”
And the one that has been quoted ever since:
“Being able to operate a tool is not the same as being a qualified penetration tester.”
The Council also made a second point that gets less attention and probably matters more. Your methodology has to be yours, documented and available for the QSA to read. Pointing at a vendor’s website and saying the tool follows their methodology does not count.
Now the caveats, because they are real.
That reply was a private email, reported second-hand on a blog. It is not an FAQ. It is not an updated Information Supplement. Eleven months later the Council still has not published anything formal.
It was also about GUI scanning tools, with Core Impact named as the example. It was not about an autonomous agent that plans and chains its own attacks under a named human’s supervision. Nobody has ruled on that configuration, in either direction.
And the guidance everyone keeps pointing at, PCI’s own Penetration Testing Guidance supplement, is version 1.1 from September 2017. It was written against PCI DSS v3.2 and has never been updated for 4.0, let alone for AI.
Its entire section on tester qualifications talks about years of experience and certifications held. In 2017 it simply never occurred to anyone that the tester might not be a person.
So the honest position is this. The strongest signal we have points firmly at requiring a human. It also has no formal standing, and it was answering a narrower question than the one you are asking.


Frameworks differ wildly on how much they specify. FedRAMP is the most prescriptive by a distance, so it is the useful reference even if you are not chasing it.
Its guidance requires the report to cover scope, the attack vectors assessed, the testing timeline, the actual tests performed and their results, findings with evidence, and access paths showing how vulnerabilities chained together.
It also requires something no other framework spells out: named personnel, including a Penetration Test Team Lead holding an industry-recognised credential.
That is the detail worth internalising. FedRAMP never says “a human must do this.” It just requires a named person with a credential to lead it, which amounts to the same thing.
Across every framework reviewed, the content list converges on the same seven things:
An autonomous agent can genuinely produce most of that list. It struggles with exactly one item, and it happens to be the one auditors care most about: someone accountable.

This is the part I find most telling, because it is the industry voting with its product design rather than its marketing.
Every AI pentesting vendor with a real compliance offering keeps a human in that specific path.
Cobalt is the most direct about it. Its own comparison page says its autonomous product “does not produce compliance attestation reports,” and points customers needing audit-ready proof at its human-led platform instead. Its own engineers wrote a piece making the same point: AI can draft a finding, it cannot decide what matters in a particular environment.
Horizon3 sells NodeZero as an autonomous platform, then sells a separate compliance service staffed by OSCP-certified humans, with NodeZero verifying fixes afterwards. I compared the two products in detail in Horizon3 vs XHack.
Synack puts a human triage team in front of every finding.
Astra, which sells autonomous testing, says in its own marketing that for PCI you should not substitute autonomous testing for human-led qualified testing, because “the human still signs the report. The auditor still asks who they are.”
Read those together. Four companies that sell AI pentesting, all building a human step into the compliance path, all saying so publicly.
On AI penetration testing compliance, that is either an industry reading the regulatory tea leaves correctly, or an industry hedging against a question nobody has answered. Both readings fit the evidence. Neither one helps you if you buy AI-only testing and your auditor says no.
Here is something that surprised me while researching this.
There is no documented, named, checkable case of an auditor accepting or rejecting an AI-run penetration test. Not one, anywhere public.
Plenty of opinion. Plenty of vendor positioning. No actual decisions on record.
That matters because the market has not set a precedent yet. Every buyer is making this call privately with their own auditor, and nobody is publishing the outcome.
If your QSA accepts an AI-led report next quarter, you will be one of the first people who knows. If you are still choosing a supplier, the questions that expose a weak provider are worth asking before this one comes up.
So this is the part where I talk about what we built, and I have tried to earn it above.
Our agent runs a seven-stage pipeline. Prompt, planner, orchestrator, executor, findings, then human triage and verification, then the report. Every stage writes to an audit log, and every finding traces back to the exact tool invocation that produced it and the reviewer who signed it off.
We built it that way before CREST published Annex B. It happens to match the model CREST landed on, for the same reason: the finding and the accountability are different jobs.
That structure exists because of everything above. The agent is fast and genuinely good at finding things. What no agent can do is be the name on the report when an auditor asks who did the work.
If you need something an auditor will accept, our researchers run the engagement with the agent alongside them.
$2,500 covers one web application or up to 50 host IPs, delivered in four days with a retest. $5,000 covers up to three applications across web, host, API and mobile. $12,000 covers larger internal and external programmes.
The scope is agreed before anyone starts, and the prices are published, which is less common in this market than it should be.
The people doing that work hold OSCP and OSCP+ certifications, which is the kind of thing the PCI guidance actually asks about.
And if you just want continuous coverage between audits, run the agent yourself from $20 a month and bring humans in when the auditor asks who is signing. That blend is what human plus agentic testing actually means in practice, and what it costs is published too.
I should also be straight that we are small. Six clients secured and 32 assessments done. If your procurement team needs a vendor with a decade of references, that is not us yet.
Possibly, because SOC 2 does not require a penetration test at all. The Trust Services Criteria never use the phrase, and the AICPA has published nothing about AI-run testing. In practice auditors ask for a pentest as evidence that your control evaluation is real, and what they will accept varies by auditor and by whether you are doing a Type I or Type II report. Ask yours directly and early, because there is no rule to appeal to if they say no. AI penetration testing is new enough that many auditors are forming a view for the first time when you ask.
Not on its own. PCI requires a “qualified internal resource or qualified external third party” with organisational independence, and the Council told a QSA in writing that operating a tool does not make someone a qualified tester. An AI agent working under a named, qualified human who reviews and signs the report is the model every PCI-facing vendor has built. The genuinely unresolved part is whether a sophisticated autonomous agent with human sign-off would satisfy the requirement in a way a simple scan tool would not, and nobody has ruled on that.
Yes, as of July 2026. CREST added a Responsible AI Use domain to its company requirements and an Annex B to its penetration testing standard, and accredited its first ten firms by September 2026. The model is that AI assists, qualified testers validate its output, and AI output does not automatically become a finding. It is a voluntary professional accreditation rather than a regulatory requirement, but it is the only formal answer that exists.
It can produce most of the content. Scope, methodology, tests performed, findings with evidence, severity, even chained access paths are all things a good agent handles. What it cannot supply is the last item on every auditor’s list, which is a named, qualified person who is accountable for the work. FedRAMP makes this explicit by requiring a credentialed Penetration Test Team Lead. Other frameworks get to the same place through convention.
Then you have genuine latitude, and ISO 27001, SOC 2 and NIS2 all fall into this category. A well-documented programme of continuous autonomous testing plus periodic human testing is defensible, particularly for a lower-risk environment with strong patch management. The expectation scales with your risk profile. If you run significant internet-facing systems handling sensitive data, expect your certification body to want independent qualified testing regardless of what the standard’s text says.
No case is documented publicly, in either direction. That is the honest state of things. The absence cuts both ways: there is no precedent saying you will be refused, and none saying you will be accepted. Until auditors start publishing decisions or a body updates its guidance, every buyer is negotiating this privately.
AI penetration testing compliance comes down to a gap between what the standards say and what auditors actually do.
No standards body has ruled that an AI agent cannot run your test. The strict frameworks got strict by naming credentialed people rather than by mentioning AI at all, and the relaxed ones never specified anything in the first place.
But every framework that genuinely requires a test expects someone accountable to put their name on it, CREST has now formalised that expectation for AI specifically, and every vendor selling into this market has built a human step into their compliance path.
So use the agent for what it is good at, which is finding things quickly and continuously and cheaply. Put a qualified human in front of the report, because that is what the person reviewing it is going to ask about.
And ask your auditor before you buy, not after. The one thing this research made unambiguous is that nobody else can answer this for you yet.
Categories
Related articles