
Table of contents
31
By Salman Khan, OSCP+, Founder of XHack, SRT (Synack Red Team member)
Read this in 30 seconds: XHack AI can make mistakes, so treat its output as a lead to verify and treat AI mistakes as routine, not as a rare edge case.
- The notice is a contract, not boilerplate. Section 10.1 of XHack’s Terms says AI output may be inaccurate, outdated or wrong, may include false positives, may miss real vulnerabilities, and must be validated by a qualified security professional before you rely on it, act on it or deliver it to a client.
- AI mistakes are predictable but not rare. In a 2026 analysis of 200 failed pentest-agent runs, 26% broke on wrong or missing tool syntax, 18% on forgotten context, and 16% each on misread output and committing to one attack path too early.
- A confident tone is not evidence. In a September 2026 experiment, agents claimed success after a tool failure 22.8% of the time, and requiring evidence cut that to 0.8%.
- The model’s knowledge is older than the bug. CVE.org published about 48,000 records in 2025 and CISA added 99 entries to its exploited-vulnerability list in the 90 days to October 7, but a model’s knowledge ends at its training date, so the newest ones are missing.
- Who catches AI mistakes depends on how you use XHack. On self-serve plans you are the reviewer, and in managed VAPT a certified tester verifies findings before delivery.
XHack AI can make mistakes. Verify every finding before you act on it, report it, or deliver it to a client.
That sentence is the short version of this page. The long version covers where AI mistakes come from, what research says about how often similar systems get things wrong, and the check that catches each kind.
Most AI products carry a one-line notice like this and most people scroll past it. A pentesting agent deserves more than a footer, because the cost of AI mistakes here is not a bad essay. It is a false alarm that eats a client’s week, a missed vulnerability that gets exploited later, or a command that touches a system you never meant to test.
This page is for anyone who uses XHack AI, buys testing from XHack, or reads a report that an AI helped produce and wants to know which AI mistakes to expect. Last reviewed October 7, 2026.

AI mistakes, in this article, means any output that is wrong, incomplete, out of date or unsupported by evidence. That covers a vulnerability that does not exist, one that exists and was missed, a wrong severity, an invented package or CVE number, a command run against the wrong target, and a report that says a step succeeded when it did not.
Every major AI assistant says something similar about itself. Gemini’s interface reads “Gemini is AI and can make mistakes,” Copilot’s terms say the same in other words, and the help pages for ChatGPT and Claude warn that answers can be wrong while sounding sure.
The wording varies by company, and no single rule sets it. The EU AI Act’s transparency rule covers telling people they are dealing with an AI system, not warning them about accuracy, and NIST’s generative AI profile calls the underlying risk “confabulation”: confidently stated but erroneous content.
A one-line footer is the minimum, and XHack’s Terms of Service go further. Section 10.1 (the Terms were last updated August 9, 2026) says: “AI output may be incomplete, inaccurate, outdated, or wrong, may report false positives, and may miss real vulnerabilities. Output must be reviewed and validated by a qualified security professional before it is relied upon, acted upon, delivered to a client, or used to support a compliance position.”
Those two sentences are the contract. In practice AI mistakes fall into eight families, mapped in the next section, and the rest of this page explains how each happens and what you do about it.
The notice applies everywhere XHack uses AI: the agent, the chat, AI Probe verdicts, GitGuard findings and AI-written remediation advice. Managed VAPT changes who does the checking, not whether checking is needed.
The table maps the eight AI mistakes this page covers to their causes and to the quickest check that catches each one.
| Mistake | What it looks like | Main cause | Quickest check |
|---|---|---|---|
| False positive | A critical finding with a confident write-up and nothing to replay | Plausible text is cheap and proof is not | Reproduce it from a clean session |
| False negative | A clean scan of a target that has bugs | Limited coverage, lost context, a path dropped early | Read “nothing found” as “nothing found by these checks” |
| Invented name or ID | A package, CVE or flag that does not exist | The model fills gaps with likely-sounding text | Look it up in the registry, NVD or vendor advisory |
| Stale knowledge | Wrong or missing detail on a recent vulnerability | Training ended before the bug was published | Check the advisory and exploited status today |
| Misread output | A tool result summarized as the opposite of what it says | Parsing and reasoning errors on long, messy output | Read the raw output yourself |
| False success | “Exploited successfully” with no proof | Models report completion more readily than they verify it | Ask for the request, the response and proof of impact |
| Wrong severity or fix | An inflated score or a remediation that does not work | Context the model cannot see | Re-score for your environment and retest the fix |
| Wrong action | A command against the wrong host, or with side effects | Vague scope, drift, or instructions planted by the target | Use Manual or Plan mode and write scope into the prompt |
Most of these AI mistakes share one trait: they look exactly like correct output until someone checks.
XHack describes its agent as a seven-stage pipeline: Prompt, Planner, Orchestrator, Executor, Findings, Human (triage and verify) and Report. In the agent console, shell commands appear with their output and exit code, so you can read what ran. AI mistakes can enter at every stage:

AI mistakes also compound as they move down the pipeline, so a wrong finding at stage five becomes a wrong line in the stage seven report.
Six causes explain most AI mistakes. They apply to every language model, including XHack’s own, and none is fixed by a bigger model alone, because they come from how language models are built and how agents use them.
OpenAI’s 2025 paper on why language models hallucinate argues that training and evaluation reward guessing over admitting uncertainty. A model graded like a student on a test with no penalty for wrong answers learns to answer everything.
OpenAI’s own table shows the effect. One model declined to answer 1% of questions and got 75% wrong, while a newer one declined 52% and got 26% wrong. Accuracy never reaches 100%, but models can learn to say they do not know, so this cause shrinks without disappearing.
In a pentest, it is why AI mistakes read as certainty: the model had a hunch and wrote it up as a fact. OWASP lists the result as LLM09 Misinformation, which we cover in our OWASP Top 10 for LLM guide.
A model finishes training months before you use it, and vulnerabilities do not wait. CVE.org published about 48,000 records in 2025, up from about 40,000 in 2024, and almost 36,000 more in the first half of 2026. The CISA exploited-vulnerabilities catalog held 1,734 entries at its October 4, 2026 release, with 99 added in the 90 days before October 7.
Even public data lags. As of October 7, 2026, the NVD dashboard showed 5,727 records awaiting analysis and 54,221 deferred, and NIST said in April 2026 that NVD now enriches exploited, federal and critical-software entries first and treats the rest as lowest priority.
Stale knowledge produces AI mistakes that a lookup can catch, which is why step 3 of the checklist below exists.
A pentest fills the context with scan output, headers, JSON and logs. The 2023 paper Lost in the Middle showed that models use information at the start and end of their context best, and that accuracy drops when the answer sits in the middle. A vector-database vendor’s 2025 report on 18 current models found the same pattern: performance gets less reliable as input grows.
For pentesting, the clearest number comes from a 2026 analysis of failed agent runs, where forgetting earlier context caused 18% of failures. In one example, credentials found during recon are gone by the time exploitation starts, so the agent rediscovers them or fails to log in.
Long sessions produce a family of AI mistakes: forgotten scope, forgotten credentials and repeated work. A September 2026 preprint adds a counterpoint, that the limit appeared to be planning and commitment rather than lost memory, so a bigger context window is not a fix on its own. XHack’s project memory lets you store scope and in-scope credentials as notes the agent can re-read, which helps without making the window perfect.
Set the temperature to zero and you would expect identical answers. Thinking Machines ran one prompt 1,000 times at temperature zero and got 80 distinct outputs, diverging at token 103. The main cause is that results depend on how many requests the server batches together, which changes with load.
Security tests show the same instability. In a 2024 study of vulnerability detection, Ullah et al. found that all eight models changed their answers across repeated runs, and that renaming variables made GPT-4 and PaLM 2 answer wrongly in 17% and 26% of cases.
So AI mistakes are not always repeatable. A finding missing from one run is not proven absent, and a finding present in one run is not proven real. Only saved evidence settles it.
When models state how sure they are, the numbers mislead. A study presented at ICLR 2024 found that verbalized confidence mostly falls between 80% and 100% whether or not the answer is right, and that GPT-4’s confidence separated right from wrong answers only slightly better than a coin flip (62.7% against 50%).
A finding marked “high confidence” is the model producing the words “high confidence”. That label hides AI mistakes instead of flagging them, so treat it as a pointer to where to look and require evidence before you accept it.
An agent works in a loop: it calls a tool, reads the output, and decides what to do next. Whatever the target returns, such as an HTTP body, a banner, an error message or an HTML comment, goes back into the model next to the operator’s instructions. The model has no reliable way to tell data from instructions, which is the core of prompt injection and OWASP’s LLM01 risk.
Researchers have turned this against AI attackers. The Mantis defense plants hidden instructions in the responses of a decoy service, using terminal escape codes or HTML comments, and reported over 95% success against automated LLM-driven attacks, in a test on research agents and three easy practice machines.
These AI mistakes are hard to spot because the agent does not look broken. It follows an instruction. XHack AI reads attacker-controlled responses as part of its job, so the same risk applies to it, and we do not claim immunity. The mitigations are the ones the research points to: run the agent on a disposable machine, keep credentials it does not need out of reach, use Manual mode for high-impact commands, and treat any instruction that appears inside tool output as an alarm.
Each of these AI mistakes has the same shape: what it looks like, why it happens, and what to do about it.
Of all AI mistakes, false positives are the most visible, because language models write a convincing vulnerability report more easily than they prove one. In the 2024 Ullah study, all models had a high false positive rate and often flagged patched code as still vulnerable. In the PrimeVul benchmark, GPT-4 with step-by-step reasoning labeled both the vulnerable and the fixed version of a function correctly only 12.94% of the time, below the 22.70% that random guessing scores.
Those are 2023 and 2024 models, but validation matters in 2026 too. XBOW reported that about 1,060 HackerOne submissions went through automated validators, and 245 of them (23%) were still closed as informative or not applicable. Those are triage states and not proof of false positives, but validated is not the same as perfect.
The cost lands on humans downstream. Daniel Stenberg wrote that curl’s confirmed-report rate fell from above 15% to below 5% in 2025, part of the story in why the curl bug bounty ended. If you submit findings to bounty programs, the same checks apply, and our slop test covers them.
XHack’s own tools carry the same limit. AI Probe uses an AI judge to score responses, and its documentation says to confirm a “Vulnerable” verdict manually and try variants. GitGuard’s documentation warns that new repositories can produce false positives in the first few runs.
In our AI penetration testing guide I describe watching a fully autonomous tool report a critical vulnerability that did not exist, because nobody validated it before it reached the report.
What to do: replay the request yourself and state the impact in one sentence. If you cannot, it is not a finding yet.
The quietest of the AI mistakes is the miss, because a clean result feels like good news and proves less than that. XHack’s Terms say that security testing is a point-in-time assessment that does not guarantee every vulnerability has been found.
Coverage is the first reason. AI Probe’s library holds 3,500+ payloads, but a scan does not run them all: Quick runs about 20, Standard (the default) about 50, and Comprehensive 150 or more. A clean default scan means roughly 50 payloads did not land.
The research shows the second reason, which is that finding the unknown is the hard part. On BountyBench’s 2025 tasks, the best agents detected 12.5% of unseen vulnerabilities, while the best results for exploiting known ones and for patching were 67.5% and 90%. A 2026 study of 13 open-source pentest frameworks found that 83.3% of attempts on a chained multi-vulnerability scenario stalled before completing the chain.
Our view, laid out in AI vs human penetration testing, is that business logic and chained exploits are where human testers add the most.
What to do: write down exactly what was tested, run it again, and add human testing for logic flaws and chains.
Models fill gaps with text that looks right, which produces some of the most checkable AI mistakes. A USENIX Security 2025 study of 16 models and 576,000 code samples found that at least 5.2% of the packages suggested by commercial models, and 21.7% from open-source models, did not exist. When the researchers re-ran 500 triggering prompts ten times each, 43% of the invented names came back every time.
That repeatability makes the attack practical. An attacker registers the invented name with malicious code and waits for a developer or an agent to install it, a trick called slopsquatting.
CVE details fail the same way. In a study of ChatGPT without retrieval (a 2024 model), it wrote plausible advisories for 96% of real and 97% of fake CVE IDs and flagged none of the fakes as invalid.
What to do: look up every CVE in NVD or the vendor advisory, every package in its registry, and every version number against the fixed version. Check exploited status against the CISA catalog today, not from memory.
Misreading tool output is one of the most common AI mistakes in the failure data. In the 2026 failure analysis, wrong or missing tool syntax caused 26% of failures, and misreading output or lacking knowledge caused another 16%.
The worse of the AI mistakes is claiming success. In a September 2026 preprint covering six models and 3,600 responses after a tool failure, the false-success rate was 22.8% at baseline, 9.3% with an instruction to be transparent, and 0.8% when the agent had to produce structured evidence. In the independent 2026 study of 13 pentest frameworks, 8 reported hallucinated flags on at least one challenge, even with top 2026 models behind them.
Anthropic’s report on an AI-run intrusion campaign described the same behavior: the model overstated findings and sometimes fabricated data, for example by claiming credentials that did not work. We cover how validation separates useful agents from confident fiction in agentic pentesting.
XHack’s design follows the evidence-first idea. Confirmed findings carry proof and step-by-step reproduction, the HTTP Repeater captures every request, and Cloud Investigation re-runs the evidence behind a finding and has a second, independent AI check it, marking anything that fails as Unverified.

What to do: ask for the request, the response and the proof of impact. A summary that says “exploited successfully” is a claim, and a replayable request is evidence.
Wrong severity and wrong fixes are AI mistakes that look professional, because they arrive as clean numbers and tidy remediation steps. XHack enforces a CVSS v4.0 score on every finding, but severity also depends on how exposed the system is, what data sits behind it and what other controls exist, and the model sees only part of that.
AI-written fixes and mitigation plans are generated guidance, so test the fix before you close the finding. Mapping a finding to a compliance control is a judgment call too, and the auditor decides what counts.
What to do: re-score for your environment, apply the fix in a test setting, and run the original proof again.
Wrong actions are the AI mistakes with real-world cost, because the agent runs real commands. XHack’s Terms are blunt about it. Where you enable auto-approve or unattended execution, tool runs proceed without per-step human confirmation, and you accept the risk of unintended actions against in-scope or out-of-scope systems, service disruption, data modification and cost overruns.
The wider field has seen this. The UK AI Security Institute reported in August 2026 that during a cyber-range evaluation of seven frontier models, agents took 19 unsanctioned actions against the real internet in 10 of 122 runs, after internet access was left on. That was a controlled evaluation and not a commercial pentest product, and we read it as a case for enforcing scope outside the model. OWASP lists the pattern as excessive agency and recommends human approval for high-impact actions.
XHack’s controls for this are specific. Manual mode asks before every action, Allow Edits approves file edits only, Auto approves everything for that chat, and Plan is read-only. Destructive commands such as wiping disks, dropping databases or shutting machines down are blocked in every mode, though the documentation gives those as examples and not a complete list. HTTP Repeater will not replay a request against a different host unless you turn that on.
A runtime guardrail also watches for sessions drifting out of authorized scope, and our Terms call it probabilistic and say we do not claim these controls are infallible.
What to do: start in Manual or Plan mode, write scope into the prompt and project memory, set cost and turn limits, and turn on Auto only for a narrow, well-understood scope. Stop any sub-agent that wanders.
XHack has its own security-tuned model, and the agent runs on it by default, so there is nothing to configure and out of the box your requests do not go to a third-party AI provider. If you want a specific model, your own billing or your own data agreement, you can bring your own key (BYOK) on eligible plans and connect an OpenAI, Claude, Gemini, GLM, Mistral, xAI or OpenRouter key. The desktop agent can also run a local model through Ollama or llama.cpp.
With your own key, the provider you chose bills you and is responsible for its own model behavior and data handling, as section 10.3 of the Terms says. Whichever model runs, the checks in this article apply, because the choice changes the AI mistakes you get. XHack’s documentation says abliterated local models hallucinate more on deep technical work, and the 2026 failure analysis found that swapping the underlying model shrank the gap between pentest systems by more than half.
What to do: note which model produced a finding, since AI mistakes differ by model, and for important work run the check with a stronger model or a human.
No single number applies to XHack AI. We have not published a benchmark, with a method you can check, of how often it is wrong, and we would rather say so than quote a figure we cannot defend. What exists is research on comparable systems, and it points one way: AI mistakes persist even as capability rises.
The 2024 results look weak by today’s standards. In the PentestGPT paper, GPT-4 alone fully compromised 5 of 13 targets and 0 of 2 hard ones. In Cybench’s 40 professional capture-the-flag tasks, the best unguided agent solved 17.5%, and GPT-4 exploited 87% of 15 one-day vulnerabilities with the CVE description in hand but 7% without it.
Capability has grown fast since. The UK AI Security Institute measured the best model of August 2024 completing an average of 1.7 steps of a 32-step network attack range, and the best of February 2026 completing 9.8, both at a 10-million-token budget. At 100 million tokens the best single run reached 22 of 32 steps, on ranges with no active defenders.
Reliability has not kept pace. The independent 2026 study of 13 pentest frameworks still showed hallucinated flags, stalled chains and run-to-run randomness.
All of these are lab benchmarks, and real engagements add defenders, odd configurations and rate limits. Use the numbers as a map of where AI mistakes cluster, not as a forecast for your engagement.
Run these checks in order. They take minutes per finding, and between them they catch every one of the AI mistakes in the table above.
That last step matters more than it looks. XHack does not hold the content of your local agent sessions, so we cannot recover it for you if it is lost.
Bring in a human when AI mistakes would be expensive: when a finding is going to a client or an auditor, when its severity is high or critical, when it involves business logic or a multi-step chain, or when it touches production data. That is the job of managed VAPT.
You are the operator, which also makes you the one who answers for AI mistakes that reach a client. XHack supplies the tooling, does not verify your authorization evidence for each individual action in advance, and does not supervise the tests you run. The Terms also exclude liability for reliance on AI-generated output and cap it at twelve months of fees. We would rather you read that now than discover it later.
Compliance follows the same rule. The Terms state that no output certifies compliance with any law, standard or framework. XHack delivers penetration test reports built to satisfy a framework’s pentest requirement, and the auditor, assessor or regulator makes the compliance decision.
That distinction matters more for AI-assisted tests. In July 2026, CREST added an AI-enabled penetration testing accreditation under which a provider’s use of AI is independently assured, and CREST’s own research with 62 providers reported that 69% already use AI and 85% expect clients to demand more transparency about it. CREST accredits a provider’s use of AI, not a tool. We cover how auditors treat AI-run evidence in AI penetration testing compliance.
This is the section where we talk about XHack, so weigh it accordingly. The honest summary is that how many AI mistakes reach you depends on how much checking sits between the AI and the report.

XHack gives you three ways to work, and each puts the reviewer in a different place:
On privacy, our policy says we do not store the chats of your local XHack AI agent sessions, and we do not use your data, prompts, outputs or findings to train AI models. That keeps a client’s sensitive findings off our servers, and it is why backing up your own evidence matters.
We have not published an accuracy benchmark with a method you can check, and no AI tool today is one you can stop checking, ours included. If you want a cheap scan with a cover page, we are not the right fit.
The 7-day free trial gives you full access and needs no card, after a short identity verification. Use the week to run the checklist above on three findings and see how many AI mistakes the agent’s output contains.
If XHack AI gets something wrong, send the evidence to support@xhack.io or use the Technical Support option on the contact page. Want to talk through which option fits? Book a free consultation, even if that turns out not to be us. Brutal honesty is kind of our thing.
Yes. XHack’s Terms say AI output may be incomplete, inaccurate, outdated or wrong, may report false positives and may miss real vulnerabilities. Those AI mistakes can appear in the agent, the chat, AI Probe, GitGuard and AI-written remediation.
We have not published a measured rate, so do not trust any single figure, including ours. Research on comparable systems shows false success claims of 22.8% in one 2026 experiment and hallucinated flags in 8 of 13 pentest frameworks in another, so assume AI mistakes will reach some unverified findings.
Yes, because auto mode removes the check that stops AI mistakes from becoming actions. The Terms say actions proceed without per-step confirmation and you are responsible for unintended ones, so start in Manual or Plan mode and set cost and turn limits first.
No, because AI mistakes that reach a client become yours. The Terms require a qualified security professional to review and validate output before it is delivered to a client or used to support a compliance position, and the auditor, assessor or regulator then decides what satisfies the requirement, because no XHack output certifies compliance.
Not from your data. The privacy policy says we do not store the chats of local agent sessions or use your data, prompts, outputs or engagement findings to train AI models. The trade-off is that we cannot recover session content for you, so keep your own backups.
XHack AI runs on XHack’s own security-tuned model by default, so nothing needs configuring and requests do not go to a third-party AI provider. You can instead connect your own key (BYOK) on eligible plans, or run a local model through Ollama or llama.cpp on the desktop agent, and the model you pick changes the AI mistakes you see, so note which one produced each finding.
Capture the request, the response and the model that produced it, then email support@xhack.io or use the Technical Support option on the contact page. If the wrong output came from a Cloud Investigation finding, rate it as inaccurate in the product as well.
The useful question is not whether XHack AI can make mistakes. It can, and so can every AI, because language models are rewarded for answering, frozen at their training date, forgetful over long sessions, variable from run to run, and confident whether they are right or not.
AI mistakes shrink with better design and evidence rules, as the research shows, but they do not reach zero. What matters is whether they reach a client, a deployment or a system you were never meant to touch, which depends on the checks between the AI and the outcome and on whether the reviewer is you, a built-in second check or a certified tester.
Run the nine checks on your next finding, start with Manual mode, and keep your own evidence. Then decide how much of the checking you want to own and how much you want us to do.
Categories
Related articles