Back to Features

AI Probe — Adversarial Security Testing for AI Systems

Test your AI systems against 350+ attack payloads across OWASP LLM Top 10 categories. Detect prompt injection vulnerabilities, data leaks, jailbreaks, and safety failures before attackers exploit them.

Find the Weaknesses in Your AI Before Attackers Do

Every AI system has blind spots. Language models can be manipulated into revealing system prompts, generating harmful content, leaking training data, or bypassing safety guardrails entirely. AI Probe automates the process of finding these weaknesses by running hundreds of adversarial attack payloads against your AI systems and reporting exactly where they fail.

Whether you are deploying a customer-facing chatbot, an internal AI assistant, a RAG-powered knowledge base, or a tool-calling agent, AI Probe tests it the way a real attacker would — systematically, persistently, and across every known category of LLM vulnerability.

Full OWASP LLM Top 10 Coverage

AI Probe ships with over 350 pre-built attack payloads organised across every category in the OWASP LLM Top 10 framework, the industry standard for AI security assessment.

Prompt Injection tests whether your AI can be tricked into following attacker-injected instructions — through direct overrides, delimiter confusion, encoding tricks like base64 and ROT13, translation bypasses, and multi-step payload splitting.

Sensitive Data Disclosure probes whether your AI can be coaxed into revealing personal data, API keys, internal configurations, or training data through extraction techniques, inference attacks, and social engineering patterns.

Insecure Output Handling tests whether your AI's responses can be used to inject XSS payloads, SQL injection strings, or other code into downstream systems that consume its output.

Excessive Agency checks whether your AI can be persuaded to take unauthorised actions, escalate its own privileges, or operate beyond its intended scope.

System Prompt Leakage uses dozens of techniques to extract the system instructions from your AI — through direct requests, roleplay, translation, encoding, completion, and logical exploitation.

RAG Poisoning tests whether document injection, metadata manipulation, or context window abuse can override your AI's behaviour through its retrieval pipeline.

Misinformation evaluates your AI's resistance to generating false citations, manufactured consensus, and plausible-sounding disinformation when prompted.

Denial of Service includes token bomb payloads, recursive expansion, infinite loops, and context window abuse that could exhaust your AI's resources or billing.

Jailbreaks covers the full spectrum of known bypass techniques — DAN variants, developer mode, roleplay escalation, hypothetical framing, encoding bypasses, skeleton key attacks, many-shot jailbreaking, and world simulation prompts.

Bias and Toxicity probes whether your AI can be led into generating hate speech, discriminatory content, stereotyping, or implicitly biased responses under various framing conditions.

Connect Any AI Target in Minutes

AI Probe works with any AI system that exposes an HTTP API. You configure a target by specifying the endpoint URL, request body template, response JSONPath, authentication headers, and timeout. Presets are available for popular providers including OpenAI, Anthropic Claude, Google Gemini, and any OpenAI-compatible API.

The body template uses a simple placeholder system. For single-turn tests, {{payload}} is replaced with the attack prompt. For multi-turn conversations, {{messages}} is replaced with the full conversation array. This means AI Probe can test REST APIs, OpenAI-compatible chat completions, and custom LLM interfaces without modification.

AI Probe — Target Configuration

Launch a Scan and Walk Away

Select your target, choose which attack categories to include, and launch. The scan runs in the background. You can close the browser, work on other things, and come back to check progress at any time.

Each payload is sent to your target, the response is captured, and an AI judge evaluates whether the target defended successfully or was compromised. Progress updates in real time — you can watch findings appear as they are discovered, or come back later to review the full results.

For targets with rate limits, the scan automatically paces itself with configurable delays between payloads to avoid triggering throttling.

AI Probe — Scan Dashboard

Multi-Turn Conversation Attacks

Simple one-shot attacks only test the surface. Real attackers use conversation. AI Probe includes multi-turn payloads that use four distinct strategies.

Crescendo attacks start with innocent questions and gradually escalate across multiple turns, testing whether your AI's defences weaken as rapport is established. Persistence attacks repeat the same request with different framings until the target gives in. Context manipulation shifts topics mid-conversation and circles back, testing whether safety context is maintained across turns. Confusion attacks send contradictory instructions to destabilise the model's coherence.

Each turn is logged and evaluated. If your AI refuses on turn one but complies on turn five, the judge catches it.

AI-Powered Evaluation

Every response is evaluated by a dedicated AI judge that has been specifically trained for security assessment. The judge does not use simple keyword matching. It understands the intent and context of both the attack and the response, and it knows the difference between a genuine refusal, a soft refusal that leaks information, and actual compliance with the attack.

Findings are classified into five severity levels — critical, high, medium, low, and informational — based on the actual impact of the response, not just whether the AI said yes or no. A response that reveals the full system prompt is critical. A response that hints at the existence of certain tools but does not name them is medium. A clean refusal with no information leakage is informational and marked as safe.

Findings Dashboard

Results are organised into three tabs — vulnerable findings, safe findings, and other (errors, timeouts, inconclusive). Each finding shows the severity, title, attack category, security score, and response time. You can expand any finding to see the full prompt sent, the complete target response, and the judge's reasoning for its classification.

The summary panel shows your overall security score (A through F), total findings by severity, and pass rate across all tested categories.

AI Probe — Findings Dashboard

Escalate — Go Deeper on Any Finding

When AI Probe discovers a vulnerability, the Escalate feature lets you take that finding and continue the attack manually. It opens an interactive chat interface pre-loaded with the original attack context — the prompt that was sent, the target's response, and the judge's assessment.

From there, you can send your own follow-up messages to the target, test variations of the attack, or try to escalate the impact. Every message and response is logged in the conversation history. You can also ask the AI bypass generator to suggest a new attack payload based on what it has observed in the conversation so far.

There is also a reset button that brings the conversation back to the original finding state, so you can try different escalation paths without losing your starting point.

AI Probe — Escalate Chat

Retest After You Fix

After patching a vulnerability, use the Retest button on any finding to re-send the same payload and have the judge evaluate the new response. The result updates in place, so you can see whether the status changed from vulnerable to safe. A Retest All button runs every finding again in one operation, giving you a full before-and-after comparison across the entire scan.

Professional Reports

Generate a complete security assessment report in HTML or PDF format at any point during or after a scan. Reports include an executive summary, security score, grade, severity breakdown chart, detailed findings with prompt and response, judge reasoning, and category-by-category results.

Reports are formatted for professional use — ready to share with engineering teams, security leads, compliance officers, or clients. They include a confidential banner, scan metadata, and clear status indicators for incomplete scans.

AI Probe — Report

Custom Payloads

Beyond the built-in library, you can create your own attack payloads tailored to your specific AI system. Custom payloads support all the same fields as system payloads — category, severity, attack mode (single or multi-turn), strategy, prompt template, expected behaviour, and custom judge criteria.

This lets you test for organisation-specific risks — proprietary data leakage, business logic abuse, or industry-specific compliance requirements that generic payloads would not cover.

Built for Security Teams

AI Probe is available on the Elite company plan and the Elite researcher plan. Administrators can manage targets, view scan history, configure payloads, and monitor scans across the organisation from the platform dashboard.

All scan data is tenant-isolated. No data is shared across organisations. API keys and authentication headers are encrypted at rest and never returned in API responses.

Ready to get started?

Experience this feature firsthand and see how it can enhance your security operations.

Start Testing
Need Help?

Our team is here to assist you with any questions or issues.

Contact Support