
Table of contents
17
By Salman Khan, OSCP+, Founder of XHack, SRT (Synack Red Team member)
Read this in 30 seconds: A bug bounty methodology is the order you work in and the output each step must produce. In 2026 the rules around each step changed: platforms now penalize unvalidated submissions, require severity on some programs, and support CVSS 4.0 in different ways.
- Validate before you submit, every time. Bugcrowd’s March 2026 rules put accounts with repeated invalid reports under review, and can suspend them or require identity verification. Apple can pause a researcher’s reports for 180 days.
- Severity is becoming a required field. HackerOne’s documentation says severity becomes required on bug bounty, vulnerability disclosure and challenge programs from September 21, 2026, though programs can opt out.
- CVSS 4.0 is now configurable on the big platforms, but each one sets it differently, so check which version the program uses.
- Automate the lists, keep the judgment. The most useful thing a methodology does is decide what gets tested and what gets reported.
- Write the report a triager can act on. Stenberg’s recipe for curl is a practical template, and it is the step that most often decides the outcome.
A methodology is not a checklist you run once. It is a loop, and in 2026 the loop has a penalty attached to every step you skip.
Most published bug bounty methodologies are long lists of tools or vulnerability types. Few say what each phase must produce, or what the platform rules now say about that output. This article is organized around outputs, and it sources each rule to the platform or standard that sets it.
We build XHack AI, which includes recon and testing phases in its agent. We are not neutral about tools. Where a claim comes from a platform’s own documentation, we link it. Where we are reading between sources, we say so.

Methodology advice written before 2026 rarely mentions how platforms respond to automated and AI-assisted reporting. The platforms now do, and their rules change what a finished step looks like.
Bugcrowd, March 10, 2026. Bugcrowd’s policy post announced several rules, which it said were initially in effect and might not be the whole list. Accounts that farm submissions face a permanent ban, with reinstatement only through an appeal that requires identity verification. Ten consecutive invalid reports trigger a review, and the account may get a 30-day suspension if the submissions look automated or AI-generated without validation. Ten invalid reports in total require identity verification before further submissions. Bugcrowd also says that sending low-detail submissions from multiple accounts to train AI triage systems violates its terms.
Bugcrowd’s post puts accountability on the human. It says the existing code of conduct already requires the researcher to be responsible for the accuracy and value of each submission. The post does not require disclosure of AI use, and it does not list specific validation steps.
Apple, Security Bounty guidelines. Apple lists “issues discovered by AI without proper validation” as ineligible, and asks researchers to avoid lengthy AI-generated descriptions. Repeatedly submitting ineligible reports may lead Apple to pause processing for 180 days, and researchers with more than two pause periods may be removed from the program. A complete report, Apple says, needs a working exploit or reliable proof of concept, plus numbered reproduction steps.
HackerOne. We could not find a published HackerOne policy on AI-generated reports in its current documentation. What we did find is its severity documentation, which is concrete. Starting September 21, 2026, severity is required for bug bounty, vulnerability disclosure and challenge programs, and programs can opt out in their settings. Researchers can pick a severity manually or use a CVSS calculator.
Put together, the 2026 rule is simple: a report that has not been validated is now a risk to your account, not just a wasted triage hour.
Output: a scope file and a written list of what you may not do.
Scope is the first input to any methodology: the set of assets a program lets you test, and the rules of engagement say how. The rules matter as much as the scope. They can set request limits, forbid automated scanning, restrict which accounts you may create, or ban certain techniques entirely.
Read the program’s policy page first, then the platform’s guidance. Intigriti’s guide on aggressive scanning says that some public programs allow automated tools only at 2 to 10 requests per second, and that breaches can lead to warnings, invalidated reports or suspension. Write the limit into your notes.
Keep the scope in a file that every later tool in your methodology reads. Most avoidable mistakes start here, and this file costs almost nothing to get right.
Output: a list of hosts and assets that belong to the target and that you are allowed to test.
Recon is the phase of your methodology most often automated, and the one where automation is most clearly useful. Passive subdomain discovery queries public sources without touching the target. Probing then shows which hosts answer and what they serve.
We covered tool choices, rate settings and a sample pipeline in recon automation for bug bounty hunters. For this methodology, the rule is passive first, filter by scope, probe at the program’s rate, then review the list by hand before any testing.
The output is a short, reviewed list of live in-scope hosts. A long list you have not reviewed is a backlog, not an output.
Output: a ranked list of targets, with a reason for each ranking.
Mapping turns a host list into a picture of the application, which is the input your methodology needs for testing. For each live host, note what it does: a login page, an API, an admin panel, a file upload, a payment flow, an old version of something. Then rank them.
Ranking is judgment, and it is where many hunters lose time. The obvious host is not always the valuable one. The host with the most user data, the most complex permissions or the most recent change often is. Write down why you ranked each target, so you can revisit that reasoning when nothing turns up.
AI can summarize what a host does and flag unusual responses. It does not know which feature handles money for this company, or which change shipped last week. That context comes from the program’s documentation and from reading the application.
Output: a set of hypotheses you tested, with the result of each, including the negative ones.
This phase is the core of the methodology, and a standard reference helps most here. The OWASP Web Security Testing Guide, version 4.2, organizes testing into these categories, in this order:
The guide is written for full assessments, so a bug bounty methodology built on it has to be trimmed to a single target, so you will not run every category on every program. Choose the categories your methodology needs for the target’s features. An API with user accounts and per-user data needs authorization and API testing. A static marketing site needs very little.
Our mapping of these categories to bounty work is our synthesis, not an OWASP statement. Authorization is where many valuable findings sit: one account reaches another account’s data, and no scanner knows which data belongs to whom. Business logic is similar: a discount that stacks, a refund issued twice, a workflow that skips a step. Those need two accounts or a careful read of the flow.
Automation fits best in categories built on known patterns, such as configuration exposure, default credentials and common input-validation mistakes. It fits worst where the bug depends on what the application is supposed to do.
Keep a record for every hypothesis your methodology produces. A negative result tells you not to test that path again, and most hunters forget to write it down.
Output: a reproduced finding, with exact steps and an honest statement of impact.
Validation is the phase where a methodology earns its keep. It means you reproduced the behavior yourself, from a clean state, and confirmed it is a vulnerability in this program’s scope. A tool’s output is a claim. Your reproduction is the evidence. In 2026 this step also protects your account: Apple requires a working exploit or reliable proof of concept, and Bugcrowd’s rules target submissions sent without validation.
Impact is the second half of the methodology’s validation work. Say what an attacker gains and under what conditions. Avoid inflated severity. Programs notice it, and it costs credibility on the next report.
On severity, the platforms now support CVSS 4.0, but in different ways:
CVSS 4.0 changes the scoring. The FIRST CVSS v4.0 specification adds Attack Requirements as a base metric, and a Supplemental group whose values do not affect the final score. A third-party bounty guide notes that the same bug can score differently under 4.0 than under 3.1, so state the version you used in the report. Score the finding against the conditions you observed, and state your assumptions. A defensible score is worth more than a high one you cannot defend.
Automation cannot do this phase. A scanner can report that a response looks unusual. It cannot tell you the response grants access the user should not have.
Output: a report the triager can act on, and a plan for the conversation that follows.
The report is the product of the whole methodology. Earlier phases prepare it, and a weak report can waste a strong finding.
Stenberg’s recipe for an excellent vulnerability report, published by the curl maintainer on June 29, 2026, is a practical template. It asks you to check the behavior is not documented, open with a short human summary of the flaw and its impact, use the program’s channel, write as a person, include a reproducer, offer a patch if you can, state the affected versions, stay available for questions, and learn from the result. We covered the full recipe in why the curl bug bounty ended.
The line that matters most for an AI-assisted methodology is the one about communication. Whatever tool helped, you write to the maintainers and answer their questions. Bugcrowd’s rule is the same in principle: the human researcher is accountable for each submission’s accuracy and value. curl’s own account of why its bug bounty ended centers on low-quality, AI-assisted reports. We are not claiming that is the whole story, only that the maintainers document it.
Follow-through matters. Triage can take days or weeks, and a fast, clear reply to a question often decides whether a report is accepted. Record the outcome, and let it change the next target’s ranking.
| Phase | AI helps with | AI does not replace |
|---|---|---|
| Scope and rules | Summarizing a long policy page | Reading the policy and deciding what it allows |
| Recon | Running and sorting tool output | Reviewing the host list and removing what is out of scope |
| Mapping | Describing what each host seems to do | Knowing which feature matters to this company |
| Testing | Generating variations and repetitive checks | Authorization and business-logic judgment |
| Validation | Suggesting reproduction steps | Reproducing the bug yourself, from a clean state |
| Reporting | Drafting structure and wording | Owning every claim and answering the triager |
AI cuts the cost of producing candidates. It does not cut the cost of deciding which candidates are real, or of standing behind the report. Platform rules now penalize the gap between the two. Our earlier analysis of AI-assisted bug bounty results makes the same point from the platform side.
XHack AI is a desktop security agent for authorized work, and bug bounty is one of the places it is used. We are not neutral about it, so here is what it does for your methodology, what it requires, and what it does not do.
What it does across each phase of the methodology. The Workflow Engine lets you lay out your own methodology as a node graph, and the agent runs it step by step with a live view of progress. The built-in browser records every request and response with its headers, cookies and body into the Repeater, so you can check the traffic behind each hypothesis, including the negative ones. Sub-agents for recon, scanning and exploitation can run in parallel, and you can pause or redirect any of them (Sub-Agents and Swarm).
Findings you can check. Confirmed issues land in a findings panel with proof, reproduction steps and the exact request, and export to PDF, HTML, Markdown or JSON (Findings and Reports). Our release notes say the newest version requires a CVSS v4.0 rating and a CWE on each finding.
Who can use it, and how access works for a methodology built on it. Access is gated by identity verification. Individual accounts upload a government ID or passport, front and back, and at least one professional credential, such as a certification like OSCP or CISSP, or a relevant education diploma. The verification page says review is usually decided within one business day, and that it happens before payment. Organization accounts follow a separate process.
What the terms require of you when you run a methodology. Section 6 of our terms, “Authorized Use Only,” says you may test only systems and networks for which you hold explicit, documented, current authorization from the owner, and that you must stay within the scope and the rules of engagement. In bug bounty terms, that means the program’s scope. Section 10.1 says AI-generated output may be incomplete, inaccurate, outdated or wrong, may report false positives, and may miss real vulnerabilities. That is why every finding in this methodology gets validated by you before it is submitted.
Pricing, as listed on the pricing page. Starter is $20 a month, which the page calls an introductory launch price. Professional is $49 a month and Elite is $150 a month. The page lists a 7-day free trial, one per customer. It does not say whether the trial requires a card, so check that before you sign up.
What it does not do. XHack AI is not a bug bounty platform. It does not choose programs for you, submit reports to them, or guarantee that a report is accepted, paid or unique. It does not replace your own methodology or reading the program’s rules, and it does not decide whether a finding is real. Those stay with you.
It is the order of phases you follow for each target, with a defined output at each phase. A useful methodology also states what validation each output needs before you move on.
Not in the rules we checked. Bugcrowd’s March 2026 post does not require disclosure of AI use, and it does not list validation steps. It does hold the human researcher accountable for each submission’s accuracy. Check each program’s policy, because programs can add their own requirements.
On Bugcrowd, ten consecutive invalid reports trigger a review, and the account can be suspended for 30 days if the submissions look automated without validation. Ten invalid reports in total require identity verification before further submissions. These are the rules in Bugcrowd’s March 10, 2026 post, which says they may not be the whole list.
HackerOne’s documentation says severity is required on bug bounty, vulnerability disclosure and challenge programs from September 21, 2026, and that programs can opt out in their settings. Check the program’s settings before you submit.
Use the version the program specifies, and state it in the report. HackerOne, Bugcrowd and Intigriti each set it differently. The same bug can score differently under CVSS 3.1 and 4.0.
Parts of it can. Recon and probing are the most automatable. Authorization and business-logic testing need a person who understands the application, and validation and reporting remain your responsibility.
A bug bounty methodology is a six-phase loop: scope, recon, mapping, testing, validation and reporting. Each phase needs a defined output, and in 2026 the platforms attach penalties to skipped validation.
Automate the lists in your methodology. Keep the judgment, the validation and the report for yourself, and write down what you tested, including what came up empty.
Categories
Related articles