Miju Labs

The security dossier

The bounty platforms

HackerOne has publicly foreclosed selling its corpus and Bugcrowd built an RL product that routes around its own researchers — the two incumbents with the best human attack data both looked at paying humans for licensable trajectories and declined.

high confidence9 minupdated 2026-08-30hackerone · bugcrowd · synack · data rights · consent

The most important finding in this research is a negative one, and it comes from the two companies best placed to have done the thing.

Between them, HackerOne and Bugcrowd hold the largest corpus of human attack trajectories that exists anywhere — millions of submitted reports across two decades, produced by professionals, against real targets. Both have looked at the question of selling that corpus to AI labs. Both declined, for different reasons, and both said so in public.

HackerOne: contractually and publicly foreclosed

HackerOne's community terms contained the line: "HackerOne may use confidential information to develop or improve its services, for example, to identify trends and to train AI models." The researcher community read that as their exploitation techniques feeding model training, and reacted accordingly.

CTO Alex Rice answered directly. HackerOne is not training LLMs on researcher submissions; the "train AI models" language predates modern LLMs and referred to a spam-classification engine built on regression analysis. Under Section 8 of the community terms, researchers retain IP. HackerOne receives a limited licence for service provision, customers get slightly broader rights to secure their own systems, and neither permits resale or redistribution. HackerOne committed to updating the terms to distinguish legacy ML from LLM training. Their own agentic PTaaS tool draws from "public benchmarks, internal benchmarks, public CVEs, and opt-in sidecar runs — not bug bounty submissions" (Critical Thinking Ep. 162).

Note what "opt-in sidecar runs" is. It is a consented, payroll-shaped data-collection mechanism bolted onto a piece-rate platform — the company reaching for exactly the structure a specialist would build, because the piece-rate corpus underneath it is unusable for the purpose.

The Finder Terms are broader than the summary suggests

HackerOne's published Finder Terms (2023) grant HackerOne a "perpetual, irrevocable, non-exclusive, transferable, sublicensable, worldwide, royalty-free license" to submissions — but "for the sole purpose of providing the Services" (HackerOne). The purpose limitation is what does the work; the adjectives before it are as broad as Gray Swan's. The customer's parallel licence carries no such purpose limit. Anyone modelling this should read the clause rather than the headline.

HackerOne does sell AI red teaming as a service: H1 AI Red Teaming, with 750+ AI-focused researchers, 1,700+ AI assets tested, 15–30 day engagements, deliverables including "full multi-turn traces" mapped to the OWASP LLM Top 10 and NIST AI RMF, with Anthropic, IBM, Snap, Adobe, Zoom and Cloudflare named as customers (HackerOne). Researcher compensation on that product is not disclosed. That is a project deliverable with a scope and an end date, not a data feed.

Bugcrowd: entered the market and routed around its own crowd

On 21 May 2026 Bugcrowd launched Reinforcement Learning Environments, aimed squarely at "large language model providers and frontier AI research teams building security-aware agents." The product offers "hundreds of thousands of training environments built from authentic open-source vulnerabilities with real source code," in which agents locate bugs, trigger them, assess exploitability and produce fixes, with objective scoring at each step. It is built on technology from Bugcrowd's November 2025 acquisition of Mayhem Security — symbolic execution and fuzzing, descended from DARPA's Cyber Grand Challenge. Frontier labs are already using it (Bugcrowd; SiliconANGLE).

And the sentence that defines the page: "no customer data or security researchers are used at any stage of the training process."

A company with 500,000+ registered researchers, selling to the exact buyer that wants attack data, chose to manufacture environments from public CVEs instead of negotiating rights with the people it already pays. CAIO David Brumley's framing of why: "You cannot train a model to be good at security by showing it what security looks like, you have to give it real problems to solve."

Read the conclusion both ways, honestly

Two incumbents with the best possible position looked at "pay humans properly for consented, licensable attack trajectories" and did not do it. There are two readings and the evidence does not settle between them.

The bearish reading: this is the strongest available evidence that the corpus cannot be assembled at a price anyone will pay. Bugcrowd had the crowd, the relationships and the buyer, and still concluded that synthesising rights-clean environments from open-source CVEs was cheaper and faster than extracting them from a crowd. If the party with every advantage routed around the problem, the problem may simply be the wrong shape.

The bullish reading: neither declined on economics. HackerOne declined on contract — its terms leave IP with researchers and it has now publicly promised not to monetise the corpus that way, which forecloses the move for as long as that promise holds. Bugcrowd declined on product shape — its RL environments need reproducibility and verifiable oracles, which crowd submissions do not have, and the Mayhem acquisition gave it a cheaper route to exactly that. Neither tested whether humans paid hourly, under consent, producing purpose-built trajectories, is a viable supply. The lane is open because the incumbents' constraints are not the entrant's constraints.

Both readings are compatible with the same facts. The honest position is that this is the largest single uncertainty in the thesis, and it is resolvable only by attempting it.

The rest of the field

PlatformResearchersPayoutsSelling data to labs?
HackerOne1.5M+ [WEAK — platform marketing]$81M (2024–25); $300M+ lifetimeNo — publicly foreclosed
Bugcrowd500K+ [WEAK]$50M+ lifetime, 1,000+ programmesYes — but explicitly not researcher data
Synack1,500+, invite-only52,000 tests/yearNo announcement found
Intigriti100K+ [WEAK]Not establishedNo announcement found
YesWeHack50K+ [WEAK]Not establishedNo announcement found
ImmunefiNot established$134M lifetime to Q1 2026; $190B TVL, 230 programmesNo announcement found

Synack is the only fully vetted population in the record — 1,500+ members behind a five-step vetting process (Synack). It is the best available proxy for "practitioners who reliably deliver professional-grade work," and it should be the number anyone uses when sizing quality supply, not the millions quoted elsewhere. Its compensation structure is also the most payroll-like on any bounty platform: vulnerabilities from $500 to several thousand, plus "missions" at $25–$50 for routine tasks with ad-hoc missions "easily exceeding $100+", plus payments for report submissions, patch verifications and mentoring. Members are 1099/W8BEN contractors, most engaging a few hours a week. Synack sells human-delivered testing; there is no AI-data initiative.

Intigriti, YesWeHack and Immunefi show no data-licensing activity that could be established. Immunefi is worth watching for a different reason: its median payout is ~$2,000 against an average of ~$52,800, a 26x mean/median gap, and one Q1 2026 bounty was 38% of all quarterly earnings. That is the sharpest illustration in the dataset of why piece-rate discovery work produces terrible expected value for the median participant — the argument Offensive supply builds the recruiting case on.

Neither has a defensive-side equivalent; the closest thing to one is on The defensive vendors, and it is a services firm in Hyderabad.

Take rates: not published by any platform. Industry convention holds that platforms charge programmes a subscription rather than clipping researcher payouts, but no source states it authoritatively. See What a rake can actually be.

The last piece is the January–February 2026 controversy, and it matters more for recruiting than for competition.

HackerOne launched an Agentic PTaaS product in January 2026, marketing that its agents "are trained and refined using proprietary exploit intelligence informed by years of testing real enterprise systems." Researchers read that as their reports. The quotes are worth keeping: "As a former H1 hunter, I hope you haven't used my reports to train your AI agents"; "We're literally training our own replacement."

CEO Kara Sprague clarified on 18 February 2026: HackerOne "does not train generative AI models, internally or through third-party providers, on researcher submissions or customer confidential data"; third-party providers cannot retain or use researcher data for their own model training; "You are not inputs to our models" (The Register).

No opt-out or consent mechanism was introduced. The controversy was resolved by denial, not by building a way for researchers to say yes.

That is the gap, and it is a recruiting asset before it is a product one. A large, vocal, professionally networked population has already organised around the question of whether its attack data is being used to train AI without consent or compensation, and the answer it received was "we are not doing it" rather than "here is how you opt in and what you get paid." Contrast Gray Swan AI, whose crowd signs an irrevocable worldwide licence at the door and does not appear to mind. A company whose proposition is we will pay you, explicitly, per artefact, for exactly that, and here is the consent document is arriving at a pre-agitated audience with a ready-made narrative. Design the consent artefact first; it is not a legal chore. See Paying the crowd and Getting cut out.

Unresolved on this page

Take rates for every platform. Researcher counts from primary sources for Intigriti, YesWeHack and Immunefi — the figures above come from an SEO comparison post. Whether any platform is quietly building a data product: nobody systematically enumerated AI/ML job postings across HackerOne, Bugcrowd, Synack, Intigriti and YesWeHack, and that scrape is cheap, fast and high-signal. Also note the two HackerOne denials come from two different executives about two different episodes; treat them as a pattern of position, not a single statement.