Miju Labs

The security dossier

Ninety days in security

A sequence with a disproof attached to every step: find the six firms first, publish one calibration benchmark, recruit against a number that beats a bug bounty hunter's entire year, and build your own targets because Hack The Box's terms ban the whole business.

medium confidence8 minupdated 2026-08-30execution · recruiting · benchmark · metrics · sequencing

Ninety days answers three questions and no others. Is there a named buyer for validated logic-bug judgement at a price above the labour cost? Can you assemble twenty reviewers who can reproduce, not just read? And does a published calibration number get returned calls? Everything below is ordered so that a no arrives as early and as cheaply as possible.

The general version of this sequence is The first ninety days. This is the security-specific one, and it differs in one respect: the buyer here publishes its vendors — and since this page was written, it has published their names — so the first ten days are spent talking to them rather than looking for them.

Days 0–10: work the six firms — they are named

This step has changed since the sequence was first written. The six external security research firms triaging Anthropic's CVD pipeline are no longer an unknown: they are Ada Logics, Anvil, Calif.io, Doyensec, Ophion Security and Trail of Bits, listed on the dashboard's own About page (Anthropic CVD — about). The ten days go on contacting them rather than finding them.

Three conversations in priority order. Calif.io first, because it is not a peer — it publishes capability evaluation for frontier labs as a service line and carries Google DeepMind on its logo wall (calif.io); the question is whether it is at capacity, and whether the calibration layer sits inside or outside what it sells. Ophion Security second, because it is the only one of the six with a published pricing page, which is the closest public price for AI-assisted vulnerability triage in existence. Ada Logics and Doyensec third, as the two most likely to be volume-constrained on open-source triage.

What it disproves: if two or more of the six say they are turning work away, this stops being a market-entry question and becomes a hiring or subcontracting plan. If Calif.io already sells severity calibration as well as capability evaluation, the wedge at The one shape that survives is occupied and the honest move is to sell into the six rather than become a seventh — see the revised room: 2 at The security read.

In the same ten days, ask two or three labs one free question during ordinary commercial discovery: what do you require of a data vendor? SOC 2 Type II, ISO 27001, nationality restrictions, air gap, personnel screening. No lab publishes its supplier security addendum, and the answer decides whether ISO 27001 is a month-three cost or a month-thirty one. See Copy Daybreak's architecture.

Days 10–30: build one thing, and make it the credential

Ship a published calibration benchmark for logic-bug severity: a set of logic vulnerabilities in published, patched open-source software, three independent expert reviewers, agreement statistics against every frontier model you can reach, and the disagreements published with the reasoning on both sides. Hold back a second set of identical construction and never release it.

Three constraints on how it gets built, each of which is load-bearing.

Published CVEs only, as the master compliance discipline. BIS's own FAQ says "[n]either the disclosure of the vulnerability nor the disclosure of the exploit code would be controlled under the rule" (BIS FAQs), and published technology falls outside the EAR entirely. Restricting to published, patched vulnerabilities simultaneously supports the export argument, removes the coordinated-disclosure obligation, removes the worst misuse case, matches the ethics position ExploitBench and Cybench already take, and is the precondition for an insurer writing the risk at all. One rule, five jobs — Exploits are free, uploading is an export.

Build your own targets. Do not touch Hack The Box content. Its Acceptable Use Policy, effective 1 April 2026, states: "You shall not use any content from the Services… to train, evaluate, fine-tune, test, benchmark, or develop any machine learning model, artificial intelligence system, large language model," and separately bans using outputs derived from the service to provide commercial services to third parties (HTB AUP). That is a categorical prohibition on this business, it is a contract rather than a criminal matter so there is no intent element to argue about, and Hack The Box also builds ranges for UK AISI — it is a competitor as well as a landlord. Read the TryHackMe, Immersive Labs and RangeForce policies in the same afternoon; nobody has. See The terms of service bite first.

Where a permissive licence exists, use it and say so. SigmaHQ's 3,000+ detection rules are under Detection Rule License 1.1, which grants rights "to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies," subject to attribution (SPDX). A commercial derivative is viable without negotiating anything. Elastic's rules are not equivalent — Elastic License v2 — so do not mix them.

What it disproves: if the frontier models already agree with expert severity at, say, 95% exact, the calibration gap is closed and the wedge is gone. That is a two-week finding and it saves a company.

Days 20–45: recruit twenty, using the number that does the work

Run recruiting in parallel with the benchmark, because the reviewers are also the authors.

The proposition, stated exactly: twenty hours at $85/hr is $1,700, which is more than the entire annual income of the median earning bug bounty researcher. HackerOne's $81M of annual payouts divided by roughly 50,000 researchers who have ever earned a bounty gives ~$1,620 a year; 35–45% of active hunters earn zero; about 40% of those who submit a report never receive a bounty (Bug Bounty Economics) [WEAK — synthesis of platform reports]. For everyone outside the top few percent, guaranteed hourly work is a category change rather than a raise. You are recruiting the 97% the platforms never monetised — see Offensive supply.

For the reviewers you actually need, pay above that. Reproduction capability is the filter and the market clears at $150–250/hr for it.

The channels, in order of leverage:

  1. The Critical Thinking — Bug Bounty Podcast — the central media property for professional bug bounty hunters, and the outlet whose coverage surfaced the HackerOne terms-of-service controversy (Ep. 162). Its audience has already organised around the question of whether their attack data trains AI without consent. Arrive with the consent document written.
  2. tl;dr sec, 90,000+ subscribers — the largest single channel into security practitioners. Establish first whether it is still open to third-party sponsorship: its creator Clint Gibler joined OpenAI in June 2026 to lead its cyber team (Benzinga).
  3. Detection Engineering Weekly, 5,200+ subscribers — small, dated, and the most precisely targeted channel to the highest-value defensive sub-specialism there is.
  4. SigmaHQ contributors and CTFtime's 38,575 teams — the highest-precision qualification signals available, because every entry is someone who has demonstrably done the thing.

Write the consent artefact before the first advert. The HackerOne episode was resolved by denial — "You are not inputs to our models" — and no opt-out or consent mechanism was built (The Register). Nobody in this market offers explicit, paid, per-artefact provenance, and one document serves three audiences: the recruit, the export authority and the buyer's counsel. Copy Mercor's warranty line verbatim — "Your work at Mercor will not involve access to confidential or proprietary information from any employer, client, or institution" — because it is the market's answer to the employer-IP problem and a buyer's lawyer will recognise it. The corpus is the target and Paying the crowd.

What it disproves: if 200 targeted approaches through those four channels do not produce 20 people who can reproduce a logic bug on a schedule, the reach score is wrong and no amount of budget fixes it. Gray Swan converted 15,000 signups into 100+ placeable professionals — 0.7% — and that is the honest base rate.

Days 45–75: publish, then convert

Publish the benchmark with the disagreements attached. Then run the conversion motion Irregular demonstrated: model-by-model assessments published continuously, each one a bid for a citation, with the private held-out version as the thing actually being sold. Irregular went from first published model assessment to $450M in six to nine months (Irregular, in full) — read that as a ceiling, not a plan.

Ask for a paid pilot, not a design partnership: a fixed number of validated reports, a price that hurts slightly, a delivery date the buyer picks. A buyer with real urgency pays to move a date; one without it offers a logo. Target three counterparties in this order: whichever of the six firms turns out to be capacity-constrained (as a subcontract, not a competition), UK AISI's evaluation team (which contracts external builders by name and is therefore applicable-to), and the labs' CVD and Safeguards functions rather than their research orgs.

Take the terms seriously on the first contract, because it sets the shape of every later one: licence rather than assignment, a statements-only hold-out on the FrontierMath model, a time-boxed exclusivity window of 12–24 months rather than a perpetual one, and no exclusivity given away as goodwill — it is worth 4–5x when sold deliberately. What you can actually sell.

Day 90: the numbers that mean continue or stop

Day 30. The six firms are named, or you know why they cannot be. At least one lab has answered the vendor-requirements question. The benchmark's first agreement statistics exist on internal data. Stop if frontier models already match expert severity above ~95% exact, or if every one of the six is a large consultancy on a framework.

Day 60. Twenty reviewers signed and background-verified, with at least ten who have reproduced a logic bug end-to-end for you under recording. The benchmark is published and has been cited, linked or discussed by someone you did not pay. Three buyer conversations have reached a price discussion. Stop if reviewer acquisition cost exceeds roughly one month of that reviewer's billings, or if no buyer will name a per-report or per-hour number even under NDA — the missing numerator is the single most important fact in this market and a buyer who will not supply it is not a buyer.

Day 90. One signed paid pilot with a delivery date and acceptance criteria. A measured throughput figure — reports per reviewer-week — against CTI-HAL's five-per-annotator-week as the benchmark to beat. A gross margin computed after rework and payout costs, not before. And a held-out set that exists as inventory rather than as an intention.

The one number that decides it

Price per validated logic-bug report, quoted by a real buyer. Everything else on this page is instrumentation. If ninety days produces a number above roughly $700 a report, the business closes on public information for the first time and the plan becomes a hiring plan. If it produces a number at or below Loginsoft's $25–49/hr labour equivalent, this is an offshore services business and should be built as one or not at all. If it produces no number, the honest conclusion is that nobody has established there is a market — which is exactly what Sizing the cyber pot says today.

What would have to be true, and what nobody has checked, is at What the dossier could not establish.