Miju Labs

The security dossier

Irregular, in full

Thirty-five people are load-bearing for four competing labs' cyber safety claims, at $450M, on a hosted service nobody can licence — and they got there in six to nine months from their first published model assessment.

high confidence9 minupdated 2026-08-30irregular · pattern labs · evals · system cards · concentration

Pattern Labs is Irregular. The rename is confirmed on the record: Sequoia's partnering post states plainly that Irregular was formerly known as Pattern Labs, and that evaluations published under the Pattern Labs name are cited in system cards for GPT-4, o3, o4-mini and GPT-5 (Sequoia). TechCrunch separately places Pattern Labs evaluation work in the documentation for Claude 3.7 Sonnet and OpenAI's o3 and o4-mini (TechCrunch). Any vendor register carrying both names is double-counting the field, as Irregular notes.

The numbers: $80M raised at a $450M valuation, announced 17 September 2025, led by Sequoia Capital and Redpoint Ventures with Wiz CEO Assaf Rappaport participating (TechCrunch; TechBuzz). Founded 2023 in Tel Aviv by Dan Lahav and Omer Nevo. Roughly 35 employees (BERI). Revenue is not disclosed; Sequoia says "already generating millions in revenue," which is a partner blog post, not an audited figure.

One vendor, four competitors' safety claims

Irregular is credited in the cyber sections of system cards across OpenAI, Anthropic, Meta and Google DeepMind. Sequoia describes the team as embedding "with research teams at Anthropic, OpenAI and Google DeepMind." Its CyScenarioBench appears by name in the Claude Opus 5 system card, and its SOLVE scoring framework is described as "widely used within the industry."

BERI's framing of what that means is the sharpest sentence written about this market: this is "not one vendor with a large market share, but one vendor whose methodology, benchmarks and infrastructure are load-bearing for multiple competitors' safety claims at once."

Dwell on that, because it is genuinely unusual. Four organisations that compete on capability, compete on safety reputation, and would not share a training corpus under any circumstances, have each bought their independent cyber assurance from the same thirty-five people. There is no equivalent in the atlas — not in Design and UI/UX, not in Law, not in Accounting, audit and tax.

Two readings are live and both are defensible.

The assurance-market reading. This is how audit, penetration testing and certification always end up: converging on a handful of accepted names, because convergence is what makes results comparable across firms. Once your framework is the shared reference, switching costs are collective rather than individual — nobody can leave alone without losing comparability. That is the strongest form of hold in the atlas, and it is not available to a vendor selling a commodity.

The single-point-of-failure reading. Thirty-five people cannot author enough hard, novel, expert-built environments to keep four labs' evaluations unsaturated. Anthropic's Opus 4.6 system card already states that the model "has saturated all of our current cyber evaluations" and that "the saturation of our evaluation infrastructure means we can no longer use current benchmarks to track capability progression." A shared vendor that cannot keep pace is a concentration risk that attracts either a regulator or an in-house replacement. See One customer is a binary event.

The headcount is the opening for anyone entering, and it is the only opening — not price, not quality, not relationships. Thirty-five.

They sell a service, and that is the finding

The most decision-relevant thing about Irregular is not the valuation. It is the shape of what they sell, because it settles the product-versus-service question that every other page in this section circles.

Irregular does not licence its evaluation set. It hosts it. Three statements from their own material make this unambiguous:

  • The platform "can directly connect to different components of AI systems, such as the underlying language model or the surrounding scaffolding / agent" — the customer plugs their model into Irregular's infrastructure, not the other way round (platform page).
  • The evaluation set "remains private to avoid contamination" (CyScenarioBench). Handing it over would destroy it. An evaluation that has leaked into the buyer's training data stops measuring anything.
  • Even at the top tier, scoring uses "automated checks and expert review" (FrontierCyber). The vendor with the strongest product story in the sector still keeps humans in the loop for open-ended work.
Contamination is the moat

Irregular's hold is the rarest structural property in the data business: the customer cannot take it in-house without destroying it. A licensed corpus is a one-time sale that erodes; a private, contamination-protected evaluation set is a subscription the buyer cannot escape without losing the measurement. That property is what a competitor should be trying to reproduce, not the benchmark itself.

The three suites underneath are worth naming because they show the ladder. Atomic Tasks are bounded technical challenges — cryptographic attacks, certificate forgery, browser exploitation, memory corruption, protocol manipulation, multi-host compromise. CyScenarioBench runs multi-stage attack scenarios in containerised network topologies with real operating systems, applications and security controls: complete web applications, authentication portals, internal wikis, corporate infrastructure, graded at task, path and campaign level. FrontierCyber is open-ended vulnerability research against physical devices — phones and routers — plus image libraries inside file-processing services, parsers behind ingestion pipelines, browsers rendering attacker-controlled pages, and databases behind application-like access patterns.

That last one is the tell. A vendor running an estate of physical phones and routers is not shipping a dataset. It is operating a laboratory, and the buyer is paying for the operation.

Six to nine months from first assessment to $450M

The timeline is the most transferable fact on this page. Irregular's first public model evaluation was Claude 3.7 Sonnet in early 2025; the $80M round closed in September 2025. That is roughly six to nine months from first published model assessment to a $450M valuation, inferred from publication dates on their research index rather than stated by the company.

The mechanism is visible on their research page: a continuous stream of model-by-model offensive-security assessments — GPT-5.4, GPT-5.5, GPT-5.6, GLM-5.2, Kimi K3, Muse Spark, Claude Sonnet 4.5, Claude 3.7 Sonnet. Publish an assessment of a frontier model, get cited in that lab's documentation, and the citation becomes the credential that sells the private engagement. The public benchmark is never the revenue line — it is what makes the private evaluation credible enough to sell.

They are also named as a UK AISI subcontractor, collaborating on advanced cyber task design alongside SpecterOps, Hack The Box and Crystal Peak Security (AISI). Government is a reference customer here as much as a revenue line.

What this proves and what it does not

Proves: labs will pay real money — millions, on the only revenue signal available — for human offensive expertise applied to pre-release frontier models. The demand is not hypothetical and the buyers are named.

Does not prove: that labs will pay for data at volume. Every observable feature of the Irregular engagement reads as high-touch consulting with deliverables that happen to be evaluations: embedded teams, private sets, expert review in the scoring loop, a physical device estate. The $450M valuation on "millions in revenue" is a bet on the category, not a validated multiple on a data business's unit economics.

The contrast with Gray Swan, in full is the cleanest structural comparison in this section: crowd versus payroll, licence-everything versus licence-nothing, $3.77 a unit versus an embedded team. Both are cited in frontier system cards. Both are valued in the hundreds of millions. Neither sells human offensive data as a product, which is the finding that runs through Offensive security and every page in this section.

What is not on the record

Revenue, ARR, gross margin, and any contract value with any of the four labs. Whether the lab relationships are recurring retainers or per-release engagements — the difference between an evals business and a consultancy. How Irregular compensates the people who author its scenarios, which is the entire supply question for anyone copying it. And the definitive "four labs" list: OpenAI, Anthropic and Google DeepMind are confirmed by name; Meta is asserted in secondary coverage rather than a primary source.