The headline number is not a typo. On the Cyber Defense Benchmark — agentic threat hunting over Windows event logs, ground truth derived from Sigma rules, scored CTF-style — Claude Opus 4.6 produced correct flags 3.8% of the time, and no model passed the ≥50%-recall bar on more than 5 of 13 ATT&CK tactics (arXiv 2604.19533). The other four models tested cleared it on zero.
The same generation of models solves 8 of 10 end-to-end attempts on a UK AISI cyber range. The offensive/defensive gap is not rhetoric; it is roughly twenty-fold between two benchmarks scored the same way. The measurement gap has the full curve.
Detection engineering is the sub-market where that gap is cheapest to attack, because the answer key is mechanical and — uniquely in security — you are allowed to sell the corpus.
What the artefact is
Not the rule. The trainable unit is the triple (behaviour description or telemetry, candidate rule, evaluation against a labelled corpus), plus the analyst's reasoning about which field to key on and where the false-positive surface lies. Finished rules are free. The process around them is not.
The licence map is why this leads the dossier on legal cleanliness, and it is more differentiated than earlier research recorded:
| Corpus | Licence | Sell a derivative dataset? |
|---|---|---|
| SigmaHQ/sigma — 3,000+ rules | DRL-1.1; the Sigma specification is public domain | Yes — DRL-1.1 grants rights to "use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies", subject to author attribution, a URI back to the rule set and retention of the licence text |
| Neo23x0/signature-base (YARA) | DRL-1.1 | Yes, same terms |
| Yara-Rules/rules | GPL-2.0 | No. Copyleft. Keep it out. |
| splunk/security_content | Apache-2.0 | Yes |
| elastic/detection-rules | Elastic License 2.0 — "You may not provide the software to third parties as a hosted or managed service" | Probably yes for a dataset sale; no as a SaaS detection service |
| OTRF/Security-Datasets — labelled attack telemetry | MIT | Yes |
| redcanaryco/atomic-red-team | MIT | Yes |
| MITRE ATT&CK | Royalty-free, commercial use permitted | Yes, with MITRE's mandated copyright designation |
Say the GPL line out loud, because it is a real and previously unrecorded hazard. Yara-Rules/rules is one of the most-cloned YARA corpora in the world and it is GPL-2.0. Mixing it into a commercial dataset contaminates the derivative. A junior engineer assembling "all the public rules" will pull it in by default and nobody will notice until diligence. Keep it in a separate tree, or out of the building. The terms of service bite first has the consolidated map.
Is there a verifier
Yes, and it is the strongest mechanical verifier in defensive security. Three legs, all automatable: the rule compiles; it fires on labelled attack telemetry; it stays silent on labelled benign telemetry. All three have permissively licensed public corpora behind them.
Three independent 2026 papers demonstrate it working:
- The Cyber Defense Benchmark scores agents against Sigma-derived ground truth over 106 real attack procedures from OTRF Security-Datasets, spanning 86 ATT&CK sub-techniques across 12 tactics, in episodes of 75,000–135,000 log records generated by a deterministic campaign simulator that time-shifts and entity-obfuscates the raw recordings (arXiv 2604.19533).
- A deterministic BAS→SIEM synthesis pipeline emits starter Sigma rules with byte-stable traceback: all 17 of 17 emitted rules parse and convert to Splunk and Elasticsearch backends, and replayed through a live OpenSearch SIEM the LLM-category rules fire on 30% of a held-out AdvBench subset and 14% of HarmBench at 7.7% false positives (arXiv 2606.05252).
- Graph-Based Structural Evaluation grades LLM-translated emulation procedures against 49 validated Sigma rules (19 Linux, 30 Windows) using normalised graph edit distance across technique, tactic, telemetry-class and logsource layers. A 29-step ALPHV/BlackCat plan scored composite 0.674 against a 0.80 deployment threshold, failing on technical realism at 0.43 against a required 0.990 (arXiv 2607.11517).
Read the third one carefully, because it is the business. Even where a verifier exists, somebody has to author the rubric layers. The verifier makes the product cheap to grade; it does not make it cheap to create.
The honest limit is the benign leg. Public benign corpora are too clean — AuditBench scored F1 1.00 on lab data and 0.45 and 0.25 on real OpTC data. A false-positive surface measured against tidy telemetry is not a false-positive surface.
Is anyone buying
Budget scores 4. There is a named buyer building precisely this product, with a band.
OpenAI's Cyber Blue Team is "an operator-led group focused on turning real defensive problems into better models", and its stated scope is to help security teams "investigate threats, create and validate detections, improve their controls, and respond with greater speed and confidence" — while explicitly disclaiming any intent to build "another SIEM or autonomous SOC" (Product Manager, Cyber Defense and Blue Team, $293K–$385K). A lab building a detection-authoring product needs detection-authoring data.
The same discipline is staffed internally at both anchor buyers: Anthropic's Security Engineer, Detection & Response ($300,000–$405,000) and Incident Manager, Detection & Response ($290,000–$365,000); OpenAI's Security Engineer, Detection and Response at $293K–$385K in San Francisco with parallel London and Sydney roles.
It is 4 and not 5 because of a specific gap: no lab has been shown buying a detection dataset. The posting proves intent to build the product. It does not prove they will not generate the data from their own SOC. That single unknown decides whether this is a business — see The labs as buyers.
What the expert costs
Cost scores 2, and this is the finding that decides the page.
| Source | Figure |
|---|---|
| ZipRecruiter, Detection Engineer (29 Aug 2026) | $156,399/yr · $75.19/hr; 25th $143,000 · 75th $172,500 · 90th $182,000 |
| Mercor, Cyber Security Experts — evaluates "alert quality, detection rules, triage decisions" (cohort filled) | $85–95/hr (listing) |
| Anthropic FTE comparator | $300,000–$405,000 |
A detection engineer's day-job hourly is $75.19. Mercor's $85–95 is a 1.1–1.3x premium. Compare Threat intelligence at 1.8x and Governance, risk and compliance at up to 2.3x. Detection engineers are the best-paid practitioners in this dossier relative to the going AI-data rate, which makes them the hardest group of the eight to recruit — and any plan using one blended security rate is wrong by a factor of about 1.6. Paying the crowd works the ladder in full.
The infrastructure is cheap by contrast: MIT-licensed attack telemetry, a replay harness and a SIEM. The money is people.
Getting to them
Reach scores 5 — the best-evidenced 5 in the dossier, because three separate channels publish counts.
- Detection Engineering Weekly — "over 16,000 subscribers" (checked 30 August 2026), a 3.1x increase on the 5,200 recorded in earlier research. The most precisely targeted channel to this sub-market anywhere.
- Blue Team Village — 10,000+ Discord members, ninth year at DEF CON, DEF CON 34 on 7–9 August 2026, six content tracks, a 501(c)(3). The first verified defensive Discord count in this research programme.
- tl;dr sec — 90,000+, with the standing complication that its creator now leads OpenAI's cyber team.
- SigmaHQ/sigma — 10.9k stars, 2.7k forks, 3,000+ rules. Every merged contributor is a person who has had a working detection rule pass QA. The exact contributor count could not be established.
- GIAC — 290,000+ certifications across 60+ types with a searchable holder directory; GCDA and GCIA are the credentials.
One cheap trick: Detection Engineering Weekly is a Gold sponsor of Blue Team Village. The two highest-precision channels are already connected, so a combined sponsorship reaches both. See Defensive supply.
Where the benchmarks sit
| Benchmark | Scale | Frontier score |
|---|---|---|
| Cyber Defense Benchmark (Apr 2026) | 106 procedures, 86 sub-techniques, 26 campaigns | Claude Opus 4.6: 3.8% correct flags; ≥50% recall on 5/13 tactics, four other models on zero |
| CTI-REALM | 37 CTI reports + detection references | 0.637 (Claude Opus 4.6) |
| AUTOSIGMA (Aug 2026) | Real APT reports and security blogs | Quantitative results could not be established |
| BAS→SIEM synthesis | 23-template library, live OpenSearch replay | 17/17 rules parse; 30% AdvBench, 14% HarmBench, 7.7% FP |
Proof scores 4: a public scoreboard at 3.8%, permissive corpora, and an evaluation harness a small team can stand up in weeks. You can be visibly the best source here faster than in any other sub-market.
What would kill it
The labour arithmetic. A 1.1–1.3x premium does not move a well-paid engineer with a day job. If the cohort has to be paid $130–150/hr to actually show up, the margin structure of the whole product changes. This is the live risk, not a theoretical one.
The lab generates its own. OpenAI's Cyber Blue Team sits inside a company with a large SOC. Detections authored against their own telemetry are free, current and uncontaminated.
SOC Prime already pays for rules — and pays badly. It runs the only established per-artefact market for security expert output: monthly royalties scaled to rule popularity, Sigma and Roota only, individuals only, no companies, authorship retained, submissions running "from 5 to 50+ per week" for active authors (SOC Prime). No dollar figure exists publicly because a usage royalty has no per-rule price by construction. A fixed, disclosed per-artefact rate would differentiate against it — but it also proves the supply is already being farmed by someone else.
Room scores 3 and defense scores 3: the rules are public and getting more public, the licence advantage is available to everyone who reads it, and the durable asset is the labelled telemetry plus the rubric layers, not the rules themselves.
Where the record is thin
No lab has been observed buying a detection dataset — the gap that decides whether this is a product line or a thesis. AUTOSIGMA's numbers are unretrieved. SigmaHQ and Elastic contributor counts are unavailable (GitHub's API is scoped off; the contributors graph is robots-disallowed), leaving fork counts as the only proxy. GIAC publishes no per-certification totals, and SOC Prime's actual payouts exist only in its contributor Discord.
Compare the two-bets-one-product argument in The security read and the sibling case at Malware reverse engineering — a YARA rule is the shared output of both, and DRL-1.1 covers it.