Miju Labs

The security dossier

Malware reverse engineering

The only security sub-market with a deterministic verifier demonstrated at 1,572-task scale and a proven cost of production a competitor cannot shortcut — 5,000 expert-hours for 262 binaries — held back by the weakest buyer evidence of the eight and a licence position worse than assumed.

buildhigh confidence8 minupdated 2026-08-30

Somebody paid for five thousand hours of expert reverse engineering and then published the result.

SRE-Bench is 262 binary instances derived from 19 programs averaging 16.9K lines of code, carrying 1,572 deterministically graded tasks, "built entirely from scratch by RE experts with over 5,000 hours" of work, including 44 in-house anti-analysis primitives to make the obfuscation realistic (arXiv 2608.11469). That is roughly 19 expert-hours per instance. At a senior RE contract rate of $150/hr it is about $2,850 per instance and ~$750,000 for the corpus [INFERENCE — the paper states hours, not money].

That number is the reason this page is ranked first. Every other sub-market in this dossier can be entered by someone with a mailing list and a spreadsheet. This one cannot. A competitor who wants to copy your corpus must find, hire and manage people who can build and defeat anti-analysis primitives, for a year, before they have anything to sell. Combine it with the fact that the frontier stalls at 31.5% full-instance completion and you have the rarest combination in expert data: a task machines cannot do, over an artefact machines can grade, with a floor under the price.

The problem is on the other side of the ledger. No frontier-lab job posting found anywhere in this research is titled or scoped for malware reverse engineering. That is the weakest buyer evidence of the eight sub-markets, and it is not a small caveat.

What the artefact is

The unit is a binary plus a set of questions whose answers are facts about that binary, graded automatically. Around it sit the pieces that are not automatically gradeable and are therefore the expert's real output: the unpacking chain, the capability inventory, the C2 extraction, the ATT&CK mapping, and — for RL — the disassembly-navigation trajectory itself.

Two properties make this shape unusual. First, freshly authored binaries are intrinsically exclusive. SRE-Bench exists because "instances must be unseen as source code" or models pattern-match instead of analysing. Contamination protection and commercial exclusivity are the same property here — exactly what The oracle decides everything found labs pay a 4–5x premium for.

Second, the published ground truth is wrong. TTPDetect recovers 85.7% of documented TTPs from real samples while discovering on average 10.5 previously unreported TTPs per sample (arXiv 2602.06325). Ten and a half missing TTPs per sample means vendor reports are incomplete ground truth. "Agreement with the published write-up" is dead as a scoring function, and the alternative — exhaustive fresh annotation — is expert labour you can sell.

Is there a verifier

Yes, for the fact-shaped half, and the split is commercially decisive.

TaskVerifierEvidence
Recover a flag, decrypt an input, identify the algorithmDeterministicCREBench: 432 challenges, 48 cryptographic algorithms, three difficulty levels, CTF-style flag recovery (arXiv 2604.03750)
Structured RE sub-tasks — function identification, control-flow claimsDeterministicSRE-Bench's 1,572 graded tasks
Unpack a packed sampleCheckable — does the unpacked artefact match?Malcat practitioner benchmark
Function-level TTP labellingConsensus against a gold setTTPDetect
"Describe this sample's capabilities and attribute the family"Expert judgement only

This is the strongest verifier position in the dossier and the reason defense scores 5. Where Vulnerability research's oracles run at 45% specificity and Threat intelligence has no oracle at all, malware RE's answer key either matches the binary or it does not.

The honest limit is the last row. Narrative capability write-ups — what a real analyst actually delivers — have no verifier, and that half prices as a service. Sell the graded half as a product and the narrative half as adjudication. See What you can actually sell.

Is anyone buying

Budget scores 2, and the schema is explicit that an argument a buyer should spend scores 2 rather than 4.

What exists is adjacent, not direct. Anthropic's nearest role is Offensive Hardware Security Engineer, Platform Security at $320,000–$405,000; OpenAI's nearest is Technical Threat Investigator, Threat Intel Engineering at $230K–$385K, which is CTI work, not binary analysis. CrowdStrike and Meta built CyberSOCEval's malware benchmark from real samples detonated in Falcon Sandbox — that is a vendor with its own sandbox building the artefact in-house rather than buying it.

And SRE-Bench itself proves somebody commissioned 5,000 expert-hours. Who funded it could not be established from the paper. That is the single most commercially important unanswered question on this page: it is either the best possible lead or the identity of an incumbent.

[UNVERIFIED — inference] The likely buyer is an EDR or sandbox vendor, or a government, rather than a frontier lab; the lab route probably runs through a general cyber evaluation suite rather than a malware-specific purchase. The labs as buyers and Government buyers set out what each has actually paid for.

What the expert costs

Cost scores 2. This is the second-most expensive pilot of the eight after Detection engineering, and the reasons compound.

SourceFigure
ZipRecruiter, Malware Analyst (Aug 2026)$86,474/yr · $41.57/hr; 25th $65,000 · 75th $100,500 · 90th $115,500
ZipRecruiter, Cyber Reverse Engineer (Aug 2026)$134,615/yr
Glassdoor, Malware Reverse Engineer [WEAK — 7 reports, most recent Feb 2024]Median total pay $181,051

The $41.57 is a trap. Three sources span $86k to $181k for nominally the same skill because "malware analyst" covers tier-1 sandbox-report readers while "reverse engineer" covers people who defeat 44 anti-analysis primitives. Price against $134k–$181k, i.e. $65–90/hr day-job equivalent, and expect $120–200/hr for contract work [UNVERIFIED — inferred from the offensive contractor bands in <a class="wikilink" href="/security/paying-the-crowd">Paying the crowd</a>].

Then add the licence bill. This is the finding that most changes the plan:

  • MalwareBazaar holds 1,128,069 confirmed samples, free under fair use, capped at 2,000 samples per IP per day — but the abuse.ch terms say use by "companies, organizations, individuals and networks with requirements likely to breach or exceed the fair use principles… may require a paid subscription service via Spamhaus, which is designed for users with commercial / for-profit requirements." A commercial data business needs a Spamhaus licence. The terms are silent on redistribution and on AI training, and the silence is itself the risk.
  • VirusShare is invitation-only, prohibits scripted downloads, and states no licence terms for the samples at all.
  • vx-underground returned HTTP 403; its terms could not be established.
  • VirusTotal redistribution terms remain unresolved.

The workaround is the same fact that makes the moat: SRE-Bench proves you do not need the public corpus. Authoring binaries from scratch fixes the contamination problem and the licence problem in one move — and costs 19 hours an instance. See The terms of service bite first.

Getting to them

Reach scores 3, the honest score for a discipline whose communities publish nothing.

The one precise instrument is the GIAC Certification Holder Directory — 290,000+ certifications issued across 60+ types, searchable, with GREM as the qualifying credential. Per-certification holder counts are not published, so the directory gives you names without a denominator.

Everything else is soft: r/ReverseEngineering at ~100K and r/Malware at ~80K [WEAK — single aggregator]; Recon, OffensiveCon and Objective by the Sea, none of which publish attendance; and the Ghidra, Binary Ninja, IDA and Malcat tool communities, real and named but counting nothing. Contrast Detection engineering, where three channels publish exact subscriber numbers.

Where the benchmarks sit

BenchmarkScaleFrontier score
SRE-Bench (Aug 2026)262 binaries, 19 programs, 1,572 graded tasks, 44 anti-analysis primitivesGPT-5.6-sol 61.4% per instance; 31.5% full-instance completion across five frontier models
CREBench (Apr 2026)432 challenges, 48 crypto algorithmsGPT-5.4 64.03/100, flag in 59% of challenges; human expert baseline 92.19
CyberSOCEval609 questions, Falcon Sandbox detonations23–34%
Malcat practitioner benchmark9 triage + 5 packed samples15.2/20 triage; 11.8/20 static unpacking

Four independent methodologies agree, which is why proof scores 4 — publish against a public scoreboard everyone is losing on and you are visibly the best source inside one paper cycle. SRE-Bench's own conclusion is the cleanest statement of the opportunity anywhere in this dossier: "strong source-code security capabilities do not yet transfer to binary analysis." More at The measurement gap.

What would kill it

Three ways this dies

The buyer never materialises. Budget 2 is not a rounding error. If the purchaser is an EDR vendor with its own sandbox and its own analysts, they build rather than buy — CrowdStrike and Meta already did.

Someone funds a second SRE-Bench and open-sources it. The moat is 5,000 hours of someone's money. It has been spent once already and the funder is unknown. A well-capitalised lab repeating it and publishing kills the cost-of-production argument overnight.

The licence trap fires late. Building on MalwareBazaar at scale without a Spamhaus subscription, or on VirusShare with no stated sample licence, produces a corpus a buyer's counsel will refuse in diligence — after you have paid for it.

Room scores 4 rather than 5 for the second of those: the position looks empty, but 5,000 expert-hours have already been spent in it by a party this research could not name.

Where the record is thin

The funder of SRE-Bench is the biggest hole and the most actionable one. Second: no live AI-data listing scoped to reverse engineering was found in this research, so the $120–200/hr contract rate is inference from adjacent offensive bands, not an observed price — see Paying the crowd. Third: GREM holder counts, Recon and OffensiveCon attendance, and vx-underground and VirusTotal redistribution terms are all unestablished.

No observed transaction, anywhere

Every other page in this dossier can point at something: a posting with a band, an invoice in a transparency file, a named vendor in a system card. Malware RE has none of those. It has the best artefact and the worst commercial record, and those two facts have to be held at the same time. Compare the ranked read in The security read and the single surviving shape in The one shape that survives.