Miju Labs

The security dossier

Copyright is the weakest thing you own

A rating is not a work in the EU — no free creative choice — so contract, trade secret and the database right do all the work. The decided cases point the same way: Bartz priced acquisition at about $3,000 a work while calling training itself exceedingly transformative, which means provenance rather than use is what generates buyer liability. And two European courts gave opposite answers on whether weights reproduce works within five days of each other.

medium confidence10 minupdated 2026-08-30copyright · bartz · getty · gema · fair use · tdm · litigation

Everyone reaching for legal protection over a taste corpus reaches for copyright first. It is the weakest of the four protections available and it does not reach the thing that is valuable.

Take the units one at a time, because the answers differ

A numeric rating or a pairwise preference is almost certainly not protected in the EU. Originality means the author's own intellectual creation reflecting free and creative choices (Infopaq, C-5/08; Cofemel, C-683/17 — IPKat on Cofemel). Choosing 4 rather than 5 on a five-point scale involves no free creative choice. A scalar rating is not a work.

The US answer is more interesting and still not usable. In CDN Inc. v. Kapes, 197 F.3d 1256 (9th Cir. 1999), wholesale coin prices were copyrightable, because they were "not mere listings of actual prices paid" but the publisher's "best estimate of the fair value," and that expert estimation "imbues the prices listed with sufficient creativity and originality" (FindLaw). In New York Mercantile Exchange v. IntercontinentalExchange (2d Cir. 2007), settlement prices were not protected: merger applied because "all possible expression takes the same form, a number," and the Second Circuit distinguished CDN on the ground that coin prices "were estimates, not discovered market facts" (FindLaw).

A design quality score is an expert estimate, not a discovered market fact, so it sits on the CDN side. That is a real argument. It is also a thin copyright that merger will be pressed hard against, in one jurisdiction, on a point two circuits disagree about. Treat it as a fallback, never as a foundation.

A written critique is protected in both systems. Two hundred words of design reasoning reflects free and creative choices in selection, emphasis and expression; Infopaq established that eleven words can qualify. This is the unit where you genuinely need a licence from the contributor, and it is why the The clause that expires your corpus in year ten has to enumerate uses rather than rely on general words.

A rubric is protected, and it is the best copyright in the corpus. A multi-axis rubric with anchored scale points, definitions and worked exemplars is a literary work by any standard — and it is authored by you or by staff whose contracts you control, not by the contributor panel. It is the one piece of IP you can own cleanly without navigating §29 UrhG or L131-3 CPI. Its version history and calibration data should also be a trade secret under Directive (EU) 2016/943 (EUR-Lex), which is the protection that actually holds.

A leaderboard as a whole may attract thin compilation copyright — Article 3 of Directive 96/9/EC protects a database whose "selection or arrangement of its contents constitute[s] the author's own intellectual creation" (EUR-Lex) — which protects the structure and not the data, and is close to worthless against a competitor who re-derives the same rankings.

What actually protects the corpus

Contract, trade secret, and — for a European operator only — the sui generis database right. In that order of practical importance. Copyright is fourth and it is the one everyone reaches for first. The corpus's value lies in facts about expert opinion, in scale, in structure and in the reliability record, and copyright reaches none of them.

What has actually been decided, and the answer is: less than the discourse implies

On the central question — is training on lawfully obtained copyrighted material fair use or within a TDM exception — no appellate court on either continent has given a final answer, and the two European courts that came closest went in opposite directions within five days.

Bartz priced acquisition, not training

Final approval of the $1.5 billion settlement in Bartz v. Anthropic (N.D. Cal.) was granted on 20 July 2026 by Judge Araceli Martínez-Olguín, at approximately $3,000 per work — four times the $750 statutory minimum. Anthropic must destroy all original files torrented or downloaded from Library Genesis and the Pirate Library Mirror, and represented that neither dataset nor any portion was in the training corpus of any commercially released model. The release covers acquisition and copying claims only, through 25 August 2025 (Authors Guild; JURIST).

Judge Alsup's June 2025 order had found training itself "exceedingly transformative," while treating the retention of a pirated library as a separate and unexcused act.

The most commercially relevant fact in US AI copyright

The money was for acquisition, not for use. A buyer's exposure runs on where the data came from, not on what was done with it — and that is precisely the exposure a documented, consented, commissioned corpus eliminates. This is the argument that converts the whole business from a data sale into a risk transfer, and it is worth more to a buyer's general counsel than any quality claim you can make about the data itself.

The five-day European split

England, 6 November 2025. Getty Images v Stability AI, [2025] EWHC 2863 (Ch). Getty abandoned its training claim mid-trial because there was no evidence the acts occurred in the UK — Stability trained overseas on AWS — and dropped the output-reproduction claim. On secondary infringement the court held that an "article" need not be tangible, so cloud-stored copies count, but that model weights are not "infringing copies" because they do not store the works: "by the end of that process they did not store any of those copyright works." Trade mark succeeded only narrowly, on watermarks appearing in outputs under ss.10(1) and 10(2) TMA 1994; the s.10(3) dilution claim failed (Mishcon de Reya; BAILII).

Germany, 11 November 2025. GEMA v OpenAI, LG München I, Case 42 O 14139/24. The court held that memorisation of song lyrics in model parameters is a reproduction, split training into phases, and applied the TDM exception only to the initial data-gathering phase — not to the embedding of content in parameters (EUIPO case note; Bird & Bird; Library of Congress).

Weights do not contain the works. Weights do contain the works. Two European jurisdictions, five days apart.

Two lessons for a vendor. Territoriality decided Getty, which means where training happens is a litigable fact under the buyer's control — expect buyers to push jurisdictional representations down onto suppliers, and expect to be asked where your contributors worked. And "the weights do not contain the works" is now a reasoned English-law holding, which cuts against any theory that a buyer's model is a derivative of your corpus. You are protected by the contract you wrote, not by residue in someone else's weights.

The German direction of travel on the input side went the other way in December 2025: the Hanseatic Higher Regional Court dismissed the appeal in Kneschke v. LAION, holding the use fell within the Article 4 / §44b commercial TDM exception and that a rights reservation in natural language is not machine-readable — the standard being that the reservation must be interpretable by machine such that an automated process does not process the reserved content (Grünecker; Kluwer Copyright Blog). [WEAK] on the appeal case number — the source carries the first-instance reference.

That matters defensively, not offensively. Article 4 is irrelevant to a buyer's use of a licensed corpus — where the data is commissioned and licensed, the contract governs and the TDM exceptions never enter the analysis. But it is entirely relevant to your own published tranche: if you publish benchmarks as marketing, a competitor can mine them freely unless your reservation is machine-actionable (robots.txt, a TDM reservation protocol file, HTTP headers, per-asset metadata) rather than a sentence in your terms of use.

The two cases still running

Thomson Reuters v. Ross Intelligence (3d Cir., No. 25-2153) is the first appellate word on AI training and fair use. Judge Bibas granted summary judgment in February 2025 on 2,834 Westlaw headnotes, finding roughly 2,430 protectable and infringed and rejecting fair use; the Third Circuit heard argument on 11 June 2026 and pressed both sides on market definition and on whether non-generative use is transformative. No decision as of 30 August 2026 (LawSites; Baker Botts). Thomson Reuters ran three theories of harm, and one of them — lost licensing to other AI companies — is directly favourable to a licensed-data vendor. If lost licensing revenue counts as market harm, the licensed market it presupposes is this market. Have a note prepared for the day it lands.

Andersen v. Stability AI, 3:23-cv-00201 (N.D. Cal.), survived two rounds of dismissal on direct infringement; a third amended complaint was filed 27 February 2026 with answers on 13 March 2026, and no class certification ruling or trial date is locatable (Mesh IP tracker; CourtListener). Three and a half years in with no merits ruling is the honest measure of how long this uncertainty persists — plan on it outlasting the company's first two funding rounds.

Two procedural facts worth carrying. In Concord Music v. Anthropic, a 71-page second amended complaint filed 22 July 2026 alleges that extraction tools stripped copyright notices, which maps onto a §1202 copyright-management-information claim (Music Business Worldwide). Stripping metadata from assets you ingest is an independent statutory wrong and is trivially avoidable — preserve CMI on every reference asset that enters a brief or reference set. And in NYT v. OpenAI, Judge Stein's court ordered production of 20 million de-identified ChatGPT logs (AI Lawsuit Tracker). Assume everything you send a lab is discoverable in litigation against that lab, and write delivery records accordingly.

Part 1 (digital replicas, July 2024) and Part 2 (copyrightability, January 2025) are final; Part 2 holds that purely AI-generated output lacking human authorship is not copyrightable while human-authored contributions to AI-assisted works are. Part 3, "Generative AI Training," was released as a pre-publication version on 9 May 2025 and Register Shira Perlmutter was removed from office days later (Center for Art Law; pre-publication PDF). Its substance: no categorical exemption for training, prima facie infringement requiring four-factor analysis, transformativeness assessed against outputs, and a novel "market dilution" theory that stylistic imitation at scale devalues human creativity without verbatim copying.

[UNVERIFIED] — a final Part 3 could not be confirmed to exist. Cite it as an indication of direction, never as settled guidance, and note that its author's removal days after publication materially reduces the weight a court is likely to give it.

What nobody has litigated

No case anywhere has been decided on a commissioned expert-judgement corpus. Every reading above is adjacent doctrine applied to a fact pattern no court has seen — which is the same position the security dossier's law page reaches from the other end of the market, and the reason both conclude that contract quality matters more than doctrine.