Contra's smartest idea is that a creative evaluation should be stratified by workflow stage rather than run flat. The popular version of the finding is the eye-catching half and the weaker half. The durable half is being under-sold by the people who discovered it.
The claim as it is usually told
Model rankings invert between ideation, mockup and refinement, which turns a leaderboard into a diagnostic. Publicly: Claude leads ideation on landing pages, Gemini dominates mockup at a 68.9% win rate, Claude reclaims refinement; Veo 3.1's 61% ideation win rate falls to 39% at refinement (contralabs.com/research).
It is a good story. A buyer who learns that the model they use for concepting is the wrong model for polish has learned something they can act on this week, and it reframes evaluation from a ranking exercise into a routing decision.
It is also the part that will not survive.
The inversion rests on 28 evaluators across 80 sessions, 93 prompts, four models per domain and 380 outputs (arXiv 2606.30561). The authors list "small evaluator pool limiting statistical power" and "limited prompt sampling without multi-iteration testing" among their own limitations. A 68.9%-versus-39% swing on a sample that size, without correction for multiple comparisons across five domains × three phases × three axes, is suggestive rather than established [WEAK].
And there is a second problem that no amount of sample size fixes. Inversions are model-specific, so they expire with each release. A diagnostic whose content changes every eight weeks is a subscription — which is commercially good — but it is not a finding you own. Anyone can re-run it on the next model generation, and the result will be different for reasons that have nothing to do with your method.
The half that generalises
On HCB ad images, Kendall's W runs 0.345 (ideation) → 0.436 (mockup) → 0.549 (refinement) (arXiv 2606.30561). The authors' own gloss is that criteria become "progressively more verifiable" as work advances.
This is a completely different kind of claim from the inversion, and it is much stronger.
It is a claim about creative work, not about any model. Rankings between models will be different next quarter; the fact that professionals agree more about a finished comp than about a napkin sketch will not be. It also survives the sample-size objection far better, because it is a within-study monotone trend on a single construct rather than a between-model difference — three points in the predicted order, on the same raters, on the same material, at successive stages.
And it is intuitive in a way that makes it saleable without a statistics argument. Early-phase work is under-determined: there are many good answers and professionals legitimately split between them. Late-phase work is constrained by everything already decided, so the remaining questions are about execution, which is the part of the discipline with shared standards. That is the same divide the aesthetics literature draws between shared and individual taste components — Vessel et al. (2018) found high shared taste for faces and landscapes and strong individual variation for architecture and artworks (Max Planck Neuroscience) — reappearing within a single piece of work as it matures.
Why the agreement gradient is worth more than the inversion
Because it is a pricing schedule, and the inversion is only a headline.
If agreement is a function of phase, then the cost per unit of reliable signal is a function of phase. Reliability of a panel mean rises with panel size; the number of raters you need to hit a target reliability scales inversely with single-rater agreement. So:
| Phase | Kendall's W (ad images) | Panel depth needed | What you sell |
|---|---|---|---|
| Ideation | 0.345 | Deepest — most raters per item | Divergence maps, concept diversity |
| Mockup | 0.436 | Middle | Standard multi-axis scoring |
| Refinement | 0.549 | Thinnest — fewest raters per item | Cheap high-confidence labels, craft signal |
Stack that on the other reliability gradient — agreement is highest on prompt adherence and lowest on visual appeal (arXiv 2606.30561) — and you get a two-dimensional capacity model: panel depth varies by axis and by phase. A vendor running a uniform panel across every axis and every phase overspends on prompt adherence at refinement and underspends on visual appeal at ideation, and cannot tell you by how much.
Publishing per-axis, per-phase reliability and varying panel depth accordingly is the thing that makes a vendor look like an instrument maker rather than a labour broker. It is the only claim in the artefact taxonomy that is simultaneously a methodological contribution, a cost advantage and a sales argument — and nobody has replicated it, which means the replication is available to whoever runs it first.
The pricing consequence is concrete. Phase-decomposed evaluation costs roughly 20–40% more than flat scoring in design and analysis overhead. If the agreement gradient lets you cut refinement panels from, say, five raters to three while deepening ideation panels from five to nine, the overhead pays for itself in the axes where it matters and the buyer gets tighter confidence intervals where their decision actually turns.
Where the decomposition maps, and where it does not
The ideation → mockup → refinement split is not universal. It maps cleanly onto:
- Copywriting — concept → draft → line edit
- Video — treatment → rough cut → grade
- Architecture — massing → scheme → detail
- Software UI — wireframe → comp → build
It maps poorly onto domains where the artefact is atomic: a single logo mark, a photograph, a typeface glyph. There is no ideation phase you can hand a panel when the deliverable is one image that either works or does not — you can decompose the process, but you cannot decompose the output into stages a rater can score separately.
It maps very well onto agentic work generally, because agent trajectories are already phase-structured — which is the connection to The trajectory moat and the reason the two products belong in the same pitch. [WEAK — this mapping is my analysis, not a sourced claim]
The honest position
No independent replication of phase-stratified creative evaluation exists anywhere. I looked; there is nothing. That absence cuts both ways: it means the finding is unconfirmed, and it means the confirmation is unclaimed.
The right posture is to sell the gradient and use the inversion. The gradient is what you put in the methodology paper, publish per-axis and per-phase, and build the pricing schedule on. The inversion is what you put in the quarterly diagnostic that a buyer renews, precisely because it expires — a routing recommendation that changes with each model release is a reason to keep paying, so long as nobody mistakes it for a law.
The failure mode is the reverse: leading with the inversion as a general truth, having a buyer's own team fail to reproduce it on next quarter's models, and losing the credibility of the gradient along with it. Contra has published both halves in the same paper and led with the wrong one. That is the seam.
The agreement gradient has been measured once, on one domain within one study — Kendall's W on ad images, 28 evaluators. Whether it holds at the same slope in typography, motion, photography or industrial design is unknown, and the slope is what sets the panel-depth ratios in the table above.
The panel-size question underneath it is unanswered for professional design evaluation generally — how many raters per axis, per phase, for a target reliability. That is the most valuable unclosed gap on the product side, it is cheap to close by re-analysing your own pilot data, and it is treated at length in The oracle problem.