Co-founder and CEO of Intelligence, the company behind Design Arena, the largest crowdsourced preference arena for AI-generated design. She claims $60M ARR selling arena data and evaluations to frontier labs, which makes her the commercial proof point, and the main counter-thesis, for Julian's lane: consumer-crowd taste instead of expert panels.
high confidenceSan FranciscoAt Design Arena (Intelligence, formerly Arcada Labs) since 2025Updated 2026-09-19
Background
Studied Computer Science and Neuroscience at Harvard and worked at Apple before founding (per YC). Index Ventures says she and Kamryn Ohly were best friends from freshman year who built a game engine at a hackathon and noticed that AI-generated visuals felt 'clunky and visually off'. TechCrunch reports they started the company a few weeks before graduating in 2025.
Design Arena launched on YC on 31 Jul 2025 (47K users in four weeks, per the launch post). By Feb 2026 she described Arcada Labs as building real-world evals for hard-to-benchmark things: design, prediction markets and social posts.
What they run now
She runs the company and its lab sales: TechCrunch quotes her saying the first major frontier-lab deal closed about a week after they recognised the demand. She is also the public voice of the leaderboard, writing most posts on notes.designarena.ai about which model wins or regresses at design (Opus 4.8, GLM-5.2, Kimi K3, Thinking Machines' open-weights model). She is widening the brand from design to a general 'realistic evals' company (Prediction Arena, audio realism benchmark).
Argued Claude Opus 4.8 regressed at single-turn web design (23rd on Web Dev non-agentic) because it was over-optimised for multi-turn agentic work. source
Real-world evals · 2026-02
Frames the company as building 'portals' that bridge AI to the real world through evals for hard-to-benchmark domains. source
Recent posts
36 posts archived · most engaged first, then the latest
WE COULD NOT BELIEVE IT WHEN WE SAW THIS
We were curious about Kimi K3's massive jump in consumed thinking tokens
After taking a closer look, we found that Kimi K3 devotes 10x more thinking tokens to coding during reasoning than any other model we've seen before
Not only that, but it has basically memorized all of the image IDs from Unsplash and, without a web search tool call, can reason about images with a higher accuracy than we've seen from any other model
The image and video below are visualizations of its (very long) thinking trace for single-turn generation
The resulting website selects images with greater relevance and accuracy than we've ever seen on DesignArena
Forgive my all-caps heading, but I am excited about this and wanted you to read it too bc this is pretty gnarly
Huge congratulations to the @Kimi_Moonshot team for this breakthrough and becoming the 1st-place model on @DesignArena!
How did @OpenAI Sol finally learn design taste?
We projected 1,000 websites by GPT-5.6 Sol into a design manifold... and discovered big holes
These holes are where GPT-5.5 previously generated outputs with bad AI "smell"
We've passed 5 million users!!
It's crazy to see the compounding value of focusing on one goal...
To give every human superintelligence and make superintelligence more human
We could not do any of it without your support. Thank you so much for being there since day 1.
We're just getting started :)
Accounts whose posts Grace has reposted recently: Prashant Shishodia, fal.
Why it matters here
She holds the one hard revenue number in design-preference data ($60M ARR claimed) and knows what frontier labs actually buy: contract shape, per-vote versus licence pricing, and turnaround. Her model is the opposite of Julian's. It uses a free consumer crowd rather than paid experts, so any pitch Julian makes to labs has to explain why expert judgement beats millions of arena votes, for example on professional standards, rationales and trajectories. She is also a possible partner or buyer if labs start asking for expert-only cuts of arena data.
How to reach
Publicly active on X (@grx_xce) and posts analysis on notes.designarena.ai; the site lists founders@ contact addresses. Best angle (speculative): offer something the arena structurally lacks, such as a vetted expert panel for validating arena results or an EU/German designer cohort, rather than pitching as a rival. A YC or Conviction mutual would be the warm path.
What we could not establish
Whether $60M ARR is gross or net, and who pays for the free model inference
Her Apple role and dates
Exact date of the rename from Arcada Labs to Intelligence
Whether expert or professional voters are segmented and priced differently