Ex-Capital One Labs product/research lead who co-founded Halluminate (YC S25), first a computer-use agent benchmark and simulation company, now selling finance RL environments built by ex-bankers and deal professionals to frontier model developers.
high confidenceSan FranciscoAt Halluminate since 2024Updated 2026-09-19
Background
Wu studied computer science and economics at Cornell (VP of Cornell Consulting Group; researched model quantization). At Capital One Labs he led product and research and launched one of the early AI agents in financial services, co-authoring three patents (YC, AI Engineer).
He founded Halluminate in 2024 with Wyatt Marshall, whom he met in freshman week at Cornell. The company's first public work was browser-agent evaluation: WebBench (5,750 tasks across 452 websites) and the 'Westworld' simulated environments, presented at AI Engineer World's Fair 2025. By 2026 the positioning had narrowed to investment banking, PE and consulting.
What they run now
Selling finance RL environments and expert-authored tasks (modeling, underwriting, quality of earnings, pitch materials, diligence) and building the contractor pipeline of finance analysts/associates who author them.
Career
2024 – presentCo-founder & CEO, Halluminate
—Led product and research, Capital One Labs — launched an early financial-services AI agent; 3 patents
—Cornell University, Computer Science & Economics
On the record
Agents need environments
The bottleneck for agents is not model intelligence but reliable training environments mirroring real software. source
Computer use
Knowledge workers spend ~90% of their day on a computer, so agents must learn to use real software. source
Read vs act · 2025
Distinguishes information retrieval from state-changing workflows; in production, auth, infra, latency and unpredictable agent behaviour decide success. source
Recent posts
34 posts archived · most engaged first, then the latest
Come train financial superintelligence with us!
Financial knowledge work is going to completely change in the coming few months. We're looking for talented bankers, investors, traders, consultants, and operators to come push this frontier with us.
- Pay: $100 - $250/hour
- Fully remote and flexible work starting at 20 hours/week
- Sponsored events and perks
If you're looking for an opportunity to work with a high-quality network of like-minded peers this summer, apply below!
The bar to build great RL envs is growing rapidly.
A great environment needs to:
- be safe and aligned
- challenge and expose realistic error modes
- have no anonymization or data leakage
- verification that doesn’t have false positives or negatives
- well engineered tools
- delivered with great research taste and service
- and more
Really hard and fun challenges. Join us if you’re interested in solving these problems in the future!
When building RL Environments the natural order of prioritization with app/tool building is:
1. Great feature rich open source app
2. Licensable and hosted closed source app (ex MSFT Office)
3. Cloned version of closed source apps
80% of the time we never make it to 2 or 3 because a great feature rich open source app has many of the capabilities needed to teach generalized understanding for RL (ex LibreOffice vs Excel).
In a world of background agents doing work, the user doesn’t care what app (closed vs open) is being used for the action space. What they care about is that the final output is economically valuable. I fully expect to see this trend continue for many software verticals (ex. Medplum for EHRs, LibreOffice for Office Suite).
Huge and unintended win for open source apps! More companies should lean in and actively make their software RL-consumable for this reason.
Jev (Noul) replacing LLM as a Judge functions in our Data Quality Pipelines for RL Environment creation.
Still collecting results on overall alignment compared to prior methods but overall seems quite strong (and blazing fast). Easy to see the potential gains this primitive can have in the entire data industry.
Hats off @CompleteSkeptic & team! The copy and paste instructions for Claude Code in the docs were also quite neat.
Although there may be some value in training on alignment specific RL Envs, my intuition is that the vast majority of "alignment" gains will come from hardening and increasing QA/QC across all RL Envs horizontally. Catching things like:
- Impossible Tasks (ex. CyberGym Attack)
- Escapable Sandboxes
- Unintended code execution
- Broken verifiers / non-comprehensive verification
Will do more in aggregate than specifically trying to train on alignment RL Envs. Its like plugging one hole while many others exist.
To put simply: this is a horizontal problem across the whole data/RL Env industry that every player needs to take super seriously. We all have a responsibility to play.
Accounts whose posts Jerry has reposted recently: Jun Liang LEE.
Why it matters here
Halluminate is a near-template for what Julian would build in design: a small team that turned generic agent environments into a vertical, expert-authored product (ex-bankers writing tasks, ground truth and weighted verifiers). Wu's pivot shows how a vertical specialist positions against generalists; his hiring posts reveal the operating model (contractor experts plus delivery processes). No creative overlap.
How to reach
Speculative: small YC founder, reachable on X (@Jerr_Wu) or jerry@halluminate.ai (listed on the company site). Founder-to-founder exchange on vertical expert recruitment and verifier design is a plausible opener.
What we could not establish
Halluminate's actual funding amount and investors' round sizes are not disclosed.