Miju Labs

← All people · Gray Swan AI

Andy Zou

Co-founder (described as CTO in an undated CMU talk listing) · Gray Swan AI

AI-safety researcher behind the universal adversarial-suffix jailbreak, representation engineering, circuit breakers and HarmBench, and co-founder of Gray Swan. His research is the technical core of the company's attack and defence products.

low confidencePittsburghAt Gray Swan AI since 2023Updated 2026-09-19

Background

BS and MS at UC Berkeley (with Dawn Song and Jacob Steinhardt), then a CMU CS PhD advised by Zico Kolter and Matt Fredrikson, conferred May 2026 with the thesis 'Improving Security and Safety of Generative Models'; Nicholas Carlini (Anthropic) sat on the committee. His adversarial-attack work was covered by The New York Times.

What they run now

Research on attacks and defences (representation engineering, circuit breakers) that feed Shade and Cygnal. Current title and scope not confirmed.

Career

  1. 2023 – presentCo-founder (CTO per an undated talk listing), Gray Swan AI
  2. UC Berkeley, BS, MS
  3. conferred May 2026Carnegie Mellon University, PhD, Computer Scienceadvisors Zico Kolter and Matt Fredrikson

Recent posts

22 posts archived · most engaged first, then the latest

Andy Zou@andyzou_jiaming·
We deployed 44 AI agents and offered the internet $170K to attack them. 1.8M attempts, 62K breaches, including data leakage and financial loss. 🚨 Concerningly, the same exploits transfer to live production agents… (example: exfiltrating emails through calendar event) 🧵
734542.2K525KView on X ↗
Andy Zou@andyzou_jiaming·
No LLM is secure! A year ago, we unveiled the first of many automated jailbreak capable of cracking all major LLMs. 🚨 But there is hope?! We introduce Short Circuiting: the first alignment technique that is adversarially robust. 🧵 📄 Paper: arxiv.org/abs/2406.04313
15120652145KView on X ↗
Andy Zou@andyzou_jiaming·
Meta: Here's a model we fine-tuned extensively to do exactly one thing (differentiating safe and unsafe content). GCG: Hold my beer...
52511447KView on X ↗
Andy Zou@andyzou_jiaming·
At @GraySwanAI, we worked with @AnthropicAI to test Sonnet 4.5's safeguards. The results were exciting: for example, Sonnet 4.5 achieved SoTA robustness against prompt injection attacks. Excited to continue partnering with Anthropic to test & strengthen security.
Andy Zou@andyzou_jiaming·
At @GraySwanAI, we worked with @OpenAI to test GPT-5's safeguards. We identified 6 universal jailbreaks on a pre-release endpoint, but overall, GPT-5 demonstrated SoTA robustness against attacks. Excited to continue partnering with OpenAI to test & strengthen security.

Interviews & talks

Connections

  • Zico KolterPhD advisor, co-founder
  • Matt FredriksonPhD advisor, co-founder
  • Dan Hendrycksco-author (Center for AI Safety), named at Gray Swan launch
Why it matters here

Little direct relevance to creative data; relevant only as the technical founder of a crowd-plus-automation business.

How to reach

Academic/X route. Low priority for Julian.

What we could not establish
  • Not listed on Gray Swan's current About page (which names only Fredrikson, Kolter, Jenks, Whitman); whether he is still in an operating role after his May 2026 PhD is unverified.
  • CTO title comes from an undated event listing only.
  • No verified LinkedIn.

Sources

  1. Personal site
  2. CMU CSD: degree conferred
  3. Luma: ML Speaker Series
  4. Zico Kolter on X: launching Gray Swan (2024-07)
  5. Gray Swan About page