About
I’m a computer science Ph.D. student at Stanford University advised by Prof. Percy Liang and Prof. Sanmi Koyejo. I’m part of Stanford AI Lab, Stanford NLP, Stanford ML, and Stanford Trustworthy AI Research (STAIR). I’m also a part-time advisor to AlphaXiv and to MLCommons.
I’m broadly interested in building, understanding, aligning, and safeguarding frontier language models. My work spans pretraining, post-training, alignment, and safety, with an emphasis on scaling laws, open model development, and rigorous evaluation.
News
- Aug 2026: Published Forecasting the Effects of Midtraining on Delphi, showing that midtraining loss and pass@k can be smoothly predicted for relevant math tasks.
- Aug 2026: The Marin team launched a new five-width MoE scaling ladder and a 535B-parameter (23B active) hero run. The model is currently training over 18T tokens on 11 GB200 NVL72 racks.
- Jan 2026: Our work on extracting copyrighted content from production LLMs was featured by The Atlantic
- Oct 2025: Released Marin 32B, the best fully open-source 32B model at the time of release, beating OLMo 2 32B Base on 14/19 standard benchmarks. Read more in our retrospective.
- July 2025: Our work on extracting copyrighted content from open-weight models was featured by The Atlantic
- May 2025: Presented our Independence Tests for Language Models paper at ICML 2025, receiving an oral presentation (top 2.6% of submissions)
- May 2025: Released Marin 8B, the best fully open-source 8B model at the time of release, outperforming Llama 3.1 8B. I was in charge of instruction tuning and decontamination.
- May 2025: Our AILuminate benchmark was covered by WIRED and Business Wire
- May 2024: Meta used our MLCommons AI Safety Benchmark v0.5 to evaluate their Llama 3 models for safety testing
- May 2023: Selected as a Knight-Hennessy Scholar (2023 cohort) — one of 84 scholars chosen from over 7,500 applications (1.1% acceptance rate)
Selected Work
Pretraining
Marin: A Fully Open-Weight Language Model
- Marin 8B outperformed Llama 3.1 8B at release; Marin 32B beat OLMo 2 32B on 14 of 19 standard benchmarks.
- The Marin team is currently running a new five-width scaling ladder and 535B-parameter MoE hero run (23B active), training over 18T tokens (~2.7 × 1024 FLOPs) on 11 GB200 NVL72 racks.
Post-training
Async RL from Scratch
- Built Marin's asynchronous LLM reinforcement-learning stack for JAX/TPUs, decoupling rollout and training to deliver a 1.21× speedup over synchronous RL and cut 32GB weight-transfer time from 29 to 14 seconds.
- Built preemption-resilient training and evaluation infrastructure, fixed an upstream vLLM-on-TPU sampling bug, and demonstrated stable 500-step RL runs.
Forecasting the Effects of Midtraining on Delphi
- Led the first controlled midtraining scaling study across Marin's Delphi ladder, spanning 3 × 1018–1022 FLOPs and nine compute-optimal model sizes from 447M to 9.7B across midtraining and post-training.
- Forecasted held-out 1022-FLOP loss within 3% and MATH-500 pass@128 within 1%; at 1022 FLOPs, midtraining raised pass@128 from 10.4% to 79.6% before post-training.
- Showed pass@128 after midtraining forecasts pass@1 after SFT across model sizes. After midtraining and SFT, the 1022-FLOP model matched Gemma 2 9B Instruct using roughly 1/44 of its pretraining FLOPs.
Alignment
Specification-Driven Alignment via Agreement-Anchored Rubric Repair
- Developed a method to mitigate underspecification in natural-language model policies using LM-judge ensemble agreement over behavioral rubrics, repairing 6 of 8 ambiguous OpenAI Model Spec statements.
- Compiled repaired specifications into 60,854 synthetic preference pairs and post-trained an open-weight model with DPO, achieving a 65.2% blinded win rate over its SFT base.
Safety
Extracting Books from Production Language Models
- Led the first study demonstrating long-form extraction of copyrighted books from production LLMs despite model- and system-level safeguards, testing 13 books and four major production model families.
- Claude 3.7 Sonnet recovered over 94% of four books, including two in-copyright works; we extracted 95.8% of Harry Potter and the Sorcerer's Stone from Claude, 76.8% from Gemini, and 70.3% from Grok.
SpecEval: Evaluating Model Adherence to Behavior Specifications
- Led the first systematic audit of frontier-model adherence to providers' published behavior specifications: 16 models, six labs, and more than 100 policy statements, uncovering consistency gaps up to 20%.
Other Publications
(*equal contribution, †alphabetical/random authorship)
![]() | Extracting books from production language models Ahmed Ahmed, A. Feder Cooper, Sanmi Koyejo, Percy Liang arXiv 2025 Website / arXiv / alphaXiv / Press Coverage / Reddit / TL;DR |
![]() | Marin: A Fully Open-Weight Language Model David Hall, Ahmed Ahmed, ..., Percy Liang 2025 8B: Blog Post / Google Blog 32B: Blog Post / Retrospective |
![]() | Extracting memorized pieces of (copyrighted) books from open-weight language models A. Feder Cooper, Aaron Gokaslan, Ahmed Ahmed, Amy B. Cyphert, Christopher De Sa, Mark A. Lemley, Daniel E. Ho, Percy Liang arXiv 2025 arXiv / alphaXiv / Press Coverage / TL;DR |
![]() | AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons Shaona Ghosh, Heather Frase, Adina Williams, ..., Ahmed Ahmed, ..., Peter Mattson, Percy Liang, Joaquin Vanschoren arXiv 2025 arXiv / alphaXiv |
![]() | Independence Tests for Language Models Sally Zhu*, Ahmed Ahmed*, Rohith Kuditipudi*, Percy Liang ICML 2025 SPOTLIGHT - TOP 2.6% arXiv / alphaXiv |
![]() | SpecEval: Evaluating Model Adherence to Behavior Specifications Ahmed Ahmed, Kevin Klyman, Yi Zeng, Sanmi Koyejo, Percy Liang arXiv 2025 arXiv / alphaXiv |
![]() | Introducing v0.5 of the AI Safety Benchmark from MLCommons Bertie Vidgen, Adarsh Agrawal, Ahmed Ahmed, ..., Percy Liang, Peter Mattson, Joaquin Vanschoren arXiv 2024 arXiv / alphaXiv |
![]() | HELM Safety: Towards Standardized Safety Evaluations of Language Models Farzaan Kaiyom, Ahmed Ahmed, Yifan Mai, Kevin Klyman, Rishi Bommasani, Percy Liang November 2024 Website |
![]() | Self-Improving Robots: End-to-End Autonomous Visuomotor Reinforcement Learning Archit Sharma, Ahmed Ahmed, Rehaan Ahmad, Chelsea Finn CoRL 2023: Conference on Robot Learning Website / Paper / alphaXiv / Code |
![]() | Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL Bogdan Mazoure, Ahmed Ahmed, Patrick MacAlpine, R Devon Hjelm, Andrey Kolobov ICLR 2022: International Conference on Learning Representations arXiv / alphaXiv / Code |
Website template adapted from Ken Liu. Thanks Ken!









