About

I’m a computer science Ph.D. student at Stanford University advised by Prof. Percy Liang and Prof. Sanmi Koyejo. I’m part of Stanford AI Lab, Stanford NLP, Stanford ML, and Stanford Trustworthy AI Research (STAIR). I’m also a part-time advisor to AlphaXiv and to MLCommons.

I’m broadly interested in building, understanding, aligning, and safeguarding frontier language models. My work spans pretraining, post-training, alignment, and safety, with an emphasis on scaling laws, open model development, and rigorous evaluation.

News

Selected Work

Pretraining

Marin: A Fully Open-Weight Language Model (2025–present)

Core contributor · Led midtraining and instruction tuning; contributed to pretraining
  • Marin 8B outperformed Llama 3.1 8B at release; Marin 32B beat OLMo 2 32B on 14 of 19 standard benchmarks.
  • The Marin team is currently running a new five-width scaling ladder and 535B-parameter MoE hero run (23B active), training over 18T tokens (~2.7 × 1024 FLOPs) on 11 GB200 NVL72 racks.
Marin 8B · Marin 32B · Live MoE report
Post-training

Async RL from Scratch (2026)

Co-first author and lead developer
  • Built Marin's asynchronous LLM reinforcement-learning stack for JAX/TPUs, decoupling rollout and training to deliver a 1.21× speedup over synchronous RL and cut 32GB weight-transfer time from 29 to 14 seconds.
  • Built preemption-resilient training and evaluation infrastructure, fixed an upstream vLLM-on-TPU sampling bug, and demonstrated stable 500-step RL runs.

Forecasting the Effects of Midtraining on Delphi (2026)

Project lead
  • Led the first controlled midtraining scaling study across Marin's Delphi ladder, spanning 3 × 1018–1022 FLOPs and nine compute-optimal model sizes from 447M to 9.7B across midtraining and post-training.
  • Forecasted held-out 1022-FLOP loss within 3% and MATH-500 pass@128 within 1%; at 1022 FLOPs, midtraining raised pass@128 from 10.4% to 79.6% before post-training.
  • Showed pass@128 after midtraining forecasts pass@1 after SFT across model sizes. After midtraining and SFT, the 1022-FLOP model matched Gemma 2 9B Instruct using roughly 1/44 of its pretraining FLOPs.
Alignment

Specification-Driven Alignment via Agreement-Anchored Rubric Repair (2026)

First author and project lead · Paper forthcoming; draft available upon request
  • Developed a method to mitigate underspecification in natural-language model policies using LM-judge ensemble agreement over behavioral rubrics, repairing 6 of 8 ambiguous OpenAI Model Spec statements.
  • Compiled repaired specifications into 60,854 synthetic preference pairs and post-trained an open-weight model with DPO, achieving a 65.2% blinded win rate over its SFT base.
Safety

Extracting Books from Production Language Models (2026)

First author and project lead
  • Led the first study demonstrating long-form extraction of copyrighted books from production LLMs despite model- and system-level safeguards, testing 13 books and four major production model families.
  • Claude 3.7 Sonnet recovered over 94% of four books, including two in-copyright works; we extracted 95.8% of Harry Potter and the Sorcerer's Stone from Claude, 76.8% from Gemini, and 70.3% from Grok.
Website · Paper · The Atlantic

SpecEval: Evaluating Model Adherence to Behavior Specifications (2025)

First author and project lead · Oral presentation, NeurIPS 2025 Workshop on Regulatable ML
  • Led the first systematic audit of frontier-model adherence to providers' published behavior specifications: 16 models, six labs, and more than 100 policy statements, uncovering consistency gaps up to 20%.

Other Publications

(*equal contribution, alphabetical/random authorship)
Marin: A Fully Open-Weight Language Model
David Hall, Ahmed Ahmed, ..., Percy Liang
2025
8B: Blog Post / Google Blog
32B: Blog Post / Retrospective
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams, ..., Ahmed Ahmed, ..., Peter Mattson, Percy Liang, Joaquin Vanschoren
arXiv 2025
arXiv / alphaXiv
SpecEval: Evaluating Model Adherence to Behavior Specifications
Ahmed Ahmed, Kevin Klyman, Yi Zeng, Sanmi Koyejo, Percy Liang
arXiv 2025
arXiv / alphaXiv
Introducing v0.5 of the AI Safety Benchmark from MLCommons
Bertie Vidgen, Adarsh Agrawal, Ahmed Ahmed, ..., Percy Liang, Peter Mattson, Joaquin Vanschoren
arXiv 2024
arXiv / alphaXiv
HELM Safety: Towards Standardized Safety Evaluations of Language Models
Farzaan Kaiyom, Ahmed Ahmed, Yifan Mai, Kevin Klyman, Rishi Bommasani, Percy Liang
November 2024
Website
Self-Improving Robots: End-to-End Autonomous Visuomotor Reinforcement Learning
Archit Sharma, Ahmed Ahmed, Rehaan Ahmad, Chelsea Finn
CoRL 2023: Conference on Robot Learning
Website / Paper / alphaXiv / Code
Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL
Bogdan Mazoure, Ahmed Ahmed, Patrick MacAlpine, R Devon Hjelm, Andrey Kolobov
ICLR 2022: International Conference on Learning Representations
arXiv / alphaXiv / Code

Website template adapted from Ken Liu. Thanks Ken!