Sayash Kapoor is a Princeton computer scientist, co-author of AI Snake Oil, and incoming assistant professor at the University of California, Berkeley’s School of Information, joining its faculty in fall 2027. His research addresses a practical question with enormous consequences for science, work, and public policy: when do impressive AI demonstrations translate into systems that people can actually trust?
Before pursuing his doctorate at Princeton, Kapoor worked on machine learning at Facebook and conducted AI-related work at Columbia University and EPFL. He subsequently developed a research agenda spanning reproducibility, accountability, and the social consequences of automation; he has also been a Mozilla senior fellow and a Porter Ogden Jacobus Fellow at Princeton.
With Princeton professor Arvind Narayanan, Kapoor wrote AI Snake Oil, distinguishing genuine advances in generative AI from exaggerated claims about predictive systems used in consequential decisions. Their subsequent AI as Normal Technology framework argues that even transformative technologies remain subject to institutional incentives, regulation, implementation constraints, and human choices.
Research and distinctive ideas
- Agent evaluation as an engineering discipline. Kapoor co-authored AI Agents That Matter, which identifies benchmark overfitting, weak holdout practices, irreproducible testing, and evaluations that ignore cost. Because agents use tools and act within environments, meaningful assessment must measure realistic tasks, standardized execution, and operational tradeoffs.
- Scientific reproducibility before scientific autonomy. His team’s CORE-Bench tests whether agents can reproduce published findings using the original code and data, spanning 270 tasks from 90 papers. The benchmark separates useful, measurable scientific assistance from unsupported claims of fully autonomous research. A subsequent study investigates what apparent benchmark saturation conceals, including shortcuts, reliability problems, weak generalization, and the importance of human collaboration.
- Cost-aware, multidimensional agent evaluation. Kapoor co-developed the Holistic Agent Leaderboard, infrastructure for comparing models and agent scaffolds under standardized conditions. Its accompanying research covers 21,730 runs across nine models and nine benchmarks. Measuring accuracy without operating cost obscures the expense of repeated model calls, recursive workflows, and growing usage.
- The capability-reliability gap. An agent that occasionally solves a problem is different from one that performs dependably in production. Kapoor’s research on agent reliability and AI Engineer Summit talk emphasize that flawed verifiers, reward hacking, and misleading benchmarks can inflate apparent performance. He treats trustworthy AI as a reliability-engineering problem: building dependable systems around inherently variable models.
His work also examines how organizational constraints shape automation: coding agents can accelerate implementation while leaving product judgment, coordination, accountability, and deployment unresolved. At Berkeley, Kapoor plans to connect technical research with policy and public communication, an emphasis he highlighted when announcing his faculty appointment.