Joel Becker is the founder and chief executive of Qally’s and an AI evaluation researcher whose work tests whether increasingly capable models make skilled people more productive. Formerly on the technical staff at METR, he helped develop human-calibrated measures of autonomous AI capabilities and co-led a randomized experiment that found experienced open-source developers worked more slowly with AI assistance.
Becker studied economics and econometrics at the University of Bristol, worked as a predoctoral fellow at Harvard and the National Bureau of Economic Research, pursued graduate economics research at New York University, and was a visiting researcher at Oxford’s Global Priorities Institute. He also co-owned a statistics consultancy serving professional soccer teams.
His research extended into genomics: he was first author of a 2021 study introducing the Polygenic Index Repository and examining bias in genetic prediction. He subsequently founded Qally’s and worked on evaluation methods at METR, applying economic ideas about incentives, counterfactuals, and measurement to advanced AI systems.
- Human-calibrated AI time horizons. Becker coauthored research measuring autonomous task completion and contributed to human baselines and data collection. The resulting metric estimates how long a task would take a skilled human if an AI system can complete it with a specified probability. A 50 percent horizon makes capability easier to interpret, but does not establish that models can handle organizational context or achieve the reliability required in production.
- The developer-productivity paradox. In a randomized study of 16 experienced open-source developers completing 246 tasks, participants took 19 percent longer when early-2025 AI tools were allowed, despite believing those tools accelerated their work. Becker attributes the gap to factors including verification costs, unreliable outputs, complex codebases, and developers’ extensive repository knowledge. The finding describes a specific population and period, not all programmers or subsequent models.
- Mergeability over benchmark passing. Becker contributed to research on agent-written pull requests indicating that approximately half of benchmark-passing changes would not have been accepted by maintainers. Automated tests cannot fully capture architectural fit, maintainability, or the judgment exercised during human review.
His research on compute slowdowns further examines how constraints on computing investment could delay projected AI capability milestones. Across these projects, Becker focuses on whether measured performance survives the practical demands of real work: ambiguous goals, accumulated context, coordination, quality control, and deciding which tasks are actually worth doing.