← All speakers

Bio, Work & Ideas

Raza Habib

Conference affiliation: Humanloop · 2024

Raza Habib is an AI researcher and entrepreneur at Anthropic who co-founded and led Humanloop, a platform for evaluating and improving AI applications. His work focuses on making language-model products dependable through rigorous testing, direct involvement from domain experts, and feedback from real users.

Habib studied physics at the University of Cambridge and completed a master’s degree in machine learning and computational statistics and a doctorate in probabilistic deep learning at University College London. He worked on speech synthesis at Google and became the founding engineer at Monolith AI, where machine-learning applications for physical engineering, including work with McLaren, demonstrated how active learning could identify the most informative experiments.

In 2020, Habib co-founded Humanloop with Peter Hayes and Jordan Burgess, alongside academic co-founders David Barber and Emine Yilmaz. The UCL research spinout initially helped businesses capture specialist knowledge through smarter data labeling and expert-defined rules. As language models advanced, Habib redirected Humanloop toward AI application development, integrating prompt experimentation, model comparison, production monitoring, and user feedback.

By its November 2024 general-availability launch, Humanloop supported versioned prompts, tools, evaluators, and application flows, with customers including Duolingo, Gusto, Vanta, and Filevine.

  • Evaluation-first development. Habib treats evaluation as an AI product’s practical specification: define useful outputs, test throughout development, and maintain regression suites when prompts or models change. He favors scorecards balancing accuracy, usefulness, latency, and cost over a single universal metric.
  • Domain experts as direct builders. Product engineers need sustained input from specialists who understand the work being automated. Duolingo linguists shaped language-learning prompts, while Filevine’s legal professionals helped develop legal AI features; both examples illustrate how experts can define behavior and evaluate results directly.
  • Simple architecture, disciplined operations. Habib describes most language-model applications through four components: a base model, instructions, selected context, and optional function calls. His AI Engineer World’s Fair talk emphasizes improving those components, comprehensive logging, and replayable traces over elaborate orchestration.
  • User feedback as ground truth. Accepted suggestions, edits, corrections, copied outputs, and regeneration requests reveal whether applications actually help people. Habib supplements those signals with narrowly scoped automated checks and human annotation, especially for sensitive or regulated work.

In August 2025, Habib, Hayes, Burgess, and much of the Humanloop team joined Anthropic. The standalone Humanloop platform closed the following month.

Read the topics behind these talks

1 conference talk

References