← All speakers

Bio, Work & Ideas

Brooke Hopkins

Conference affiliation: Coval · 2025

Brooke Hopkins is the founder and chief executive of Coval, which builds simulation and evaluation infrastructure for autonomous voice and chat agents. She applies lessons from self-driving cars to a central challenge of conversational AI: making systems reliable when every interaction changes what happens next.

Hopkins studied computer science and mathematics at New York University and began her career developing voice technology for Google Assistant. At Waymo, she built systems for managing simulation datasets and led the evaluation-job infrastructure team, creating developer tools for running autonomous-driving simulations.

She founded Coval in 2024. The company joined Y Combinator’s Summer 2024 batch and raised $3.3 million led by MaC Venture Capital, with participation from General Catalyst and Y Combinator.

How Hopkins approaches reliable AI agents

  • Probabilistic evaluation: Instead of forcing agents into rigid decision trees, Hopkins simulates many plausible conversations and measures how consistently they complete tasks. Her reference-free evaluation judges overall outcomes without prescribing every conversational turn. Repeating failures distinguishes rare anomalies from systematic weaknesses.
  • Purpose-built simulation fidelity: Text-based tests can expose problems with workflows, tool calls, and instruction following; voice interactions reveal interruptions and latency; realistic accents, background noise, and degraded audio matter when those specific conditions are being evaluated. Hopkins treats reliability metrics as product-specific: appointment booking, refunds, and outbound sales require different standards.
  • Human-calibrated automated judgment: Coval’s Metric Studio aligns automated language-model evaluations with human-labeled conversations before applying those judgments at scale. Hopkins also coauthored research on scripted comparative evaluation, which fixes one side of a conversation to isolate model behavior. Coval’s open voice-model benchmarks measure speech-recognition and speech-generation latency and accuracy.

For Hopkins, dependable autonomy depends on continuously measuring real-world uncertainty, translating failures into better tests, and preserving the flexibility that makes conversational agents useful.

Read the topics behind these talks

1 conference talk

References