← All speakers

Bio, Work & Ideas

Julia Neagu

Conference affiliation: Quotient · 2025

Julia Neagu is a member of technical staff at Databricks and the co-founder and former chief executive of Quotient AI, the agent-evaluation company Databricks acquired in March 2026. She builds systems that identify failures in deployed AI agents and turn those failures into improvements.

Neagu studied physics at Princeton and Harvard, earning an AB, MA, and PhD. She developed quantitative models at Aon, led analytics at Tamr, and subsequently directed data and evaluation work for GitHub Copilot. In 2023, she and former GitHub colleague Freddie Vargus founded Quotient to help developers test AI applications against actual product requirements. Her announcement introducing Quotient emphasized realistic, domain-specific experimentation over subjective judgments and generic benchmarks.

  • Production evaluation over static benchmarks. Search agents confront changing websites, unpredictable questions, retrieval failures, hallucinations, and reasoning errors simultaneously. Neagu argues that fixed datasets and conventional monitoring cannot capture these shifting, interconnected production conditions, as she outlined in a collaborative AI-search evaluation session.
  • Reference-free failure detection. Quotient developed evaluators that identify problems in live agent interactions without waiting for labeled answers or human feedback. Its systems examine production traces for failures involving grounding, retrieval, reasoning, and tool use.
  • Continuous agent improvement. Neagu envisions agents recognizing unreliable sources, outdated information, and emerging hallucinations, then using those signals to improve subsequent behavior. Databricks acquired Quotient to bring production-derived evaluation datasets and reward signals into enterprise products including Genie, Genie Code, and Agent Bricks.
  • Practical adoption of open models. Her analysis of open and proprietary language models weighs customization, infrastructure, operating costs, and domain-specific performance; fine-tuning on specialized data can favor open models, while proprietary systems can simplify deployment.

Read the topics behind these talks

1 conference talk

References