
Building Metrics That Actually Work — David Karam, Pi Labs
A hands-on Pi Labs workshop applies lessons from Google Search to evaluating stochastic LLM applications. The presenters discuss defining application-specific metrics, combining natural-language questions with generated Python checks, creating synthetic examples, calibrating scores against data, and using an evaluation copilot, spreadsheets, an SDK, and…
David Karam
Evals · Coding and developer tools · RAG, context, and search
