
How to look at your data; what to look for, how to measure
Jeff Huber and Jason Liu present a two-part approach to improving AI applications by examining both retrieval inputs and application outputs. Huber explains how generated and real queries, application-specific embedding evaluations, and recall@10 can reveal performance differences that broad MTEB rankings obscure, illustrated with a Weights & Biases…
Jeff Huber · Jason Liu
Other / unclassified · RAG, context, and search · Agent engineering