
Your Evals Are Meaningless (And Here’s How to Fix Them)
HoneyHive co-founder Mohak Sharma explains why developer-written test sets, generic metrics, and uncalibrated LLM judges can produce evaluation scores that fail to predict real production performance. He recommends domain-expert-defined criteria and edge cases, continuously incorporating problematic production queries into evaluation datasets, and tracking…
Mohak Sharma
Evals · Safety and governance