
RAG Evaluation Is Broken! Here's Why (And How to Fix It)
AI21 Labs presenters Yuval Belfer and Niv Granot argue that conventional RAG benchmarks overreward questions answerable from individual chunks while neglecting realistic aggregation across documents. Using financial examples and a 22-document FIFA World Cup corpus, they report baseline accuracies of 5% and 11% and describe a structured-RAG alternative that…
Yuval Belfer · Niv Granot
Finance · RAG, context, and search · Evals
