← All speakers

Bio, Work & Ideas

Michael Aaron

Conference affiliation: Google DeepMind · 2026

Michael Aaron is a Kaggle software engineer building more transparent, reproducible ways to evaluate AI models and agents. His work includes Kaggle Community Benchmarks, the FACTS Grounding benchmark, and Kaggle Game Arena—systems that test grounded answers, invite domain experts to create evaluations, and compare models through strategic games.

Aaron spent an extended period at Google before focusing on evaluations and benchmarks at Kaggle. He contributed ideas, data collection, and early experiments to FACTS Grounding, which measures whether models produce long-form responses faithful to supplied documents and user instructions.

In January 2026, he and Kaggle product lead Meg Risdal introduced Kaggle Community Benchmarks, enabling contributors to assemble evaluation tasks, inspect model interactions, and compare results on shared leaderboards. The associated open-source benchmark SDK supports tests involving coding, tools, multimodal inputs, and multiturn conversations.

  • Community-created evaluations: Domain specialists can design tests reflecting professional expertise that standard research datasets miss. Aaron focuses on making those evaluations reproducible while acknowledging that good tasks require substantial effort and sustained contributor incentives.
  • Strategic games as dynamic benchmarks: Kaggle Game Arena compares models through poker, chess, and Werewolf, revealing differences in risk-taking, deception, and strategy. His engineering concerns include fair prompts, OpenSpiel-based harnesses, simulation infrastructure, public game visualizations, and Elo-style rankings.
  • Statistical rigor at practical cost: Bradley–Terry pairwise comparisons help limit expensive matchups while preserving meaningful rankings. Changing model endpoints and disappearing older releases complicate comparisons over time.
  • Separating models from their harnesses: Agent performance reflects the surrounding prompts, tools, orchestration, and execution environment as well as the underlying model. Aaron treats that distinction as essential to interpreting leaderboard results fairly.

At AI Engineer Europe 2026, Aaron appeared alongside Nicholas Kang under a Google DeepMind affiliation, concentrating on game-based benchmarks, community participation, and the operational challenges of evaluating rapidly changing AI systems.

Read the topics behind these talks

1 conference talk

References