Nicholas Kang is a Kaggle product manager leading Kaggle Benchmarks, a platform for creating, running, and sharing evaluations of AI systems. Originally from Singapore, he focuses on opening AI assessment to domain specialists and making autonomous agents easier to test before deployment.
Kang developed Community Benchmarks, which lets participants build evaluation tasks from their professional expertise, combine those tasks into benchmarks, and compare models. He subsequently coauthored a Google announcement introducing local benchmark development, bringing evaluation authoring into editors, command-line workflows, and AI coding agents.
His defining contributions include:
- Community-created AI evaluations. Kang argues that researchers cannot anticipate every consequential application of AI. He highlights a wastewater-treatment engineer who created a benchmark rooted in operational safety and specialist experience—a concrete example of expertise that conventional research benchmarks can miss.
- Transparent evaluation methodology. Identical benchmarks can produce conflicting scores when model providers change their evaluation harness or context compaction settings. Kang considers those implementation choices essential context for interpreting leaderboards, alongside practical measures such as cost and speed.
- Standardized Agent Exams. Kang authored the launch of Kaggle’s agent assessments, which let an autonomous agent take an exam and receive a comparative score without custom infrastructure. The initial assessment examines multistep reasoning and adversarial safety; Kang has emphasized the importance of testing consumer agents before they manage sensitive or consequential tasks.
- Cognitive-abilities benchmarks. He contributed to and judged Measuring Progress Toward AGI: Cognitive Abilities, a community initiative involving Google DeepMind researchers that invites benchmarks for capabilities including learning, attention, executive function, metacognition, and social cognition.
At AI Engineer Europe 2026, Kang appeared with Kaggle engineer Michael Aaron under a Google DeepMind conference affiliation, outlining his approach to community benchmarks, accessible agent exams, and broader participation in AI evaluation.