← All speakers

Bio, Work & Ideas

Greg Kamradt

Conference affiliation: ARC Prize Foundation · 2025

Greg Kamradt is president and a board member of the ARC Prize Foundation, where he builds benchmarks and open competitions to test whether artificial intelligence can learn unfamiliar skills. He also created Needle In A Haystack, an open-source evaluation exposing the difference between a language model’s advertised context window and its ability to retrieve information within it.

Kamradt worked in finance at NetApp, then in strategy, growth, and predictive analytics at Salesforce before leading operations and product at Digits, an AI-focused accounting company. He founded Leverage, an AI product and education studio, and taught developers to build language-model applications through tutorials, open-source material, and practical demonstrations.

His Needle In A Haystack evaluation places specific facts at varying depths inside increasingly long inputs, then measures whether models can recover them. The project now includes multiple-fact retrieval and chained reasoning, making long-context performance testable beyond headline token limits.

In early 2024, Zapier co-founder Mike Knoop recruited Kamradt to help organize ARC Prize, an open competition based on François Chollet’s Abstraction and Reasoning Corpus. Working with Knoop, Chollet, and Bryan Landers, Kamradt shifted from building AI applications toward evaluating whether models genuinely acquire new capabilities. He subsequently co-authored the 2024 and 2025 ARC Prize technical reports.

What his benchmarks are designed to reveal

  • Skill-acquisition efficiency: Intelligence requires learning an unfamiliar task and applying its rules effectively. Success at chess, Pokémon, or other familiar games cannot establish that ability when training data, developer knowledge, or repeated intervention may explain the result.
  • Human-calibrated generalization: ARC tasks draw on basic concepts such as objects, geometry, counting, and other agents, using human performance as a reference point. Problems ordinary people can solve but frontier models cannot reveal missing capabilities more directly than specialized expert-level trivia.
  • Interactive reasoning benchmarks: Static puzzles cannot fully evaluate exploration, goal discovery, memory, planning, and adaptation. Hidden game-like environments force agents to infer unfamiliar rules without instructions, while private evaluation sets prevent developers from preparing systems for the exact tasks.
  • Diagnostic evaluation: Action counts measure learning efficiency, and replaying model decisions exposes specific failures: mistaken world models, inappropriate analogies to familiar games, and solutions that do not transfer between levels.

In March 2026, Kamradt launched ARC-AGI-3, the foundation’s first fully interactive benchmark, alongside competitions offering more than $2 million in prizes. His subsequent analysis of frontier-model reasoning traces documented how capable systems can recognize isolated effects without understanding an environment. Among early competition winners, he highlighted agents that wrote and executed Python to investigate unfamiliar games instead of depending on elaborate predefined scaffolding.

Kamradt also continues building practical AI products, including agnts.sh and ReadSide: tools for serving agent-readable web content and discussing specific passages while reading.

Read the topics behind these talks

1 conference talk

References