Alex Duffy is the co-founder and chief executive of Good Start Labs, which uses games to evaluate and train artificial-intelligence models. His work tests whether interactive environments can develop abilities that conventional benchmarks struggle to measure: negotiation, trust, deception, and strategic judgment.
Duffy co-founded AI Camp, where students built working AI products, and served as head of product. He later became vice president of AI at Salt AI, leading researchers working on drug discovery and developing tools for emerging AI workflows. His 2024 introduction to Salt’s platform described combining language, images, video, and other media into collaborative software.
At Every, he led AI training and consulting for organizations in journalism, finance, construction, and technology. With Tyler Marques, he developed AI Diplomacy, an experiment placing language models inside a strategy game where success requires private negotiation, coalition-building, and anticipating betrayal. Their work became Good Start Labs, which spun out of Every with $3.6 million in funding in October 2025.
- Benchmarks shape model behavior. Once an evaluation becomes influential, developers optimize for it; poorly designed tests can therefore reward shallow proxies or harmful behavior. Duffy identifies conversational sycophancy as one consequence of feedback systems favoring agreement over judgment, and advocates human-centered model evaluation grounded in practical situations.
- Games expose social and strategic reasoning. AI Diplomacy makes models navigate incomplete information and conflicting incentives. In one observed game, Gemini 2.5 Pro gained an early advantage before OpenAI o3 prevailed through shifting alliances, while Claude Opus proved vulnerable to manipulation. These results describe particular game conditions, not fixed model personalities. Duffy develops this perspective in his AI Engineer World’s Fair talk.
- Training through play can test real-world transfer. Good Start Labs investigates whether reinforcement-learning data generated in games improves practical work. Duffy’s research on the railroad-finance game 1830 reported apparent improvements in research involving corporate filings—an early investigation into transferable strategic reasoning, not proof of universal transfer.
Duffy also argues that teachers, journalists, financial professionals, and other domain specialists should help define worthwhile AI behavior. Their judgments can make evaluation more relevant to the people expected to trust and use these systems.