← All speakers
  • SWE Hiring is Cookednickheiner.com
  • Nick Heiner's Substack | Substack

    Independent benchmarks, essays on the future of work, and dispatches from someone building AI products and testing AI agents every day at Surge AI. Click to read Nick Heiner's Substack, a Substack publication. Nick Heiner's Substack

    nickheiner.com

Bio, Work & Ideas

Nick Heiner

Conference affiliation: VP of RL Environments · Surge AI · 2026

Nick Heiner leads reinforcement-learning environments at Surge AI, developing realistic simulations that test whether AI agents can complete demanding professional work. A former Netflix engineer and founding engineer at Fixie, he focuses on the widening gap between impressive benchmark scores and practical usefulness.

Heiner studied at Cornell University and worked at Opower and the United States Digital Service before joining Netflix, where he worked on user-interface platforms. After ChatGPT accelerated his interest in AI, he joined Fixie as a founding engineer and subsequently moved to Surge. The company has also identified him as its vice president of product in an assessment of Claude’s coding behavior.

He coauthored Corecraft, a simulated customer-support workplace that measures agents against multistep tasks, realistic operational tools, and expert-authored grading criteria. The research found that leading models struggled when required to satisfy every criterion, while training in high-fidelity environments improved performance on held-out tasks and other evaluations.

  • Benchmark contamination obscures actual capability. Public test material can enter training data, making memorization resemble progress. In his AI Engineer World’s Fair talk, Heiner examines SWE-bench Verified and argues for private holdout sets, contamination disclosures, and scrutiny of what models have already encountered.
  • Prompt-verifier alignment makes evaluations trustworthy. Graders should check everything a task requests, and tasks should disclose everything graders enforce. Otherwise, contradictory instructions, rigid formatting requirements, and Unicode-based reward hacking can produce misleading scores regardless of an agent’s real competence.
  • Expert judgment is essential for sophisticated work. Heiner helped develop Hemingway-bench, which uses blind comparisons by professional writers to evaluate qualities such as originality, coherence, and taste. He argues that automated judges and mechanical checklists cannot reliably capture nuanced writing quality.

Read the topics behind these talks

1 conference talk

References