← All speakers

Bio, Work & Ideas

Peter Gostev

Conference affiliation: Arena.ai · 2026

Peter Gostev is AI Capability Lead at Arena and the creator of BullshitBench, an open-source benchmark measuring whether language models recognize nonsensical questions or confidently accept their premises. His work exposes a consequential weakness behind rising model scores: systems that perform impressively on conventional tests can still lack the judgment to challenge incoherent requests.

Gostev began his career in strategy and analytics, including at Accenture, before leading AI strategy at NatWest Group. There, he explored applications including retrieval-based access to internal human-resources policies and encountered the security, coordination, and deployment constraints facing large regulated organizations. As Head of AI at Moonpig, he helped apply generative AI to customer service and visual models to greeting-card tagging and discovery. His approach to enterprise AI strategy combined immediate operational improvements with longer-term experiments, while encouraging contributions from employees outside specialist engineering teams.

At Arena, Gostev applies that practical perspective to model evaluation, combining targeted benchmarks with real users’ assessments of competing systems.

  • Helpfulness without automatic agreement. Gostev distinguishes factual hallucination from a subtler failure: producing coherent explanations for questions that make no sense. An invented relationship between deployment frequency, code indentation, and variable-name length should prompt skepticism, not elaborate causal analysis. His account of BullshitBench argues that trustworthy assistants must challenge invalid premises instead of rewarding users with plausible nonsense.
  • Reasoning does not guarantee judgment. Gostev found that additional reasoning sometimes worsened performance: models could recognize a flawed premise, then spend extensive effort answering anyway. Comparisons of open-source systems likewise showed no clear relationship between parameter count and performance on this task.
  • User dissatisfaction signal. Arena lets users reject both answers in anonymous head-to-head comparisons. Gostev uses those votes to track persistent weaknesses across professional tasks, finding stronger gains in quantitative work than in creative writing, medicine, finance, law, and some software applications. His AI Engineer Europe presentation identifies game development as a revealing case: generating functional code does not ensure coherent mechanics or engaging design. Because prompts and expectations change over time, dissatisfaction measures evolving user experience rather than a fixed benchmark.

Read the topics behind these talks

1 conference talk

References