← All speakers

Bio, Work & Ideas

Šimon Podhajský

Conference affiliation: Waypoint AI · 2026

Šimon Podhajský is a Principal AI Engineer at Curebase, founder of the Prague evaluation community Evals.cz, and creator of the open-source Personal Intelligence Kit. He builds systems that test whether AI applications work reliably and help people understand their behavior without surrendering control over consequential decisions.

Podhajský studied cognitive science at Yale and conducted neuroeconomics and neuroscience research before working in data science at SRI International and data engineering at Pure Storage. He introduced scientists to Git and later rebuilt an unwieldy genomics workflow, applying software-engineering discipline to scientific work. At Waypoint AI, he held an AI leadership role focused on language-model applications and retrieval-augmented generation before joining Curebase.

  • Evaluation grounded in production failures. Podhajský favors an evaluation flywheel that starts with observing live application behavior, analyzes concrete errors, and converts those failures into regression tests. He distinguishes evaluations from benchmarks and guardrails, questions unvalidated synthetic test data, and cautions that formal evaluations can become premature when an application is still changing rapidly.
  • Read-only personal AI. The Personal Intelligence Kit combines browser history, journals, notes, tasks, email, and contacts inside an analytical workspace while writing results only to a separate destination. Using Claude Code, Python, SQLite, and an Obsidian vault, Podhajský builds weekly reflections and identifies contacts who might appreciate articles he has read. His AI Engineer Europe talk argues that an AI observer should inform human decisions without sending messages or modifying source records.
  • Cognitive exhaust and informed consent to risk. Everyday digital traces can reveal attention drift, neglected relationships, and gaps between intentions and actions that individual applications cannot see. Keeping those records untouched preserves evidence of actual human behavior, but Podhajský acknowledges that concentrated personal data, shell access, and external model providers still create serious privacy and exfiltration risks.
  • Debate as a reasoning benchmark. His open-source DebateFlow tests whether language models can judge extended arguments and identify deliberately introduced weaknesses across multiple turns.

Podhajský organizes Evals.cz with Veronika Pešková, teaches through Czechitas, contributes to the Czech-language Data Talk and AI ta Krajta podcasts, and serves on the Czech Debate Association’s board.

Read the topics behind these talks

1 conference talk

References