← All speakers

Bio, Work & Ideas

Benedikt Sanftl

Conference affiliation: Mutagent · 2026

Benedikt Sanftl is the co-founder and chief executive of Mutagent, which builds software agents that diagnose, evaluate, and improve other AI agents. His central concern is practical: production failures multiply faster than engineers can inspect traces, identify causes, and safely test fixes.

Sanftl studied electrical engineering at the Technical University of Munich, completing master’s work in high-frequency technology while working at Airbus. He subsequently earned an engineering doctorate and became chief product officer at Beam AI, developing AI products for complex enterprise processes. After approximately two and a half years of manually optimizing agents, he launched Mutagent with two co-founders, including chief technology officer Burak Cemil Özafşar.

How he approaches reliable AI agents

  • The agentic AI engineer: Sanftl and Özafşar envision specialized agents collaborating across specification, implementation, testing, monitoring, diagnosis, and improvement. Their AI Engineer World’s Fair session connected an offline development loop with an online loop that converts production failures into better tests and targeted fixes.
  • Production-trace diagnostics: Sanftl distinguishes identifying a poor outcome from determining why it occurred. His writing on evaluation and diagnosis emphasizes grouping recurring failures and tracing them to actionable problems, including incomplete context, faulty instructions, or broken tools.
  • Scenario-based agent evaluation: He advocates realistic scenarios over fixed evaluation targets susceptible to gaming. Evaluation suites should grow as production uncovers unfamiliar cases, while unreliable model-generated scores require careful calibration.
  • Diagnose, mutate, validate: His AI engineering framework links root-cause analysis to a specific change and subsequent verification. Mutagent applies this approach through coordinated evaluation and diagnostic agents, existing coding and observability integrations, and human-reviewed workflows.

Read the topics behind these talks

1 conference talk

References