← All speakers

Bio, Work & Ideas

Ali Khial

Conference affiliation: Head of AI/ML · G2i · 2026

Ali Khial is a software-engineering leader working to make AI coding-agent evaluation reflect the demands of production development. His focus spans software engineering benchmarks, frontier-model evaluation, agentic workflows, and the human judgment required to assess whether autonomous systems can handle commercially meaningful work.

Khial studied at the University of Sciences and Technologies of Algiers and built his career in backend engineering, technical leadership, and digital commerce, including work involving ButcherBox and ConnectiveRx. His production-software background informs a practical skepticism toward benchmark scores divorced from business requirements. At the 2026 AI Engineer World’s Fair, he represented G2i in an AI/ML leadership role.

Building benchmarks engineers can trust

Khial treats a benchmark as an interconnected system of instructions, agents, execution environments, graders, and recorded trajectories. His analysis of coding-agent evaluation identifies four essential improvements:

  • Human-authored benchmark instructions: Describe goals and constraints without leaking test files, prescribing implementation interfaces, or rewarding mechanical compliance.
  • Behavioral grading: Evaluate observable outcomes instead of rejecting sound solutions over arbitrary variable names; apply rigorous unit, integration, and end-to-end tests where security or business logic requires precision.
  • Economically meaningful evaluation tasks: Measure work engineering teams actually need, not difficult exercises whose completion says little about production usefulness.
  • Contamination-resistant evaluation: Use novel tasks, private holdout sets, and controlled environments to limit repository leakage and reward hacking, while publishing failure patterns and execution trajectories alongside scores.

For Khial, trustworthy evaluation depends on experienced software engineers helping define realistic tasks, sound tests, and evidence that an agent’s capabilities transfer beyond the leaderboard.

Read the topics behind these talks

1 conference talk

References