← All speakers

Bio, Work & Ideas

Dmitry Kuchin

Conference affiliation: Multinear · 2025

On this page

Dmitry “Dima” Kuchin is an AI engineering leader at Mirable and co-founder of Multinear, an open-source platform for building and evaluating reliable generative AI applications. He specializes in turning promising prototypes into production systems whose behavior can be measured against actual customer needs.

From enterprise data to reliable AI

Kuchin’s career spans startup technology leadership and engineering or data work involving Meta, Lemonade, ZoomInfo, and WeWork. He was part of Unomy’s team when WeWork acquired the company in 2017.

By 2023, he led Sphere Partners’ data and AI practice and wrote about applying emerging generative AI systems to business problems. A 2024 accounting-industry report identified him as Sphere’s chief data officer; his approach emphasized combining AI-assisted work with professional human judgment.

He subsequently co-founded Multinear with Asaf Bord. The open-source Multinear platform supports application-specific evaluation workflows, with examples covering bank customer support and Text-to-SQL. At Mirable, he addresses production failures including inconsistent answers, hallucinations, drifting agents, and inadequate measurement.

How Kuchin approaches AI reliability

  • Business-specific AI evaluations: Judge applications by whether they accomplish the user’s goal. A support response can be factually accurate yet still fail if it omits essential instructions or forces escalation to a human.
  • Evaluation from the first prototype: Build tests alongside the earliest application, inspect individual failures, refine prompts, models, logic, or source data, and rerun evaluations to detect regressions.
  • Synthetic scenarios grounded in real workflows: Generate realistic questions and explicit success criteria from reference materials, then vary wording and customer personas while checking that essential details remain intact.
  • Benchmarks for architectural decisions: Use a stable evaluation baseline to compare GPT-4o with GPT-4o mini, GraphRAG with simpler retrieval, and agentic workflows with lower-latency alternatives. His AI Engineer World’s Fair talk connects these choices directly to quality and inference cost.
  • Text-to-SQL mock-database testing: Match evaluation methods to the application: customer-support systems can use an LLM judge, while schema-matched mock databases and controlled records make generated SQL results directly testable.

More recently, Kuchin has applied the same rigor to AI-assisted programming with Cursor, Claude Code, and Codex, combining active supervision with domain-specific end-to-end evaluations.

Read the topics behind these talks

1 conference talk

References