← All speakers

Bio, Work & Ideas

Ziv Ilan

Conference affiliation: NVIDIA · 2026

Ziv Ilan is a solutions architect on NVIDIA’s Paris-based AI Labs team, specializing in the evaluation, optimization, and deployment of generative AI. He works with frontier-model developers to make language, image, and video models faster, more efficient, and practical to run in production.

Before joining NVIDIA, Ilan worked at Deci, subsequently acquired by NVIDIA, on neural architecture search and model compression for generative AI and computer vision. He holds a bachelor’s degree in electrical and electronics engineering from Tel Aviv University and an MBA from HEC Paris.

His work spans the full operational lifecycle of advanced models. He coauthored an account of Domyn’s Colosseum 355B development, covering distributed infrastructure, continued pretraining, alignment, and evaluation for regulated industries. His writing on production-grade LLM evaluation describes an Amdocs architecture combining fine-tuning, GitOps, automated regression testing, application-specific grading, and human review.

  • Continuous evaluation as production infrastructure. General benchmarks, domain-specific tests, automated grading, and expert review serve distinct purposes: catching regressions, measuring business-task performance, and determining whether customized models merit deployment.
  • Diffusion acceleration without sacrificing fidelity. Ilan combines quantization, caching, and distillation to reduce the repeated denoising passes that make image and video generation expensive. His AI Engineer Europe session emphasizes applying these optimizations incrementally and measuring their impact on visual quality.
  • Dynamic quantization for FLUX.2. Work with Black Forest Labs illustrates how lower-precision computation can reduce memory requirements and improve inference speed. Because diffusion models are attention-heavy and sensitive to numerical changes, Ilan distinguishes post-training quantization from quantization-aware training and highlights techniques that adjust numerical ranges during inference.
  • Chunk-based caching and step distillation. Spatial caching avoids recomputing regions that barely change between denoising steps; overly aggressive reuse can visibly degrade results. Step distillation teaches a model to produce comparable images in fewer iterations without necessarily reducing its parameter count. Ilan contrasts trajectory-based and distribution-matching methods and points to NVIDIA Research’s FastGen framework for distributed post-training.

Ilan also helps convene the Paris developer community, including an NVIDIA gathering featuring Mistral and Hugging Face.

Read the topics behind these talks

1 conference talk

References