← All speakers

Bio, Work & Ideas

Bertrand Charpentier

Conference affiliation: Pruna AI · 2026

Bertrand Charpentier is the cofounder, president, and chief scientist of Pruna AI, which makes machine-learning models faster, cheaper, and more energy-efficient. His research spans uncertainty estimation, neural-network pruning, and generative-model optimization; his central technical argument is that production models should be selected for specific applications, not their position on a general-purpose leaderboard.

During doctoral work at the Technical University of Munich, Charpentier investigated how models recognize the limits of their predictions. His projects included uncertainty in asynchronous event prediction, Posterior-Network, which estimates uncertainty without out-of-distribution samples, and Natural Posterior Networks, which extend Bayesian uncertainty estimation to exponential-family distributions. He also worked on differentiable sampling of directed acyclic graphs.

His research subsequently focused on reducing computation. An early-pruning project explored identifying useful sparse subnetworks sooner; the 2024 paper Structurally Prune Anything, written with Xun Wang, John Rachwan, and Stephan Günnemann, developed architecture-independent structured pruning.

Charpentier founded Pruna AI with Rachwan, Günnemann, and Rayan Nait Mazi, who became chief executive. The company initially built an automated pipeline for compressing publicly available Hugging Face models and comparing their performance before and after optimization. Charpentier subsequently announced $6.5 million in funding.

  • Application-specific model evaluation: Image-model rankings change between benchmarks and tasks such as object removal, background editing, and text rendering. Charpentier favors representative prompts, multiple evaluators, and metrics aligned with actual product requirements over aggregate rankings or isolated CLIP scores.
  • Pareto-efficient model selection: He evaluates quality against inference latency, price, and energy consumption.
  • Composable model compression: Pruna’s open-source optimization framework combines pruning, quantization, compilation, distillation, and caching. Different model components can require different quantization strategies, while caching and distillation reduce repeated denoising work in image and video generation.

Read the topics behind these talks

1 conference talk

References