← All speakers

Bio, Work & Ideas

Stephen Batifol

Conference affiliation: Black Forest Labs · 2026

Stephen Batifol is a developer advocate at Black Forest Labs and a coauthor of the research behind FLUX.1 Kontext, an image-generation and editing model designed to maintain visual consistency through successive revisions. His career bridges production machine learning, open-source vector databases, and the development of responsive creative tools.

Batifol began as an Android developer before moving into data science and machine-learning engineering. At Wolt, he built infrastructure supporting production models, including operational predictions such as delivery-time estimates. His account of Wolt’s machine-learning platform describes consolidating fragmented deployment practices around Kubernetes, Flyte, MLflow, and Seldon Core.

He also founded the Berlin MLOps Community meetup. In 2024, he joined Zilliz as its first developer advocate in Europe, working with the Milvus open-source vector database. His local retrieval application combines Ollama, LangChain, Milvus, and Llama 3 to make French and Berlin parliamentary documents searchable through local retrieval-augmented generation.

Batifol subsequently joined Black Forest Labs, where he contributed to the FLUX.1 Kontext research paper and published FLUX model adaptations.

  • Consistency makes visual generation usable. Kontext preserves recognizable characters and products across successive edits; multi-reference image editing combines distinct images into coherent outfits, interiors, storyboards, and product compositions.
  • Speed changes creative software. Batifol champions interactive visual generation through FLUX.2 [klein], whose low-latency generation and editing enable users to adjust images continuously instead of waiting through disconnected rendering cycles.
  • Generation and representation should develop together. He explains Black Forest LabsSelf-Flow research, authored by other company researchers, as an alternative to fixed external encoders that can constrain scaling and complicate multimodal systems. Its student-teacher approach jointly develops generation and internal representations across images, video, and audio.
  • Visual models can extend into physical action. His AI Engineer Europe presentation connects world models and robotics to experimental action prediction: systems that learn spatial relationships and anticipate how objects move, extending visual generation toward embodied AI.

Read the topics behind these talks

1 conference talk

References