← All speakers

Bio, Work & Ideas

Soumya Gupta

Conference affiliation: ML Engineer · Uber · 2026

Soumya Gupta is an applied AI technical lead at Uber building computer-vision and generative systems for production. Her work on closed-loop multimodal evaluation helps Uber Eats improve restaurant photographs without inventing ingredients, misrepresenting portions, or diluting merchants’ visual identities.

Earlier, at Schlumberger, Gupta developed recurrent autoencoders for sensor validation in oilfield operations, research credited in a NeurIPS 2019 conference program. By 2021, she had joined Uber and was advocating common evaluation standards and benchmark datasets for industrial computer vision. Her professional background also includes study at the University of California, Berkeley.

At Uber, Gupta and Jai Chopra developed a multimodal food-photography pipeline that routes images for editing, evaluates proposed changes, and blocks unreliable results. Her contributions emphasize four practical principles:

  • Human-aligned evaluation: Representative datasets spanning geographies, dishes, and image quality establish a consistent human-labeled standard; repeated production sampling catches drift that static benchmarks miss.
  • Recall-sensitive image routing: Unnecessary edits waste computing resources and can degrade good photographs. Missing discrepancies between photographs and menu descriptions creates greater risks, including fabricated ingredients or incorrect portions.
  • Configuration-driven agent tuning: A diagnostic agent identifies mismatches with human labels, while reflection and synthesis agents propose improved configurations. Updates must pass existing benchmarks before deployment, with observability and rollback providing safeguards.
  • Faithfulness-preserving image enhancement: Image-specific prompts and multidimensional quality checks assess plating, color, portions, and fidelity. Failed images receive targeted feedback and limited regeneration attempts, measured through pass-at-K; images that remain unfaithful are withheld.

Gupta’s conference presentation positions human judgment, continual evaluation, and bounded automation as operational requirements for trustworthy generative imagery.

Read the topics behind these talks

1 conference talk

References