← All speakers

Bio, Work & Ideas

Benjamin Cowen

Conference affiliation: Modal · 2026

Benjamin Cowen is a forward-deployed machine learning engineer at Modal who helps companies decide when specialized AI products need their own models—and how to train and deploy them without managing dedicated computing clusters. His work builds on a research career spanning mathematical optimization, scientific imaging, acoustic sensing, and fault-tolerant computing.

Cowen studied applied mathematics at Case Western Reserve University, completing his bachelor’s degree in 2014 and master’s degree in 2015, before earning a doctorate in electrical engineering from New York University in 2019. His research explored numerical optimization and signal processing, including methods for reconstructing useful information from incomplete measurements; during his doctoral training, he also worked on autonomous vehicles at NVIDIA.

At NYU, he co-developed LSALSA, which transforms an iterative sparse-coding algorithm into a trainable neural architecture for faster signal separation. He also coauthored an alternative to backpropagation that divides neural-network training into smaller optimization problems through online alternating minimization.

He subsequently worked as a senior scientist for a U.S. Department of Energy contractor and became an assistant research professor at Penn State’s Applied Research Laboratory. His projects included piezoelectric ceramics research and AirSAS, an in-air synthetic-aperture sonar apparatus designed to produce reproducible acoustic datasets and test whether models respond to meaningful physical differences instead of environmental artifacts.

At Modal, Cowen works on performance engineering and infrastructure for demanding machine-learning applications. He also coauthored research on fault-tolerant training with idle GPUs, investigating how distributed fleets can redirect unused accelerators toward workloads that tolerate interruptions.

  • Fine-tuning as a product decision: Specialized models become attractive when API costs erode margins, enterprise customers require tighter latency or throughput, or application-specific evaluations stop improving through prompting.
  • Evaluation data before model training: Production feedback, credible evaluations, and existing agent harnesses can supply the foundations for supervised fine-tuning or reinforcement learning; without them, customization lacks a reliable target.
  • Serverless training and inference: On-demand GPU containers support parallel experimentation and hyperparameter searches, while autoscaling tools such as vLLM, SGLang, and NVIDIA Triton Inference Server make customized models practical to serve.
  • Sandboxed reinforcement-learning rollouts: Isolated execution environments let agents practice application-specific tasks concurrently, connecting training feedback directly to real product behavior.

Cowen’s AI Engineer Europe session crystallized his practical priority: give application teams control over model behavior and economics while preserving the rapid iteration that made general-purpose APIs appealing.

Read the topics behind these talks

1 conference talk

References