← All speakers

Bio, Work & Ideas

Lukas Biewald

Conference affiliation: Weights & Biases · 2024

Lukas Biewald is senior vice president of AI initiatives at CoreWeave and co-founder and former chief executive of Weights & Biases, the AI developer platform CoreWeave acquired in 2025. His companies have addressed two foundational problems in practical machine learning: obtaining reliable training data and making model development reproducible.

Biewald studied mathematics and computer science at Stanford before working on search at Yahoo. In 2007, he co-founded CrowdFlower with Chris Van Pelt, building a human-in-the-loop data labeling platform that combined software and human judgment to prepare machine-learning datasets. The company became Figure Eight and was acquired by Appen in 2019.

As deep learning advanced, Biewald returned to hands-on experimentation with neural networks, taught deep learning, and spent time at OpenAI. In 2017, he co-founded Weights & Biases with Van Pelt and Shawn Lewis to give machine-learning teams better tools for tracking experiments and understanding differences in model behavior. His account of founding both companies traces the shift from solving training-data shortages to improving how engineers build and evaluate AI systems.

Weights & Biases grew from experiment tracking into a platform for model development and generative-AI applications. CoreWeave completed its acquisition in May 2025, integrating AI developer tooling with its computing infrastructure.

  • The experimental record is intellectual property. Failed prompts, discarded configurations, and unexpected model behavior contain knowledge that disappears when teams preserve only the final application. Passive experiment tracking captures that history automatically, improving reproducibility, collaboration, and iteration speed.
  • Layered evaluation makes releases dependable. Biewald favors combining mandatory checks for unacceptable failures, fast tests for everyday development, and larger evaluation suites run less frequently. Metrics should reflect actual user experience, not abstract benchmark performance.
  • Production improvements compound across techniques. For an open-source voice-assistant experiment, Biewald combined speech recognition, prompt engineering, model comparisons, synthetic training examples, and QLoRA fine-tuning on consumer hardware. His central engineering challenge was translating speech into valid function calls accurately and quickly; measured results determined which combination worked.
  • W&B Weave applies these principles to tracing, evaluating, and improving generative-AI applications. Biewald also hosts Gradient Dissent, interviewing researchers, founders, and engineers about the technical decisions behind modern AI systems.

Read the topics behind these talks

1 conference talk

References