Sander Dieleman is a Google DeepMind research scientist who helped develop WaveNet and contributed to Lyria, Imagen, and Veo, generative models for music, images, and video. His work examines how neural networks represent sound, language, and visual information—and how those representations determine what machines can create.
At Ghent University, Dieleman studied neural-network representations of musical audio, completing his doctorate in 2016. A 2014 internship at Spotify applied those ideas to music recommendation. He also became a lead developer of Lasagne, a Theano-based neural-network library, and developed a winning Galaxy Zoo classification system that became published astronomy research.
At DeepMind, he contributed to AlphaGo and co-authored WaveNet, which generated speech and music directly from raw audio waveforms. Subsequent work included musical-note synthesis, the MAESTRO piano dataset, and Piano Genie, an interactive system for improvising music with an eight-button controller. His research portfolio now includes Lyria, Imagen 2 and 3, and Veo, extending his focus from audio to image and video generation.
- Learned latent representations: Dieleman argues that useful compression must preserve the spatial relationships and semantic structure generative models need. Learned autoencoders reduce image and video data to manageable latent grids without imposing the destructive compromises of conventional codecs; his writing on latent-space generation examines the resulting architectural tradeoffs.
- Diffusion as spectral autoregression: Natural images concentrate more energy at lower frequencies, while Gaussian noise spreads energy across frequencies. Denoising therefore tends to reconstruct broad composition before fine detail, giving image generation a coarse-to-fine ordering without flattening pictures into arbitrary pixel sequences. His analysis of noise schedules explores how training decisions affect this process.
- Guidance, diversity, and control: Stronger guidance improves adherence to prompts while reducing output diversity and potentially oversaturating images. Dieleman also emphasizes curated training data, distributed training, fewer-step sampling through distillation, and richer controls such as visual references and camera movement, while distinguishing general technical possibilities from specialized capabilities he does not personally develop. His AI Engineer Europe talk connects these choices to large-scale image and video systems.