← All speakers

Bio, Work & Ideas

Neil Zeghidour

Conference affiliation: Co-founder & CEO · Gradium AI · 2026

On this page

Neil Zeghidour is co-founder and chief executive of Gradium, which builds speech models and infrastructure for real-time voice applications. His research helped establish the technical foundations of modern generative audio, from Google’s SoundStream and AudioLM to Kyutai’s Moshi, a conversational model capable of listening and speaking simultaneously.

From audio compression to voice agents

Zeghidour earned a doctorate in machine learning from École Normale Supérieure in Paris, following master’s degrees in machine learning and quantitative finance. He spent three years at Facebook AI Research before joining Google, where he built and led a generative-audio research team.

At Google, he was lead author of SoundStream, a neural codec for compressing speech, music, and other audio. He subsequently co-authored AudioLM, which generates coherent speech and music using discrete audio representations, and SoundStorm, which accelerated audio generation.

He later helped create Kyutai, an open-research laboratory in Paris, and co-authored Moshi and Hibiki, a simultaneous speech-to-speech translation system. His recollection of an early Moshi demonstration highlights two central challenges: audio instruction data and multistream architectures.

In September 2025, Zeghidour co-founded Gradium with Olivier Teboul, Laurent Mazaré, and Alexandre Défossez. The company launched publicly that December with $70 million in seed funding and a focus on voice infrastructure, including streaming speech recognition, synthesis, and voice cloning.

What natural voice AI actually requires

  • Full-duplex speech: Human conversations include interruptions, overlapping voices, and brief acknowledgments. Moshi can listen and respond concurrently, avoiding the rigid conversational turns that cause many voice assistants to stumble.
  • Production-ready conversational agents: Natural speech alone cannot replace dependable tool use, safety controls, observability, and personalization. Zeghidour now acknowledges that conventional speech-recognition, language-model, and speech-synthesis pipelines remain easier to inspect and deploy than fully integrated speech-to-speech systems.
  • Application-level latency: External searches and tool calls can delay conversations more severely than speech synthesis. His approach keeps an assistant speaking naturally while information is retrieved, then incorporates the result without an awkward pause.
  • Paralinguistic understanding: Hesitation, discomfort, tone, and other vocal signals convey information that transcripts discard. Speech models must be trained to recognize and act on these cues, not merely receive audio inputs.
  • Gradium Phonon: His on-device text-to-speech model runs on a smartphone CPU, making voice applications less dependent on cloud inference while keeping sensitive interactions local. The goal is practical consumer-scale voice AI without prohibitive synthesis costs.

Read the topics behind these talks

1 conference talk

References