← All speakers

Bio, Work & Ideas

Karan Goel

Conference affiliation: Cartesia · 2024

Karan Goel is the co-founder and chief executive of Cartesia, where he develops real-time voice and multimodal intelligence using alternatives to transformer-based architectures. He co-created the S4 sequence-modeling architecture, a formative advance in state space models that helped establish his approach to continuous, efficient machine learning.

Raised in Delhi, Goel studied electrical engineering at IIT Delhi, completed a machine-learning master’s degree at Carnegie Mellon in 2018, and earned his Stanford computer-science doctorate in 2023 under Christopher Ré. He also worked with Snorkel AI and Salesforce AI Research, contributing to Robustness Gym, model auditing, data augmentation, and Meerkat, a system for exploring machine-learning datasets.

At Stanford, Goel collaborated with Albert Gu and Ré on structured state spaces for long sequences, developing S4 to process extended streams without prohibitive memory and computational demands. He subsequently investigated simpler diagonal state-space architectures and led research applying state-space models to raw-audio generation, bringing sequence-modeling theory closer to practical speech systems.

Goel co-founded Cartesia in 2023 with fellow Stanford researchers. The company introduced Sonic in 2024, translating that research into low-latency generative voice, and later expanded into Ink streaming speech recognition and Line, its code-first voice-agent platform.

  • Streaming intelligence, not batch responses. Goel argues that conversational assistants, robotics, interactive worlds, and on-device applications need models that absorb information incrementally and respond immediately. Their practical constraints include latency, power consumption, memory, and the cost of continuously processing audio, video, and sensor data.
  • Compression as working memory. State-space models update an internal representation as each token arrives instead of repeatedly inspecting an expanding context. Goel considers selective retention especially valuable for long, noisy inputs such as security footage, while acknowledging that compression can discard information important in shorter contexts. His AI Engineer World’s Fair talk connects that tradeoff to long-lived multimodal systems.
  • Voice as a full-stack systems problem. Natural conversation depends on timing as much as output quality. Goel’s work combines speech generation, recognition, model serving, deployment, debugging, latency measurement, and evaluation; Sonic 2.0 and Cartesia’s Series A extended that focus to controllability, efficient inference, and richer streaming architectures.
  • Developer-controlled voice agents. Line gives engineers application logic, integrations, background reasoning, local development, conversation-level debugging, and production feedback loops for building agents that can handle situations beyond rigid scripted workflows.

Read the topics behind these talks

1 conference talk

References