← All speakers

Bio, Work & Ideas

Shivam Verma

Conference affiliation: Staff Machine Learning Engineer · Spotify · 2026

Shivam Verma is a staff machine learning engineer at Spotify who builds systems that translate listening history into personalized recommendations. A former Twitter engineer, he has led work on user representations within Spotify’s AI Foundation organization, combining listener embeddings, catalog knowledge, and large language models to make music and podcast recommendations more responsive to individual preferences.

In 2025, Verma coauthored Cross-modal Adaptive Mixture-of-Experts (CAMoE) with Vivian Chen and Darren Mei. The modality-aware advertising architecture uses specialized experts, separate audio and video prediction heads, and selective training updates to prevent signals from one advertising format from overwhelming another. Its deployment increased audio-ad click-through rates by 14.5 percent and video-ad click-through rates by 1.3 percent.

His subsequent research on cold-start podcast discovery combined advertising and in-app promotions within a shared model trained on streams, clicks, likes, and follows. For less-listened-to creators, online experiments produced approximately 24 percent higher impression-to-stream rates, 27 percent more streams, and 22 percent lower effective cost per stream.

How Verma approaches personalization

  • Foundational user representations: Cross-session listening histories become embeddings that serve multiple recommendation and search systems. Representing listeners, songs, and podcast episodes in a shared space allows models to learn relationships across content formats.
  • Hierarchical Semantic IDs: Compact token sequences encode catalog items with shared broad characteristics and finer distinctions. Adapted language models can then generate recommendations as actual songs or podcast episodes, while combining proprietary catalog knowledge with their existing understanding of the world.
  • Personalized soft-token projections: Listener embeddings are projected into a language model’s representation space, giving it persistent individual context without retraining a separate model for every user. Verma describes user representations, Semantic IDs, and soft tokens as complementary components of personalized, catalog-aware systems.
  • User-steerable generative recommendation: Natural-language prompts and editable taste profiles let listeners influence recommendations directly. His AI Engineer conference presentation connects these capabilities to products including DJ, prompted playlists, and podcast recommendations.

Read the topics behind these talks

1 conference talk

References