← All speakers

Bio, Work & Ideas

Tom Shapland, PhD

Conference affiliation: Livekit · 2025

Tom Shapland, PhD, is a founder and product leader who cofounded Tule, an agricultural technology company acquired by CropX, and Canonical AI, a voice-agent analytics startup. His work spans agricultural science, machine learning, and the challenge of making conversational AI understand when people have finished speaking.

At the University of California, Davis, Shapland’s doctoral research improved ways to measure actual evapotranspiration, the water passing from crops and soil into the atmosphere. He cofounded Tule with Jeff LaBarge to turn field-scale crop-water measurements into irrigation recommendations. A member of Y Combinator’s Summer 2014 cohort, Tule expanded into computer-vision products that estimated plant water stress from smartphone imagery. CropX acquired the company in January 2023, after which Shapland initially led its Tule business unit.

Shapland subsequently founded Canonical AI with Adrian Cowham, Tule’s former chief technology officer. They initially explored retrieval-augmented generation for specialized publishers, then shifted to semantic caching after encountering language-model latency, cost, and rate limits. Their account of Canonical’s founding traces that evolution toward AI infrastructure. Canonical later developed voice-agent observability tools for analyzing caller journeys, conversation stages, and production failures.

By AI Engineer World’s Fair 2025, Shapland had joined LiveKit as a product manager. His work on conversational interruptions concentrates on how speech systems interpret hesitation, context, and competing attempts to speak.

What makes conversations work

  • Meaning predicts the end of a turn. Conventional voice activity detection often treats silence as evidence that someone has finished speaking. Shapland argues that semantic content is a stronger predictor, supported by sentence structure, tone, and conversational context.
  • Semantic turn detection needs both sides of the conversation. He described a LiveKit model that evaluates the preceding four conversational turns and extends the silence threshold when an utterance appears incomplete. Speech-recognition systems that hear only the user can miss context supplied by the agent’s preceding question.
  • Production voice systems need controllability. Full-duplex models can listen and generate simultaneously, enabling more natural timing and acknowledgments. Shapland nevertheless emphasizes instruction following, predictable behavior, and pronunciation control; he expects improved cascaded speech pipelines and contextual turn detection to remain valuable in commercial applications.

Read the topics behind these talks

1 conference talk

References