Devendra Singh Chaplot is an artificial-intelligence researcher and adviser to Sarvam who helped build Mistral AI’s open-weight language models and Thinking Machines Lab’s training infrastructure. A founding-team member at both companies, he has worked on Mistral 7B, Mixtral 8x7B, Mistral Large, Pixtral, and Tinker, pursuing models that deliver greater capability without prohibitive training or inference costs.
From embodied intelligence to frontier models
Chaplot graduated from IIT Bombay in 2014, studying computer science and engineering with a minor in applied statistics, and completed his machine-learning doctorate at Carnegie Mellon. His doctoral research on intelligent autonomous navigation combined perception, spatial memory, language, planning, and decision-making; his open-source projects include Active Neural SLAM and object-goal navigation. Carnegie Mellon named him a 2020 Facebook Fellow in computer vision.
After completing his doctorate in 2021, Chaplot worked as a research scientist at Facebook AI Research. He joined Mistral’s founding team in 2023 and helped train Mistral 7B, Mixtral 8x7B, and Mistral Large. He subsequently led the multimodal research team behind Pixtral 12B and Pixtral Large and established Mistral’s Palo Alto office.
In 2025, he joined Thinking Machines Lab’s founding team and worked on Tinker, an API that lets researchers fine-tune models without managing distributed training infrastructure. He announced a move to xAI and SpaceX in March 2026 and subsequently became an adviser to Sarvam.
Technical convictions
Sparse mixture-of-experts models: Mixtral activates two of eight experts per token, drawing on a larger parameter pool while limiting inference costs. Chaplot prioritizes useful performance per unit of active compute over raw parameter counts.
Open models as commercial infrastructure: Public weights encourage experimentation, private deployment, customization, and community-built integrations while generating demand for hosted services and stronger commercial models.
Data quality and targeted fine-tuning: Filtering, deduplication, and dataset composition can matter more than indiscriminately adding training data. He recommends prototyping with capable general-purpose models, then fine-tuning smaller open models when production economics warrant it.
Domestic frontier-AI capability: At Sarvam, Chaplot advocates developing India’s own model-training expertise, computing infrastructure, and technical talent so systems can accommodate local languages and requirements.