Anish Agarwal is the co-founder and chief executive of Traversal, which builds an AI site reliability engineer for diagnosing failures in complex production software. He is also an assistant professor at Columbia University, applying his research in causal machine learning to a problem that conventional observability tools struggle to solve: distinguishing what actually broke from everything that broke alongside it.
From causal inference to automated troubleshooting
Agarwal worked at Boston Consulting Group before earning a PhD in electrical engineering and computer science at MIT, where his advisers included Alberto Abadie, Munther Dahleh, and Devavrat Shah. His subsequent work included Microsoft Research, a postdoctoral position with Amazon Core AI, a fellowship at Berkeley’s Simons Institute, and consulting on experimental design and causal inference for Uber and TauRx Therapeutics.
With Shah and Dennis Shen, he developed synthetic interventions, extending synthetic-control methods to estimate the effects of multiple possible treatments from incomplete observations. His subsequent work on synthetic combinations extended causal analysis to combinations of interventions.
Agarwal founded Traversal with Ahmed Lone, Raaz Dwivedi, and Raj Agrawal; Matthew Schoenbauer, initially a student in Agarwal’s causal-AI class, became its first employee. In 2025, Traversal emerged from stealth with $48 million in seed and Series A funding led by Sequoia Capital and Kleiner Perkins. Agarwal framed the company’s mission around combining causal machine learning with enterprise incident-response agents in his public launch announcement.
The ideas behind Traversal
AI-generated code widens the operational context gap.Coding assistants accelerate development while leaving engineers responsible for increasingly complex systems they understand less thoroughly, shifting work toward incident response and on-call troubleshooting.
Root causes require causal reasoning. Production failures cascade across interconnected services; anomaly detectors identify correlated symptoms, while causal machine learning helps isolate the initiating change or dependency.
Telemetry needs statistical and semantic interpretation. Numerical signals reveal behavior that language models cannot reliably infer alone, while logs, code, and metadata supply context that purely statistical systems miss.
Parallel agent swarms investigate unfamiliar incidents. Coordinated agents test hypotheses and search observability data concurrently, avoiding the bottlenecks of sequential workflows and obsolete runbooks. Traversal’s work with DigitalOcean illustrates this approach to autonomous production troubleshooting, combining causal analysis, operational context, and evidence-backed incident investigation.