← All speakers

Bio, Work & Ideas

Shreya Shankar

Conference affiliation: UC Berkeley · 2024

Shreya Shankar is a computer scientist, visiting assistant professor at Carnegie Mellon University, and creator of DocETL, an open-source system for analyzing unstructured documents with language models. Her research combines database systems, human-computer interaction, and AI evaluation to make automated analysis more reliable and responsive to human judgment. In 2027, she will join Carnegie Mellon’s computer science faculty as a tenure-track assistant professor, with a courtesy appointment in its Human-Computer Interaction Institute.

Shankar earned bachelor’s and master’s degrees in computer science at Stanford and researched machine-learning security at Google Brain. She subsequently became the first machine-learning engineer at Viaduct, building infrastructure for connected-vehicle data. At the University of California, Berkeley, where she completed her doctorate in 2026, she studied how engineers actually maintain deployed models: her research on production machine-learning workflows, based on interviews with 18 practitioners, documented the practical demands of validation, monitoring, versioning, and continuous iteration. Her subsequent work on interactive debugging for retrieval-augmented generation received a Best Paper award at ACM CHI 2026.

  • Agentic document processing. Shankar built DocETL after Berkeley journalists needed to analyze extensive collections of public records. Its agentic query rewriting breaks difficult document-processing tasks into smaller operations, tests alternative execution plans, and checks their results. Public defenders, climate scientists, and journalists have used the system to investigate document collections that resist conventional analysis. In a public DocETL demonstration, she showed how an agent-oriented workflow could analyze years of Hacker News discussions and produce a dashboard.
  • Data quality assertions from prompt revisions. Her SPADE research converts developers’ prompt-editing histories into automated checks: repeated attempts to prevent particular mistakes reveal requirements that applications should enforce explicitly. The research reported deployment through LangSmith across more than 2,000 pipelines.
  • Criteria drift in AI evaluation. Through EvalGen, Shankar identified a central difficulty of judging model outputs: people often discover what quality means only after examining examples, and their standards evolve during that process. Her approach combines generated evaluators with human feedback so automated judgments can track changing preferences.
  • Human-steerable data systems. Her DocWrangler research examines how people refine document-analysis workflows by inspecting intermediate outputs, revising instructions, and converting broad questions into tractable tasks. She emphasizes examining real production outputs, tracing failures to specific prompt and model versions, and recalibrating automated judges as applications change.

With Hamel Husain, Shankar teaches AI Evals for Engineers & PMs and is coauthoring an O’Reilly book on evaluation for AI engineers.

Read the topics behind these talks

1 conference talk

References