← All speakers

Bio, Work & Ideas

Vikas Paruchuri

Conference affiliation: Datalab · 2025

On this page

Vikas “Vik” Paruchuri is the founder and chief executive of Datalab, the document-intelligence company behind Marker, Surya, and Chandra. He builds open-source systems that transform PDFs, scans, forms, tables, handwriting, and other difficult documents into structured information usable by software and AI models.

From diplomacy to document intelligence

Paruchuri studied American history, worked at UPS and Pepsi, and served as a U.S. diplomat before an interest in financial-market prediction drew him into programming and machine learning. Kaggle competitions sharpened his practical skills, and he became a machine-learning engineer at edX. Seeing the limitations of lecture-driven online courses, he founded Dataquest, developing an approach to project-based technical education that emphasized solving meaningful problems over memorizing syntax.

He led Dataquest for eight years. Pandemic-era expansion and subsequent layoffs convinced him that specialization, management layers, and excessive meetings could reduce productivity even as headcount increased. In March 2024, he became chairman, appointed a new chief executive, and concentrated on document-processing research.

Paruchuri had taught himself deep learning by implementing neural-network architectures, studying foundational papers, and building Zero to GPT. While assembling training datasets, he encountered valuable material trapped in PDFs and began developing Marker and specialized document models. His account of that transition traces his progression into a research role at Answer.AI, where working with Jeremy Howard reinforced his preference for generalists, practical infrastructure, and small teams. He subsequently founded Datalab to commercialize his document-intelligence work.

Defining projects and convictions

  • Integrated document intelligence. Marker converts documents into Markdown, JSON, and HTML while preserving structure such as tables and equations. Surya handles multilingual OCR, layout analysis, text detection, and tables; Chandra targets complex forms, handwriting, and visually demanding documents. Accurate extraction depends on coordinating layout, recognition, reading order, structured output, and deployment.
  • PDF grounding over unnecessary generation. Many PDFs already contain machine-readable text. Paruchuri uses that embedded information where possible, reserving OCR and generation for material that actually requires them. Surya extends the approach through line-level grounding against PDF text and character-level bounding boxes, reducing opportunities for transcription errors and wasted computation.
  • Small generalist teams with end-to-end ownership. Developing Surya OCR with research engineer Tarun Ram Menta required handling customer discovery, datasets, architecture, training, inference, and product integration. Paruchuri’s AI Engineer presentation makes the organizational case: fewer handoffs preserve customer context and accelerate feedback. At Datalab, he favors reusable components across hosted and on-premises deployments, modular code, and server-rendered interfaces using HTMX and Alpine.
  • OCR evaluation should measure meaning. In his critique of OCR benchmarks, Paruchuri argues that exact-string matching and edit-distance metrics can penalize mathematically equivalent equations and harmless formatting differences. Document systems should be judged by whether they preserve the underlying information, not whether every character matches an arbitrarily formatted reference.

Read the topics behind these talks

1 conference talk

References