Arjun Bansal is the co-founder and chief executive of Everest by Log10, which develops artificial-intelligence systems for clinical documentation, regulatory submissions, and other demanding life-sciences workflows. Previously, he co-founded Nervana Systems, the deep-learning company Intel acquired in 2016, and led AI software and research at Intel.
From neuroscience to enterprise AI
Bansal studied computer science at Caltech, earned a neuroscience doctorate at Brown University, and completed postdoctoral training at Boston Children’s Hospital and Harvard Medical School. His research combined machine learning with neurophysiological data collection and analysis, grounding his technical career in both computational methods and biological experimentation.
At Nervana Systems, Bansal led algorithms, machine-learning software, and data science as the company developed an integrated deep-learning hardware and software stack. Following its acquisition, he became an Intel vice president overseeing AI software and research, including healthcare partnerships with GE, Siemens, and Philips.
He subsequently co-founded and led XOKind, whose Una travel-planning assistant combined collaborative planning with personalized recommendations. He later co-founded Log10 to tackle a fundamental obstacle to deploying AI applications: measuring whether their outputs actually meet professional standards.
Teaching AI systems to recognize quality
Bansal’s work centers on human-calibrated automated evaluation: preserving expert judgment while making quality assessment fast enough for production systems. His approach has several distinctive components:
Reviewer-specific evaluation models. Generic AI judges can favor longer responses, their own outputs, or whichever answer appears first. In research co-authored with Ansup Babu, Bansal developed evaluators trained against explicit grading criteria and individual reviewers’ judgments; reported summarization experiments reduced absolute error by approximately 44 percent.
Synthetic bootstrapping from scarce expert feedback. Those experiments also used small sets of human-reviewed examples to generate additional training data, allowing evaluators initialized with roughly 25 to 50 expert labels to approach the performance of systems trained on substantially larger labeled datasets.
AutoFeedback as production infrastructure.Log10’s open-source integration and evaluation tooling connects application traces, grading rubrics, automated assessments, and model comparisons. At AI Engineer World’s Fair 2024, Bansal explained how these feedback loops can detect hallucinations, prioritize human review, curate training data, and improve prompts or fine-tuned models.
Through Everest, Bansal now applies this evaluation-first approach to clinical-study reports, safety documentation, and regulatory submissions—workflows where specialized expertise, traceability, and reliable review directly determine whether generated material is usable.