← All speakers

Bio, Work & Ideas

Nishant Gupta

Conference affiliation: Software Engineer, Tech Lead · Meta · 2026

Nishant Gupta is a distributed-systems engineer focused on making autonomous AI agents reliable, observable, and safe in production. In 2026, he was a software engineering technical lead at Meta, working on training and inference infrastructure associated with Meta Superintelligence Labs.

Gupta earned a master’s degree in computer science from the University of California, Los Angeles, in 2019. In 2024, he was the first-listed author of research on safely oversubscribing Meta’s datacenter capacity. The work explored dynamic idle resource leasing: assigning unused computing capacity to additional workloads while protecting reliability and service-level guarantees. That background in elastic scheduling and resource allocation informs his approach to agents whose compute requirements shift with reasoning, tool use, and workload complexity.

  • Deterministic execution boundaries. Models should propose actions while independent infrastructure validates requests, enforces policies, and controls access to production systems. Gupta connects agent reliability architecture to established distributed-systems safeguards, including controlled retries, circuit breakers, resource quotas, and isolated execution.
  • Agentic control planes. Scheduling, workload routing, shared memory, policy enforcement, observability, and evaluation need a coordinated operational layer. Gupta identifies runaway retry loops, escalating compute costs, stale reads, and conflicting updates as infrastructure failures that can masquerade as poor model reasoning.
  • Scenario-driven production evaluation. Agent performance depends on complete workflows: planning, tool execution, failure recovery, task completion, latency, cost, safety, and escalation. His approach to evaluating production agent behavior uses realistic scenarios, distributed traces, and live operational signals to expose failures that isolated answer-quality benchmarks miss.
  • Targeted human oversight. Human reviewers resolve ambiguous failures, assess consequential decisions, and calibrate automated evaluations. Continuous telemetry and feedback help direct their attention toward situations where autonomous systems cannot reliably govern themselves.

Read the topics behind these talks

2 conference talks

References