← All speakers

Bio, Work & Ideas

Varun Krishna

Conference affiliation: SambaNova Systems · 2024

Varun Badrinath Krishna is a senior principal AI solutions engineer at SambaNova Systems whose work spans enterprise AI infrastructure, cybersecurity, and specialized language models. He develops practical approaches to making AI systems more accurate on proprietary information while reducing the hardware, latency, and operational costs of deploying them.

Krishna studied engineering at the National University of Singapore before earning a master’s degree and doctorate in computer engineering at the University of Illinois Urbana-Champaign. His research applied machine learning to smart-grid security, including detecting electricity theft, compromised meters, and anomalous consumption data. That work received recognition at QEST and CRITIS in 2015, and he became a Siebel Scholar in the class of 2018.

After research placements at ABB, IBM Research, and Cisco, Krishna worked at C3 AI, advancing from data science into technical management and deploying machine-learning applications for enterprises. At SambaNova, he helped support a 2024 AI Engineer workshop on Llama inference and enterprise retrieval, alongside Rachelle Mattern and Petro Milan.

  • AttackQA and cybersecurity evaluation. Krishna authored AttackQA, a cybersecurity question-answering dataset containing 25,335 examples with rationales. Its pipeline uses language models to generate questions, filter weak examples, and assess answers for security-operations applications.
  • Domain-specific fine-tuning for retrieval-augmented generation. His cybersecurity-focused Llama research combines specialized embeddings, model adaptation, and retrieved evidence. In the reported task-specific evaluation, the resulting Llama 3 8B system outperformed a GPT-4o-based alternative; the comparison does not establish superiority beyond that particular application.
  • Model bundling for multi-model agents. Krishna argues that assigning every model its own hardware node becomes inefficient when agents invoke multiple systems for retrieval, reasoning, validation, and synthesis. His case for co-locating models on shared infrastructure emphasizes memory management, higher utilization, fewer network handoffs, and lower deployment costs.

His healthcare-focused enterprise agent demonstration extends these concerns to graph-based retrieval, multiple language models, and sensitive clinical information.

Read the topics behind these talks

1 conference talk

References