Rachelle Mattern is an enterprise AI solutions-engineering leader focused on deploying specialized language models without sacrificing data privacy, operational control, or cost predictability. In 2024, she was director of solutions engineering at SambaNova Systems, serving the company’s global customers.
Her approach confronts a practical enterprise dilemma: general-purpose models offer straightforward integration, while smaller models tailored to legal, finance, human-resources, and coding tasks provide stronger customization but create operational complexity. With SambaNova’s Samba-1 platform, Mattern articulated how Composition of Experts addresses that tradeoff by placing multiple specialized models behind a single secure model endpoint. Requests can be routed to appropriate experts, directed to individual models, or chained across multiple systems.
She emphasized the governance required to make that architecture workable: model-level access controls restrict which teams and applications can use specific models, while independently scheduled fine-tuning accommodates different update cycles for financial policies and software development.
Mattern also connected model orchestration to SN40L’s three-tier memory architecture, which combines on-chip, high-bandwidth, and DDR memory to support serving multiple models. At AI Engineer World’s Fair 2024, she framed SambaNova’s advertised Llama 3 inference speed within those broader requirements of privacy, infrastructure, and manageability. Her collaborative workshop with Petro Milan and Varun Krishna extended that framework into enterprise retrieval-augmented generation, document search, embeddings, and application development.