Val Bercovici is chief AI officer at WEKA, focused on the infrastructure economics that determine whether AI agents can operate at scale. He argues that persistent context, cache efficiency, and memory architecture increasingly dictate the cost and responsiveness of production inference.
Bercovici held senior technical and cloud-strategy roles at NetApp, including chief technology officer-at-large, and subsequently became chief technology officer of NetApp/SolidFire. He chaired the Storage Networking Industry Association’s Cloud Storage Initiative, contributing to the Cloud Data Management Interface, and represented NetApp on the Cloud Native Computing Foundation governing board.
He later founded and led PencilDATA, applying blockchain to enterprise data integrity, and also led cybersecurity company Chainkit. At WEKA, he directs AI and open-source strategy, extending his background in storage and distributed infrastructure into the economics of agentic inference.
- AI token economics: Bercovici links token costs and inference throughput directly to storage latency, context caching, and accelerator utilization. His public writing connects these efficiencies to practical AI deployment across business and science.
- The AI memory wall: Persistent, multi-turn agents can outgrow high-bandwidth GPU memory, forcing systems to reconstruct context and spend additional compute. Bercovici treats this memory bottleneck as an architectural challenge requiring faster, larger memory tiers.
- Context-platform engineering: At AI Engineer Code 2025, Bercovici and Callan Fox introduced an open-source toolkit for modeling agent swarms and translating application-level service agreements into infrastructure objectives. Bercovici articulated the platform strategy and token economics; Fox developed the load generator and presented WEKA Labs’ research and benchmarks.
- KV-cache reuse and persistent context: Bercovici argues that short cache lifetimes and repeated prefills inflate inference costs and exhaust subscription limits. His analysis of AI unit economics connects these problems to Augmented Memory Grid, WEKA’s approach to extending inference memory through persistent, high-performance storage.