State of the Union: Why Local, Why Now
Nader Khalil · Alex Cheema · Matthew Berman · Ahmad Osman · Joseph Nelson
AI Engineer World's Fair 2026 · 44:29
Local AI and distributed inference
EXO Labs builds exo, open-source software that connects Macs and workstations into a local AI inference cluster. Developers and organizations can pool device memory to run models too large for one machine, load models from Hugging Face, and manage their cluster through a dashboard. Automatic device discovery and model partitioning reduce manual setup, while OpenAI-, Claude- and Ollama-compatible APIs let users connect existing clients to locally hosted models.
Founded in 2024 and based in London, the company was co-founded by Alex Cheema, its CEO, and Mohamed Baioumy. Its engineering work includes separating prompt processing from token generation across different hardware. In its DGX Spark and Mac Studio implementation, EXO assigns these phases to devices with different compute and memory-bandwidth strengths, streaming the model’s KV cache between them while computation continues. This approach coordinates both memory capacity and the work performed during an inference request.
EXO Labs also produces the speed and evaluation data behind local.ai, a reference for choosing local AI setups. It compares combinations of models, hardware, agent harnesses, inference engines and configurations using task quality, completion time and cost. This helps people buying hardware or configuring existing machines assess an entire setup rather than relying on token throughput alone.
Nader Khalil · Alex Cheema · Matthew Berman · Ahmad Osman · Joseph Nelson
AI Engineer World's Fair 2026 · 44:29
Affiliations reflect their AIE appearances, not necessarily current employment.