Dylan Patel is the founder, chief executive, and chief analyst of SemiAnalysis, the semiconductor and AI-infrastructure research firm he started building in 2020. He analyzes the physical and financial foundations of advanced artificial intelligence: accelerator architectures, memory bandwidth, networking, chip manufacturing, electricity, and the economics of operating enormous computing clusters.
An unconventional route into chips
Patel grew up around his family’s small businesses and developed an early fascination with computer hardware after repairing an Xbox. He immersed himself in online communities focused on smartphones, graphics processors, and personal computers, then spent two years as a quantitative analyst at a risk-focused financial firm before turning his independent semiconductor research into SemiAnalysis. What began as a solo publication grew into a research and advisory business covering the entire chain from chip manufacturing to cloud infrastructure. His account of that career trajectory helps explain his close attention to both engineering constraints and commercial incentives.
The infrastructure questions that define his work
Prefill and decode demand different machines. Processing a prompt is compute-intensive; generating successive tokens depends heavily on memory bandwidth. Patel applies that distinction to frontier-model serving, including continuous batching, vLLM, TensorRT-LLM, and the practical cost of deploying large open models.
Disaggregated prefill changes inference economics. Assigning prompt processing and token generation to separate accelerators reduces contention between customers and enables hardware tailored to each workload. His research on Nvidia’s Rubin CPX connects specialized accelerators, memory choices, rack architecture, and operating costs.
Context caching makes long documents affordable. Legal and contract-review applications can waste money repeatedly processing identical source material. Reusing previously computed model state reduces that expense, although caches consume substantial memory and may need to move between accelerators, host systems, and storage.
Cluster performance depends on reliability and power. Optical failures, slower individual chips, software immaturity, and electricity availability can erase theoretical hardware advantages. Patel’s comparison of H100 and GB200 systems evaluates training performance alongside downtime, energy consumption, and total ownership costs.
AI competition is industrial and geopolitical. His research on Huawei Ascend production identifies high-bandwidth memory as a potential constraint on Chinese accelerator manufacturing. His analysis of global AI infrastructure extends to export controls, Middle Eastern data-center financing, competing rack-scale architectures, and American electricity shortages.
Patel advocates hardware-software co-design: optimizing models, kernels, memory, networking, and silicon together. As inference clusters increasingly support reinforcement-learning workloads, he treats computing capacity as a shared resource constrained by manufacturing, financing, reliability, and power.