← All speakers

Bio, Work & Ideas

Mozhgan Kabiri Chimeh

Conference affiliation: NVIDIA · 2026

Mozhgan Kabiri Chimeh is an NVIDIA developer relations manager specializing in high-performance computing, research software, and practical AI deployment. She helps developers adopt GPU-accelerated tools, from scientific simulation and CUDA debugging to running large language models locally.

Kabiri Chimeh earned a doctorate in computer science from the University of Glasgow in 2016, researching faster logic-circuit simulation across GPUs and other parallel architectures. She also holds a Glasgow master’s degree in information technology and worked on GPU-accelerated stereo vision for the European CLoPeMa robotics project.

At the University of Sheffield, she worked as a research associate and research software engineer. She contributed software development, validation, and manuscript review to FLAME GPU 2, an open-source framework for high-performance agent-based simulation. As a Software Sustainability Institute fellow, she advocated reproducible research software and supported mentorship and wider participation through Women in HPC.

She joined NVIDIA in January 2020, leading GPU hackathons and boot camps across Europe, the Middle East, and Africa. She later became a developer relations manager, supporting AI software adoption across the United Kingdom and Ireland.

Her technical work centers on two complementary concerns:

  • Making GPU software reliable. Her coauthored NVIDIA Compute Sanitizer series explains how developers can detect CUDA memory errors, uninitialized data, synchronization failures, and build custom debugging tools.
  • Measuring local inference as users experience it. Her DGX Spark benchmarking used Qwen models, vLLM, Docker, warm-up runs, GPU telemetry, and streaming instrumentation to compare throughput with time to first token. Her tests found that NVFP4 quantization made a 14-billion-parameter model deliver its first token 3.4 times faster than its unoptimized counterpart. The practical distinction is crucial: unified memory determines whether a model fits, while memory bandwidth and numerical precision shape how quickly it responds.

Read the topics behind these talks

1 conference talk

References