Myungjong Kim is a deep learning scientist at NVIDIA developing speech-recognition systems that identify individual speakers, handle overlapping conversations, and deliver multilingual transcriptions in real time. His work spans assistive technology for speech disorders, commercial speaker diarization, and the infrastructure behind customizable enterprise speech AI.
From accessible speech to NVIDIA
Kim earned a doctorate in electrical engineering from the Korea Advanced Institute of Science and Technology in 2016, focusing on dysarthric speech recognition for people whose motor speech disorders make conventional transcription unreliable. As a postdoctoral researcher at the University of Texas at Dallas from 2016 to 2018, he investigated silent speech recognition from articulatory movements and helped study speech intelligibility among people with amyotrophic lateral sclerosis.
He subsequently worked as a speech scientist at Avoma before joining Samsung Research America, where he developed approaches to speaker diarization: determining who spoke and when, even during overlapping or fragmented conversations. His research on multi-scale speaker embeddings strengthened speaker identification from short audio segments, while Samsung’s Bixby diarization system combined overlap detection, speech separation, clustering, and corrective postprocessing.
Kim joined NVIDIA in 2022. His contributions extend from multi-speaker transcription research to practical guidance for adapting large multilingual speech models. He also participated in a joint NVIDIA session at the 2025 AI Engineer World’s Fair on enterprise speech recognition, customization, and deployment.
- META-CAT and target-speaker transcription. Kim coauthored META-CAT, which combines diarization-derived speaker information with speech-recognition representations so one architecture can transcribe multiple people or isolate a chosen speaker without a separate audio-masking pipeline.
- Cache-aware FastConformer streaming. His coauthored Nemotron 3.5 ASR fine-tuning guide describes a 600-million-parameter model supporting 40 language-locales. Its encoder reuses previously computed states, while an RNN-T decoder emits text incrementally and configurable attention controls the latency–accuracy tradeoff.
- Multilingual adaptation without catastrophic forgetting. Kim’s recent work addresses specialization for particular languages, accents, and domains while preserving broader multilingual performance, including replaying examples from other languages during fine-tuning and integrating punctuation and capitalization into transcription output.