
AI Engineer World's Fair 202516:24
The End of Awkward AI Transcriptions
NVIDIA presenters explain how Riva speech-AI systems combine FastConformer encoders with CTC, RNN-T, and TDT decoding across Parakeet and Canary model families. They describe Sortformer-based diarization and speaker-kernel integration for target-speaker and multitalker transcription, alongside word boosting, text normalization, voice activity detection,…
Travis Bartley · Myungjong Kim · Byungjoong · Jaehan
Speech and audio · Architecture