Shijia Liao is the co-founder and chief scientist of Fish Audio, where he builds expressive, instruction-controlled voice synthesis. His work gives synthesized speech recognizable character, emotional range, and conversational context while making powerful voice models available to developers.
After studying at the University of Maryland, Liao contributed to multimodal AI research at NVIDIA, including LITA, which helps video-language models locate events in time, and Eagle, which combines visual encoders to improve multimodal understanding.
Dissatisfied with the flat synthetic voices in VTuber and anime content, he began training speech models on a single consumer GPU. The experiment became Fish Speech, an open-source project that developed into Fish Audio, which he co-founded with chief executive Rissa Cao. In July 2026, the company reported a $52 million seed round, $21 million in annual recurring revenue, and more than eight million users.
Liao was first author of the 2024 Fish Speech paper, which introduced a dual-autoregressive speech architecture supporting multilingual generation and voice cloning. At AI Engineer World’s Fair 2025, he demonstrated OpenAudio S1’s ability to control vocal emphasis and emotional delivery. As the first-listed core contributor to Fish Audio S2, he extended that work into natural-language direction, multi-speaker dialogue, and streaming inference, with publicly released model weights and fine-tuning tools.
His contributions center on three connected challenges:
- Directed vocal performance: Giving creators precise control over character, emotion, emphasis, and delivery.
- Open-source voice infrastructure: Releasing speech models, training tools, and serving infrastructure for practical experimentation and deployment.
- Production-grade speech inference: Using custom CUDA kernels, FP8 quantization, continuous batching, and GPU scheduling to reduce serving costs while checking that optimizations preserve audio quality and speaker identity.