← All speakers

Bio, Work & Ideas

Max Ryabinin

Conference affiliation: Together AI · 2026

Max Ryabinin is vice president of model shaping at Together AI, developing infrastructure that makes advanced language models easier to train, customize, and deploy. His work ranges from decentralized open-source systems for sharing computational resources to transformer training across five-million-token contexts.

After machine-learning internships at Replika and Yandex Translate, Ryabinin became a senior research scientist at Yandex Research. He completed a doctorate in decentralized deep learning at HSE University in 2023 and taught efficient deep-learning systems at HSE and the Yandex School of Data Analysis.

He co-created Hivemind, an open-source PyTorch library for decentralized deep learning across heterogeneous machines and unreliable networks. Related projects, including DeDLOC and SWARM Parallelism, addressed slow connections, uneven hardware, and participants joining or leaving during training.

In 2021 and 2022, Ryabinin chaired BigScience’s engineering and scaling working group, contributing to the collaboration behind the multilingual BLOOM model. He also helped develop Petals, which distributes large-model inference and fine-tuning across internet-connected computers. His research on distributed inference details fault tolerance and adaptive workload allocation across geographically dispersed hardware.

Ryabinin joined Together AI as a distinguished research scientist before moving into research-and-development leadership. He helped lead the company’s engineering and research operations in Amsterdam and now oversees model shaping.

  • Training across five-million-token contexts. Ryabinin combines established techniques—including fully sharded data parallelism, DeepSpeed-Ulysses, FlashAttention, activation checkpointing, CPU offloading, and chunked computation—to address memory bottlenecks throughout the transformer. His AI Engineer Europe presentation distinguishes attention’s quadratic computational demands from the activation-memory growth that can independently derail training.
  • UPipe and headwise chunking. With Ravi Ghadia, Maksim Abraham, and Sergei Vorobyov, he coauthored Untied Ulysses, introducing UPipe to process attention heads in smaller groups and reuse intermediate buffers. The approach balances memory consumption against throughput and enabled Llama3-8B training with five-million-token contexts on eight H100 GPUs.
  • Production-ready model customization. Ryabinin helped develop Together AI’s fine-tuning platform, supporting long-context training, conversational data, preference optimization, continued training, and customer control over model weights.
  • Learning without reliable answer checkers. His coauthored Relativistic Adversarial Reasoning Optimization trains a policy and comparative critic together, allowing models to improve from expert examples when task-specific correctness cannot be mechanically verified.

Read the topics behind these talks

1 conference talk

References