← All speakers

Bio, Work & Ideas

Andres Marafioti

Conference affiliation: Hugging Face · 2026

Andrés Marafioti leads multimodal research at Hugging Face, where he develops open models and robots that run on accessible hardware. He led development of SmolVLM, helped create SmolDocling, and builds conversational systems for Reachy Mini.

Marafioti holds a doctorate in applied machine learning focused on generative speech and music. His early research included reconstructing missing sections of musical recordings. Before joining Hugging Face, he was a senior machine-learning engineer at Unity. He subsequently coauthored research on Idefics3 and openly available training datasets before leading efforts to shrink vision-language models without sacrificing practical usefulness.

Introduced in 2024, SmolVLM began as a two-billion-parameter model family and expanded to 500-million- and 256-million-parameter versions. The SmolVLM technical paper, led by Marafioti, reports that its smallest model requires less than one gigabyte of GPU memory for inference while outperforming an earlier, substantially larger Idefics model. He also contributed to SmolVLM2’s video capabilities and nanoVLM, a readable PyTorch implementation for training compact vision-language models. In 2026, he coauthored the O’Reilly book Vision Language Models.

  • Affordable, repairable conversational robots. Reachy Mini is an expressive, deliberately nonhumanoid robot intended for experimentation. Its approximately $300 configuration and $450 wireless version—with a Raspberry Pi and battery—arrive unassembled, helping owners understand, repair and modify their hardware. Marafioti argues that students and independent developers should shape human-robot interaction instead of leaving it to companies selling expensive humanoids.
  • Speech-to-speech systems designed end to end. His work on Hugging Face’s open speech-to-speech framework combines voice-activity detection, speech recognition, language-model inference and speech synthesis. Reachy Mini adds echo cancellation, camera input, movement and independently scaled inference services; experienced latency includes network and orchestration delays, not merely model speed.
  • Streaming speech synthesis for real interactions. Marafioti optimized Faster Qwen3-TTS using streaming generation, a static key-value cache and CUDA graph capture. His demonstration reported first-token latency below 200 milliseconds and approximately four-times-real-time synthesis.
  • Locally controlled embodied AI. His work on running Reachy Mini conversations locally makes the conversational stack available on users’ own hardware, including Apple Silicon, so developers can inspect and replace the components governing how a robot hears, speaks and responds.

Read the topics behind these talks

1 conference talk

References