Ari Heljakka is the founder and chief executive of Scorable, formerly Root Signals, which builds infrastructure for evaluating and improving AI agents. His work focuses on whether autonomous systems understand their environment, follow organizational policies and take actions that advance their assigned goals.
From generative models to dependable agents
Before founding Root Signals in 2023, Heljakka co-founded the enterprise software company Dream Broker and pursued machine-learning research at Aalto University. His doctoral research on generative neural networks, published in 2020, investigated how models could produce convincing images while preserving control over their visual characteristics.
With Arno Solin and Juho Kannala, he developed PIONEER, a progressively growing generative autoencoder designed to reconstruct images and navigate their latent representations. His subsequent Deep Automodulators research, published at NeurIPS 2020, explored reconstructing images and combining characteristics from multiple examples.
Evaluate understanding and action separately. Heljakka’s agent-evaluation framework distinguishes semantic quality—grounding, relevance, consistency and policy alignment—from behavioral performance, including tool selection, valid API calls, error handling and progress toward a goal. Faithfulness to retrieved documents, he emphasizes, does not guarantee factual accuracy about the wider world.
Maintain the judges as carefully as the agents. His concept of EvalOps treats the agent workflow and its judgment workflow as separate operational systems. Automated judges introduce their own costs, latency, uncertainty and failure modes, requiring calibration, versioned evaluation suites and continual refinement.
Put evaluation inside the feedback loop. Through the Model Context Protocol, agents can consult evaluators while they work and use explanatory feedback to improve their responses. In one demonstration, a hotel-reservation agent stops recommending a competing property after connecting to an evaluator enforcing the hotel’s booking policy.
Build tests from actual failures. Heljakka advocates production-grounded evaluation: converting execution traces, observed failure patterns and human annotations into evaluators that govern future releases. Observability establishes what an agent did; evaluation determines whether that behavior met the organization’s requirements.