Rashi Agrawal is Head of Agentic AI and ML at Hinge Health, building conversational healthcare systems that must protect sensitive information, recognize emergencies, and remain accountable to clinicians. Her defining principle is that decisions with potentially serious medical consequences need enforceable software controls, not faith in a language model’s judgment.
Agrawal began her career at Accenture before spending nearly a decade at Yahoo, where she advanced from intern to engineering leader and worked on advertising technology, including Yahoo Gemini. She subsequently led AI engineering at GoodLeap, applying machine learning to sustainable-energy financing, before joining Hinge Health. Her career in engineering and AI leadership spans consumer-scale software, financial technology, and regulated healthcare.
She holds a master’s degree from San José State University, pursued executive education at Stanford Graduate School of Business, and founded Women In Tech AI. At Hinge Health, she brings machine-learning scientists and software engineers together to build personalized healthcare conversations and proactive communications.
How she engineers safer healthcare AI
- Privacy enforced by architecture. Protected health information should be removed during ingestion, before reaching analytics infrastructure. Separate production and development environments, together with role-based and geographic access restrictions, prevent sensitive data from spreading beyond authorized systems.
- Deterministic guardrails above the model. Emergency escalation, identity verification, and consequential clinical routing belong in application code executed before the model responds. Prompts cannot serve as authentication boundaries or dependable defenses against prompt injection.
- Continuous clinical evaluation. Automated assessments of accuracy, safety, escalation, relevance, and drift should complement member feedback and human-reviewed conversation traces. High-risk interactions merit particular scrutiny, while newly discovered failures should trigger additional monitoring.
- Validate the evaluator. Automated judges can misclassify medically appropriate guidance as unsafe. Before changing an agent after a quality score drops, teams must determine whether the response failed or the evaluator misunderstood its clinical context.
- Worst-case risk governs release decisions. Severity follows the most serious plausible harm, independent of development capacity or launch schedules. Safety defects warrant conservative decisions; minor polish issues generally do not. Meaningful accepted risks require explicit approval.
Her guardrails-first healthcare AI framework also identifies a persistent operational constraint: automated monitoring can generate signals faster than qualified people can interpret and act on them. Safe deployment therefore depends on sufficient human oversight alongside technical safeguards.