Contents
  1. I. Purpose and development
    1. Healthcare work and settings
    2. Intended contribution
    3. Continuing lines of development
  2. II. Applications and completed work
    1. Documentation and total work
    2. Administration and access
    3. Detection and clinical response
  3. III. Records, exchange, and permitted use
    1. The record as an account of care
    2. Preserve meaning through FHIR
    3. Access and permitted use
  4. IV. Patient-specific conclusions
    1. Applicability beyond relevance
  5. V. Responsibility and integration
    1. Meaningful clinical review
    2. Accepted clinical handoffs
    3. Commitment at the point of care
  6. VI. Evidence for use
    1. Validate the task and population
    2. Errors, probabilities, and workload
    3. Evaluate the complete intervention
  7. VII. Operation and recovery
    1. Monitor changing conditions
    2. Recovery and affected work
  8. Check understanding
  9. Open questions
  10. Selected talks
  11. References
  12. Talk library
← All topics

AI in Healthcare

AI can help healthcare teams prepare documentation, find relevant information, identify concerns, and coordinate work. The engineering task is to connect that capability to a service people can actually use. A fluent note, an accurate prediction, a successful record exchange, and an accepted handoff each accomplish something different. This chapter follows those distinctions from intended use through integration, evaluation, and recovery.

I. Purpose and development

Healthcare work and settings

Healthcare involves assessing people, delivering care, documenting what happened, and coordinating what happens next. A clinical workflow comprises the physical and mental tasks people perform within and across care environments. Some tasks occur together; others wait for information or another person. Care coordination organizes those responsibilities and communications across participants. Improving one task helps only if the surrounding work still reaches its intended result.

A care setting is the environment in which care occurs. A patient population is the group an application is intended to serve. These definitions become requirements: which people can use the interface, what information is available, and who can respond when assistance is insufficient. The following comparison identifies design considerations, not universal staffing arrangements or response deadlines.

SettingWork and informationResponsibility to establish
Outpatient visitsCare without hospital admission; preparation, consultation, and later follow-up may occur separately.Ownership of results and unfinished work after the visit.
Inpatient careCare during hospital admission; observations and assessments accumulate across a team.Who reviews a concern and who assesses the patient.
Emergency accessTriage prioritizes attention by urgency rather than supplying a complete diagnosis.The route to an appropriately urgent clinical response.
Remote follow-upInformation arrives through a mediated interaction rather than an in-person examination.Eligibility for the interaction and access to clinician-delivered care.

Population limits can be concrete. A 2024 prospective study of the Dora R1 telephone assistant examined follow-up after routine cataract surgery, with an ophthalmologist supervising calls in real time. It excluded non-English speakers and people with hearing or cognitive difficulties. That study therefore cannot establish suitability for everyone who might receive a telephone call. Likewise, domain expertise must match the task: experience as a physician does not necessarily include experience performing medical coding.

Intended contribution

Intended use connects a specific task to its users, patient population, setting, available information, output, permitted actions, and exclusions. Start there rather than with a model or an agent framework. Clinical decision support supplies information at useful points in care to assist decisions. It can take the form of reference information, reminders, summaries, templates, or rules; it need not involve AI. The existing workflow or a simpler rule is therefore a meaningful alternative.

A learned prediction is an estimate produced by behavior fitted from examples rather than entirely authored rules; Machine Learning Fundamentals explains that distinction. A generated answer is a composed response, not necessarily a passage retrieved intact. In retrieval-augmented generation, selected records or references supply information for that composition. Neither mechanism determines who may act on the output.

For example, a bounded result-review assistant might prepare a source-linked summary for authorized clinicians using records available for the selected patient and encounter. Its output remains a draft; missing required information is surfaced for review; it cannot independently place orders. This specification identifies both useful work and a limit on delegation. Displaying a concern, moving it into a review queue, and committing a change to a record are separate operations, even when they originate from the same model output.

Medical-purpose terminology also follows intended use rather than product branding. The International Medical Device Regulators Forum’s Software as a Medical Device definition concerns software performing an intended medical purpose without being part of a hardware medical device. Calling an application a copilot does not settle its local regulatory status, risk class, or authorization requirements.

Continuing lines of development

Healthcare AI developed along several lines because reasoning, information delivery, pattern recognition, and documentation solve different problems. These turning points explain the continuing division of work.

LandmarkContribution
Ledley–Lusted framework — July 3, 1959Robert Ledley and Lee Lusted’s diagnostic-reasoning paper separated logic, probability, and treatment value judgments: foundations for assistance, not a deployed diagnostic service.
HELP — around 1970; 1975At LDS Hospital and the University of Utah, a common clinical database and then a medical-decision language connected decision support to incoming patient information.
MYCIN — 1972 beginnings; 1974 implementationStanford’s rule-based consultation project, implemented in Edward Shortliffe’s dissertation, made expert knowledge explicit, inspectable, and modifiable.
Learned retinal-image detection — November 29, 2016Varun Gulshan, Lily Peng, and collaborators demonstrated detection from learned visual features, leaving clinical use and patient benefit as further questions.

Contemporary ambient documentation applies language assistance to producing records for later care and billing. It addresses another persistent task rather than replacing diagnosis or hospital decision support. Explicit rules, timely information, learned detection, and documentation assistance remain useful for different parts of healthcare work.

II. Applications and completed work

Documentation and total work

A clinical note records information about care, such as the history discussed, findings, assessment, and plan. It becomes part of the electronic health record, or EHR: the digital record accumulated across care interactions. Ambient documentation captures a care conversation and generates a draft for professional checking and correction. The task ends with usable documentation, not with generated text.

Checking must cover both unsupported additions and consequential omissions. A sentence can sound plausible while introducing something the encounter did not establish; a concise draft can omit information needed by the next clinician. Notes also support billing, so errors can travel into both clinical and administrative work. From Ambient Documentation to Clinical Intelligence explains these downstream uses.

Inspectability can reduce the distance between a disputed sentence and its evidence. Abridge’s linked-evidence interface connects selected note text to transcript passages and original audio. The transcript and note remain derived representations: a link helps a clinician inspect the source but does not establish accurate speech recognition or valid clinical interpretation.

This fictional encounter separates generation of a transcript and an unfinalized draft from reviewer navigation back to the same source segment. Source links enable checking; they do not certify transcription or interpretation.

A 24-week randomized trial involving 66 ambulatory practitioners in one health system introduced the documentation-only tool at staggered, randomly assigned stages—a stepped-wedge design. Compared with usual practice, it reported less time spent on notes and lower work exhaustion/interpersonal disengagement. Participants were voluntary early adopters and knew when they were using the tool. Unedited-note quality was assessed by a model judge, an AI model scoring the notes, rather than comprehensive independent clinical review. These findings concern documentation and practitioner outcomes, not improved diagnosis or patient health.

Administration and access

Administrative assistance organizes access and payment rather than determining a diagnosis. A referral requests evaluation or treatment from another clinician or service. Intake assembles the clinical question, relevant history, and records for specialty review. The reviewer may advise, schedule care, or request workup—additional information or investigations needed to assess the patient. The referring team obtains missing material, watches for requested results, and resubmits the referral for review. When a visit is needed, scheduling and follow-up must remain tracked: an appointment does not establish attendance or completed care.

Within this process, a proposed AI intake assistant could organize supplied records and flag missing requested material for the referring team. Its value should be compared with the existing workflow. An AHRQ-supported non-AI referral evaluation found shorter initial waits in most studied medical clinics, but interviews identified shifted work and appointment-selection difficulties. Faster intake alone would not establish better access to completed care.

Payment work has similar boundaries. A payer is the organization financing covered care; prior authorization is a process for reviewing a proposed service against coverage requirements. A denial appeal seeks reconsideration of a rejected request or claim. Preparing a clinical appeal requires reconciling patient records, care guidelines, and payer policies under a deadline—not merely generating a persuasive letter. In the announced Anterior–Stellarus workflow, AI assists research and policy application while licensed clinicians make final determinations. That division of work does not itself demonstrate better access.

Detection and clinical response

Screening looks for a target condition in an intended population. Diagnosis determines what condition is present. Prognosis estimates a future outcome. Triage prioritizes attention by urgency, also called acuity. These tasks produce different outputs, but all require clarity about the response that follows. An urgency category is not a complete diagnosis, and identifying risk does not establish which intervention will help.

Hospital deterioration monitoring makes the response dependency visible. Kaiser Permanente describes Advance Alert Monitor as a coordinated service: threshold crossings prompt off-site nurse review, a reviewing nurse contacts the bedside rapid-response nurse, and bedside assessment informs a care-team decision that considers patient preferences. This operating account explains responsibilities; it is not a model-only intervention or, by itself, a causal estimate of patient benefit.

Outpatient screening must connect image interpretation to care beyond the screening visit. Gulshan, Peng, and collaborators’ 2016 retinal-image study demonstrated detection using learned features rather than explicitly programmed lesion detectors. It tested different operating thresholds against ophthalmologist judgments and called for clinical-setting and patient-outcome studies. IDx-DR’s 2018 pivotal study, led by Michael Abràmoff and colleagues, examined a bounded diagnostic workflow in ten primary-care offices. Trained office staff acquired images, and the system supplied image-quality guidance and a diagnostic result without specialist interpretation of every screening image. This prospective study tested more than classification of an existing image collection, but did not establish completed follow-up or prevention of vision loss.

ACCESS, reported by Risa Wolf and colleagues in 2024, separates screening from subsequent care. At two pediatric diabetes sites, point-of-care AI screening at the visit was compared with referral and scripted education. Follow-up among intervention participants with positive results used a different denominator from screening in either arm. Screening completion and indicated follow-up are different outcomes.

Screening completion and indicated follow-up

ACCESS: youth with diabetes at two pediatric sites. Positive-result follow-up belongs to a subgroup of the intervention cohort.

Screening and follow-up have different denominators81 of 81 intervention participants screened at the visit, including 25 with positive results. Of those 25, 16 attended eye care within six months. Separately,18 of 82 analyzed controls completed screening within six months.One participant scale for every bar; each denominator is stated explicitly.Intervention · screening at the visit81/81 screened25 positive results within the 81 screenedFollow this subgroup onlyPositive-result subgroup · eye-care attendance within six months16/25 attended9 without recorded attendance in this window; this does not establish disease absence.Separate analyzed controls · screening within six months18/82 screenedReferral and scripted education; this is not a positive-result follow-up rate.
ACCESS studied youth with diabetes; the device was not labeled for this pediatric population, and specialists overread all study images. Screening, eye-care attendance, completed treatment and preservation of vision are different endpoints; the study did not measure vision preservation.

More generally, high risk does not imply large treatment benefit: an adverse outcome might remain likely with or without the proposed intervention. Prediction and causal effect are separate claims. Patient-facing assistance needs the same discipline. A remote follow-up conversation can collect information and request clinical involvement, but nuanced symptoms and access to clinician-delivered care remain part of the task.

III. Records, exchange, and permitted use

The record as an account of care

An EHR contains accounts of observations and care, not a complete state of the patient. An encounter is a particular interaction with a healthcare service. Patient identity establishes whose information is represented; encounter identity distinguishes which care interaction it concerns. Notes describe events and interpretations, orders request activities, and results report findings. Combining them without preserving those roles changes their meaning.

Medication records illustrate the difference. A prescription contains an order and instructions; dispensing records a supplied quantity; administration records that medicine was given; a medication statement reports that it was taken or given. A prescription is not proof of administration. Similarly, a statement that a symptom was denied, a statement that it occurred previously, and no statement about it are not interchangeable.

Provenance records where information came from and how it was produced or revised. Preserve it alongside the information; Data Quality and Curation explains the general principle. For a clinical result, distinguish when the observation applies, when a result version became available to providers, and when the application received it. A result collected at 09:00 but issued at 10:15 cannot support a prediction made at 09:30. Later corrections must not silently replace the historical account of what was available.

Collection, issue, receipt and retained snapshots

One investigation, with example timestamps. All three application snapshots remain available as historical records.

A receipt enables its own decision snapshotCollected 09:00. Version 1 issued 10:15 and received 10:20. Correction version 2 received 11:05, issue time unspecified. Snapshot 09:30 has no result; snapshot 10:20 version 1; snapshot 11:05 version 2. Earlier snapshots are retained.Clinical eventProvider issueApplication receipt09:00 collected10:15 version 110:20 version 111:05 version 2Preserved snapshots09:30No result available10:20Version 1 retained11:05Version 2 availableHorizontal position follows time. Correction issue time is unspecified. Collection alone enables no result snapshot.
Collection, provider issue and application receipt are different events. Arrows connect receipts only to the snapshots they enable. The correction’s issue time is unspecified; its arrival changes new work without changing what earlier decisions knew.

Status and absence also carry information. A report may be preliminary, final, or corrected. A missing measurement might be unperformed, not asked about, temporarily unknown, withheld, or lost through an error. FHIR supplies distinct data-absent-reason codes for such cases. If no reason is recorded, the application cannot invent one. Information may also exist only in narrative text or another institution’s records; an absent smoking-history code, for example, does not identify a nonsmoker.

Preserve meaning through FHIR

Interoperability means exchanging information that recipients can use with its intended meaning. Fast Healthcare Interoperability Resources, or FHIR, provides typed building blocks called resources, connected through references. A profile constrains a resource for a particular use—for example, by specifying required content or terminology. The FHIR R5 overview describes these mechanisms. Actual integrations must establish their deployed version, profiles, and supported behavior; shared syntax alone does not establish complete or clinically consistent records.

Consider a result-review application. A ServiceRequest describes a requested investigation. A DiagnosticReport can reference that request, its encounter, and Observation resources containing individual results. Patient references identify whose information is represented. These are links between distinct records, not interchangeable copies of one document. A review draft should retain its source-result version so that the basis of the summary remains inspectable.

A result retains its context

Example

A review draft derives from a result version, while resource references preserve the investigation and patient context.

Example resource links, not universally mandatory fields. Solid reference arrows point to the referenced resource; the dashed derivation arrow points from the selected result version to its application draft. The draft is not a required FHIR resource type.
Read the diagram as text
  • ServiceRequest. Requested investigation.
  • DiagnosticReport. Report connecting the investigation and its results.
  • Observation, selected version. Identified result version with applicable time, status and typed content.
  • Patient. Whose information is represented.
  • Encounter. The particular care interaction.
  • Review draft. Application draft retaining its source-version identity.
  • DiagnosticReportServiceRequest: Reference: basedOn.
  • DiagnosticReportEncounter: Reference: encounter.
  • DiagnosticReportPatient: Reference: subject.
  • DiagnosticReportObservation, selected version: Reference: result.
  • Observation, selected versionPatient: Reference: subject.
  • Observation, selected versionReview draft: Derivation: summarized into.

An Observation separates what was observed from its result and optional interpretation. Results can use different datatypes or appear in components; a missing top-level number is not a normal finding. LOINC, Logical Observation Identifiers Names and Codes, supplies shared identifiers for observations. Maintained by the Regenstrief Institute, it helps recipients interpret tests originally named by local codes. LOINC identifies the observation; FHIR transports its representation.

Preserve a code’s terminology system and version; matching display labels do not establish equivalent meanings. UCUM, the Unified Code for Units of Measure, supplies coded units that can support compatible conversions. Preserve comparators too: a result below a value is not exactly that value. Matching units alone does not establish the same measurement, specimen, or method. The FHIR datatype definitions make these representation distinctions explicit.

SMART on FHIR adds application-launch and authorization capabilities. An EHR launch can provide selected patient and encounter context, while declared server capabilities describe what is supported. This is separate from the resource representation. Existing HL7 Version 2 feeds exchange messages for activities such as laboratory orders and patient transfers. Converting them to FHIR still requires local mapping because implementations differ in field use and placement. A new API does not eliminate that interpretation work.

Access and permitted use

Purpose limitation constrains processing to an approved purpose. Data minimization limits information, recipients, and retention to what that purpose requires. Secondary use means another use, such as developing models from material collected for care. Reading an EHR, recording a conversation, sending it to a service, retaining a draft, and reusing it for development require separate consideration. Approve actual service paths explains why the product, feature, endpoint, recipients, and settings matter more than a provider name.

For each artifact, record the approved purpose, recipients, permission basis, retention rule, and accountable owner. Do not assign one lifetime to the entire encounter.
ArtifactDecision to document
Source recordWhich patient data and operations the application may access.
Encounter audioRecording arrangements, processing recipients, and disposition.
Draft noteAuthorized reviewers, record-entry process, and draft retention.
Service requestFields sent, actual processing path, logs, and reuse terms.
Audit entryAction, actor, authority, source/version references, and protected evidence access.

NHS England’s ambient-scribing guidance calls for explaining the tool and respecting patient objections. In that setting, retain a workable documentation route without optional recording. This is scoped guidance, not a universal consent rule.

Removing obvious identifiers does not settle reuse or eliminate identification risk. HHS guidance describes HIPAA’s Safe Harbor and Expert Determination methods, including identifiers embedded in free text, and acknowledges residual risk even after proper de-identification. A name-stripping function is therefore not a compliance determination. Similarly, useful auditability does not require copying every sensitive payload into ordinary logs: preserve attributable actions and references, with evidence access and retention governed separately.

IV. Patient-specific conclusions

Applicability beyond relevance

A clinical guideline offers recommendations for specified circumstances. Finding a relevant passage is not the same as establishing that its circumstances hold for this patient. Evidence sufficiency concerns whether the available information permits the requested conclusion. Here that requires matching patient identity, current status, relevant history, exclusions, source dates, and local use conditions.

The HL7 clinical-practice-guideline guide separates describing a patient from proposing an activity. It also separates eligibility for a pathway from enrollment: preferences or other clinical considerations can prevent participation even when eligibility holds. An application should therefore keep required premises visible instead of compressing a guideline and a record into an unexplained yes.

This schematic comparison illustrates evidence states without specifying medical eligibility criteria.
Premise neededRecord evidencePermitted interpretation
A particular investigation has a resultA linked result is present.Summarize that result with its status and source.
The result has the required final statusThe available report is preliminary.The final-status requirement is not established.
A relevant historical condition is absentThat history is undocumented.The condition remains unknown, not absent.

Clinical meaning can fail even when terminology matches. In Lovejoy’s prior-authorization example, an answer treated imaging findings as establishing suspected disease, while the speaker interpreted the existing confirmed diagnosis as changing the guideline answer. The lesson is to evaluate the actual criterion with the history, not to generalize that interpretation to every use of the word suspected. Source identity matters too: a patient assertion and a clinician’s documented finding should not become indistinguishable merely because both entered a summary.

Missing premises can guide the next question. MYCIN worked backward from a goal through expert-authored rules and asked for facts it could not derive. Its explanation subsystem exposed supporting rules and reasons for questions. Its certainty factors represented degrees of belief, not conditional probabilities. The medical knowledge was separate from the execution machinery: removing that knowledge base later produced EMYCIN, a framework for building expert systems in other domains. Neither an inspectable rule trace nor a modern source-linked answer establishes that the underlying medical interpretation is correct. Abstention withholds an unsupported claim; it can still preserve a useful factual summary and identify the information or clinical review required before answering further.

V. Responsibility and integration

Meaningful clinical review

Review is work, not a safety property conferred by an approval button. A reviewer needs suitable expertise, the original evidence, time to inspect it, and authority to reject the output. Expertise should match the workflow and its failure modes. The person operating the software, the person assessing clinical content, and the person authorized to change care need not be the same person.

Assign these responsibilities explicitly, even when one person holds several roles.
RoleResponsibility
Service operatorMaintain functioning interfaces, access controls, and incident response.
Clinical assessorEvaluate content against patient evidence and the clinical task.
Action authorizerApprove the particular consequential operation within their authority.
Unresolved-work ownerEnsure pending review or follow-up receives an accepted disposition.

Automation bias is inappropriate reliance on automated advice. In a web-based experiment with 223 dermatologists, participants assessed 24 cases before and after AI advice, including five deliberately incorrect recommendations. Some initially correct answers became wrong after erroneous advice; useful advice was also rejected. This controlled task did not measure patient outcomes, but it shows why oversight assessment must distinguish accepting helpful advice from rejecting harmful advice.

Alert fatigue concerns reduced responsiveness associated with burdensome alerting. A retrospective primary-care study associated more reminders per encounter and repeated reminders for the same patient with lower acceptance. It did not establish a general workload effect or assess whether overrides were appropriate. Lowering override rates is therefore not automatically a safety improvement.

Make corrections diagnostically useful. A review surface can place the source record, guideline, and requested conclusion together, then collect why an answer is wrong rather than only an incorrect label. Such critiques help distinguish missing evidence, misunderstood criteria, and unsupported additions. Review before consequential commitment develops the interface mechanics. Patients and staff also need a route to report disputed content to an accountable reviewer and learn its disposition.

Accepted clinical handoffs

A clinical handoff transfers information, responsibility, and authority. Under AHRQ TeamSTEPPS guidance, the sender remains responsible until the receiver acknowledges understanding and acceptance.

For application tracking, distinguish requested, received, accepted, and resolved. Assign local deadlines and overdue handling, including after-hours coverage. Workflow Automation develops this coordination contract.

A patient-facing voice demonstration makes the boundary concrete: after additional symptom information, the assistant recommends immediate nurse involvement. The recording demonstrates a request, not completion of a transfer. Likewise, a code-controlled routing layer can ensure that a selected clinical path runs before ordinary conversational generation, but deterministic execution does not establish correct recognition of urgency. Both detection and accepted takeover need assessment.

Commitment at the point of care

A workflow trigger starts consideration of work: a clinician request, incoming result, or scheduled review. Point-of-care integration determines where assistance appears, who receives it, and how long it remains relevant. A draft is not finalized documentation; a suggestion is not an order; delivery of an alert is not a completed response.

SEIPS, the Systems Engineering Initiative for Patient Safety model introduced by Pascale Carayon and colleagues in 2006, connects people, tasks, tools, organizational conditions, and physical surroundings to care processes and outcomes. Its work-system perspective explains why a technically correct feature can still create interruptions or transfer workload. Review time and staffing belong in the design. HELP’s earlier data-triggered and time-based support illustrates the enduring importance of delivering information within actual work.

Acceptance must also account for change. A clinician may review a draft before a new result arrives or another user edits the target record. Refresh clinical context when relevant premises change; separately check authority and target-record concurrency. SMART distinguishes selected patient context from granted resource operations. Approval in the application does not create an EHR permission.

FHIR R4’s version-aware updates use ETag and If-Match to reject overwriting a resource whose version changed. Conditional creation can avoid another resource when specified matching criteria identify an existing one. Server support must be checked, and these protections do not prove clinical appropriateness or freshness of every source used by a draft. Authorization at the protected operation and real integration tests cover the underlying mechanics. Test wrong-patient context, intervening changes, rejection, and uncertain submissions as part of the clinical workflow.

Reviewed is not yet committed

Example

Clinical validity, current authority, and target-record concurrency are independent acceptance conditions.

A proposed path to a confirmed record change, not completed care. Source freshness is checked separately from the target resource version; a version check cannot certify every clinical premise. Server support is required for version-aware updates. Uncertain submissions go to reconciliation.
Read the diagram as text
  • Reviewed proposal. Proposed content and identified source versions.
  • Current clinical context. Confirm patient and encounter; reassess relevant source changes.
  • Current write authority. Check the requested operation against server permissions.
  • Target version matches. Where supported, apply version-aware update protection.
  • Refresh and review. No commitment of this proposal under stale or conflicting conditions.
  • Denied. Do not perform an unauthorized write.
  • Commit confirmed. Authoritative result confirms the accepted record change.
  • Reviewed proposalCurrent clinical context: Submission requested.
  • Current clinical contextRefresh and review: Context invalid or materially changed.
  • Current clinical contextCurrent write authority: Context remains applicable.
  • Current write authorityDenied: Permission absent.
  • Current write authorityTarget version matches: Permission granted.
  • Target version matchesRefresh and review: Version conflict.
  • Target version matchesCommit confirmed: Match and successful update.

VI. Evidence for use

Validate the task and population

Clinical validation assesses clinically relevant performance for the intended task, population, and care context. IMDRF’s clinical-evaluation framework distinguishes a valid clinical association between output and condition, reliable production of the intended output from inputs, and achievement of the intended clinical purpose. Technical correctness, clinical performance, workflow behavior, and patient outcomes remain separate claims; Evals and Benchmarks develops that general distinction.

A reference standard is the procedure used to establish the comparison answer. It may involve independent expert assessment, adjudication of disagreement, or later outcome ascertainment. Agreement with a comparator that does not establish condition status cannot automatically be called diagnostic sensitivity or specificity. Report counts and sampling uncertainty, not percentages alone. More observations reduce sampling uncertainty but do not remove a biased reference procedure.

The distinction predates modern models. Yu and colleagues’ 1979 MYCIN evaluation asked eight outside specialists to assess anonymized prescriptions for ten deliberately challenging retrospective meningitis cases. The study assessed expert acceptability, with disagreement among evaluators. It did not establish routine-care outcomes, cost effectiveness, or deployment readiness.

The target itself can be wrong. Obermeyer and colleagues’ 2019 study examined a system using predicted healthcare spending as a proxy for need. Black patients were sicker than White patients at the same score, with lower spending at comparable illness burden reflecting unequal access and utilization. Accurate cost prediction could therefore allocate help poorly. The paper illustrates target mismatch, not merely a model fitting error.

A suitable target can still fail to transfer. A 2021 external validation of a widely deployed sepsis predictor found poor performance in one academic center and examined calibration, alert burden, and added value relative to existing practice. This was a historical evaluation of a specified model and cohort, not a judgment about every version or hospital. Language, age, comorbidity, equipment, and access patterns can all define materially different use conditions.

Match separation to the intended claim, retaining only information available when the system would run. The detailed methods belong in Data Quality and Curation.
Evaluation boundaryMain transfer questionRemaining limitation
Patient separationPerformance on unseen people.Does not establish future or cross-site performance.
Later periodPerformance as time and practice change.Returning patients may still overlap.
Another institutionPerformance under another site's conditions.One external site does not represent every setting.

Errors, probabilities, and workload

Sensitivity is the fraction of actual cases detected. Specificity is the fraction of non-cases correctly left unflagged. Positive predictive value is the fraction of alerts that are actual cases. Prevalence is the proportion of the evaluated population with the target condition or event. Sensitivity therefore does not tell a reviewer how often an alert is correct.

Hold sensitivity and specificity fixed to isolate prevalence. Real population changes can also alter those rates.

Prevalence changes the alert workload

Hypothetical cohort of 1,000; sensitivity 80% and specificity 90% stay fixed. Change only prevalence to isolate its effect on alerts.

170 alerts: 80 actual cases + 90 false alerts.
Positive predictive value: 80/170 = 47.1%. 20 actual cases are missed.
Alert counts with fixed sensitivity and specificity80 detected cases and 90 false alerts out of 170 alerts. Same 0 to 250 alert scale at every prevalence.050100150200250Detected cases: 80False alerts: 90All alerts: 170 · counts, not review minutes or treatment benefit
Actual casesDetectedMissedFalse alertsCorrectly unflaggedCohort
1008020908101,000
Reference comparison · available without changing controls
PrevalenceDetectedMissedFalse alertsAll alertsPPV
1%82991077.5%
10%80209017047.1%
The calculation holds sensitivity and specificity fixed, not a clinical threshold or a treatment policy. Lower prevalence changes both alert composition and volume. Actual review time, response delays and clinical benefit must be measured separately.

Lower prevalence leaves fewer real cases while false alerts still arise from the much larger non-case population. Review demand must therefore be assessed alongside detection: count alerts and missed cases, then measure the actual work and response delays they create. Changing an action threshold changes that workload and its errors; available capacity is a constraint, not evidence that an arbitrary higher threshold is clinically acceptable.

Calibration concerns absolute probabilities. Among comparable patients assigned approximately 10% risk for a defined event and horizon, approximately 10% should experience it. Ranking higher-risk patients well does not establish calibrated probabilities, and calibration at one institution may not transfer. Nor does calibration select an appropriate intervention. See Interpret probability forecasts. A confidence interval for measured sensitivity answers another question: uncertainty in a population performance estimate, not an individual patient’s event probability.

For notes and conversations, measure the errors that matter to that task rather than forcing every output into a diagnostic score. Clinicians can identify unsupported additions, omissions, incorrect intervention categories, or intervention at the wrong conversational turn. In Evals Driven-Development for a Mental Health AI Coach, clinician annotations become typed regression cases with conversation input, an expected observation, category, and replay turn. Both missed and inappropriate interventions matter. If a model grades these cases, validate the judge against independent clinical assessment; a generated critique is not its own reference standard.

Evaluate the complete intervention

Retrospective evaluation uses previously collected cases. Prospective evaluation is planned around newly occurring work. Silent operation runs on current inputs while withholding outputs from care decisions. It can reveal input and performance problems, but cannot show how clinicians respond to advice they never see. Usual care is the actual existing alternative, including its staffing, tools, and follow-up—not simply the absence of a model.

Choose the design for the remaining claim; these are not mandatory steps in a universal approval ladder.
DesignExposureWhat it can investigate
Retrospective testingPreviously collected cases.Specified output performance under reconstructed conditions.
Silent prospective operationNew inputs; outputs withheld from care.Current data availability and prospective performance.
Supervised live evaluationUsers see and may act on outputs.Reliance, usability, workflow changes, and observed failures.
Comparative outcome studyDefined assisted and comparison workflows.Differences in prespecified outcomes under the study design.

DECIDE-AI, published in 2022, provides reporting guidance for early live evaluation, including actual use, overrides, workflow, and harm. CONSORT-AI, published in 2020, extends clinical-trial reporting to describe algorithm version, inputs, users, outputs, and effects on decisions. Both make the human–AI intervention more inspectable; neither proves effectiveness or supplies authorization to deploy.

Random assignment uses an unpredictable chance process rather than prognosis or preference to allocate an intervention. Concealing the next assignment prevents recruitment from being steered toward a particular arm. Assignment may be by patient, clinician, or care unit: shared staff and practices can let one arm affect another, so the unit must fit the workflow. Concurrent observation alone does not randomize exposure. Choose the live experiment develops these design choices. Clinical randomized studies are not categorically prohibited; exposure and safeguards require an appropriate study and institutional review.

Keep endpoints separate. The ambient-documentation trial measured practitioner and documentation outcomes; ACCESS distinguished screening from indicated follow-up. Neither established the patient-health claims those intermediate outcomes might suggest. A surrogate endpoint substitutes another measure for an outcome directly relevant to patients. FDA–NIH BEST emphasizes that correlation with health outcomes is insufficient: intervention-induced changes in the surrogate must reliably predict benefit in the specified context. An efficiency gain, a clinical benefit, a safety assessment, and permission to operate therefore require different support.

VII. Operation and recovery

Monitor changing conditions

Distribution shift means operating conditions differ from those evaluated. Patient mix, measurement practices, interfaces, protocols, staffing, and model versions can change different parts of the system. A detected change warrants investigation; it does not alone establish degraded clinical performance. Machine Learning Fundamentals explains the broader concept. Monitoring should preserve separate views of input availability, meaning, output quality, response workload, and outcomes.

Assign clinical and technical owners according to the changed relationship.
Observed changeInitial investigationWhat remains unproven
Fewer results arriveInspect source and interface availability.Whether affected patients have normal findings.
A field or unit changesReview mapping with source-system and clinical owners.Whether prior interpretation remains valid.
More alerts remain pendingInspect case mix, review effort, staffing, and acceptance.Whether detector performance itself deteriorated.
Recent outcomes are unavailableTrack follow-up maturity and ascertainment.Clinical performance in that incomplete cohort.

Missing documentation, delayed outcomes, and selective follow-up are different problems. An absent historical entry might never have been recorded. A recent outcome may not yet be observable. Follow-up confined to selected patients can leave the others without comparable labels. Do not classify these unknowns as successes or negative outcomes; incomplete feedback requires explicit accounting.

Actions can also change the outcome being monitored. In the clinical AI quality-improvement framework’s hypothetical early-warning example, an alert prompts treatment and the predicted event does not occur. The record alone cannot distinguish a false prediction from successful prevention: the outcome without treatment is unobserved. Simple prediction-versus-outcome scoring can therefore misdescribe the intervention’s effect.

A non-event leaves the treatment effect unresolved

Same patient, one observed course. The alternative is not a second patient or a measured study arm.

Observed course

  1. Alert prompts treatment.
  2. Treatment is followed by no observed event.

Counterfactual · unobserved

Without that treatment, what would have happened to this same patient?

Outcome unknown
In this hypothetical example, no event is observed after treatment. That sequence alone cannot distinguish an incorrect prediction from successful prevention: the same patient’s outcome without treatment is unobserved.

Retain source and model versions, relevant policy and workflow versions, observed effects, and reviewer decisions under defined access and retention controls. Sensitive telemetry needs its own governance. Clinical and technical owners should agree on conditions for reassessment, restricted use, and suspension. A model or prompt change requires renewed evaluation of affected behavior; a healthy request pipeline is not a substitute.

Recovery and affected work

Recovery begins by identifying what failed. Degraded operation deliberately provides a smaller supported service. Reconciliation determines what actually happened and resolves discrepancies against authoritative records. A record amendment corrects or qualifies documentation while preserving attributable history. These responses solve different problems; restoring a model version addresses future execution, not every effect already produced.

FailureRequired response
Assistance unavailableContinue through an established non-AI workflow; retain ownership of pending work.
Required source records unavailableWithhold dependent conclusions or clearly bound any partial information.
Submission outcome uncertainInspect authoritative status and use supported duplicate-protection mechanisms before repetition.
Accepted output known to be wrongReview affected work, correct through the applicable record process, and assign necessary communication.

The SAFER contingency-planning guide recommends tested downtime procedures, named clinical and technical leadership, communication independent of failed infrastructure, and alternative documentation. Restored systems must reconcile information captured during downtime. Suspending assistance must not silently abandon urgent, overdue, or unaccepted work.

An uncertain write is not a confirmed failure, and resource-level duplicate protection is not proof that a clinical action occurred exactly once. A known wrong note has another remedy: correction with history. The historical CMS amendment guidance illustrates preserving distinguishable original and changed content with authorship and dates. It is not a determination of current local policy. Nor can a record correction undo an utterance already heard or an action already taken.

Restoring software without erasing effects requires separate completion evidence. Before resuming the affected use, verify repaired dependencies and an evaluated operating mode, reconcile pending effects, and confirm ownership of unfinished work. Long-running corrective work may continue under an accountable plan. The final measure of recovery is a workable care service with understood obligations—not merely a responding endpoint.

Open questions

  1. Detection can expand faster than follow-up capacity. Establishing useful expansion requires studying completed indicated care, delays, and workload across services—not screening volume alone. Progress would demonstrate that additional detection reaches appropriate follow-up without leaving less visible groups behind.

  2. Effective oversight must remain effective under routine workload. Reviewers can accept incorrect advice and reject useful advice, while repeated alerts change response behavior. Progress would show sustained error detection and correction in real workflows, with review effort and downstream consequences measured together.

  3. Clinical monitoring becomes difficult when outcomes arrive late or treatment changes them. A non-event after an alert may reflect successful intervention rather than a false alarm. Progress requires evaluation designs that separate prediction quality, response behavior, and intervention effects while keeping unavailable outcomes explicit.

  4. Local adaptation must accommodate specialty and organizational differences without silently changing clinical authority. Decentralized expert review can improve contextual fit, but safe update methods and cross-setting regression evidence remain necessary. Progress would make each local change attributable, reviewable, and independently assessable.

Follow the curated reading path through the speakers and demonstrations behind this entry.

19 min

AI Engineer World's Fair 2026 · 2026

Shipping AI to a Million Patients Without an A/B Test

Jared Joselowitz

Cited in this entry

Explains hazard-driven testing for patient conversations and why rollback cannot retract delivered speech. Read its deployment argument as motivation for pre-exposure evidence, not a blanket prohibition on clinical randomized studies.

Watch talk

Explore more talks

The rest of the library, beyond the curated path. Cited talks support this entry; reviewed transcripts were processed in full. Metadata candidates have not been reviewed as sources or verified as topic members.

41 matching talks

Every catalogued talk on this subject: Healthcare

TalkSpeakerEventYear
Vivek MuppallaAI Engineer World's Fair 20262026
Rashi AgrawalAI Engineer World's Fair 20262026
Vasant KearneyAI Engineer World's Fair 20262026
Frank CoyleAI Engineer World's Fair 20262026
Hamed Firooz, Maziar SanjabiAI Engineer World's Fair 20252025
Yu SuAI Engineer World's Fair 20262026
Apoorva JoshiAI Engineer World's Fair 20262026
Roy DerksAI Engineer Summit 20252025
Adam TerlsonAI Engineer Summit 20252025
How to Build Trustworthy AI

Transcript reviewed

Allie HoweAI Engineer World's Fair 20252025
Angel Ortmann LeeAI Engineer World's Fair 20262026
Ian ButlerAI Engineer World's Fair 20252025
Jun Yu TanAI Engineer World's Fair 20252025
Jesse HuAI Engineer Code 20252025
Angus J. McLeanAI Engineer Europe 20262026
Christopher LovejoyAI Engineer Europe 20262026
AI’s Jurassic Park Period

Transcript reviewed

Aaron StanleyAI Engineer World's Fair 20262026
Jeremy Silva, Chris HernandezAI Engineer World's Fair 20252025
Mohak SharmaAI Engineer Summit 20252025
Anna Marie BenzonAI Engineer World's Fair 20262026
Giran Moodley, Mayan Soni, Oussama Hafferssas, Mayank SoniAI Engineer Europe 20262026
Sandipan BhaumikAI Engineer Europe 20262026
Ayush BhardwajAI Engineer World's Fair 20262026
Philip RathleAI Engineer World's Fair 20242024
Christopher Lovejoy, Saul HowardAI Engineer World's Fair 20262026
Dan MasonAI Engineer World's Fair 20252025
Sandipan BhaumikAI Engineer Europe 20262026
Agents Need Feature Flags

Transcript reviewed

Sachin GuptaAI Engineer World's Fair 20262026
Dan FengAI Engineer World's Fair 20262026
Clay Cockrell, Tony FabrikantAI Engineer World's Fair 20262026
Denys LinkovAI Engineer World's Fair 20262026
Anant ShankhdharAI Engineer World's Fair 20262026
Stephen ChinAI Engineer Europe 20262026
Andreas Kollegger, Zaid ZaimAI Engineer Europe 20262026
Don't be data poor

Metadata candidate

Anuj IravaneAI Engineer World's Fair 20262026
Rossella Blatt Vital, Deepsha MenghaniAI Engineer World's Fair 20252025
Mike BursellAI Engineer World's Fair 20252025
Kwindla Hultman KramerAI Engineer World's Fair 20242024
Christopher LovejoyAI Engineer World's Fair 20252025
Kwindla Kramer, Shrestha Basu MallickAI Engineer World's Fair 20252025
Nik CaryotakisAI Engineer Summit 20252025

References

Coverage and source review
Processed transcripts
34 processed in full · 6 in the curated path
Automated source review
Passed
Metadata candidates
13 unreviewed; not verified topic membership
Corpus version
1bd8e407b26a07b33815594e1b2db5f41827119a2b3cb6fbf240f9fc571fc767

Automated review checks source support; it is not publication approval.

A synthesis of selected conference talks and technical references. Citations link to the source material; they do not imply that every talk on this subject is included.

  1. AHRQ: What is workflow?

    Clinical workflow comprises physical and mental tasks performed by people within and across care environments. Tasks can occur sequentially or simultaneously. AHRQ's medication-order example includes clinician reasoning, patient communication, prescription entry, transmission, and pharmacy fulfillment. The toolkit recommends assigning responsibility for assessing current and anticipated workflows, beginning before implementation and continuing afterward.

  2. IHI: Closing the Loop—A Guide to Safer Ambulatory Referrals in the EHR Era

    A specialty referral asks a specialist to evaluate or treat a patient. IHI describes referral safety as a closed communication loop: relevant information must reach the correct person through appropriate channels and at the right time. The organization identifies unclear roles, communication failures, workloads, and differing specialist requirements as obstacles to completing that loop.

  3. TRIPOD+AI expanded checklist: timing, labels, and evaluation separation

    Specify the intended decision point t0, eligible population, predictor definitions and measurement times, and outcome assessment procedure. For prognosis, define the horizon h, such as an event within 28 days after t0. For diagnosis, identify the reference standard used to determine condition status. Predictors should be measured before or when the model is intended to run; assessment and blinding procedures must avoid outcome leakage. Report participant overlap, including repeated records, across training, tuning, and evaluation. Engineering interpretation: patient-disjoint evaluation asks about unseen people; a later-period evaluation asks about performance as time and practice change; evaluation at another institution asks about transport across settings. These dimensions can be combined, and their dates, populations, and differences must be explicit.

  4. Use of an Electronic Referral System to Improve the Outpatient Primary Care–Specialty Care Interface

    Referral intake included relevant history, the referring clinician’s question, and linked EHR information. A designated specialty reviewer selected routine scheduling, expedited overbooking, or further information/workup. Referring clinicians had to watch for requested results and resubmit referrals. No-shows appeared on worklists without automatic messages; clinics maintained supplementary appointment and notification logs. Scheduling after review could miss patient availability, and one reviewer reported no integration with the appointment system. Cardiology waiting times increased alongside reduced staffing. The report defined never-scheduled referrals using a 180-day appointment window, not confirmed resolution of patient need.

  5. Kaiser Permanente: Early Warning System for Hospitalized Patients

    Kaiser Permanente describes Advance Alert Monitor as a coordinated response workflow. When a patient's risk score crosses a threshold, an off-site team of specially trained nurses reviews the record. A reviewing nurse contacts the bedside rapid-response nurse, who works with the physician and other professionals to adjust the care plan while considering patient preferences. The implementation assigns separate responsibilities for identifying risk, reviewing the alert, assessing the patient, and deciding what to do.

  6. WHO: Interagency Integrated Triage Tool

    Acuity-based triage sorts patients by estimated urgency of intervention: who needs immediate attention and who can wait. It can occur at access points including ambulance services, outpatient clinics, and hospitals. WHO, ICRC, and Médecins Sans Frontières developed the Interagency Integrated Triage Tool for facility-based emergency-unit triage, with distinct materials for adults and children. Its categories communicate urgency rather than providing a complete diagnosis.

  7. Accuracy and safety of an autonomous artificial intelligence clinical assistant conducting telemedicine follow-up assessment for cataract surgery

    Dora R1 asked symptom-based questions by telephone approximately three weeks after routine cataract surgery. In this prospective two-hospital study, an ophthalmologist supervised calls in real time; investigators compared independently made recommendations about symptoms and further review. The study also examined subsequent management, usability, and potential staff costs. Transcript analysis identified difficulty capturing nuanced symptoms with a binary question. The authors emphasized explaining the respective roles of automation and clinicians and preserving access to clinician-delivered care.

  8. How to Leverage Domain Expertise — Chris Lovejoy, Notius Labs

    Hire or empower someone with direct experience of the target workflow, rather than treating a broad professional credential as sufficient.

  9. eCQI Resource Center: Clinical Decision Support

    Clinical decision support supplies general or patient-specific information at useful points in care to assist decisions. It includes alerts, reminders, order sets, patient summaries, documentation templates, diagnostic support, and contextual reference information. The guidance emphasizes matching information, recipients, format, delivery channel, and workflow timing, while preserving clinician judgment.

  10. AI System Design: From Idea to Production

    Use the four-phase framework of product requirements, system design, evaluation and monitoring, and optimization; begin with a measurable, solution-agnostic user problem.

  11. Google: Supervised learning foundations

    Features are input values x; a label y is the target recorded for an example. Parameters θ are learned numbers defining a model fθ. Training repeatedly compares predictions with labels using a loss and adjusts these numbers to improve agreement. Inference evaluates the trained function on a new input without needing its label. Generalization means performing well on examples outside training; evaluation therefore compares predictions against withheld labels. Programmer-oriented illustration: training fits w and b in f(x)=wx+b from examples, whereas an authored rule such as if x>c executes a programmer-specified condition. Both execute code at inference, but the origin of their decision behavior differs.

  12. AI System Design: From Idea to Production

    A structured claims process can use RAG, control flow, and human-in-the-loop review without giving an agent end-to-end autonomy.

  13. FDA, Health Canada, and MHRA: Good Machine Learning Practice Guiding Principles

    The 2021 principles connect intended clinical use to data quality, representative populations, independent test data, clinically relevant evaluation, and monitoring after deployment. They emphasize the performance of the human-AI team rather than only the model. The resulting engineering question is not merely whether a predictor is accurate, but whether its inputs, users, workflow, and response to errors support the intended clinical task. Retraining also requires evaluation because changing the learned system can change its behavior.

  14. SMART App Launch: Patient context and resource permissions

    SMART separates patient context, which identifies whose record is in view, from scopes, which specify allowed resources and operations. For example, patient/Observation.rs requests reading and searching observations for the contextual patient; patient/Observation.c requests creation and does not imply reading. launch/patient requests context, not unrestricted record access. Granted scopes may differ from requested scopes and remain constrained by underlying system policies and user permissions. Engineering implication: approving a proposed clinical action in an application's review interface does not mint authorization to read or write an EHR. Execution must independently satisfy the server's authorization and patient-context checks.

  15. IMDRF: SaMD definition and intended medical purpose

    IMDRF defines SaMD by software's intended medical purpose and its ability to perform that purpose without being part of a hardware medical device. Medical purposes include diagnosis, prevention, monitoring, and treatment. Running on a phone, server, or general-purpose computer does not exclude SaMD status. Software intended to drive a hardware medical device falls outside this SaMD definition, although that does not make it unregulated. The document also excludes software for making or maintaining devices from its concept of medical-purpose software. Engineering implication: describe intended users, medical task, and use conditions before reasoning about regulatory status; an implementation label such as AI or app is insufficient.

  16. Ledley and Lusted: Reasoning Foundations of Medical Diagnosis

    Robert Ledley and Lee Lusted's July 3, 1959 paper examined how computers might assist diagnostic reasoning. It distinguished logical relationships between findings and possible diseases, probabilities for evaluating alternatives, and value judgments involved in treatment choices. Computers could organize information and bring overlooked diagnostic possibilities to attention. Simplified clinical cases illustrated the proposed methods; the authors explicitly presented foundations for future practical procedures rather than an immediately deployable diagnostic service.

  17. Pryor, Gardner, Clayton, and Warner: The HELP System

    The HELP developers describe a different emphasis from a stand-alone consultation: integrating clinical information and decision support into hospital work. At LDS Hospital and the University of Utah, previously separate clinical systems began moving to a common database around 1970. Development of a medical-decision language followed in 1975. Incoming patient data could trigger relevant decision logic automatically, while time-based processing supported periodic checks. Outputs could reach patient records, nursing units, or diagnostic departments. Requests for additional information were limited to what a particular decision needed, reducing redundant entry.

  18. Buchanan and Shortliffe: Preface to Rule-Based Expert Systems

    Bruce Buchanan and Edward Shortliffe trace MYCIN to discussions between Stanford medical and computer-science researchers in spring 1972 and Shortliffe's 1974 dissertation implementation. The project investigated whether explicit rules could support expert consultation while remaining understandable and modifiable. Its experiments extended through 1982. A concrete downstream development was EMYCIN: removing the medical knowledge base produced a framework for constructing expert systems in other domains.

  19. Gulshan and colleagues: Development and Validation of a Deep Learning Algorithm for Detection of Diabetic Retinopathy

    Varun Gulshan, Lily Peng, and collaborators reported a retinal-image classifier that learned predictive features without explicitly programmed lesion detectors. Published online November 29, 2016, the study evaluated two operating thresholds selected during development. On 8,788 fully gradable EyePACS-1 images, the high-sensitivity threshold achieved 97.5% sensitivity and 93.4% specificity against ophthalmologist majority judgments. The contribution was demonstrated image-based detection with adjustable screening tradeoffs, not autonomous delivery of comprehensive eye care. The authors explicitly called for clinical-setting and patient-outcome studies.

  20. From Ambient Documentation to Clinical Intelligence

    Clinical notes affect both billing and future clinical context, so documentation errors can propagate beyond the original encounter.

  21. eCQI Resource Center: Electronic health record

    An electronic health record is a digital record of patient health information accumulated across visits and care settings. Its contents can include progress notes, diagnoses, medications, allergies, vital signs, laboratory results, treatments, and imaging.

  22. NHS England: Ambient scribing guidance for health and care professionals

    Ambient scribes capture a care conversation and produce notes for professional review and editing. NHS guidance assigns responsibility for information entered into records to the professional using the tool. It calls for meaningful review every time, correction before inclusion in the record, additional validation where translation is involved, and consideration of whether third-party information should be retained. Patients should receive an explanation of the tool, and their objections should be respected.

  23. Abridge: Verify a Note With Linked Evidence

    Abridge links selected text in a generated clinical note to corresponding transcript passages and allows playback of the original audio. This is an auditability mechanism: a clinician can inspect the evidence behind a note. Linking does not by itself establish that the transcript is correct or the clinical interpretation is valid.

  24. A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being

    A 24-week stepped-wedge randomized trial enrolled 66 ambulatory practitioners in one health system across two states. The documentation-only system generated transcripts and draft notes from conversations, using both private Epic APIs and FHIR R4 APIs; private APIs carried audio and returned notes. Compared with usual practice, the study reported reduced work exhaustion/interpersonal disengagement and reduced time spent on notes. Recording required patient agreement. Unedited-note quality was assessed using an LLM judge rather than comprehensive independent clinical adjudication.

  25. Use of an Electronic Referral System to Improve the Outpatient Primary Care–Specialty Care Interface

    This evaluation of a non-AI electronic referral system combined system logs, surveys, interviews, and process simulation. Specialist review could request further primary-care workup, provide advice without an appointment, or schedule care. Logs showed initial waiting-time reductions in most studied medical clinics. Interviews also identified administrative work shifting to referring clinicians and difficulty involving patients in appointment selection, particularly for homeless and limited-English-speaking populations. Simulated labor costs differed between medical and surgical clinics because reviewer roles and staffing costs differed.

  26. AI That Pays: Lessons from Revenue Cycle

    A clinical appeal requires reconciling patient evidence, care guidelines, and payer coverage policies under a submission deadline.

  27. Anterior and Stellarus prior-authorization workflow

    The announced joint workflow uses AI for case research, documentation and policy application in prior authorization. It can recommend approval or route a case for review; licensed clinicians make final determinations. This is the scope of the announced partnership, not every Anterior deployment or evidence of patient benefit. Exclude vendor accuracy and population claims.

  28. Abràmoff and colleagues: Pivotal Trial of an Autonomous AI-Based Diagnostic System

    Michael Abràmoff and colleagues' August 28, 2018 report evaluated IDx-DR in ten primary-care offices. Trained office personnel acquired retinal images; the system supplied image-quality guidance and a diagnostic output without requiring a specialist to interpret every screening image. Among 819 participants with fully analyzable data, sensitivity was 87.2% and specificity 90.7% for the study's more-than-mild diabetic-retinopathy target against an expert reading-center reference. The study tested a bounded primary-care diagnostic workflow, extending beyond retrospective classification of an existing image collection.

  29. Wolf and colleagues: The ACCESS Randomized Control Trial

    Risa Wolf and colleagues' 2024 ACCESS report compared point-of-care AI eye screening with referral plus scripted education in youth with diabetes. Screening was completed by 81 of 81 intervention participants at the study visit versus 18 of 82 analyzed controls within six months. However, only 16 of the 25 intervention participants with positive AI results attended an eye-care provider within six months. Completing screening and completing indicated follow-up were therefore distinct workflow outcomes.

  30. Hernán and Robins: Counterfactual outcomes and confounding

    Y(1) and Y(0) denote outcomes under specified treatment and no-treatment strategies; only the outcome under the received strategy is observed. For an adverse outcome, average benefit within group X=x can be defined as E[Y(0)−Y(1)|X=x]. High untreated risk E[Y(0)|X=x] does not determine this difference: risk can remain high under both strategies. Illustrative risks of 0.8 versus 0.8 imply no reduction, while 0.3 versus 0.1 imply a 0.2 reduction. Confounding arises when treatment groups differ in causes of outcome, such as severity influencing both treatment selection and prognosis. Their observed outcome difference then mixes treatment effects with baseline differences.

  31. HL7 FHIR R5: DiagnosticReport

    FHIR distinguishes a requested investigation, its report, and individual results. DiagnosticReport can link to a ServiceRequest describing what was requested, an Encounter providing care-event context, and Observation resources containing individual results. It can also carry interpretation and an attached report. Panels group related observations rather than turning every measurement into an unrelated document. Report status distinguishes preliminary, final, amended, corrected, and other states. These relationships preserve how a result fits into a clinical investigation.

  32. HL7 FHIR: MedicationAdministration

    FHIR distinguishes a medication order and administration instructions from dispensing a supply and recording actual administration. MedicationStatement instead represents a report from a patient or clinician that medication was taken or given. These records describe different evidence: a prescription is not itself a record that the patient received the medicine.

  33. What Every Reader Should Know About Studies Using Electronic Health Record Data but May Be Afraid to Ask

    Available EHR fields are not necessarily routinely populated. The authors explain that smoking-history codes are underused, so their absence cannot identify a nonsmoker. Relevant information may instead appear inconsistently in narrative notes, be fragmented across fields, or reside in another institution’s record. Their consortium encountered transferred patients whose earlier clinical course was unavailable. Thus missing data can reflect documentation practices and care boundaries rather than a negative clinical finding.

  34. FHIR R5: Observation effective and issued times

    Observation.effective[x] records when the observed value is asserted to apply, commonly specimen collection or procedure time. Observation.issued records when this version became available to providers, typically following review and verification. Worked illustration: a specimen collected at 09:00 has a result issued at 10:15; a predictor running at 09:30 cannot use that result merely because its effective time precedes prediction. Engineering inference: reconstruct the version actually accessible at prediction time, including ingestion delays and later corrections, rather than joining the latest result by clinical event time alone.

  35. HL7 FHIR: Provenance

    FHIR Provenance records the activity that created or updated resources, participating actors, and source entities. It distinguishes when an activity occurred from when it was recorded. Targets can identify particular resource versions; references must include version information when different versions need distinguishing. Provenance describes generation and revision, while AuditEvent covers usage and other activities.

  36. FHIR R5 data-absent-reason value set

    The data-absent-reason terminology distinguishes unknown, temporarily unknown, not asked, masked, error and not performed. In particular, not-performed means the value is unavailable because the procedure or test was not performed; masked means information is withheld for privacy or security. These are different explanations for unavailable data, not normal results. Engineering implication: retain the recorded reason rather than collapsing every absence into a healthy or zero-valued measurement.

  37. HL7 FHIR R5 Overview

    FHIR represents exchangeable healthcare content as resources with common data structures and metadata, combined through references. Profiles and other conformance artifacts tailor those resources to specific uses, including terminology and cardinality constraints. This gives programmers a structured way to exchange clinical information, but interpreting a longitudinal record still requires the meaning and context of the underlying observations. A syntactically valid resource is not proof that an observation is clinically correct or comparable with another institution’s data.

  38. FHIR R5 Observation: results, interpretation and absent data

    Observation.code identifies what is observed; value[x] contains its actual result, which may be a Quantity or another permitted datatype. Interpretation is a separate optional qualitative assessment such as normal, high or low. dataAbsentReason explains a missing expected result and is allowed only when value[x] is absent. Components can carry their own results, interpretations and absent-data reasons. Engineering implication: a missing top-level numeric value is not a normal result; inspect the observation code, result datatype, components and available absence explanation before constructing a feature.

  39. Regenstrief Institute: LOINC Users' Guide—Introduction

    LOINC supplies shared names and identifiers for laboratory tests and clinical observations, such as serum potassium or vital signs. It addresses a practical exchange problem: a receiving system cannot interpret another laboratory's private test codes without mappings. LOINC identifiers can be carried within messages, documents, or APIs; identifying the observation is distinct from transporting its result. The guide distinguishes individual measurements from panels that enumerate multiple measurements.

  40. HL7 FHIR R5 Datatypes: Coding and CodeableConcept

    Coding retains a code, its terminology system, an optional version, and display text. The system identifies where the code’s definition comes from; display text is not intended for computation. Identical system, version and code establish the same coded meaning. When these differ or version information is absent, relationships require consulting terminology definitions and available mappings. CodeableConcept can contain multiple codings whose granularity differs. Constructed example: receiving the same display label from two systems does not establish that their codes mean the same thing.

  41. FHIR R5 Quantity: coded units and comparison semantics

    Quantity separates numeric value, comparator, human-readable unit, coded unit and coding system. A unit code requires its system; display text must not be assumed to be a valid UCUM expression. UCUM coding can support conversion into canonical values for comparison. A comparator such as less-than or greater-than changes the meaning of the number and cannot be ignored. Engineering implication: confirm that quantities describe a comparable measurement, validate compatible coded units and convert them before numeric comparison, preserving bounds rather than treating them as exact values.

  42. SMART App Launch: Conformance

    SMART on FHIR specifies application-launch and authorization capabilities. During EHR launch, a server can pass existing patient and encounter context to an application. Patient and encounter context are separately declared capabilities, as are patient-level and user-level permissions. The server publishes authorization endpoints and supported capabilities through its SMART configuration document.

  43. HL7 FHIR R5: Comparison with HL7 Version 2

    HL7 Version 2 exchanges messages assembled from reusable segments for activities such as laboratory orders and patient transfers. FHIR additionally supports independently identifiable resources and multiple exchange styles, including REST, documents, and messaging. HL7 warns that Version 2 implementations vary in supported fields and their placement, so conversion mappings generally require implementation-specific work. Conversion must also determine when repeated references identify the same entity and how conflicting historical and current values should be represented.

  44. Why Your Enterprise Tech Stack Isn't Ready for AI Agents - And What to Build Instead

    Design the audit record to preserve actions, data access, and the authorization behind each action, rather than treating diagnostic logs as sufficient.

  45. HHS: De-identification methods and residual identification risk

    HHS describes two HIPAA de-identification methods. Safe Harbor removes specified identifiers, including relevant dates and geographic detail, and requires absence of actual knowledge that remaining information can identify a person. Expert Determination documents that identification risk is very small in the anticipated recipient context; techniques can suppress values, generalize detail, or perturb data. Identifiers in free text count as well as structured fields. HHS explicitly states that properly de-identified data retain nonzero identification risk because linkage can reconnect records with people. Removing names and record numbers alone therefore does not establish de-identification, anonymity, or protection against all privacy threats.

  46. HL7 Clinical Practice Guidelines: CPGPlan

    The guideline implementation guide separates logic describing a patient's condition, risk, or severity from logic proposing what to do. Recommendations specify triggering events, applicability conditions, and proposed activities, while preserving supporting evidence and recommendation strength. Eligibility identifies whether a patient falls within a guideline's scope. Enrollment is separate: patient preferences or other clinical factors can mean an eligible patient should not enter that pathway. Finding a relevant guideline passage is therefore different from establishing applicability and participation in a patient-specific care plan.

  47. Mission-Critical Evals at Scale: Learnings from 100,000 Medical Decisions

    An answer can cite relevant clinical findings yet misinterpret a distinction that changes the guideline answer.

  48. Citation Needed: Provenance for LLM-Built Knowledge Graphs

    A synthesized fact can hide both its original wording and the authority of its actual source, so retain verbatim inputs and explicit links to derived artifacts.

  49. William van Melle: The Structure of the MYCIN System

    MYCIN separated consultation, explanation, and knowledge acquisition. Consultation applied expert-authored rules to patient information, working backward from a goal to the facts needed to establish it and asking the user when those facts could not be derived. Explanations exposed which rules or supplied facts supported a conclusion and why a question was being asked. Knowledge acquisition supported changes to the expert knowledge base. Its certainty factors represented degrees of belief and were not conditional probabilities.

  50. Mission-Critical Evals at Scale: Learnings from 100,000 Medical Decisions

    Give reviewers accessible source context and collect explanations of errors that can support reusable reference answers.

  51. FDA, Health Canada, and MHRA: Transparency for Machine Learning-Enabled Medical Devices

    Transparency means communicating relevant information about intended use, development, performance, and, when available, the basis of an output to the appropriate audience. The guidance treats explainability as one aspect of that broader obligation and emphasizes the human-AI team. Information must support the user’s actual decision: a technical explanation that does not clarify limitations, use conditions, or the required response may not be sufficient. Interface design and user comprehension therefore belong in evaluation.

  52. AHRQ TeamSTEPPS: Handoff

    A clinical handoff transfers information together with responsibility and authority during a transition in care. AHRQ includes uncertainty, recent changes, response to treatment, and contingency plans among the information to communicate. The sender remains responsible until the receiving person acknowledges understanding and acceptance. Electronic delivery alone is insufficient: the sender cannot assume that a message was read or understood. The receiver also needs an opportunity to ask questions.

  53. Psychological Factors Influencing Appropriate Reliance on AI-enabled Clinical Decision Support Systems: Experimental Web-Based Study Among Dermatologists

    In this web experiment, 223 dermatologists classified 24 cases before and after seeing AI advice; five recommendations were deliberately incorrect. The analysis distinguished accepting useful advice, rejecting incorrect advice, and changing a previously correct answer into an incorrect answer. Some initially correct judgments became wrong after erroneous advice, while correct advice was also frequently rejected. This provides a concrete way to test oversight behavior rather than counting approvals alone.

  54. Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system

    This retrospective study examined alert acceptance among ambulatory primary-care clinicians. More reminders per encounter and repeated reminders for the same patient were associated with lower acceptance. The study did not find support for a general workload effect or declining acceptance over time after introducing a new alert. It therefore supports investigating alert burden and repetition without treating all reduced responsiveness as simple desensitization.

  55. NIST AI RMF 1.0: accountability, appeals and override

    NIST calls for documented responsibilities and communication, defined human oversight, feedback channels that let affected people report problems and appeal outcomes, and monitoring that includes appeal and override. It also calls for documenting responses to errors; its human-AI discussion identifies override frequency and rationale as useful evidence. Engineering application: attach a dispute to its decision ID and original input/model/policy versions; route it to an accountable reviewer empowered to inspect evidence and authorize correction or override. Record the reviewer, evidence, rationale, revised disposition and downstream correction status, and communicate the result to the affected person. Feed recurring errors into evaluation and remediation. A trace alone does not provide this correction path.

  56. 200 Million Patient Interactions Later: What the Generic Voice Stack Misses

    The demonstration gathers additional symptom information and then requests immediate nurse involvement; medication uncertainty is referred to the patient's doctor.

  57. Guardrails First: Engineering Member-Facing Health AI

    Place emergency escalation and identity verification in code that runs before the model on every turn.

  58. Carayon and colleagues: Work System Design for Patient Safety—The SEIPS Model

    Pascale Carayon and colleagues' 2006 SEIPS model connects people, tasks, tools and technologies, physical surroundings, and organizational conditions. Their interactions shape care processes and outcomes for patients, workers, and organizations. The model extends a structure–process–outcome view of healthcare quality with explicit attention to work-system design. A technological change can therefore alter workload, communication, and safety through its interaction with staffing and existing tasks, rather than through the tool alone.

  59. HL7 FHIR HTTP: Version-aware updates and conditional creation

    FHIR version-aware updates use ETag and If-Match to prevent overwriting a resource changed since it was read; a mismatched version returns HTTP 412. Conditional creation uses If-None-Exist search criteria: no match creates a resource, one match returns the existing outcome without creating another, and multiple matches fail. This supports protecting reviewed work from concurrent edits and avoiding duplicate records when equivalence is defined appropriately.

  60. IMDRF: Software as a Medical Device—Clinical Evaluation

    IMDRF separates valid clinical association, analytical validation, and clinical validation. Clinical association concerns whether the output has a well-founded relationship to the targeted condition. Analytical validation concerns correctly and reliably producing the intended output from inputs. Clinical validation concerns achieving the intended purpose in the target population and care context with clinically meaningful outputs. Clinical evaluation comprises ongoing assessment before and after distribution.

  61. FDA: Sampling uncertainty in diagnostic test performance

    Sensitivity estimates the positive-test fraction among people with the target condition; specificity estimates the negative-test fraction among those without it. Different samples can produce different estimates, so confidence intervals quantify uncertainty arising from sampling. FDA recommends reporting numerator/denominator counts alongside percentages and two-sided 95% intervals. More observations generally reduce sampling uncertainty, but do not remove systematic bias. A reference standard establishes condition status independently of the candidate test; agreement against a non-reference comparator cannot automatically be labeled sensitivity or specificity.

  62. Yu and colleagues: An Evaluation of MYCIN's Advice

    Victor Yu and colleagues evaluated antimicrobial recommendations for ten deliberately challenging meningitis cases drawn from retrospective records. Eight outside infectious-disease specialists assessed standardized prescriptions without knowing their authors. MYCIN received acceptable ratings in 52 of 80 assessments, or 65%; individual Stanford faculty prescribers ranged from 42.5% to 62.5%. Evaluators disagreed about some appropriate therapies. The authors concluded that clinical acceptability, effects on patient care, and costs still required investigation.

  63. Obermeyer et al.: Healthcare cost as a biased proxy for need

    The studied algorithm ranked patients for extra care using predicted healthcare spending as a proxy for health need. Black patients were sicker than White patients at the same score. The authors linked this mismatch to lower spending on Black patients at comparable illness burden, reflecting unequal access and utilization. Predicting the chosen cost label could therefore work while allocation by that score underserved patients with greater need. Their estimated correction increased the Black share selected for additional help from 17.7% to 46.5%. The mechanism is target mismatch: accurate prediction of resource consumption does not establish accurate measurement of need.

  64. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients

    This retrospective external validation evaluated discrimination, calibration, alert burden, and potential added value relative to existing clinical practice. The studied sepsis model performed poorly on that institution’s cohort, despite its prior adoption elsewhere. Threshold selection changed both missed cases and how many patients would trigger alerts. The study illustrates why ranking risk, estimating calibrated probability, and improving a clinical workflow are different questions; a widely deployed model still needs evaluation in its intended setting.

  65. Classification: Accuracy, Recall, Precision and Related Metrics

    Recall is the fraction of actual positives detected; precision is the fraction of positive predictions that are true positives. Accuracy combines correct positives and negatives, so it can look high when the important positive class is rare. These denominators answer different questions: sensitivity is not the probability that a flagged patient has disease. At fixed sensitivity and specificity, reducing prevalence increases the share of false alarms among positive results; this follows by applying the confusion-matrix counts to a different population.

  66. Calibration: the Achilles Heel of Predictive Analytics

    Discrimination measures separation of outcomes, while calibration asks whether estimated risks agree with observed event frequencies. A calibration curve compares predicted risks with observed proportions; intercept and slope diagnose systematic level and spread errors but cannot establish every local part of the curve. Different disease incidence, case mix and practice patterns can invalidate calibration at a new site. Ranking patients well is therefore insufficient when a clinical action depends on an absolute risk threshold.

  67. NIST: What a confidence level means

    A 95% confidence-interval procedure is designed so that, over repeated samples under its assumptions, approximately 95% of the resulting intervals contain the fixed population parameter. It does not assign a 95% probability to an individual patient's outcome. Applied to diagnostic metrics, the parameter is a population sensitivity, specificity, or predictive value; an individual risk prediction instead estimates an event probability conditional on that patient's inputs. These answer different questions: uncertainty in a measured performance rate versus predicted outcome risk.

  68. Evals Driven-Development: Engineering a Mental Health AI Coach Ethically & Safely

    The team's learning loop converts clinician annotations on traces into typed evaluations committed to CI.

  69. Evals Driven-Development: Engineering a Mental Health AI Coach Ethically & Safely

    Score false positives, false negatives, category correctness, and intervention timing under a clinical expert's definition of good.

  70. From Ambient Documentation to Clinical Intelligence

    Abridge encodes clinician judgment into LLM judges to give non-clinician engineers a domain-informed feedback loop.

  71. DECIDE-AI: Early clinical workflow, safety, and human-factors evidence

    DECIDE-AI recommends reporting of users, training, input handling, interface, workflow timing, and who reaches the final decision. Reports should include actual use and adherence, workflow changes, user agreement or overrides, usability, learning curves, errors, indirect harm, and prespecified outcomes. Operational distinction: silent prospective evaluation runs on newly arriving clinical data while withholding outputs from care decisions; live evaluation exposes users to outputs that can influence care. Consequently, silent performance cannot establish how clinicians respond, whether the interface disrupts work, or whether AI-supported decisions improve patient outcomes.

  72. CONSORT pragmatic-trial extension: Applicability to usual care

    A pragmatic randomized trial aims to inform choices between care options under ordinary practice conditions. Pragmatism is a continuum across design decisions, not a binary label. Broadly relevant participants and clinicians, ordinary settings and resources, flexible intervention delivery, and a comparator reflecting commonly used care make results more applicable to similar practice. Reports must describe what usual care actually comprised, recruitment and exclusions, delivery requirements, and outcomes important to patients and decision makers. Highly selected participants, exceptional implementation support, or a comparator unlike local care can limit transfer even when the trial's internal comparison is valid.

  73. CONSORT-AI: reporting clinical trials of AI interventions

    CONSORT-AI asks clinical trial reports to identify the algorithm version, input acquisition and quality handling, deployment setting, required user expertise, output format, and how outputs affected decisions. It also calls for analysis of performance errors. These details make the intervention reproducible and explain the link between a model prediction and downstream care. A trial of a human-AI workflow evaluates more than the model function alone, so results cannot automatically transfer to a different interface, threshold, or use protocol.

  74. CONSORT 2010 explanation: random assignment and concealment

    Random assignment uses an unpredictable chance process rather than prognosis or clinician preference to allocate treatment. Proper implementation avoids systematic baseline selection on known and unknown prognostic factors; chance imbalance can still occur. Allocation concealment hides the next assignment until enrollment and consent, preventing recruiters from steering particular participants toward an arm. A concealed sequence and unpredictable generation are separate requirements. Applied to AI-supported workflow versus usual care, concurrent observation alone does not randomize assignment: clinician, site or patient selection can still confound outcome differences. Concealment differs from blinding after assignment, which addresses performance and outcome-assessment bias.

  75. FDA–NIH BEST: Validated surrogate endpoints

    A clinical endpoint measures an outcome of direct patient relevance, such as symptoms, function, or survival. A surrogate endpoint substitutes another measure, often a biomarker, for that outcome. Validation requires evidence that intervention-induced changes in the surrogate reliably predict clinical benefit in the specified context; correlation between the biomarker and outcomes across individuals is insufficient. BEST describes antiarrhythmic drugs that reduced premature ventricular beats while increasing mortality, illustrating why improving a plausible surrogate can still harm patients. Supporting evidence generally combines mechanistic reasoning with treatment-effect evidence, often from multiple randomized trials.

  76. Clinical artificial intelligence quality improvement: towards continual monitoring and updating of AI algorithms in healthcare

    The authors propose integrating clinical AI monitoring with hospital quality-improvement work and multidisciplinary responsibility. Monitoring should distinguish changing inputs, outcomes, and input–outcome relationships. Their hypothetical early-warning example explains intervention-induced confounding: when an alert triggers treatment and the event does not occur, observed data alone cannot determine whether the prediction was wrong or the response prevented the event. The outcome without that intervention is unobserved. Monitoring can therefore require more than comparing predictions with subsequent recorded outcomes.

  77. NIST AI RMF: Monitoring, incident response, and recovery

    NIST calls for ongoing risk tracking, user feedback, documented incident response and recovery, change management, and assigned authority to override or deactivate systems inconsistent with intended use. It explicitly considers tracking risks when suitable metrics are not yet available. Engineering application: define escalation owners and suspension criteria, preserve incident and version records, restore an evaluated safe operating mode, and evaluate changed behavior before resuming. When clinical labels arrive late, monitor input availability, operational failures, and reported incidents immediately, then calculate outcome metrics on cohorts whose follow-up has matured. Pending labels must remain unknown rather than being counted as negative outcomes or successful predictions.

  78. SAFER Guide: Contingency Planning

    The SAFER guide recommends trained and tested downtime procedures, communication independent of the failed EHR infrastructure, named clinical and technical leadership, and orderly stopping and restarting of interfaces. Alternative documentation must support continued care. Information captured on paper during downtime should be entered and reconciled after restoration, including coded orders and relevant documents. Backups require restoration testing, not merely successful creation.

  79. CMS Transmittal 442: Amendments, Corrections and Delayed Entries in Medical Documentation

    This CMS recordkeeping guidance requires amendments and corrections to be identified, dated, and attributed while preserving identifiable original content. For electronic records, original and modified content and each modification's authorship and date must remain distinguishable. It provides a concrete example of record correction through an amendment history rather than silent replacement.

  80. Shipping AI to a Million Patients Without an A/B Test

    An already-delivered clinical utterance cannot be undone, so reactive rollout monitoring cannot substitute for evidence gathered before exposure.

  81. NIST SP 800-61r3: incident response and verified recovery

    Triage validates an incident report and estimates severity and urgency; response priority considers impact, scope, and available resources. Containment limits expansion, while eradication removes persistence mechanisms and entry points. Investigators preserve the integrity and provenance of evidence and action records, with restricted access and defined retention. Recovery can begin during response under explicit criteria. Verify restoration assets before use, check restored systems for compromise, remediate root causes, and verify restoration before production use. Confirm service restoration with owners and monitor its adequacy. For an AI deployment-state example, specify affected components, containment, preserved evidence, authorized recovery actions, trusted restoration inputs, verification criteria, and who confirms resumed service; these are application choices implementing the guidance.

  82. Understanding AI Stakes to Break Production Code

    A behavioral-health deployment uses a copilot workflow with traceable inputs, editable paragraph-level summaries, and multiple evaluation stages.

  83. Biases in Electronic Health Record Data Due to Processes Within the Healthcare System

    EHR data reflect physiology and the process of observing and recording it. A diagnosis-code date can mark recognition rather than disease onset; an abnormal laboratory value remains unobserved unless someone orders the test. Clinician decisions, provider and payer workflows, reimbursement, and changing care practices affect what is recorded. This retrospective study separates laboratory values from ordering time and frequency: care-process variables themselves predict later survival. Missing measurements and repeated testing can encode clinical concern rather than random noise. Such signals may fail to transfer when practices change; association with ordering time does not establish that changing it improves survival.

  84. Why Your Enterprise Tech Stack Isn't Ready for AI Agents - And What to Build Instead

    Use a unified, immutable event log as the source of truth, accepting more complex reads in exchange for reconstructable history.

  85. The Target Trial Framework for Causal Inference From Observational Data: Why and When Is It Helpful?

    A causal question asks how outcomes would differ under alternative interventions, not merely which patients are likely to have an outcome. The target-trial framework first specifies the hypothetical randomized trial that would answer that question and then attempts to emulate it using observational data. This makes the comparison explicit and can prevent design-induced bias. It cannot repair inadequate data or make unmeasured confounding disappear; effect identification still requires assumptions such as measured baseline confounders.

  86. Demystifying evals for AI agents

    An agent evaluation separates a task and its success criteria from repeated trials, execution transcripts, graders, and final environment outcomes. A booking claim in a transcript is different from an actual reservation in the database. The system under test includes both model and agent harness. Code-based checks suit precise state or test assertions; model graders cover more open-ended properties but require calibration; human review helps establish the standard. Capability suites explore difficult behavior, while regression suites protect behavior that already works.

  87. Case Study + Deep Dive: Telemedicine Support Agents with LangGraph/MCP

    The workflow aims to turn routine operators into agent supervisors while preserving escalation for situations outside the approved blueprint.

  88. How to Leverage Domain Expertise — Chris Lovejoy, Notius Labs

    Tandem is described as scaling a single medical Oracle into a decentralized Oracle model with doctors responsible for particular customer or workflow subsets.

  89. How to Leverage Domain Expertise — Chris Lovejoy, Notius Labs

    The prior-authorization case reports that scaling assessment did not solve the engineering iteration bottleneck created by differing organizational interpretations of policies.

  90. Citation Needed: Provenance for LLM-Built Knowledge Graphs

    Lineage must survive graph mutation: entity merges retain both source sets, and invalidation records the new evidence responsible for the change.

  91. Mission-Critical Evals at Scale: Learnings from 100,000 Medical Decisions

    With fixed reviewer throughput, maintaining a review fraction makes staffing requirements grow with case volume; reducing the fraction leaves more outcomes unverified.

  92. Shipping AI to a Million Patients Without an A/B Test

    Start with concrete patient harms and turn them into scenarios grounded in the intended clinical workflow.