I. Purpose and development
Healthcare work and settings
Healthcare involves assessing people, delivering care, documenting what happened, and coordinating what happens next. A clinical workflow comprises the physical and mental tasks people perform within and across care environments. Some tasks occur together; others wait for information or another person. Care coordination organizes those responsibilities and communications across participants. Improving one task helps only if the surrounding work still reaches its intended result.
A care setting is the environment in which care occurs. A patient population is the group an application is intended to serve. These definitions become requirements: which people can use the interface, what information is available, and who can respond when assistance is insufficient. The following comparison identifies design considerations, not universal staffing arrangements or response deadlines.
| Setting | Work and information | Responsibility to establish |
|---|---|---|
| Outpatient visits | Care without hospital admission; preparation, consultation, and later follow-up may occur separately. | Ownership of results and unfinished work after the visit. |
| Inpatient care | Care during hospital admission; observations and assessments accumulate across a team. | Who reviews a concern and who assesses the patient. |
| Emergency access | Triage prioritizes attention by urgency rather than supplying a complete diagnosis. | The route to an appropriately urgent clinical response. |
| Remote follow-up | Information arrives through a mediated interaction rather than an in-person examination. | Eligibility for the interaction and access to clinician-delivered care. |
Population limits can be concrete. A 2024 prospective study of the Dora R1 telephone assistant examined follow-up after routine cataract surgery, with an ophthalmologist supervising calls in real time. It excluded non-English speakers and people with hearing or cognitive difficulties. That study therefore cannot establish suitability for everyone who might receive a telephone call. Likewise, domain expertise must match the task: experience as a physician does not necessarily include experience performing medical coding.
Intended contribution
Intended use connects a specific task to its users, patient population, setting, available information, output, permitted actions, and exclusions. Start there rather than with a model or an agent framework. Clinical decision support supplies information at useful points in care to assist decisions. It can take the form of reference information, reminders, summaries, templates, or rules; it need not involve AI. The existing workflow or a simpler rule is therefore a meaningful alternative.
A learned prediction is an estimate produced by behavior fitted from examples rather than entirely authored rules; Machine Learning Fundamentals explains that distinction. A generated answer is a composed response, not necessarily a passage retrieved intact. In retrieval-augmented generation, selected records or references supply information for that composition. Neither mechanism determines who may act on the output.
For example, a bounded result-review assistant might prepare a source-linked summary for authorized clinicians using records available for the selected patient and encounter. Its output remains a draft; missing required information is surfaced for review; it cannot independently place orders. This specification identifies both useful work and a limit on delegation. Displaying a concern, moving it into a review queue, and committing a change to a record are separate operations, even when they originate from the same model output.
Medical-purpose terminology also follows intended use rather than product branding. The International Medical Device Regulators Forum’s Software as a Medical Device definition concerns software performing an intended medical purpose without being part of a hardware medical device. Calling an application a copilot does not settle its local regulatory status, risk class, or authorization requirements.
Continuing lines of development
Healthcare AI developed along several lines because reasoning, information delivery, pattern recognition, and documentation solve different problems. These turning points explain the continuing division of work.
| Landmark | Contribution |
|---|---|
| Ledley–Lusted framework — July 3, 1959 | Robert Ledley and Lee Lusted’s diagnostic-reasoning paper separated logic, probability, and treatment value judgments: foundations for assistance, not a deployed diagnostic service. |
| HELP — around 1970; 1975 | At LDS Hospital and the University of Utah, a common clinical database and then a medical-decision language connected decision support to incoming patient information. |
| MYCIN — 1972 beginnings; 1974 implementation | Stanford’s rule-based consultation project, implemented in Edward Shortliffe’s dissertation, made expert knowledge explicit, inspectable, and modifiable. |
| Learned retinal-image detection — November 29, 2016 | Varun Gulshan, Lily Peng, and collaborators demonstrated detection from learned visual features, leaving clinical use and patient benefit as further questions. |
Contemporary ambient documentation applies language assistance to producing records for later care and billing. It addresses another persistent task rather than replacing diagnosis or hospital decision support. Explicit rules, timely information, learned detection, and documentation assistance remain useful for different parts of healthcare work.
II. Applications and completed work
Documentation and total work
A clinical note records information about care, such as the history discussed, findings, assessment, and plan. It becomes part of the electronic health record, or EHR: the digital record accumulated across care interactions. Ambient documentation captures a care conversation and generates a draft for professional checking and correction. The task ends with usable documentation, not with generated text.
Checking must cover both unsupported additions and consequential omissions. A sentence can sound plausible while introducing something the encounter did not establish; a concise draft can omit information needed by the next clinician. Notes also support billing, so errors can travel into both clinical and administrative work. From Ambient Documentation to Clinical Intelligence explains these downstream uses.
Inspectability can reduce the distance between a disputed sentence and its evidence. Abridge’s linked-evidence interface connects selected note text to transcript passages and original audio. The transcript and note remain derived representations: a link helps a clinician inspect the source but does not establish accurate speech recognition or valid clinical interpretation.
A 24-week randomized trial involving 66 ambulatory practitioners in one health system introduced the documentation-only tool at staggered, randomly assigned stages—a stepped-wedge design. Compared with usual practice, it reported less time spent on notes and lower work exhaustion/interpersonal disengagement. Participants were voluntary early adopters and knew when they were using the tool. Unedited-note quality was assessed by a model judge, an AI model scoring the notes, rather than comprehensive independent clinical review. These findings concern documentation and practitioner outcomes, not improved diagnosis or patient health.
Administration and access
Administrative assistance organizes access and payment rather than determining a diagnosis. A referral requests evaluation or treatment from another clinician or service. Intake assembles the clinical question, relevant history, and records for specialty review. The reviewer may advise, schedule care, or request workup—additional information or investigations needed to assess the patient. The referring team obtains missing material, watches for requested results, and resubmits the referral for review. When a visit is needed, scheduling and follow-up must remain tracked: an appointment does not establish attendance or completed care.
Within this process, a proposed AI intake assistant could organize supplied records and flag missing requested material for the referring team. Its value should be compared with the existing workflow. An AHRQ-supported non-AI referral evaluation found shorter initial waits in most studied medical clinics, but interviews identified shifted work and appointment-selection difficulties. Faster intake alone would not establish better access to completed care.
Payment work has similar boundaries. A payer is the organization financing covered care; prior authorization is a process for reviewing a proposed service against coverage requirements. A denial appeal seeks reconsideration of a rejected request or claim. Preparing a clinical appeal requires reconciling patient records, care guidelines, and payer policies under a deadline—not merely generating a persuasive letter. In the announced Anterior–Stellarus workflow, AI assists research and policy application while licensed clinicians make final determinations. That division of work does not itself demonstrate better access.
Detection and clinical response
Screening looks for a target condition in an intended population. Diagnosis determines what condition is present. Prognosis estimates a future outcome. Triage prioritizes attention by urgency, also called acuity. These tasks produce different outputs, but all require clarity about the response that follows. An urgency category is not a complete diagnosis, and identifying risk does not establish which intervention will help.
Hospital deterioration monitoring makes the response dependency visible. Kaiser Permanente describes Advance Alert Monitor as a coordinated service: threshold crossings prompt off-site nurse review, a reviewing nurse contacts the bedside rapid-response nurse, and bedside assessment informs a care-team decision that considers patient preferences. This operating account explains responsibilities; it is not a model-only intervention or, by itself, a causal estimate of patient benefit.
Outpatient screening must connect image interpretation to care beyond the screening visit. Gulshan, Peng, and collaborators’ 2016 retinal-image study demonstrated detection using learned features rather than explicitly programmed lesion detectors. It tested different operating thresholds against ophthalmologist judgments and called for clinical-setting and patient-outcome studies. IDx-DR’s 2018 pivotal study, led by Michael Abràmoff and colleagues, examined a bounded diagnostic workflow in ten primary-care offices. Trained office staff acquired images, and the system supplied image-quality guidance and a diagnostic result without specialist interpretation of every screening image. This prospective study tested more than classification of an existing image collection, but did not establish completed follow-up or prevention of vision loss.
ACCESS, reported by Risa Wolf and colleagues in 2024, separates screening from subsequent care. At two pediatric diabetes sites, point-of-care AI screening at the visit was compared with referral and scripted education. Follow-up among intervention participants with positive results used a different denominator from screening in either arm. Screening completion and indicated follow-up are different outcomes.
Screening completion and indicated follow-up
ACCESS: youth with diabetes at two pediatric sites. Positive-result follow-up belongs to a subgroup of the intervention cohort.
More generally, high risk does not imply large treatment benefit: an adverse outcome might remain likely with or without the proposed intervention. Prediction and causal effect are separate claims. Patient-facing assistance needs the same discipline. A remote follow-up conversation can collect information and request clinical involvement, but nuanced symptoms and access to clinician-delivered care remain part of the task.
III. Records, exchange, and permitted use
The record as an account of care
An EHR contains accounts of observations and care, not a complete state of the patient. An encounter is a particular interaction with a healthcare service. Patient identity establishes whose information is represented; encounter identity distinguishes which care interaction it concerns. Notes describe events and interpretations, orders request activities, and results report findings. Combining them without preserving those roles changes their meaning.
Medication records illustrate the difference. A prescription contains an order and instructions; dispensing records a supplied quantity; administration records that medicine was given; a medication statement reports that it was taken or given. A prescription is not proof of administration. Similarly, a statement that a symptom was denied, a statement that it occurred previously, and no statement about it are not interchangeable.
Provenance records where information came from and how it was produced or revised. Preserve it alongside the information; Data Quality and Curation explains the general principle. For a clinical result, distinguish when the observation applies, when a result version became available to providers, and when the application received it. A result collected at 09:00 but issued at 10:15 cannot support a prediction made at 09:30. Later corrections must not silently replace the historical account of what was available.
Collection, issue, receipt and retained snapshots
One investigation, with example timestamps. All three application snapshots remain available as historical records.
Status and absence also carry information. A report may be preliminary, final, or corrected. A missing measurement might be unperformed, not asked about, temporarily unknown, withheld, or lost through an error. FHIR supplies distinct data-absent-reason codes for such cases. If no reason is recorded, the application cannot invent one. Information may also exist only in narrative text or another institution’s records; an absent smoking-history code, for example, does not identify a nonsmoker.
Preserve meaning through FHIR
Interoperability means exchanging information that recipients can use with its intended meaning. Fast Healthcare Interoperability Resources, or FHIR, provides typed building blocks called resources, connected through references. A profile constrains a resource for a particular use—for example, by specifying required content or terminology. The FHIR R5 overview describes these mechanisms. Actual integrations must establish their deployed version, profiles, and supported behavior; shared syntax alone does not establish complete or clinically consistent records.
Consider a result-review application. A ServiceRequest describes a requested investigation. A DiagnosticReport can reference that request, its encounter, and Observation resources containing individual results. Patient references identify whose information is represented. These are links between distinct records, not interchangeable copies of one document. A review draft should retain its source-result version so that the basis of the summary remains inspectable.
A result retains its context
ExampleA review draft derives from a result version, while resource references preserve the investigation and patient context.
Read the diagram as text
- ServiceRequest. Requested investigation.
- DiagnosticReport. Report connecting the investigation and its results.
- Observation, selected version. Identified result version with applicable time, status and typed content.
- Patient. Whose information is represented.
- Encounter. The particular care interaction.
- Review draft. Application draft retaining its source-version identity.
- DiagnosticReport → ServiceRequest: Reference: basedOn.
- DiagnosticReport → Encounter: Reference: encounter.
- DiagnosticReport → Patient: Reference: subject.
- DiagnosticReport → Observation, selected version: Reference: result.
- Observation, selected version → Patient: Reference: subject.
- Observation, selected version → Review draft: Derivation: summarized into.
An Observation separates what was observed from its result and optional interpretation. Results can use different datatypes or appear in components; a missing top-level number is not a normal finding. LOINC, Logical Observation Identifiers Names and Codes, supplies shared identifiers for observations. Maintained by the Regenstrief Institute, it helps recipients interpret tests originally named by local codes. LOINC identifies the observation; FHIR transports its representation.
Preserve a code’s terminology system and version; matching display labels do not establish equivalent meanings. UCUM, the Unified Code for Units of Measure, supplies coded units that can support compatible conversions. Preserve comparators too: a result below a value is not exactly that value. Matching units alone does not establish the same measurement, specimen, or method. The FHIR datatype definitions make these representation distinctions explicit.
SMART on FHIR adds application-launch and authorization capabilities. An EHR launch can provide selected patient and encounter context, while declared server capabilities describe what is supported. This is separate from the resource representation. Existing HL7 Version 2 feeds exchange messages for activities such as laboratory orders and patient transfers. Converting them to FHIR still requires local mapping because implementations differ in field use and placement. A new API does not eliminate that interpretation work.
Access and permitted use
Purpose limitation constrains processing to an approved purpose. Data minimization limits information, recipients, and retention to what that purpose requires. Secondary use means another use, such as developing models from material collected for care. Reading an EHR, recording a conversation, sending it to a service, retaining a draft, and reusing it for development require separate consideration. Approve actual service paths explains why the product, feature, endpoint, recipients, and settings matter more than a provider name.
| Artifact | Decision to document |
|---|---|
| Source record | Which patient data and operations the application may access. |
| Encounter audio | Recording arrangements, processing recipients, and disposition. |
| Draft note | Authorized reviewers, record-entry process, and draft retention. |
| Service request | Fields sent, actual processing path, logs, and reuse terms. |
| Audit entry | Action, actor, authority, source/version references, and protected evidence access. |
NHS England’s ambient-scribing guidance calls for explaining the tool and respecting patient objections. In that setting, retain a workable documentation route without optional recording. This is scoped guidance, not a universal consent rule.
Removing obvious identifiers does not settle reuse or eliminate identification risk. HHS guidance describes HIPAA’s Safe Harbor and Expert Determination methods, including identifiers embedded in free text, and acknowledges residual risk even after proper de-identification. A name-stripping function is therefore not a compliance determination. Similarly, useful auditability does not require copying every sensitive payload into ordinary logs: preserve attributable actions and references, with evidence access and retention governed separately.
IV. Patient-specific conclusions
Applicability beyond relevance
A clinical guideline offers recommendations for specified circumstances. Finding a relevant passage is not the same as establishing that its circumstances hold for this patient. Evidence sufficiency concerns whether the available information permits the requested conclusion. Here that requires matching patient identity, current status, relevant history, exclusions, source dates, and local use conditions.
The HL7 clinical-practice-guideline guide separates describing a patient from proposing an activity. It also separates eligibility for a pathway from enrollment: preferences or other clinical considerations can prevent participation even when eligibility holds. An application should therefore keep required premises visible instead of compressing a guideline and a record into an unexplained yes.
| Premise needed | Record evidence | Permitted interpretation |
|---|---|---|
| A particular investigation has a result | A linked result is present. | Summarize that result with its status and source. |
| The result has the required final status | The available report is preliminary. | The final-status requirement is not established. |
| A relevant historical condition is absent | That history is undocumented. | The condition remains unknown, not absent. |
Clinical meaning can fail even when terminology matches. In Lovejoy’s prior-authorization example, an answer treated imaging findings as establishing suspected disease, while the speaker interpreted the existing confirmed diagnosis as changing the guideline answer. The lesson is to evaluate the actual criterion with the history, not to generalize that interpretation to every use of the word suspected. Source identity matters too: a patient assertion and a clinician’s documented finding should not become indistinguishable merely because both entered a summary.
Missing premises can guide the next question. MYCIN worked backward from a goal through expert-authored rules and asked for facts it could not derive. Its explanation subsystem exposed supporting rules and reasons for questions. Its certainty factors represented degrees of belief, not conditional probabilities. The medical knowledge was separate from the execution machinery: removing that knowledge base later produced EMYCIN, a framework for building expert systems in other domains. Neither an inspectable rule trace nor a modern source-linked answer establishes that the underlying medical interpretation is correct. Abstention withholds an unsupported claim; it can still preserve a useful factual summary and identify the information or clinical review required before answering further.
V. Responsibility and integration
Meaningful clinical review
Review is work, not a safety property conferred by an approval button. A reviewer needs suitable expertise, the original evidence, time to inspect it, and authority to reject the output. Expertise should match the workflow and its failure modes. The person operating the software, the person assessing clinical content, and the person authorized to change care need not be the same person.
| Role | Responsibility |
|---|---|
| Service operator | Maintain functioning interfaces, access controls, and incident response. |
| Clinical assessor | Evaluate content against patient evidence and the clinical task. |
| Action authorizer | Approve the particular consequential operation within their authority. |
| Unresolved-work owner | Ensure pending review or follow-up receives an accepted disposition. |
Automation bias is inappropriate reliance on automated advice. In a web-based experiment with 223 dermatologists, participants assessed 24 cases before and after AI advice, including five deliberately incorrect recommendations. Some initially correct answers became wrong after erroneous advice; useful advice was also rejected. This controlled task did not measure patient outcomes, but it shows why oversight assessment must distinguish accepting helpful advice from rejecting harmful advice.
Alert fatigue concerns reduced responsiveness associated with burdensome alerting. A retrospective primary-care study associated more reminders per encounter and repeated reminders for the same patient with lower acceptance. It did not establish a general workload effect or assess whether overrides were appropriate. Lowering override rates is therefore not automatically a safety improvement.
Make corrections diagnostically useful. A review surface can place the source record, guideline, and requested conclusion together, then collect why an answer is wrong rather than only an incorrect label. Such critiques help distinguish missing evidence, misunderstood criteria, and unsupported additions. Review before consequential commitment develops the interface mechanics. Patients and staff also need a route to report disputed content to an accountable reviewer and learn its disposition.
Accepted clinical handoffs
A clinical handoff transfers information, responsibility, and authority. Under AHRQ TeamSTEPPS guidance, the sender remains responsible until the receiver acknowledges understanding and acceptance.
For application tracking, distinguish requested, received, accepted, and resolved. Assign local deadlines and overdue handling, including after-hours coverage. Workflow Automation develops this coordination contract.
A patient-facing voice demonstration makes the boundary concrete: after additional symptom information, the assistant recommends immediate nurse involvement. The recording demonstrates a request, not completion of a transfer. Likewise, a code-controlled routing layer can ensure that a selected clinical path runs before ordinary conversational generation, but deterministic execution does not establish correct recognition of urgency. Both detection and accepted takeover need assessment.
Commitment at the point of care
A workflow trigger starts consideration of work: a clinician request, incoming result, or scheduled review. Point-of-care integration determines where assistance appears, who receives it, and how long it remains relevant. A draft is not finalized documentation; a suggestion is not an order; delivery of an alert is not a completed response.
SEIPS, the Systems Engineering Initiative for Patient Safety model introduced by Pascale Carayon and colleagues in 2006, connects people, tasks, tools, organizational conditions, and physical surroundings to care processes and outcomes. Its work-system perspective explains why a technically correct feature can still create interruptions or transfer workload. Review time and staffing belong in the design. HELP’s earlier data-triggered and time-based support illustrates the enduring importance of delivering information within actual work.
Acceptance must also account for change. A clinician may review a draft before a new result arrives or another user edits the target record. Refresh clinical context when relevant premises change; separately check authority and target-record concurrency. SMART distinguishes selected patient context from granted resource operations. Approval in the application does not create an EHR permission.
FHIR R4’s version-aware updates use ETag and If-Match to reject overwriting a resource whose version changed. Conditional creation can avoid another resource when specified matching criteria identify an existing one. Server support must be checked, and these protections do not prove clinical appropriateness or freshness of every source used by a draft. Authorization at the protected operation and real integration tests cover the underlying mechanics. Test wrong-patient context, intervening changes, rejection, and uncertain submissions as part of the clinical workflow.
Reviewed is not yet committed
ExampleClinical validity, current authority, and target-record concurrency are independent acceptance conditions.
Read the diagram as text
- Reviewed proposal. Proposed content and identified source versions.
- Current clinical context. Confirm patient and encounter; reassess relevant source changes.
- Current write authority. Check the requested operation against server permissions.
- Target version matches. Where supported, apply version-aware update protection.
- Refresh and review. No commitment of this proposal under stale or conflicting conditions.
- Denied. Do not perform an unauthorized write.
- Commit confirmed. Authoritative result confirms the accepted record change.
- Reviewed proposal → Current clinical context: Submission requested.
- Current clinical context → Refresh and review: Context invalid or materially changed.
- Current clinical context → Current write authority: Context remains applicable.
- Current write authority → Denied: Permission absent.
- Current write authority → Target version matches: Permission granted.
- Target version matches → Refresh and review: Version conflict.
- Target version matches → Commit confirmed: Match and successful update.
VI. Evidence for use
Validate the task and population
Clinical validation assesses clinically relevant performance for the intended task, population, and care context. IMDRF’s clinical-evaluation framework distinguishes a valid clinical association between output and condition, reliable production of the intended output from inputs, and achievement of the intended clinical purpose. Technical correctness, clinical performance, workflow behavior, and patient outcomes remain separate claims; Evals and Benchmarks develops that general distinction.
A reference standard is the procedure used to establish the comparison answer. It may involve independent expert assessment, adjudication of disagreement, or later outcome ascertainment. Agreement with a comparator that does not establish condition status cannot automatically be called diagnostic sensitivity or specificity. Report counts and sampling uncertainty, not percentages alone. More observations reduce sampling uncertainty but do not remove a biased reference procedure.
The distinction predates modern models. Yu and colleagues’ 1979 MYCIN evaluation asked eight outside specialists to assess anonymized prescriptions for ten deliberately challenging retrospective meningitis cases. The study assessed expert acceptability, with disagreement among evaluators. It did not establish routine-care outcomes, cost effectiveness, or deployment readiness.
The target itself can be wrong. Obermeyer and colleagues’ 2019 study examined a system using predicted healthcare spending as a proxy for need. Black patients were sicker than White patients at the same score, with lower spending at comparable illness burden reflecting unequal access and utilization. Accurate cost prediction could therefore allocate help poorly. The paper illustrates target mismatch, not merely a model fitting error.
A suitable target can still fail to transfer. A 2021 external validation of a widely deployed sepsis predictor found poor performance in one academic center and examined calibration, alert burden, and added value relative to existing practice. This was a historical evaluation of a specified model and cohort, not a judgment about every version or hospital. Language, age, comorbidity, equipment, and access patterns can all define materially different use conditions.
| Evaluation boundary | Main transfer question | Remaining limitation |
|---|---|---|
| Patient separation | Performance on unseen people. | Does not establish future or cross-site performance. |
| Later period | Performance as time and practice change. | Returning patients may still overlap. |
| Another institution | Performance under another site's conditions. | One external site does not represent every setting. |
Errors, probabilities, and workload
Sensitivity is the fraction of actual cases detected. Specificity is the fraction of non-cases correctly left unflagged. Positive predictive value is the fraction of alerts that are actual cases. Prevalence is the proportion of the evaluated population with the target condition or event. Sensitivity therefore does not tell a reviewer how often an alert is correct.
Hold sensitivity and specificity fixed to isolate prevalence. Real population changes can also alter those rates.
Prevalence changes the alert workload
Hypothetical cohort of 1,000; sensitivity 80% and specificity 90% stay fixed. Change only prevalence to isolate its effect on alerts.
Positive predictive value: 80/170 = 47.1%. 20 actual cases are missed.
Reference comparison · available without changing controls
Lower prevalence leaves fewer real cases while false alerts still arise from the much larger non-case population. Review demand must therefore be assessed alongside detection: count alerts and missed cases, then measure the actual work and response delays they create. Changing an action threshold changes that workload and its errors; available capacity is a constraint, not evidence that an arbitrary higher threshold is clinically acceptable.
Calibration concerns absolute probabilities. Among comparable patients assigned approximately 10% risk for a defined event and horizon, approximately 10% should experience it. Ranking higher-risk patients well does not establish calibrated probabilities, and calibration at one institution may not transfer. Nor does calibration select an appropriate intervention. See Interpret probability forecasts. A confidence interval for measured sensitivity answers another question: uncertainty in a population performance estimate, not an individual patient’s event probability.
For notes and conversations, measure the errors that matter to that task rather than forcing every output into a diagnostic score. Clinicians can identify unsupported additions, omissions, incorrect intervention categories, or intervention at the wrong conversational turn. In Evals Driven-Development for a Mental Health AI Coach, clinician annotations become typed regression cases with conversation input, an expected observation, category, and replay turn. Both missed and inappropriate interventions matter. If a model grades these cases, validate the judge against independent clinical assessment; a generated critique is not its own reference standard.
Evaluate the complete intervention
Retrospective evaluation uses previously collected cases. Prospective evaluation is planned around newly occurring work. Silent operation runs on current inputs while withholding outputs from care decisions. It can reveal input and performance problems, but cannot show how clinicians respond to advice they never see. Usual care is the actual existing alternative, including its staffing, tools, and follow-up—not simply the absence of a model.
| Design | Exposure | What it can investigate |
|---|---|---|
| Retrospective testing | Previously collected cases. | Specified output performance under reconstructed conditions. |
| Silent prospective operation | New inputs; outputs withheld from care. | Current data availability and prospective performance. |
| Supervised live evaluation | Users see and may act on outputs. | Reliance, usability, workflow changes, and observed failures. |
| Comparative outcome study | Defined assisted and comparison workflows. | Differences in prespecified outcomes under the study design. |
DECIDE-AI, published in 2022, provides reporting guidance for early live evaluation, including actual use, overrides, workflow, and harm. CONSORT-AI, published in 2020, extends clinical-trial reporting to describe algorithm version, inputs, users, outputs, and effects on decisions. Both make the human–AI intervention more inspectable; neither proves effectiveness or supplies authorization to deploy.
Random assignment uses an unpredictable chance process rather than prognosis or preference to allocate an intervention. Concealing the next assignment prevents recruitment from being steered toward a particular arm. Assignment may be by patient, clinician, or care unit: shared staff and practices can let one arm affect another, so the unit must fit the workflow. Concurrent observation alone does not randomize exposure. Choose the live experiment develops these design choices. Clinical randomized studies are not categorically prohibited; exposure and safeguards require an appropriate study and institutional review.
Keep endpoints separate. The ambient-documentation trial measured practitioner and documentation outcomes; ACCESS distinguished screening from indicated follow-up. Neither established the patient-health claims those intermediate outcomes might suggest. A surrogate endpoint substitutes another measure for an outcome directly relevant to patients. FDA–NIH BEST emphasizes that correlation with health outcomes is insufficient: intervention-induced changes in the surrogate must reliably predict benefit in the specified context. An efficiency gain, a clinical benefit, a safety assessment, and permission to operate therefore require different support.
VII. Operation and recovery
Monitor changing conditions
Distribution shift means operating conditions differ from those evaluated. Patient mix, measurement practices, interfaces, protocols, staffing, and model versions can change different parts of the system. A detected change warrants investigation; it does not alone establish degraded clinical performance. Machine Learning Fundamentals explains the broader concept. Monitoring should preserve separate views of input availability, meaning, output quality, response workload, and outcomes.
| Observed change | Initial investigation | What remains unproven |
|---|---|---|
| Fewer results arrive | Inspect source and interface availability. | Whether affected patients have normal findings. |
| A field or unit changes | Review mapping with source-system and clinical owners. | Whether prior interpretation remains valid. |
| More alerts remain pending | Inspect case mix, review effort, staffing, and acceptance. | Whether detector performance itself deteriorated. |
| Recent outcomes are unavailable | Track follow-up maturity and ascertainment. | Clinical performance in that incomplete cohort. |
Missing documentation, delayed outcomes, and selective follow-up are different problems. An absent historical entry might never have been recorded. A recent outcome may not yet be observable. Follow-up confined to selected patients can leave the others without comparable labels. Do not classify these unknowns as successes or negative outcomes; incomplete feedback requires explicit accounting.
Actions can also change the outcome being monitored. In the clinical AI quality-improvement framework’s hypothetical early-warning example, an alert prompts treatment and the predicted event does not occur. The record alone cannot distinguish a false prediction from successful prevention: the outcome without treatment is unobserved. Simple prediction-versus-outcome scoring can therefore misdescribe the intervention’s effect.
A non-event leaves the treatment effect unresolved
Same patient, one observed course. The alternative is not a second patient or a measured study arm.
Observed course
- Alert prompts treatment.
- Treatment is followed by no observed event.
Counterfactual · unobserved
Without that treatment, what would have happened to this same patient?
Outcome unknownRetain source and model versions, relevant policy and workflow versions, observed effects, and reviewer decisions under defined access and retention controls. Sensitive telemetry needs its own governance. Clinical and technical owners should agree on conditions for reassessment, restricted use, and suspension. A model or prompt change requires renewed evaluation of affected behavior; a healthy request pipeline is not a substitute.
Recovery and affected work
Recovery begins by identifying what failed. Degraded operation deliberately provides a smaller supported service. Reconciliation determines what actually happened and resolves discrepancies against authoritative records. A record amendment corrects or qualifies documentation while preserving attributable history. These responses solve different problems; restoring a model version addresses future execution, not every effect already produced.
| Failure | Required response |
|---|---|
| Assistance unavailable | Continue through an established non-AI workflow; retain ownership of pending work. |
| Required source records unavailable | Withhold dependent conclusions or clearly bound any partial information. |
| Submission outcome uncertain | Inspect authoritative status and use supported duplicate-protection mechanisms before repetition. |
| Accepted output known to be wrong | Review affected work, correct through the applicable record process, and assign necessary communication. |
The SAFER contingency-planning guide recommends tested downtime procedures, named clinical and technical leadership, communication independent of failed infrastructure, and alternative documentation. Restored systems must reconcile information captured during downtime. Suspending assistance must not silently abandon urgent, overdue, or unaccepted work.
An uncertain write is not a confirmed failure, and resource-level duplicate protection is not proof that a clinical action occurred exactly once. A known wrong note has another remedy: correction with history. The historical CMS amendment guidance illustrates preserving distinguishable original and changed content with authorship and dates. It is not a determination of current local policy. Nor can a record correction undo an utterance already heard or an action already taken.
Restoring software without erasing effects requires separate completion evidence. Before resuming the affected use, verify repaired dependencies and an evaluated operating mode, reconcile pending effects, and confirm ownership of unfinished work. Long-running corrective work may continue under an accountable plan. The final measure of recovery is a workable care service with understood obligations—not merely a responding endpoint.
Open questions
Detection can expand faster than follow-up capacity. Establishing useful expansion requires studying completed indicated care, delays, and workload across services—not screening volume alone. Progress would demonstrate that additional detection reaches appropriate follow-up without leaving less visible groups behind.
Effective oversight must remain effective under routine workload. Reviewers can accept incorrect advice and reject useful advice, while repeated alerts change response behavior. Progress would show sustained error detection and correction in real workflows, with review effort and downstream consequences measured together.
Clinical monitoring becomes difficult when outcomes arrive late or treatment changes them. A non-event after an alert may reflect successful intervention rather than a false alarm. Progress requires evaluation designs that separate prediction quality, response behavior, and intervention effects while keeping unavailable outcomes explicit.
Local adaptation must accommodate specialty and organizational differences without silently changing clinical authority. Decentralized expert review can improve contextual fit, but safe update methods and cross-setting regression evidence remain necessary. Progress would make each local change attributable, reviewable, and independently assessable.














































