Financial work and its development
Four kinds of financial work
Financial institutions and finance teams work with rights and obligations. A financial instrument is a contract that creates a financial asset for one party and a financial liability or equity instrument for another. A loan, for example, connects a lender's right to receive payments with a borrower's obligation to make them. This contractual meaning matters: describing an obligation, recording it, and fulfilling it are different activities. AASB 132 supplies the formal definition.
Financial operations maintain financial records and carry transactions through their required stages. Decision support prepares information for a consequential choice. Neither automatically grants decision authority: permission and responsibility to approve or execute that choice. Research, operations, customer service, and decision support can share document-reading capabilities while requiring different completion checks.
| Work | Inputs and useful output | Responsibility and possible effect |
|---|---|---|
| Research | Disclosures and records → checked analysis | An analyst reviews the conclusion; a separate decision may commit capital. |
| Operations | Statements and internal records → resolved discrepancies | Operations staff control corrections and completion; errors can misstate books or misdirect payments. |
| Customer service | Customer request, policy, and account facts → resolution or escalation | Service staff act within account permissions; answers can affect access to help and financial choices. |
| Decision support | Applicant or transaction information → assessment and proposed disposition | Authorized decision-makers determine credit or payment treatment; affected parties include borrowers and legitimate customers. |
This breadth is visible in the Bank of England and FCA's 2024 survey of 118 firms. Respondents described operational, fraud, risk, and other applications; the survey separately classified automation and materiality. Adoption therefore did not mean autonomous authority. The reported benefits were firms' perceptions, not controlled measurements of productivity.
Scoring, rules, and financial language
Financial AI draws on several continuing traditions. Credit scoring estimates credit-related outcomes from applicant or account information. An expert system applies explicitly represented domain rules. Financial language processing interprets text whose meaning depends on accounting and business context. Their histories explain why a language assistant can extend an existing workflow without replacing its numerical models or controls.
Contributions that still coexist
1941Credit-Rating FormulaeCombines borrower attributes statistically to support credit investigation.
Contributors: David Durand; 1941 NBER chapter.
What changed: Examined statistically derived weights as an alternative to intuitive weighting. Durand treated the formulae as supplements to judgment and warned that previously approved loans did not represent the full applicant population.
1988 deploymentAuthorizer’s AssistantCombines account records with credit and fraud policies for authorization.
Contributors: American Express and Inference Corporation; reported in 1989 by Dzierzanowski, Chrisman, MacKinnon, and Klahr.
What changed: Handled selected U.S. Personal and Gold card transactions automatically and advised human authorizers on others. Integration, acceptance testing, and transfer of business control to operations were part of deployment.
1992FalconApplies neural-network modeling to payment-fraud detection.
Contributors: FICO; date attributed to its retrospective corporate history.
What changed: Addresses suspicious payment patterns, distinct from estimating repayment risk or processing credit applications. FICO’s retrospective dates Falcon to 1992.
February 2011Loughran–McDonaldMeasures financial text with domain-appropriate word categories.
Contributors: Tim Loughran and Bill McDonald; When Is a Liability Not a Liability? Textual Analysis, Dictionaries, and 10-Ks.
What changed: General negative-word lists mischaracterized financial filings. The contribution was better measurement of financial meaning, not simply a larger vocabulary.
March 2023BloombergGPTUses mixed financial and general text for broad financial-language tasks.
Contributors: Shijie Wu, Ozan Irsoy, and colleagues; March 2023 introduction.
What changed: The 50-billion-parameter model supports tasks including entity recognition, entity linking, and question answering. Its benchmark findings concern language capabilities, not autonomous investment decisions or completed transactions.
These developments did not eliminate one another. A financial application can use language processing to interpret documents, a predictive model to prioritize cases, explicit rules to enforce policy, and people to investigate exceptions. The useful architectural question is which responsibility each component can support.
Financial records and their use
Preserve what a financial fact means
An issuer creates an instrument; an account records a financial relationship and activity; a holding, or position, gives the instrument quantity held there. A transaction records activity; a balance gives an amount at a time. Keep these identities distinct when joining records: an issuer does not identify an instrument, holding, or payment.
Different records answer different questions. Reference data supplies identifying and contractual attributes; a security master is an institution's maintained collection of such instrument records. Transaction data records activity. Financial statements summarize position or performance. A balance sheet reports assets, liabilities, and equity at an instant: resources, obligations, and the owners' residual interest. An income statement reports revenue and expenses over a period; a cash-flow statement reports cash movements. Profit therefore differs from cash received. Before comparing extracted amounts, preserve the entity, measure, period, currency, scale, and accounting context. The U.S. Securities and Exchange Commission's financial-statements guide explains these distinctions.
Standards preserve different parts of this meaning. A Legal Entity Identifier (LEI) identifies a legal entity; an International Securities Identification Number (ISIN) identifies an instrument. Their issuer mapping does not identify a customer's holding or payment, and missing coverage is not proof of nonexistence. eXtensible Business Reporting Language (XBRL) represents reporting facts with context rather than bare values.
| Field | Meaning to preserve |
|---|---|
| Entity and concept | Whose assets the fact describes, and which reporting concept it uses. |
| Reporting instant | 2020-01-01 at midnight represents the end of December 31, 2019. |
| Value and unit | EUR 1,230,000; numerical strings avoid representation-related precision loss. |
| Precision and qualifications | Declared precision, applicable dimensions, and footnotes qualify interpretation; they do not certify truth. |
Extraction must retain these relationships even when they appear in a table heading or footnote rather than beside the number. See Recover table meaning. A standardized record is easier to process, but still needs semantic checking against its source.
Separate reporting time from known time
Provenance records information's origin and history. Distinguish reporting period, publication, application receipt, and correction. As-of information is eligible at a cutoff. A restatement revises earlier reported amounts; the IAS 8 overview describes retrospective correction of material prior-period errors, subject to its conditions.
FRED's real-time periods retrieve historical information; its current view includes revisions. Date-level vintages establish neither intraday publication nor ingestion. Application replay must respect later receipt.
Aggregation can obscure this boundary. EDGAR frames select the last-filed fact fitting a calendar period; that historical period is not a historical availability guarantee. Its aggregate XBRL APIs also exclude some custom and segment disclosures. Preserve the selected source version rather than assuming an API's convenient view is the required information set.
In a hypothetical example, both versions describe the same company’s quarterly revenue in USD millions. The application receives the original value of 100 before its decision cutoff. A correction to 96 is published before the cutoff but reaches the application afterward. Historical replay retains 100; the later corrected view uses 96. See Origins, times and permitted use for provenance and Applicable sources and unresolved conflicts for alignment before synthesis.
Published does not mean received
Both versions describe Company C’s quarter Q revenue, in USD millions. Event order runs left to right; spacing does not measure elapsed time.
At cutoff: USD 100 million
The original was received; the published correction had not reached the application.
After revision receipt: USD 96 million
Corrected reporting uses the revision. It does not rewrite the historical input.
Check permission for the actual use
Governance determines and enforces appropriate information use. Access is only one part of permission: a source may be readable internally but restricted from external processing or redistribution. Source entitlements specify permitted access and uses; redistribution passes information to another recipient. The general distinction among legal, contractual, organizational, and technical permission is developed in Establish permitted uses.
Public visibility does not remove contractual restrictions. CME's inspected website market-data terms prohibit specified AI uses of covered data and metadata, including training and generating outputs. That example does not describe every exchange or a separately negotiated license. It shows why a programmer must check the actual agreement before putting a visible price into an AI pipeline.
Some information requires restrictions because of its significance to investors, independently of a data license. In U.S. securities terminology, material nonpublic information is information important to a reasonable investor that has not been disseminated so investors generally can access it. The SEC's adopting release explains these concepts; confidentiality alone does not establish materiality. Information barriers restrict flows between functions, including flows to investment decision-makers. Putting restricted records into a broadly searchable internal index can defeat that separation even if the records never leave the institution.
Review each proposed path separately: internal analysis, processing by a model provider, and disclosure to a customer. Name the purpose, recipient, fields, permitted retention, and output use. Combining eligible inputs does not settle whether the combined output may be disclosed. Provider approval must cover the actual service path, not merely a familiar company name; see Approve actual service paths.
Applying AI to financial workflows
Build checked financial research
Financial research turns eligible information into an analysis someone can inspect. A disclosure communicates financial or business information, such as a company filing. An earnings call is a discussion of reported results and business conditions, often including management commentary and questions. A research workflow begins with a specific question, locates relevant sources, aligns comparable facts, performs necessary calculations, and presents a conclusion for analyst review.
Keep the kinds of statement visible. Reported revenue is a source fact; management's expectation is an attributed statement; a ratio is a derived quantity; next year's estimate is a forecast; an assessment of business quality is a judgment. Grounding connects a claim to applicable supporting material. It does not make a forecast true or prove an analyst's interpretation. Check support, not just references explains the general checking task.
Financial vocabulary also changes what counts as support. Loughran and McDonald's 2011 study found that 73.8% of general-dictionary negative-word occurrences in its 10-K corpus came from words typically not negative in financial usage. A liability can be an ordinary accounting category rather than evidence of deteriorating conditions. The lesson is to validate the meaning being measured, not just recognize financial-looking words.
Calculation and interpretation
For numerical analysis, separate deciding what to compute from executing the calculation. The model can identify the needed facts and propose operations; tools retrieve the values and perform the arithmetic. Preserve the selected inputs and operations so the result can be checked. Exact execution still cannot establish that the chosen inputs or formula answer the question. FinQA, introduced in 2021, made this separation explicit by pairing financial questions with executable reasoning programs over report text and tables. Kepler's financial-services approach assigns planning to the model and extraction and arithmetic to tools, retaining each number's derivation.
For example, suppose compatible records report quarterly net income of USD 8 million and revenue of USD 100 million. Under the selected definition, net margin is 8 ÷ 100 = 8%. The calculation service can reproduce that result. The analyst must still establish that both figures concern the same entity and period and that the selected definition answers the question. Concluding that this margin is sustainable introduces a further judgment the division cannot verify.
Comparisons need equal care. Non-GAAP measures use adjustments outside the relevant generally accepted accounting principles. The SEC's interpretations describe how changed adjustments and misleading labels undermine comparison. Similarly named measures need not be equivalent across firms. Preserve definitions and adjustment histories alongside values. A useful research assistant reduces total inspection and correction effort while leaving these decisions visible.
Reconcile records and resolve exceptions
Financial operations must distinguish recording an obligation from fulfilling it. A ledger is an accounting record; posting enters an accounting transaction into it. Settlement completes the agreed transfer of cash or assets. In double-entry bookkeeping, an entry has debit and credit sides whose totals must match. A balanced entry need not move cash: providing a service on credit records revenue and a receivable, an amount the customer owes. Collecting payment later increases cash and reduces the receivable without recording revenue again. The entries balance in both cases, but only collection brings in cash. Balance alone does not prove correct accounts, amounts, or dates.
Reconciliation compares independently maintained records and investigates differences. An exception is an item requiring handling outside the routine path: an unmatched payment, an ambiguous correspondence, or a discrepancy needing explanation. AI can help interpret payment descriptions and assemble likely supporting records. The proposed match must still satisfy the financial correspondence, not merely resemble another description.
Consider a statement line for USD 500 and two internal records for USD 500. Amount equality does not establish which record corresponds to the payment. A transaction reference, counterparty, date, or further investigation must resolve the ambiguity. Dynamics 365's matching documentation makes this concrete: default rules can take the first qualifying transaction, while configuration can require manual matching when multiple documents match on amount. A successful rule execution is not evidence that its selection was correct.
Keep completion states equally precise. In Dynamics' reconciliation workflow, unmatched items can carry forward after a statement is marked reconciled, and completion can trigger correction postings. Separate a proposed match, an approved correction, a confirmed posting, and an unresolved item. Otherwise, an automation-rate dashboard can conceal both open discrepancies and unintended accounting changes.
Resolve the customer's financial problem
Account servicing helps customers understand and manage an existing financial relationship. Agent assist supplies information or drafts to a human service representative. DBS's July 2024 announcement described an assistant combining call transcription, knowledge-base search, summaries, and prefilled service-request fields. Its pilots began in October 2023. This is a concrete assistance workflow; the announcement's projected handling-time savings were not demonstrated end-to-end results.
General product information and account-specific information require different access. Authentication establishes control of an account identity; authorization determines permitted operations. Neither a recognized customer nor a retrieved policy automatically authorizes an account change. See Distinguish people and actors. The service design should identify which record supplies the balance, transaction status, applicable fee, and effective policy rather than letting the model blend them.
A pending transaction has not finished processing or posting. Its amount can change, as with a restaurant tip; it is not interchangeable with a posted entry. Product and channel also matter. Chase's published dispute process requires a credit-card charge to post before opening a dispute. It allows a pending debit-card transaction to be disputed by telephone, while online debit disputes require posting. Applying one rule to every pending transaction gives incorrect instructions.
Dispute handling addresses a contested transaction, not merely a request to explain its status. For covered U.S. electronic fund transfers, the Regulation E interpretation distinguishes checking whether a transfer occurred from alleging an error. Reporting a lost access device together with possible unauthorized use changes the servicing path. These are scoped examples, not universal rules for every payment product.
The assistant should say what is established, avoid promising an unconfirmed refund or completion date, and transfer unresolved facts or out-of-authority requests to a reachable owner. The CFPB's chatbot report describes generic answers and difficult human offramps leaving financial problems unresolved. Count verified resolution, repeat contact, and successful escalation—not just conversations that ended without a person.
Keep scores separate from decisions
Credit underwriting assesses a proposed loan and its terms. Default means failure to meet the relevant contractual obligation. Fraud detection identifies activity that may be deceptive or unauthorized. Both can use a predictive model—behavior fitted from examples, as introduced in Machine Learning Fundamentals—but neither task reduces to producing a score.
An additive credit scorecard assigns points to applicable borrower attributes and sums them. In the SAS implementation, grouped attributes feed a fitted statistical model that is scaled into points. Grouping, fitting, scaling, and the subsequent approval policy are separate choices. A convenient score direction or point total does not itself establish calibration or an appropriate lending threshold.
A decision policy maps information to actions under objectives and constraints. Ranking cases orders attention; a calibrated probability estimates how often a defined outcome occurs among similarly scored cases. Even a meaningful probability does not choose the action. Changing the cost of blocking a legitimate payment relative to missing fraud can change a threshold without changing the prediction. From probabilities to actions develops this distinction.
The Worldline-informed fraud study separates authorization checks from later investigation using rules, learned scores, records, and cardholder contact. An alert is a reason to investigate, not a finding of fraud. In credit review, document synthesis can prepare an assessment, while policy and authorized judgment determine the disposition. Missed losses and unjustified denial or blocking must remain separate consequences in both workflows.
Preserve the actual decision path when explaining an outcome. For covered adverse-action statements, Regulation B requires specific principal reasons reflecting factors actually considered or scored. A plausible explanation generated afterward cannot substitute for those reasons.
Identify the exposure and scenario
Risk exposure is a position or obligation through which an adverse event can cause loss. A counterparty is another party to the financial arrangement. Before choosing a model or explanation, identify the loss mechanism and the time horizon. Several mechanisms can affect the same transaction.
| Risk | Mechanism | Relevant evidence |
|---|---|---|
| Credit | A borrower or counterparty fails to meet obligations. | Obligation, due time, outstanding exposure, and recoverable value. |
| Market | Market-price movements reduce a position's value. | Positions, price sensitivities, and adverse market assumptions. |
| Liquidity | Cash or collateral is unavailable when needed, or selling a position materially moves its price. | Cash timing, funding access, market depth, and exit horizon. |
| Operational | A failed process, person, system, or external event causes loss. | Routing, approvals, execution records, and failed controls. |
Expected credit loss depends on the likelihood of default, the amount exposed at default, and the portion lost after recovery. A ranking of borrowers addresses only part of that assessment. Funding liquidity concerns meeting cash and collateral needs; market liquidity concerns trading without unacceptable price effects. An institution can have assets exceeding liabilities yet lack cash at the required time.
A scenario asks what follows if specified conditions hold. It might assume delayed receipts, reduced recoveries, or difficulty selling assets. Its output is conditional analysis, not a guaranteed forecast. Expected loss also differs from extreme-loss severity: a typical outcome or percentile cutoff can conceal worse losses beyond it. Keep scenario assumptions and the exposure they act on visible so reviewers can judge whether the analysis covers the decision.
Authority and authoritative effects
Authorize the actual financial action
Reading a record, drafting a recommendation, approving it, and submitting an instruction are different permissions. Segregation of duties separates responsibilities that should not be exercised unchecked by one actor. Maker-checker control separates preparation from independent approval; calling a second model does not create that institutional independence. Basel Principle 26 addresses delegation, approval limits, reconciliation, and separation among committing the bank, paying funds, and accounting for assets and liabilities. These supervisory standards are not automatically local law.
Delegated authority should identify who may act for whom, on which accounts or instruments, for what amounts and recipients, and for how long. Enforce these restrictions outside generated content. The model may propose an action; a protected operation must determine whether that action is permitted. Authorization belongs at the protected operation explains the general interface boundary.
Approval must remain attached to the actual proposal through delays and retries. A useful financial precedent is Article 5 of EU Regulation 2018/389: within its strong-customer-authentication scope, the authentication code binds to the agreed amount and payee, and changing either invalidates it. As an engineering rule, changed material arguments, expired approval, or invalidated source conditions should stop submission and require the appropriate renewed decision.
Limits must also cover accumulated exposure. The SEC staff's market-access FAQ describes aggregate credit or capital thresholds and checks before orders enter the market for covered broker-dealers. A per-order limit alone cannot establish an aggregate bound. Test concurrent submissions against shared exposure, including relevant pending commitments; the required property is that individually acceptable requests cannot jointly bypass the limit.
Customer communications deserve explicit authority too. Morgan Stanley's June 2024 Debrief announcement describes drafting an email that an advisor may edit and send at their discretion. It separately describes saving a note into Salesforce. A review boundary for one output does not establish the boundary for every write the application performs.
Connect terms, entitlement, and receipt
A system of record is the designated authority for a particular business fact. One source may own instrument terms, another accounting entries, another external transaction status, and another investigation status. An assistant's summary is a derived view of these records. Name the authoritative records explains the general ownership principle; a financial integration must also preserve identifiers, exact monetary values, units, relevant dates, and evidence of completed effects.
Consider bond-interest reconciliation as a teaching example. A bond is debt issued to investors. Principal, also called face value, is the amount repayable under its terms; interest is the contractual payment for borrowing. These are promised payments, not guaranteed receipts, and market value can differ from principal. The SEC's bond guide introduces this vocabulary.
A custodian holds or administers clients' assets. Entitlement is the payment due under event rules. Review identified, dated terms, eligible holdings, calculations, and external and internal records. Today's holding times a displayed rate is insufficient: eligibility and contractual conventions matter. Keep licensed terms and restricted holdings on permitted processing paths.
The Depository Trust Company's (DTC) Distributions Service Guide separates announcements, entitlements, and allocations of received funds. Payable dates do not prove receipt; position or rate corrections can generate credits or debits. Participant allocation does not prove bank receipt by the beneficial owner, the investor whose assets are held. Reconcile corresponding external allocations and internal postings, retaining unresolved differences.
Resolve uncertain execution
A lost response changes what the application knows, not necessarily what the provider did. Keep unknown outcome separate from confirmed failure without an effect and confirmed completion. Preserve the logical operation identity while consulting authoritative records. A fresh payment request can duplicate an effect that already occurred.
Idempotency is a receiver-defined contract for repeated requests. Stripe's contract reuses the first executed request's stored status and body when the same key and parameters recur, including stored failures. Parameter changes are rejected. Keys can be pruned once at least 24 hours old; reuse after pruning creates a new request. Validation failures and concurrent execution conflicts do not create saved results. This is bounded duplicate protection, not permanent exactly-once settlement.
Stripe's error guidance recommends retrying network failures with the same key and parameters using backoff. A server error can remain indeterminate despite a repeatable cached response. Provider reconciliation and subsequent events may reveal resulting objects; a local operation identifier in metadata helps correlate them. Preserve pending state until authoritative evidence resolves it rather than changing the key to escape uncertainty.
The effect and knowledge of it can diverge
One possible run of logical operation P. Time runs downward; spacing does not encode duration. The request arrives and creates an effect, but its response is lost.
Other possible resolutions: confirmed effect, confirmed no effect, or still unknown. These are alternatives, not three later events in this run.
Recovery can create new financial work. A compensating entry records a correction instead of erasing the original accounting event; Dynamics documents reversal through a new transaction rather than editing an already reconciled statement. Track the original effect and the remedy separately, including whether the remedy completed. See Recover from the effects that occurred for the wider recovery pattern.
Evidence of useful performance
Measure the financial workflow
An evaluation assesses specified behavior against an intended use. A baseline is the credible alternative: existing rules, software, models, and human work—not necessarily the newest competing model. Define the assessment unit before choosing a metric. A correct answer, a correct match, and a resolved case describe different units and support different claims. What an evaluation establishes and Define worthwhile improvement provide the general framework.
| Workflow and unit | Checks | Complete-work comparison |
|---|---|---|
| Research: claim and analyst task | Source support, numerical correctness, omitted material, and justified interpretation. | Coverage and total analyst time, including checking and correction. |
| Operations: match and reconciliation case | Correct correspondence, confirmed postings, unresolved amounts, and deadlines. | Complete handling effort, remaining exceptions, and correction burden. |
| Service: customer problem | Applicable information, authorized actions, resolution, and usable escalation. | Repeat contact, waiting, access to help, and customer consequences. |
| Decision support: decision and affected population | Policy adherence, decision-relevant errors, and outcomes after sufficient follow-up. | Losses, unjustified restrictions, group differences, and investigation effort. |
FinanceBench, introduced by Pranab Islam and colleagues in November 2023, paired financial questions with answers and supporting passages. Its experiments assessed 2,400 outputs across 16 configurations on 150 cases. Providing the correct evidence pages improved results but did not eliminate reasoning errors; qualitative review also found valid answers differing from references. It is useful evidence about the separation of retrieval, reasoning, and assessment, not certification for a new institution's workflow.
For an investigation queue, precision is the fraction of flagged cases confirmed to have the specified condition. The condition matters: an anomaly is not necessarily fraud. In her cross-document compliance presentation, Varsha Shah reported approximately 91% precision in an evaluation involving roughly three million records over five years and four jurisdictions. Confirmation concerned genuine anomalies, not uniformly established fraud. The report did not supply the labeling process or evaluation split details. This result concerns the quality of flagged cases; it does not establish how much fraud went unflagged, how long investigations took, or whether the workflow produced a net financial benefit.
Respect the historical information boundary
Backtesting assesses behavior on historical situations. Look-ahead bias occurs when later information enters a decision represented as historical. A chronological split helps, but only if source selection, preprocessing, and model selection respect each cutoff. Rolling-origin evaluation repeats this process as the decision date advances and scores the forecast horizon actually needed. See Validation without leakage.
| Distortion | What it changes | Evidence to retain |
|---|---|---|
| Later source revisions | Later values enter historical decisions. | Versions and availability cutoffs. |
| Later knowledge in the model | Historical source restrictions do not restrict everything learned during training. | Model lineage and tests for prohibited later information. |
| Survivorship bias | Today's survivors replace the population eligible historically. | Historical membership, identifier intervals, distributions, and exits. |
| Selection overfitting | The reported winner was chosen from many historical experiments. | Alternatives tried and independent assessment of the selected configuration. |
The model itself can cross the cutoff. Lookahead Bias in Pretrained Language Models found later pandemic information in risk descriptions generated from pre-pandemic earnings calls. Instructions to ignore later facts did not eliminate the demonstrated leakage. Restricting retrieved documents therefore supports only the retrieval boundary, not a complete historical-information guarantee.
Survivorship bias conditions inclusion on a later outcome. Removing securities that subsequently disappeared can omit investments available at the time, including failures and their losses; not every delisting is a failure, since mergers also cause exits. Separately, repeatedly selecting the strongest historical configuration can overfit the selection process. An attractive winning run does not disclose how many alternatives lost.
Keep representative later-period assessment separate from stress testing, which deliberately examines severe but plausible conditions. Stress scenarios reveal vulnerabilities; their frequency in a test set is not an estimate of how often they occur in ordinary work. Both are useful, but they answer different deployment questions.
Keep missing and delayed outcomes visible
A label is the specified outcome or judgment used for assessment. Selective labels arise when earlier decisions determine whose outcomes become observable. Repayment on an institution's loans is observed only after it grants those loans; rejected applicants do not become successful repayment examples. Evaluating solely on approved borrowers can misrepresent a replacement policy serving a different population. The selective-labels research formalized this problem in another decision domain; lending has the same observation structure.
This limitation predates modern AI. Durand's 1941 credit study explicitly noted that its records contained previously approved loans and omitted important information. More sophisticated modeling does not manufacture the missing outcomes. Account for incomplete feedback explains the general problem.
Selection and delay are separate. An outcome horizon is the follow-up period required by a claim. For default within twelve months, an account observed without default for only three months remains unresolved. It is right-censored: observation ended before the eventual event time became known. Treating it as a twelve-month negative silently changes the target.
Approval selects whose repayment is observed
ExampleHistorical rejection and insufficient follow-up create different gaps in outcome evidence.
Read the diagram as text
- Historical applications. The population encountered by the prior decision policy.
- Granted loans. Credit was extended and repayment can subsequently be observed.
- Rejected applications. No loan was extended by this institution.
- Loan outcome unobserved. Not a repayment success or failure label for the proposed loan.
- Outcome follow-up. Compare observed events and observation duration with the specified target.
- Target outcome assessable. The target event occurred, or sufficient event-free follow-up completed.
- Outcome still unresolved. No event observed, but follow-up ends before the target horizon.
- Historical applications → Granted loans: historical policy: grant.
- Historical applications → Rejected applications: historical policy: reject.
- Rejected applications → Loan outcome unobserved: no corresponding loan outcome.
- Granted loans → Outcome follow-up: observe subsequent records.
- Outcome follow-up → Target outcome assessable: event observed or horizon completed.
- Outcome follow-up → Outcome still unresolved: event-free observation ends early.
Payment outcomes require equally explicit definitions. Stripe's dispute lifecycle distinguishes payment time, later notification, evidence review, and changing case status. No observed dispute at a cutoff is not confirmed non-fraud. Its fraudulent dispute category records an allegation of unauthorized payment, which can include a legitimate charge the customer did not recognize. A reported fraud dispute, a won or lost case, and independently investigated fraud are different labels.
Investigation capacity also selects feedback. The Worldline-informed study assessed precision among the limited number of cards investigators could inspect, rather than treating multiple transactions on one card as independent investigations. Retain both the review-selection rule and outcome maturity. Better results on the selected queue do not establish performance on everything outside it.
Institutional responsibility and intervention
Assign ownership and independent challenge
Model risk concerns harm from an incorrect or misused model. Its intended use defines the decisions, population, and conditions for which it is accepted. Validation assesses suitability for that use; effective challenge requires expertise, objectivity, and enough influence to change the decision. Documentation without the power to require remediation is not equivalent to challenge.
| Responsibility | Role in deployment |
|---|---|
| Business management | Own day-to-day risks and operate controls. |
| Risk and compliance specialists | Support, challenge, and monitor the business's handling of risk. |
| Internal audit | Provide independent assurance on controls, with functional reporting to the board. |
Independent assurance assesses whether controls are appropriate and effective; it is not another execution stage inside the model's workflow. A payment team checking duplicates operates a control. Internal audit examining that control performs a different responsibility. Supplier involvement does not dissolve the institution's own decisions about use, monitoring, or response.
Scope each obligation before translating it into a release gate. The Federal Reserve's April 2026 SR 26-2 model-risk guidance is nonbinding, risk-based guidance primarily relevant to banks above its stated asset threshold, with exceptions. Its scope explicitly excludes generative and agentic AI. Applying its validation and monitoring practices to those systems is an engineering analogy, not a GenAI mandate; historical SR 11-7 should not be presented as the current guidance.
Other sources have different authority. FINRA Notice 24-09 reminds member firms that existing technology-neutral obligations apply to generative AI, including third-party tools; it creates no new requirements. Covered Regulation B notices must give actual principal decision reasons. Data licenses create contractual restrictions, while an institution can impose stricter internal approval policies. Record which category supports each requirement, its applicability, owner, and enforcement evidence. See Verify and revisit decisions.
Make review capable of changing the outcome
Human participation matters only if the person can make the needed intervention. Factual review checks evidence and calculations. Policy review determines whether the proposed treatment fits applicable rules. Authorization grants permission for an effect. Takeover assigns someone responsibility for unresolved work. A reviewer may be qualified for one role but not another.
Provide a review packet suited to that decision: applicable source versions, calculation inputs and definition, unresolved discrepancies, proposed effects, and approval limits. For bond reconciliation, distinguish the expected entitlement, external allocation evidence, and internal posting. A summary should help the reviewer reach those records, not replace them. Review before consequential commitment develops the interface choices.
Automation bias is inappropriate reliance on automated output. In a Duolingo false-alert experiment, skilled reviewers accepted half of fabricated cheating alerts inserted into legitimate historical sessions. Customers were not affected by the experiment. This is cross-domain evidence that approval is not a correctness measure, not an estimate of financial-review error. Test local reviewers with independently assessed correct and erroneous proposals, measuring whether they detect and correct consequential errors before commitment.
Review capacity constrains the permitted workload. Track required handling effort, qualified staffing, aging, and the time remaining before financial or customer deadlines. When demand exceeds capacity, invoking human review does not supply an answer: admission, prioritization, escalation, or the operating mode must change. Urgent work still needs an accountable owner. Operate the exception workload explains queue management.
Correction also needs a route after the initial decision. NIST's voluntary AI Risk Management Framework supports feedback, appeal, and override mechanisms. A practical design links a complaint to the decision and its versions, assigns a reviewer empowered to change the disposition, records the rationale, and tracks downstream correction and communication. A trace explains recorded events; it does not itself give an affected person a remedy.
Deployment and continued control
Choose a bounded operating mode
Deployment is a choice about permitted effects, not a race toward autonomy. Integration tests establish interface behavior. Historical assessment examines past cases under a defined information boundary. Shadow operation processes current inputs without applying candidate effects. Assisted use introduces the candidate into human work. Limited live action exposes a bounded set of real operations. Each answers a different remaining question.
| Mode | Evidence it can add | What remains unestablished |
|---|---|---|
| Historical assessment | Behavior on reconstructed past information and cases. | Performance under today's work and consequences of applying outputs. |
| Shadow | Current-input compatibility, proposed outputs, and resource demand. | Effects on customers, reviewers, and financial outcomes when proposals are used. |
| Assisted use | Reviewer handling, corrected results, and actual use of suggestions. | Safety or usefulness without that human intervention. |
| Bounded live action | Authorized effects and recovery within the exposed scope. | Suitability for broader populations, higher limits, or new activities. |
Shadow execution must technically deny or isolate writes, notifications, and financial actions; discarding the final answer does not undo tool effects. The live-experiment chapter explains this boundary. Define eligible work, allowed sources and recipients, permitted effects, aggregate exposure, intervention capacity, acceptance conditions, and stopping rules before exposure. Read-only assistance can remain the appropriate final mode.
Compare the complete process with its incumbent: source licensing, integration, validation, model calls, review, correction, support, and waiting. Faster case preparation can be valuable while failing to reduce total handling effort. Price accepted outcomes develops that accounting. The historical Authorizer's Assistant report treated integration, acceptance testing, and transfer of business control to operations as deployment work—not as consequences automatically supplied by a working model.
Reconstruct, monitor, and reassess
An audit trail is attributable evidence connecting information, decisions, approvals, and observed effects. For financial AI, preserve the relevant source and instrument versions, as-of conditions, model and policy versions, checked calculations, reviewers, authority, execution attempts, and resulting external identifiers. Generated reasoning is not a substitute for records of what was read, approved, and done. Record meaning-changing boundaries explains the diagnostic design.
The SEC's 2022 electronic-recordkeeping amendments provide a scoped example: an audit-trail alternative preserves modifications, deletions, timestamps, applicable identities, and information needed to recreate original and intermediate records. It is not a universal requirement to retain every prompt. The NIST Privacy Framework also calls for minimizing audit data. Preserve necessary evidence with controlled access and defined lifetimes; do not indiscriminately copy customer payloads into logs. See Set artifact lifetimes.
| Observed change | Reassessment |
|---|---|
| Stale sources, corrections, or changed rights | Whether inputs remain applicable and permitted, and which outputs need review. |
| Growing unresolved amounts or older cases | Whether reconciliation and intervention capacity still meet the operating contract. |
| Complaints, overrides, or changing outcomes | Whether service and decision quality remain acceptable in the affected population. |
| New provider, product, policy, or authority | Whether previous tests and approvals cover the changed system. |
| Failed limits or shutdown activation | What further work stopped, what already executed, and who must reconcile it. |
Containment and recovery are separate. A runtime kill switch can stop future decisions or tool calls when active work next checks it; a flag read only at session creation misses ongoing work. Stopping does not reverse payments. Identify completed and uncertain effects, reconcile authoritative records, authorize corrections, test recovery, and obtain approval to restore the affected scope. Record who changed controls and whether mitigation actually worked.
Finally, examine dependencies beyond one application. The FSB's 2024 assessment identifies provider concentration, correlated behavior, cyber risk, and model risk as possible amplification channels. Local usefulness does not remove a shared failure dependency. Continued permission to operate rests on maintained evidence, workable intervention, and tested recovery—not the age of the launch approval.
Open questions
End-to-end financial productivity remains difficult to establish because faster preparation can shift work into verification, corrections, and support. Useful progress would compare complete analyst or reconciliation tasks against the incumbent process, preserving quality and unresolved-work measures alongside total effort.
Effective financial oversight needs evidence that qualified reviewers detect consequential errors under realistic workload and deadline pressure. Approval rates cannot supply that evidence. Progress would demonstrate correct intervention before commitment, including overload and unavailable-reviewer conditions.
Historical language-model evaluation faces an information boundary that document filtering cannot fully enforce: later facts may already be encoded in the model. Progress would establish auditable historical-knowledge constraints or evaluation designs that do not depend on pretending those facts are absent.
Outcome learning must handle decisions that hide alternatives and labels that mature late. Progress would make observation rules, unresolved cases, and deployment-population coverage explicit, rather than treating increasingly large operational datasets as automatically representative.
Local controls do not resolve shared infrastructure risk. Institutions need ways to test substitution and recovery without silently changing data permissions or financial behavior. Progress would demonstrate continuity across actual dependencies, not merely configure an alternate provider name.










































































