Purpose and choices
Leadership owns the commitment
AI engineering leadership takes responsibility for turning systems that use learned model behavior into useful, supportable organizational capabilities. A technical capability performs an operation, such as drafting an explanation. An application connects that operation to records, interfaces, and users. An organizational capability is the repeatable ability to deliver the result, supported by people, knowledge, tools, access, and time. Integration, assessment, support, and eventual retirement remain work after the demonstration succeeds.
Leaders connect decisions that otherwise drift apart: the outcome worth pursuing, the resources committed, the people able to deliver, the authority to permit use, and the conditions for continuation. Automation can perform more implementation and checking without taking over accountability for what reaches users. Addy Osmani's account of engineering ownership makes this distinction explicit: generated work supplies evidence; accountable owners decide whether it is sufficient and defend the resulting commitment.
Tool access is only one such commitment. In its February 2026 presentation, Thomson Reuters described supplementing broad experimentation and learning opportunities with leadership oversight, three-year roadmaps, dedicated change expertise, and process mapping. These were reported organizational arrangements, not proof of their financial effect. Enterprise AI explains how adoption becomes workable; this chapter addresses which efforts receive resources and who answers for their consequences.
Foundations of investment and learning
The organizational problems surrounding AI did not begin with generative models. Economic research examined where coordination should sit; organizational research examined why experience does not always change decisions. Machine-learning engineering later made the maintenance obligations of learned systems explicit. These contributions address different problems and remain useful together.
| Contribution | Date | Problem clarified |
|---|---|---|
| The Nature of the Firm — R. H. Coase | November 1937 | Market contracting and internal coordination both cost resources; neither organization nor purchasing is universally preferable. |
| Organizational Learning and Management Information Systems — Chris Argyris | 1977 | Better reporting can correct deviations while leaving governing assumptions and defensive management relationships unchallenged. |
| Exploration and Exploitation in Organizational Learning — James G. March | February 1991 | Discovering possibilities and refining existing competence compete for scarce resources, with different timing and uncertainty of returns. |
| Hidden Technical Debt in Machine Learning Systems — D. Sculley and colleagues at Google | 2015 | Useful models bring system-wide maintenance obligations through data dependencies, feedback, configuration, and integration—not just model code. |
Together, these ideas explain why cheaper implementation does not settle investment choices. An organization still has to choose a useful objective, coordinate delivery, learn whether the objective remains sensible, and maintain what others depend on.
Choose the outcome and alternatives
Strategy connects a chosen outcome to activities and resource commitments, including explicit exclusions. Complementary activities reinforce one another: an improved service may depend on specialist knowledge, reliable delivery, and customer support together. The Harvard Institute for Strategy and Competitiveness places differentiation in this activity system. Acquiring a model or dataset alone does not establish an advantage. Some investments instead provide capability parity: work the organization needs to remain a credible option, without making its offer distinctive.
An investment thesis states why a particular intervention should improve an important outcome and what would overturn that expectation. Record the present baseline, intended users, scope and exclusions, required organizational changes, and disconfirming observations. Compare alternatives before selecting the technology. Opportunity cost is the value forgone by choosing one feasible use of resources over another; already-employed specialists still have alternative work.
| Alternative | Proposed mechanism | Commitment to examine |
|---|---|---|
| Process change | Remove an unnecessary approval | Authority to change the procedure |
| Conventional software | Route requests using explicit rules | Rule maintenance and exception handling |
| Purchased functionality | Use an existing supported workflow | Supplier fit and integration duties |
| AI assistance | Interpret varied requests and propose responses | Assessment, review, and correction capacity |
| Maintain the current arrangement | Avoid transition work | Continuing delay and displaced opportunities |
The comparison keeps the outcome fixed while changing how it might be achieved. For a customer-facing offer, customer validation tests whether people make the commitments needed to use or purchase it; enthusiasm for a demonstration is weaker than supplying real work or releasing a budget. AI Startups and Small Teams develops that distinction. Leadership uses the findings to choose scope and investment, rather than treating technical feasibility as sufficient demand.
Investment economics and commitment
Turn improvement into benefit
Benefits realization converts improvement into an organizational outcome. Released time means less effort on an activity; usable capacity means that time can perform other work; realized savings mean expenditure actually falls. A benefit owner must arrange the conversion with finance and operational teams, preserve required outcomes, and avoid transferring costs or counting the same benefit twice. Enterprise AI explains these distinctions in everyday operation.
The first obligation is to establish whether the task improved. In METR's early-2025 developer study, Joel Becker, Nate Rush, Beth Barnes, and David Rein randomized AI access across 246 tasks performed by 16 experienced open-source developers in familiar repositories during February–June 2025. AI access increased completion time by 19%, although participants afterward estimated roughly 20% acceleration. This is a bounded result for those workers, tasks, and tools—not a general estimate of AI productivity.
A later result cannot simply replace that estimate if the comparison changes. METR's February 24, 2026 update reported that participation, task selection, compensation changes, and simultaneous agent use complicated its follow-up. It did not provide a reliable current effect estimate. The leadership lesson is to distinguish perceived usefulness from measured improvement while keeping the comparison's conditions visible.
Even a demonstrated gain needs a conversion plan. Factory's outcome-chain recommendation connects engineering changes to fewer defects, customer satisfaction, and business results. Each connection is a hypothesis to measure. Assign an owner who can change staffing assignments or spending, fund the necessary transition, and specify when the downstream result should become observable. Faster drafting alone does not establish any of those later outcomes.
Improvement needs a conversion decision
ExampleReleased time can support different benefits, but neither follows automatically.
Read the diagram as text
- Measured task improvement.
- Account for remaining work. Review, correction, and transition.
- Released time.
- Additional useful output.
- Reduced expenditure.
- Measured task improvement → Account for remaining work: Check the complete work.
- Account for remaining work → Released time: Net effort falls; quality holds.
- Released time → Additional useful output: Usable capacity + demand + funded redeployment.
- Released time → Reduced expenditure: Avoidable spend + budget authority.
Compare complete costs and timing
Return on investment, or ROI, expresses net monetary benefit relative to cost. Use an explicit convention: let be monetary benefits attributable to the program over a declared period, and its included program costs. The ROI Institute methodology subtracts costs before dividing; alone is a benefit-cost ratio.
For an arithmetic example, benefits of $150,000 against $100,000 of included costs produce $50,000 net benefit and 50% ROI, not 150%. Whether those benefits are attributable, collectible, or merely valued capacity is a separate claim. For increased sales, count the monetary contribution left after the costs of supplying those sales, not their full revenue. Declare the benefit period rather than combining several years of benefits with an unexplained cost window.
Total cost of ownership includes acquisition and continuing operation. Bring integration, specialist time, assessment, change work, maintenance, support, and replacement into the stated boundary. Cost and Performance Engineering explains how to price the required service. Keep a dated ledger of initial outlays, recurring work, contractual commitments, transition costs, and expected benefits. An internal allocation identifies who carries a cost; an avoidable cost is expenditure changed by the decision. These are different quantities.
Payback is the point when cumulative net cash benefits recover the investment; it may never occur. A positive aggregate ROI does not establish affordable cash timing. Sensitivity analysis varies consequential assumptions, such as uptake, review effort, or benefit delay; a switching value identifies a change large enough to reverse the choice. Keep nonmonetary requirements visible instead of forcing every consequence into the ratio.
The Productivity J-Curve, Erik Brynjolfsson, Daniel Rock, and Chad Syverson's October 2018 working paper, explains why new general-purpose technologies require complementary investments in processes and human capital. When these investments are poorly measured, productivity can initially appear understated and later overstated as benefits arrive. This explains a measurement problem, not a guaranteed cash-return trajectory. Necessary preparation deserves funding; disappointing results still require a prospective case for further expenditure.
Choose a feasible portfolio
A portfolio manages investments together because they share objectives, resources, dependencies, or exposure. APM's portfolio guidance emphasizes strategic fit and delivery capacity. Separate project budgets do not create separate domain experts, integration teams, or support staff. Select combinations and sequences that can actually be delivered, including existing operating duties, rather than simply funding proposals in descending ROI order.
Consider a small planning example. After existing service duties, a domain-review team has two days a week available. A response-assistance pilot requires both days; a classification pilot also requires both. Both depend on the same model service and need a shared evaluation service to check their behavior. A third proposal builds that evaluation service using separately available platform capacity. Assume funding is sufficient and these estimates hold. Leaders can fund the shared service and either pilot, but must have the evaluation capability ready when the pilot needs it. They cannot run both pilots together without exceeding review capacity. Choosing one delays learning about the other; funding alone does not remove the conflict.
Funding, review capacity, and readiness are separate
Funding is sufficient. Two domain-review days per week remain after existing duties; each pilot needs both days. Platform capacity is separately available.
Evaluation gate: not ready.
Review capacity fits. The selected pilot remains conditional on evaluation capability becoming ready.
Evaluation work is selected and uses separate platform capacity. Both pilots retain the same model-service dependency.
Response pilotrequires →2 review days/week · evaluation capability · shared model service
Classification pilotrequires →2 review days/week · evaluation capability · the same model service
Evaluation proposalrequires →Separately available platform capacity
Arrows show requirements, not execution order. Either pilot uses 2/2 review days; both use 4/2 even when evaluation is ready. These checks do not establish model-service availability or operating authorization.
Concentration risk arises when several commitments depend on the same vulnerable resource or supplier. The Bank of England's April 2025 analysis explains how externally supplied models, cloud infrastructure, and common components can create shared operational exposure. Internally built applications and different product brands do not establish independent dependencies. Investigate the actual upstream resources and whether substitution is practical before treating diversity as protection.
Required service or control obligations also consume portfolio capacity. They need an economical delivery choice, not invented revenue to compete with growth proposals. Compare discretionary additions only after accounting for those duties. Shared capabilities may enable several applications, but their expense belongs somewhere, and the same released staff time cannot be claimed independently by every project that helped release it.
Fund the next uncertainty
Funding learning differs from funding delivery. An evaluation systematically assesses behavior against intended use; Evals and Benchmarks explains what it can establish. A working prototype may resolve technical feasibility while leaving usefulness, economics, and permission to operate unsettled. A bounded experiment is worthwhile when its possible findings could change the next commitment. Repeating an already-settled demonstration buys little decision-relevant information.
March called investigating new possibilities exploration, and refining established competence exploitation. His organizational models explain a persistent tension: dependable near-term improvements can crowd out uncertain learning, while perpetual experimentation can fail to develop usable competence. Neither implies a fixed budget percentage. Allocate learning resources according to the uncertainties that constrain important decisions, while preserving the capacity to operate what already works.
A real option preserves the ability, without the obligation, to invest later. In The Value of Waiting to Invest, Robert McDonald and Daniel Siegel's November 1982 working paper examined irreversible investment under uncertainty. Committing now sacrifices the alternative of acting after more information arrives, so positive expected net benefits need not justify immediate commitment. Waiting can nevertheless lose opportunities, and maintaining reversibility can cost money. The theory supplies a comparison, not a reason to delay everything.
Make the next funding agreement specific: identify the uncertainty, bounded money and staff time, observations that distinguish alternatives, decision owner, and review deadline. State what each result would justify and which duties survive stopping. The enterprise portfolio-of-experiments argument is useful when scope and value emerge through work, but it does not establish venture-capital return distributions or excuse unlimited experiments. A negative finding can make a test successful by preventing a larger mistaken commitment.
Sourcing and dependence
Source capabilities and duties
Build versus buy allocates continuing work between the organization and suppliers. A partnership adds coordinated development or delivery. As AI Startups and Small Teams explains, buying transfers specified duties, not the complete product obligation. At portfolio scale, separate application ownership, model supply, operation, and domain expertise before comparing arrangements. A delivery-model assessment can apply to individual components rather than forcing one sourcing choice onto the whole service.
Coase's coordination argument explains the tradeoff. Purchasing incurs the work of specifying, negotiating, and adapting transactions. Internal organization can avoid some of that work, but managing more activities creates its own costs and mistakes. Building is attractive only when the internal arrangement is preferable to available alternatives—not simply because it offers nominal ownership.
Transaction-Cost Economics, Oliver Williamson's October 1979 article, develops how uncertainty, recurrence, and relationship-specific investments affect governance. Asset specificity means an investment loses substantial value in alternative uses. Specialized training and accumulated supplier knowledge can make replacement difficult even when the initial market was competitive. Standardized inputs can be easier to substitute. The implication is to compare adaptation and retained competence alongside price; elaborate partnership arrangements also impose costs when simple purchasing would suffice.
| Capability or duty | Possible arrangement | Work retained internally |
|---|---|---|
| Model execution | Purchase a hosted service | Specify required behavior and assess changes |
| Domain interpretation | Develop with a specialist partner | Resolve requirements and retain usable knowledge |
| Application integration | Build internally | Maintain interfaces and coordinate source owners |
| Service operation | Split duties by an explicit agreement | Own the end-to-end outcome and unresolved handoffs |
Partnership agreements should make joint work governable: milestones, funding, responsibilities, work-product rights, access, reporting, knowledge transfer, disagreement, and termination. UKRI's collaboration guidance provides a concrete research-partnership example of these continuing obligations. The appropriate terms depend on the relationship. Agreement on a deliverable does not demonstrate that either partner can operate or replace it.
Preserve a practical exit
Vendor lock-in is the difficulty or expense of changing suppliers. It includes integration changes, behavior revalidation, information movement, transition work, and lost expertise. Intuit reported that even moving between models from the same supplier required substantial evaluation. Prompts—instructions and other input supplied when invoking a model—can encode expectations that do not transfer smoothly. Prompt transfer boundaries explain the mechanism. Compatible request syntax reduces one kind of switching work; it does not establish equivalent application behavior.
When Your LLM Reaches End-of-Life, an April 29, 2026 Verint preprint by Emma Casey, David Roberts, David Sim, and Ian Beaver, illustrates application-level replacement. Candidates first faced internal vetting, licensing, cost, and regional-availability constraints. Comparisons then examined correctness, refusals, structure, style, and response time against the incumbent. Nova 2 Lite and Qwen3-32B were suitable candidates under those tests; covering all required regions and modalities could require multiple models. The English-only main analysis used two test sets with limited human calibration. Candidate selection did not establish a completed worldwide migration, full transition expenditure, or realized savings.
| Obligation | What must be available |
|---|---|
| Change notification | Enough notice to assess and prepare |
| Artifacts and access | Usable records, interfaces, and necessary information |
| Knowledge transfer | People able to maintain the replacement |
| Transition support | Funded assistance and clear handoffs |
| Acceptance | Receiving owners and explicit service conditions |
An exit clause establishes an obligation, not demonstrated transition capability. Preserve time and expertise to exercise it. Review permissions for the actual service path, including the selected feature and deployment arrangement, rather than assuming a supplier-level approval covers the replacement. Also investigate shared upstream dependencies: two contracts may still rely on one critical resource.
Capabilities and organization
Close the actual capability gap
Translate the chosen portfolio into work the organization must repeatedly perform: domain judgment, application engineering, data stewardship, assessment, operation, and supplier management. One person may cover several responsibilities; some need specialist support. The UK Government AI Playbook describes capability needs across the service lifecycle and allows hiring, contractors, and internal development. Choose among them by time to competence, retained knowledge, and continuing coverage—not job-title completeness.
Diagnose the missing condition before purchasing a remedy. Someone may know how to review outputs but have no available time. Another person may have capacity but lack access to the source records. A named owner may understand the problem yet lack authority to change it. These require workload allocation, access decisions, or a mandate—not the same training course.
Intuit's tax-explanation work illustrates complementary capabilities. Tax analysts contributed domain knowledge, prompts, and initial judgments, while data-science and machine-learning staff concentrated on quality measures and test datasets. The division connects expertise to reusable assessment infrastructure. It does not establish an optimal staffing ratio or specify who could authorize release.
Absorptive capacity is the ability to understand and apply incoming knowledge; Enterprise AI develops its role in transferring practices. Fund protected learning time and practical demonstrations, not attendance alone. Require receiving teams to perform the work with documentation and support, and provide backups so expertise survives absence or departure. APM's transition guidance treats skills, operational acceptance, and continuing support as delivery work rather than a final administrative handoff.
Place decisions near the work
An operating model arranges the work, people, information, suppliers, management processes, and decision rights needed to deliver a service. The Operating Model Canvas makes these arrangements explicit. Reporting lines alone are insufficient: identify who defines domain requirements, accepts application behavior, operates shared services, and decides exceptions. Then examine the handoffs those choices create.
| Arrangement | Placement | Coordination consequence |
|---|---|---|
| Centralized | A common team delivers applications and shared capabilities | Pool specialists; maintain access to domain owners and manage the central queue |
| Embedded | Business units deliver their applications | Keep domain work close; address duplicated expertise and infrastructure |
| Federated | Local delivery combines with explicit central services and constraints | Preserve local judgment; define shared-service duties and exception authority |
A center of excellence concentrates expertise, but the name does not determine its mandate. It might teach teams, deliver difficult work, maintain shared capabilities, or approve defined exceptions. Those roles need different capacity and authority. If every ordinary change requires central approval, that queue becomes part of delivery. If application teams can bypass shared responsibilities without agreement, federation has not resolved ownership.
Scale the arrangement to the work. Oleve's small-team example separates product-outcome ownership from cross-product automation without implying a large departmental hierarchy. Shared code, templates, and operated services carry different service responsibilities. Team Topologies provides a vocabulary for team responsibilities and interactions; its home discussion explains collaboration, service consumption, and enabling work. Use these boundaries to decide where resources belong, not to copy an organization chart.
Shared funding
Fund the continuing service
A project budget funds a bounded change; a service continues creating obligations. Sculley and colleagues located maintenance liabilities in data dependencies, configuration, integration, and hidden consumers as well as model code. Taking on such debt can be strategically reasonable, but it leaves work to finance. Assign continuing review, support, maintenance, and replacement funding before declaring initial delivery complete.
Shared costs may remain centrally funded or be allocated using fixed shares, consumption, or another stated proxy. FinOps allocation guidance recommends an explicit policy for each category and allows mixed approaches. The policy should expose who benefits and who can influence demand; changing the allocation does not itself change the underlying expenditure.
| Arrangement | Accountability effect | Tradeoff |
|---|---|---|
| Central funding | A shared owner carries expenditure | Demand needs explicit review |
| Showback | Consumers see attributed costs | Visibility does not transfer budget responsibility |
| Chargeback | Consumers carry assigned expenses | Adds accounting work; allocation rules influence choices |
A forecast is an agreed expectation of future spending and value based on workload, pricing, timing, and lifecycle assumptions. FinOps forecasting brings engineering, product, finance, and leadership into that model; budget owners must meet it or seek additional funding. Review consumption alongside usefulness. A forecast is not proof of realized benefit, and an unused service does not become worthwhile because every team has been assigned a share.
Validated commonality means that different users genuinely share a requirement. It is a stronger basis for common investment than repeated requests for similarly named features. Forward Deployed Engineering explains how to establish it. Fund the shared capability together with its support boundary, and keep exceptional local work explicitly owned rather than hiding it in the platform budget.
Authority and assurance
Give ownership real authority
Responsibility identifies who performs work; accountability identifies who answers for an outcome; decision rights specify who may approve, change, or stop something. Enterprise AI develops the distinction. An ownership chart cannot grant missing authority. Separate the application outcome owner, technical operator, risk owner, and people who challenge claims, even when a small organization assigns several duties to one person.
A risk owner ensures that a specified risk is managed and monitored, with authority to act or escalate; the owner need not perform every mitigating action. Residual risk is exposure remaining after controls. Risk appetite describes the types and amounts of exposure an organization is willing to accept in pursuing objectives. These concepts require explicit limits, not a general declaration that leadership accepts risk. The historical Orange Book supplies the ownership and residual-risk vocabulary.
| Decision | Mandate to specify |
|---|---|
| Commit resources | Budget scope, available capacity, and escalation limit |
| Interrupt operation | Who can stop which uses and obtain necessary access |
| Require remediation | Who assigns work and funds the remedy |
| Accept remaining exposure | Authorized scope and obligations outside that authority |
| Resume or expand | Required findings, decision maker, and absence cover |
DWP's 2025 account of its arrangements illustrates written delegation and escalation when mitigation fails, exposure exceeds limits, or the current owner lacks control. The transferable principle is an effective route to greater authority, not its particular committee structure. NIST's AI Risk Management Framework similarly assigns continuing executive risk decisions and authority to disengage systems. Funding approval does not grant permission for every resulting use.
Automation need not eliminate ownership. In Intercom's April 2026 review account, an automated-approval pilot excluded changes deemed too broad or large and allowed human review requests. Engineers retained responsibility for observing their changes in production and rolling them back. That is a concrete allocation of duties, not general proof that automated approval improves safety or ROI.
Commission the right assurance
Assurance supports justified confidence through examination suited to a claim; it is not a guarantee. The UK introduction to AI assurance emphasizes measuring, evaluating, and communicating trustworthiness, including limitations and mitigations. Begin with the disputed claim and commission the missing examination. More evidence of task accuracy does not resolve an unexamined permission, operating, or benefit claim.
Independent challenge requires more than a second checking pass. The IIA Three Lines Model distinguishes delivery and risk-management roles from independent internal audit. Its September 2024 revision of the 2020 model connects independence to governing-body accountability, resources, information access, and freedom from interference. Specialist challenge can remain part of management. The lines describe concurrent responsibilities, not three mandatory departments or serial approval gates. An evaluation team whose findings can be suppressed by the delivery sponsor is not independent merely because it has a different name.
Independent challenge needs protected relationships
ExampleInternal audit has direct governing-body accountability and access to management information; specialist challenge within management remains a different role.
Management challenge and independent audit
Delivery and specialist risk challenge sit inside management. Independent internal audit sits outside that boundary. The governing body supplies audit resources and mandate; audit returns findings and accountability directly and has access to relevant management people and information. Positions carry no quantitative meaning.
- 1. Management boundary
- 2. Governing body supplies resources and mandate
- 3. Audit reports findings and accountability directly
- 4. Audit accesses management people and information
Read coordinates and regions as data
X: 0–16 conceptual units; Y: 0–10 conceptual units, increasing up. Equal scale on both axes.
(0.5, 1); (7, 1); (7, 6); (0.5, 6)
(10, 8); (10, 4.5)
(13, 4.5); (13, 8)
(10, 3.2); (6.5, 2.2)
Management: (3.75, 5.4)
Delivery / operations: (3.75, 4.4)
Risk support / challenge: (3.75, 3.4)
People and information: (3.75, 2)
Governing body: (11.5, 8.7)
Independent internal audit: (11.5, 3.9)
Resources and mandate: (9.7, 6.6)
Findings and: (13.3, 6.7)
accountability: (13.3, 6)
Access to people and information: (8.4, 1.6)
| Claim | Relevant examination | What remains separate |
|---|---|---|
| The task is performed acceptably | Representative cases and competent judgments | Benefit under everyday use |
| The organization benefits | Baseline comparison including remaining human work | Permission for wider exposure |
| Operations can receive the service | Acceptance criteria and readiness with operational users | Every future operating condition |
| A supplier owes a control | Applicable agreement or attestation | Actual configuration and enforcement |
Separate prompts or models for implementation and testing can add scrutiny, but they do not establish independent errors or organizational independence. Similarly, an AI-generated safety argument identifies claims to inspect rather than proving its own completeness. Leaders must provide reviewers with competence, time, access, and a route to the authority that can resolve disagreement.
The disposition can be conditional acceptance, restricted use, further examination, or withheld authorization. Record the supported scope, remaining uncertainty, responsible decision maker, and changes that trigger reconsideration. NIST's framework treats residual-risk decisions and post-deployment monitoring as continuing responsibilities. A passing test and a funded project do not combine into unrestricted permission to operate.
Organizational learning
Turn field findings into investment
Forward deployed engineering brings engineers close to customer work to learn the domain and implement useful solutions; its home chapter explains those responsibilities. Leadership decides what the findings justify beyond the engagement. Productization makes a capability repeatable and supported. Repeated requests are a reason to investigate shared requirements, not proof that one implementation or operating arrangement fits them all.
Kevin Bai's platform-versus-customer distinction keeps unique behavior local while using deployment discoveries to identify generalizable capabilities. The leadership choices are broader than accepting or rejecting a feature: fund local adaptation, create a supported shared capability, offer explicitly staffed service work, or decline the request. Each choice needs an owner and a maintenance boundary.
| Requirement | Potentially shared | Still locally owned |
|---|---|---|
| Read the same case system | A supported integration | Access approval for each use |
| Produce a summary | Common interface and reusable implementation | Different domain definitions of an acceptable summary |
| Keep the service useful | Shared technical support | Distinct review, escalation, and staffing duties |
The common integration can deserve investment without centralizing every policy decision. CNCF's platform guidance prioritizes common needs while allowing capabilities outside the platform. Before expanding, preserve the conditions behind the original result: users, domain rules, data access, review effort, support, and beneficiaries. Establish which recur in the receiving setting. A success produced by unusually intensive specialist support does not by itself justify expansion without that support.
Change assumptions and incentives
Revise the governing choice
Single-loop learning corrects action while preserving governing goals and rules. Double-loop learning reconsiders those goals or rules and then changes action. Argyris's 2002 explanation makes the distinction explicit. Reducing review effort to meet an existing savings target is single-loop improvement. Reconsidering whether that target remains justified after discovering necessary review work changes the governing assumption.
More reporting does not automatically produce that reconsideration. In Argyris's 1977 management-information analysis, tighter financial information helped address some deviations while defensive relationships concealed harder organizational problems. A dashboard can make a target's shortfall visible without making it acceptable to challenge the target. Leadership must permit the investment thesis itself—not just execution quality—to be examined.
Escalation of commitment means increasing resources devoted to an existing course despite adverse feedback. In Knee-Deep in the Big Muddy, Barry Staw's 1976 experiment assigned 240 business students hypothetical investment decisions. Commitment was greatest when participants owned the earlier allocation and received negative results. This bounded experiment illustrates a danger of defending one's prior choice, not a universal explanation of executive behavior. Continued investment can be rational when future benefits justify it; prior personal responsibility is not itself such a benefit.
Make truthful reporting workable
Psychological safety is a shared belief that interpersonal risk-taking is safe. Amy Edmondson's 1999 study of 51 manufacturing teams associated it with learning behavior and considered contextual support and leader coaching alongside interpersonal beliefs. The observational study does not establish an AI-specific causal effect. Its relevance is practical: admitting uncertainty and seeking help need conditions beyond a reporting channel. Safety to challenge is compatible with demanding standards for the work.
Incentives are consequences that influence behavior, including recognition, workload, and professional standing. Rewarding licenses used, demonstrations delivered, or output generated can favor activity without establishing benefit. KCS guidance recommends outcome goals while using activity counts diagnostically. Recognize accurate reporting and justified stopping; Enterprise AI explains implementation with affected teams.
Workforce commitments need the same honesty. State changed duties, learning resources, support coverage, and employment expectations instead of assuming saved task time warrants fewer staff. The OECD's workplace survey report found consultation associated with more favorable worker reports in finance and manufacturing, but not a causal productivity estimate. Skills were discussed more often than job loss or wages. Leadership should resource and address the whole transition; change-management practice belongs with the teams doing it.
Renewal and retirement
Reallocate without abandoning duties
Review commitments when material assumptions change, before renewal deadlines, and often enough to address consequential uncertainty. Compare observed benefits with expectations, then examine remaining expenditure, operating duties, scarce capabilities, and feasible alternatives. The available dispositions are different decisions: continue within limits, expand, redirect, narrow, pause, or retire. A project status does not specify the resources or authority each requires.
Sunk costs are past expenditure that cannot be recovered. They should not determine the next choice. Future exit costs, unavoidable commitments, reusable assets, and scarce staff time still matter. Cancelling a project's allocated share does not save money if total expenditure remains unchanged.
Keep a short decision record that makes reallocation explicit.
- Changed basis — Observations that changed the thesis, including unresolved benefits.
- Remaining commitment — Prospective spending, service duties, and constrained capacity.
- Alternatives — Feasible uses of released resources and what each displaces.
- Disposition — Decision, reasons, authorized scope, and accountable owner.
- Reconsideration — Next deadline or observation that requires another decision.
Retirement is funded transition work. Government Digital Service guidance, first published in 2016, requires attention to continuing user needs, knowledge transfer, communication, API consumers, and protection of retained information. Assign those duties and finance them even when new development stops. Enterprise AI develops the operating practice. Ending a roadmap does not end obligations to people already relying on the service.
A narrower service or staged replacement can preserve demonstrated value while releasing capacity from unsupported expansion. That is neither automatic retreat nor a defense of the original plan. It is the central leadership discipline: commit resources to work that remains worthwhile, make the authority match the consequences, and preserve responsibility for what the organization has already undertaken.
Open questions
Measuring productivity becomes harder when tools change who participates, which tasks are attempted, and how concurrent work is timed. Progress requires credible comparisons that represent changed working practices without mistaking preference for measured benefit.
Practical supplier independence remains difficult to value. Behavioral compatibility, retained expertise, transition time, and common upstream resources all matter. More informative replacement studies would report completed transitions and their full costs, not only candidate performance.
Organizations still need workable ways to scale scrutiny without making approval the limiting resource or weakening challenge. Progress would show that reviewers retain access and influence while consequential decisions remain timely and accountable.























