Contents
  1. Purpose and choices
    1. Leadership owns the commitment
    2. Foundations of investment and learning
    3. Choose the outcome and alternatives
  2. Investment economics and commitment
    1. Turn improvement into benefit
    2. Compare complete costs and timing
    3. Choose a feasible portfolio
    4. Fund the next uncertainty
  3. Sourcing and dependence
    1. Source capabilities and duties
    2. Preserve a practical exit
  4. Capabilities and organization
    1. Close the actual capability gap
    2. Place decisions near the work
  5. Shared funding
    1. Fund the continuing service
  6. Authority and assurance
    1. Give ownership real authority
    2. Commission the right assurance
  7. Organizational learning
    1. Turn field findings into investment
    2. Change assumptions and incentives
      1. Revise the governing choice
      2. Make truthful reporting workable
  8. Renewal and retirement
    1. Reallocate without abandoning duties
  9. Check understanding
  10. Open questions
  11. Selected talks
  12. References
  13. Talk library
← All topics

AI Engineering Leadership

AI engineering leadership decides which technical possibilities deserve organizational commitment. A system can work without being worth operating, and a worthwhile application can fail because the people responsible lack time, knowledge, or authority. The leadership task is to connect intended outcomes to resources and responsibilities—and to change those commitments when experience changes the case.

Purpose and choices

Leadership owns the commitment

AI engineering leadership takes responsibility for turning systems that use learned model behavior into useful, supportable organizational capabilities. A technical capability performs an operation, such as drafting an explanation. An application connects that operation to records, interfaces, and users. An organizational capability is the repeatable ability to deliver the result, supported by people, knowledge, tools, access, and time. Integration, assessment, support, and eventual retirement remain work after the demonstration succeeds.

Leaders connect decisions that otherwise drift apart: the outcome worth pursuing, the resources committed, the people able to deliver, the authority to permit use, and the conditions for continuation. Automation can perform more implementation and checking without taking over accountability for what reaches users. Addy Osmani's account of engineering ownership makes this distinction explicit: generated work supplies evidence; accountable owners decide whether it is sufficient and defend the resulting commitment.

Tool access is only one such commitment. In its February 2026 presentation, Thomson Reuters described supplementing broad experimentation and learning opportunities with leadership oversight, three-year roadmaps, dedicated change expertise, and process mapping. These were reported organizational arrangements, not proof of their financial effect. Enterprise AI explains how adoption becomes workable; this chapter addresses which efforts receive resources and who answers for their consequences.

Foundations of investment and learning

The organizational problems surrounding AI did not begin with generative models. Economic research examined where coordination should sit; organizational research examined why experience does not always change decisions. Machine-learning engineering later made the maintenance obligations of learned systems explicit. These contributions address different problems and remain useful together.

ContributionDateProblem clarified
The Nature of the Firm — R. H. CoaseNovember 1937Market contracting and internal coordination both cost resources; neither organization nor purchasing is universally preferable.
Organizational Learning and Management Information Systems — Chris Argyris1977Better reporting can correct deviations while leaving governing assumptions and defensive management relationships unchallenged.
Exploration and Exploitation in Organizational Learning — James G. MarchFebruary 1991Discovering possibilities and refining existing competence compete for scarce resources, with different timing and uncertainty of returns.
Hidden Technical Debt in Machine Learning Systems — D. Sculley and colleagues at Google2015Useful models bring system-wide maintenance obligations through data dependencies, feedback, configuration, and integration—not just model code.

Together, these ideas explain why cheaper implementation does not settle investment choices. An organization still has to choose a useful objective, coordinate delivery, learn whether the objective remains sensible, and maintain what others depend on.

Choose the outcome and alternatives

Strategy connects a chosen outcome to activities and resource commitments, including explicit exclusions. Complementary activities reinforce one another: an improved service may depend on specialist knowledge, reliable delivery, and customer support together. The Harvard Institute for Strategy and Competitiveness places differentiation in this activity system. Acquiring a model or dataset alone does not establish an advantage. Some investments instead provide capability parity: work the organization needs to remain a credible option, without making its offer distinctive.

An investment thesis states why a particular intervention should improve an important outcome and what would overturn that expectation. Record the present baseline, intended users, scope and exclusions, required organizational changes, and disconfirming observations. Compare alternatives before selecting the technology. Opportunity cost is the value forgone by choosing one feasible use of resources over another; already-employed specialists still have alternative work.

For example, consider the outcome of reducing the time needed to resolve routine internal requests. These are alternative proposals, not measured results.
AlternativeProposed mechanismCommitment to examine
Process changeRemove an unnecessary approvalAuthority to change the procedure
Conventional softwareRoute requests using explicit rulesRule maintenance and exception handling
Purchased functionalityUse an existing supported workflowSupplier fit and integration duties
AI assistanceInterpret varied requests and propose responsesAssessment, review, and correction capacity
Maintain the current arrangementAvoid transition workContinuing delay and displaced opportunities

The comparison keeps the outcome fixed while changing how it might be achieved. For a customer-facing offer, customer validation tests whether people make the commitments needed to use or purchase it; enthusiasm for a demonstration is weaker than supplying real work or releasing a budget. AI Startups and Small Teams develops that distinction. Leadership uses the findings to choose scope and investment, rather than treating technical feasibility as sufficient demand.

Investment economics and commitment

Turn improvement into benefit

Benefits realization converts improvement into an organizational outcome. Released time means less effort on an activity; usable capacity means that time can perform other work; realized savings mean expenditure actually falls. A benefit owner must arrange the conversion with finance and operational teams, preserve required outcomes, and avoid transferring costs or counting the same benefit twice. Enterprise AI explains these distinctions in everyday operation.

The first obligation is to establish whether the task improved. In METR's early-2025 developer study, Joel Becker, Nate Rush, Beth Barnes, and David Rein randomized AI access across 246 tasks performed by 16 experienced open-source developers in familiar repositories during February–June 2025. AI access increased completion time by 19%, although participants afterward estimated roughly 20% acceleration. This is a bounded result for those workers, tasks, and tools—not a general estimate of AI productivity.

A later result cannot simply replace that estimate if the comparison changes. METR's February 24, 2026 update reported that participation, task selection, compensation changes, and simultaneous agent use complicated its follow-up. It did not provide a reliable current effect estimate. The leadership lesson is to distinguish perceived usefulness from measured improvement while keeping the comparison's conditions visible.

Even a demonstrated gain needs a conversion plan. Factory's outcome-chain recommendation connects engineering changes to fewer defects, customer satisfaction, and business results. Each connection is a hypothesis to measure. Assign an owner who can change staffing assignments or spending, fund the necessary transition, and specify when the downstream result should become observable. Faster drafting alone does not establish any of those later outcomes.

Improvement needs a conversion decision

Example

Released time can support different benefits, but neither follows automatically.

Conditional uses of released capacity, not two benefits automatically earned from the same hours. Account for remaining and transition work within the measurement period, preserve quality, and establish the demand, redeployment or spending authority needed for the chosen benefit.
Read the diagram as text
  • Measured task improvement.
  • Account for remaining work. Review, correction, and transition.
  • Released time.
  • Additional useful output.
  • Reduced expenditure.
  • Measured task improvementAccount for remaining work: Check the complete work.
  • Account for remaining workReleased time: Net effort falls; quality holds.
  • Released timeAdditional useful output: Usable capacity + demand + funded redeployment.
  • Released timeReduced expenditure: Avoidable spend + budget authority.

Compare complete costs and timing

Return on investment, or ROI, expresses net monetary benefit relative to cost. Use an explicit convention: let BB be monetary benefits attributable to the program over a declared period, and C>0C>0 its included program costs. The ROI Institute methodology subtracts costs before dividing; B/CB/C alone is a benefit-cost ratio.

ROI=100×BCC%\mathrm{ROI}=100\times\frac{B-C}{C}\%

For an arithmetic example, benefits of $150,000 against $100,000 of included costs produce $50,000 net benefit and 50% ROI, not 150%. Whether those benefits are attributable, collectible, or merely valued capacity is a separate claim. For increased sales, count the monetary contribution left after the costs of supplying those sales, not their full revenue. Declare the benefit period rather than combining several years of benefits with an unexplained cost window.

Total cost of ownership includes acquisition and continuing operation. Bring integration, specialist time, assessment, change work, maintenance, support, and replacement into the stated boundary. Cost and Performance Engineering explains how to price the required service. Keep a dated ledger of initial outlays, recurring work, contractual commitments, transition costs, and expected benefits. An internal allocation identifies who carries a cost; an avoidable cost is expenditure changed by the decision. These are different quantities.

Payback is the point when cumulative net cash benefits recover the investment; it may never occur. A positive aggregate ROI does not establish affordable cash timing. Sensitivity analysis varies consequential assumptions, such as uptake, review effort, or benefit delay; a switching value identifies a change large enough to reverse the choice. Keep nonmonetary requirements visible instead of forcing every consequence into the ratio.

The Productivity J-Curve, Erik Brynjolfsson, Daniel Rock, and Chad Syverson's October 2018 working paper, explains why new general-purpose technologies require complementary investments in processes and human capital. When these investments are poorly measured, productivity can initially appear understated and later overstated as benefits arrive. This explains a measurement problem, not a guaranteed cash-return trajectory. Necessary preparation deserves funding; disappointing results still require a prospective case for further expenditure.

Choose a feasible portfolio

A portfolio manages investments together because they share objectives, resources, dependencies, or exposure. APM's portfolio guidance emphasizes strategic fit and delivery capacity. Separate project budgets do not create separate domain experts, integration teams, or support staff. Select combinations and sequences that can actually be delivered, including existing operating duties, rather than simply funding proposals in descending ROI order.

Consider a small planning example. After existing service duties, a domain-review team has two days a week available. A response-assistance pilot requires both days; a classification pilot also requires both. Both depend on the same model service and need a shared evaluation service to check their behavior. A third proposal builds that evaluation service using separately available platform capacity. Assume funding is sufficient and these estimates hold. Leaders can fund the shared service and either pilot, but must have the evaluation capability ready when the pilot needs it. They cannot run both pilots together without exceeding review capacity. Choosing one delays learning about the other; funding alone does not remove the conflict.

Funding, review capacity, and readiness are separate

Funding is sufficient. Two domain-review days per week remain after existing duties; each pilot needs both days. Platform capacity is separately available.

Select funded proposals
Check prerequisite readiness

Funding the evaluation proposal does not complete it. Readiness is a separate condition.

Review demand: 2/2 days per week — capacity fits

Evaluation gate: not ready.

Review capacity fits. The selected pilot remains conditional on evaluation capability becoming ready.

Evaluation work is selected and uses separate platform capacity. Both pilots retain the same model-service dependency.

Response pilotrequires →2 review days/week · evaluation capability · shared model service

Classification pilotrequires →2 review days/week · evaluation capability · the same model service

Evaluation proposalrequires →Separately available platform capacity

Arrows show requirements, not execution order. Either pilot uses 2/2 review days; both use 4/2 even when evaluation is ready. These checks do not establish model-service availability or operating authorization.

Selecting a funded proposal does not supply extra specialist time or complete its prerequisites.

Concentration risk arises when several commitments depend on the same vulnerable resource or supplier. The Bank of England's April 2025 analysis explains how externally supplied models, cloud infrastructure, and common components can create shared operational exposure. Internally built applications and different product brands do not establish independent dependencies. Investigate the actual upstream resources and whether substitution is practical before treating diversity as protection.

Required service or control obligations also consume portfolio capacity. They need an economical delivery choice, not invented revenue to compete with growth proposals. Compare discretionary additions only after accounting for those duties. Shared capabilities may enable several applications, but their expense belongs somewhere, and the same released staff time cannot be claimed independently by every project that helped release it.

Fund the next uncertainty

Funding learning differs from funding delivery. An evaluation systematically assesses behavior against intended use; Evals and Benchmarks explains what it can establish. A working prototype may resolve technical feasibility while leaving usefulness, economics, and permission to operate unsettled. A bounded experiment is worthwhile when its possible findings could change the next commitment. Repeating an already-settled demonstration buys little decision-relevant information.

March called investigating new possibilities exploration, and refining established competence exploitation. His organizational models explain a persistent tension: dependable near-term improvements can crowd out uncertain learning, while perpetual experimentation can fail to develop usable competence. Neither implies a fixed budget percentage. Allocate learning resources according to the uncertainties that constrain important decisions, while preserving the capacity to operate what already works.

A real option preserves the ability, without the obligation, to invest later. In The Value of Waiting to Invest, Robert McDonald and Daniel Siegel's November 1982 working paper examined irreversible investment under uncertainty. Committing now sacrifices the alternative of acting after more information arrives, so positive expected net benefits need not justify immediate commitment. Waiting can nevertheless lose opportunities, and maintaining reversibility can cost money. The theory supplies a comparison, not a reason to delay everything.

Make the next funding agreement specific: identify the uncertainty, bounded money and staff time, observations that distinguish alternatives, decision owner, and review deadline. State what each result would justify and which duties survive stopping. The enterprise portfolio-of-experiments argument is useful when scope and value emerge through work, but it does not establish venture-capital return distributions or excuse unlimited experiments. A negative finding can make a test successful by preventing a larger mistaken commitment.

Sourcing and dependence

Source capabilities and duties

Build versus buy allocates continuing work between the organization and suppliers. A partnership adds coordinated development or delivery. As AI Startups and Small Teams explains, buying transfers specified duties, not the complete product obligation. At portfolio scale, separate application ownership, model supply, operation, and domain expertise before comparing arrangements. A delivery-model assessment can apply to individual components rather than forcing one sourcing choice onto the whole service.

Coase's coordination argument explains the tradeoff. Purchasing incurs the work of specifying, negotiating, and adapting transactions. Internal organization can avoid some of that work, but managing more activities creates its own costs and mistakes. Building is attractive only when the internal arrangement is preferable to available alternatives—not simply because it offers nominal ownership.

Transaction-Cost Economics, Oliver Williamson's October 1979 article, develops how uncertainty, recurrence, and relationship-specific investments affect governance. Asset specificity means an investment loses substantial value in alternative uses. Specialized training and accumulated supplier knowledge can make replacement difficult even when the initial market was competitive. Standardized inputs can be easier to substitute. The implication is to compare adaptation and retained competence alongside price; elaborate partnership arrangements also impose costs when simple purchasing would suffice.

This illustrative allocation shows how one application can combine several sourcing arrangements.
Capability or dutyPossible arrangementWork retained internally
Model executionPurchase a hosted serviceSpecify required behavior and assess changes
Domain interpretationDevelop with a specialist partnerResolve requirements and retain usable knowledge
Application integrationBuild internallyMaintain interfaces and coordinate source owners
Service operationSplit duties by an explicit agreementOwn the end-to-end outcome and unresolved handoffs

Partnership agreements should make joint work governable: milestones, funding, responsibilities, work-product rights, access, reporting, knowledge transfer, disagreement, and termination. UKRI's collaboration guidance provides a concrete research-partnership example of these continuing obligations. The appropriate terms depend on the relationship. Agreement on a deliverable does not demonstrate that either partner can operate or replace it.

Preserve a practical exit

Vendor lock-in is the difficulty or expense of changing suppliers. It includes integration changes, behavior revalidation, information movement, transition work, and lost expertise. Intuit reported that even moving between models from the same supplier required substantial evaluation. Prompts—instructions and other input supplied when invoking a model—can encode expectations that do not transfer smoothly. Prompt transfer boundaries explain the mechanism. Compatible request syntax reduces one kind of switching work; it does not establish equivalent application behavior.

When Your LLM Reaches End-of-Life, an April 29, 2026 Verint preprint by Emma Casey, David Roberts, David Sim, and Ian Beaver, illustrates application-level replacement. Candidates first faced internal vetting, licensing, cost, and regional-availability constraints. Comparisons then examined correctness, refusals, structure, style, and response time against the incumbent. Nova 2 Lite and Qwen3-32B were suitable candidates under those tests; covering all required regions and modalities could require multiple models. The English-only main analysis used two test sets with limited human calibration. Candidate selection did not establish a completed worldwide migration, full transition expenditure, or realized savings.

Conceptual application cutaway. The incumbent remains live while a candidate undergoes behavioral comparison and operational qualification. A common request shape does not establish correctness, refusal, structure, style or response-time fit; licensing, region, modality, support and receiving ownership also require assessment.
An exit plan connects outgoing duties to receiving-team readiness.
ObligationWhat must be available
Change notificationEnough notice to assess and prepare
Artifacts and accessUsable records, interfaces, and necessary information
Knowledge transferPeople able to maintain the replacement
Transition supportFunded assistance and clear handoffs
AcceptanceReceiving owners and explicit service conditions

An exit clause establishes an obligation, not demonstrated transition capability. Preserve time and expertise to exercise it. Review permissions for the actual service path, including the selected feature and deployment arrangement, rather than assuming a supplier-level approval covers the replacement. Also investigate shared upstream dependencies: two contracts may still rely on one critical resource.

Capabilities and organization

Close the actual capability gap

Translate the chosen portfolio into work the organization must repeatedly perform: domain judgment, application engineering, data stewardship, assessment, operation, and supplier management. One person may cover several responsibilities; some need specialist support. The UK Government AI Playbook describes capability needs across the service lifecycle and allows hiring, contractors, and internal development. Choose among them by time to competence, retained knowledge, and continuing coverage—not job-title completeness.

Diagnose the missing condition before purchasing a remedy. Someone may know how to review outputs but have no available time. Another person may have capacity but lack access to the source records. A named owner may understand the problem yet lack authority to change it. These require workload allocation, access decisions, or a mandate—not the same training course.

Intuit's tax-explanation work illustrates complementary capabilities. Tax analysts contributed domain knowledge, prompts, and initial judgments, while data-science and machine-learning staff concentrated on quality measures and test datasets. The division connects expertise to reusable assessment infrastructure. It does not establish an optimal staffing ratio or specify who could authorize release.

Absorptive capacity is the ability to understand and apply incoming knowledge; Enterprise AI develops its role in transferring practices. Fund protected learning time and practical demonstrations, not attendance alone. Require receiving teams to perform the work with documentation and support, and provide backups so expertise survives absence or departure. APM's transition guidance treats skills, operational acceptance, and continuing support as delivery work rather than a final administrative handoff.

Place decisions near the work

An operating model arranges the work, people, information, suppliers, management processes, and decision rights needed to deliver a service. The Operating Model Canvas makes these arrangements explicit. Reporting lines alone are insufficient: identify who defines domain requirements, accepts application behavior, operates shared services, and decides exceptions. Then examine the handoffs those choices create.

AWS's operating-model guidance distinguishes three arrangements. These are design alternatives, not a maturity ranking.
ArrangementPlacementCoordination consequence
CentralizedA common team delivers applications and shared capabilitiesPool specialists; maintain access to domain owners and manage the central queue
EmbeddedBusiness units deliver their applicationsKeep domain work close; address duplicated expertise and infrastructure
FederatedLocal delivery combines with explicit central services and constraintsPreserve local judgment; define shared-service duties and exception authority

A center of excellence concentrates expertise, but the name does not determine its mandate. It might teach teams, deliver difficult work, maintain shared capabilities, or approve defined exceptions. Those roles need different capacity and authority. If every ordinary change requires central approval, that queue becomes part of delivery. If application teams can bypass shared responsibilities without agreement, federation has not resolved ownership.

Scale the arrangement to the work. Oleve's small-team example separates product-outcome ownership from cross-product automation without implying a large departmental hierarchy. Shared code, templates, and operated services carry different service responsibilities. Team Topologies provides a vocabulary for team responsibilities and interactions; its home discussion explains collaboration, service consumption, and enabling work. Use these boundaries to decide where resources belong, not to copy an organization chart.

Shared funding

Fund the continuing service

A project budget funds a bounded change; a service continues creating obligations. Sculley and colleagues located maintenance liabilities in data dependencies, configuration, integration, and hidden consumers as well as model code. Taking on such debt can be strategically reasonable, but it leaves work to finance. Assign continuing review, support, maintenance, and replacement funding before declaring initial delivery complete.

Shared costs may remain centrally funded or be allocated using fixed shares, consumption, or another stated proxy. FinOps allocation guidance recommends an explicit policy for each category and allows mixed approaches. The policy should expose who benefits and who can influence demand; changing the allocation does not itself change the underlying expenditure.

Showback reports attributed costs. Chargeback assigns those expenses to official budgets. FinOps does not treat chargeback as inherently more mature.
ArrangementAccountability effectTradeoff
Central fundingA shared owner carries expenditureDemand needs explicit review
ShowbackConsumers see attributed costsVisibility does not transfer budget responsibility
ChargebackConsumers carry assigned expensesAdds accounting work; allocation rules influence choices

A forecast is an agreed expectation of future spending and value based on workload, pricing, timing, and lifecycle assumptions. FinOps forecasting brings engineering, product, finance, and leadership into that model; budget owners must meet it or seek additional funding. Review consumption alongside usefulness. A forecast is not proof of realized benefit, and an unused service does not become worthwhile because every team has been assigned a share.

Validated commonality means that different users genuinely share a requirement. It is a stronger basis for common investment than repeated requests for similarly named features. Forward Deployed Engineering explains how to establish it. Fund the shared capability together with its support boundary, and keep exceptional local work explicitly owned rather than hiding it in the platform budget.

Authority and assurance

Give ownership real authority

Responsibility identifies who performs work; accountability identifies who answers for an outcome; decision rights specify who may approve, change, or stop something. Enterprise AI develops the distinction. An ownership chart cannot grant missing authority. Separate the application outcome owner, technical operator, risk owner, and people who challenge claims, even when a small organization assigns several duties to one person.

A risk owner ensures that a specified risk is managed and monitored, with authority to act or escalate; the owner need not perform every mitigating action. Residual risk is exposure remaining after controls. Risk appetite describes the types and amounts of exposure an organization is willing to accept in pursuing objectives. These concepts require explicit limits, not a general declaration that leadership accepts risk. The historical Orange Book supplies the ownership and residual-risk vocabulary.

Use a practical authority record rather than a name alone.
DecisionMandate to specify
Commit resourcesBudget scope, available capacity, and escalation limit
Interrupt operationWho can stop which uses and obtain necessary access
Require remediationWho assigns work and funds the remedy
Accept remaining exposureAuthorized scope and obligations outside that authority
Resume or expandRequired findings, decision maker, and absence cover

DWP's 2025 account of its arrangements illustrates written delegation and escalation when mitigation fails, exposure exceeds limits, or the current owner lacks control. The transferable principle is an effective route to greater authority, not its particular committee structure. NIST's AI Risk Management Framework similarly assigns continuing executive risk decisions and authority to disengage systems. Funding approval does not grant permission for every resulting use.

Automation need not eliminate ownership. In Intercom's April 2026 review account, an automated-approval pilot excluded changes deemed too broad or large and allowed human review requests. Engineers retained responsibility for observing their changes in production and rolling them back. That is a concrete allocation of duties, not general proof that automated approval improves safety or ROI.

Commission the right assurance

Assurance supports justified confidence through examination suited to a claim; it is not a guarantee. The UK introduction to AI assurance emphasizes measuring, evaluating, and communicating trustworthiness, including limitations and mitigations. Begin with the disputed claim and commission the missing examination. More evidence of task accuracy does not resolve an unexamined permission, operating, or benefit claim.

Independent challenge requires more than a second checking pass. The IIA Three Lines Model distinguishes delivery and risk-management roles from independent internal audit. Its September 2024 revision of the 2020 model connects independence to governing-body accountability, resources, information access, and freedom from interference. Specialist challenge can remain part of management. The lines describe concurrent responsibilities, not three mandatory departments or serial approval gates. An evaluation team whose findings can be suppressed by the delivery sponsor is not independent merely because it has a different name.

Independent challenge needs protected relationships

Example

Internal audit has direct governing-body accountability and access to management information; specialist challenge within management remains a different role.

Management challenge and independent audit

Delivery and specialist risk challenge sit inside management. Independent internal audit sits outside that boundary. The governing body supplies audit resources and mandate; audit returns findings and accountability directly and has access to relevant management people and information. Positions carry no quantitative meaning.

048121602.557.510Role placement (conceptual units)Relationship placement (conceptual units)Management boundaryGoverning body supplies resources and mandateAudit reports findings and accountability directlyAudit accesses management people and informationManagementDelivery / operationsRisk support / challengePeople and informationGoverning bodyIndependent internal auditResources and mandateFindings andaccountabilityAccess to people and information
  • 1. Management boundary
  • 2. Governing body supplies resources and mandate
  • 3. Audit reports findings and accountability directly
  • 4. Audit accesses management people and information
Read coordinates and regions as data

X: 016 conceptual units; Y: 010 conceptual units, increasing up. Equal scale on both axes.

Management boundary (polygon)

(0.5, 1); (7, 1); (7, 6); (0.5, 6)

Governing body supplies resources and mandate (polyline)

(10, 8); (10, 4.5)

Audit reports findings and accountability directly (polyline)

(13, 4.5); (13, 8)

Audit accesses management people and information (polyline)

(10, 3.2); (6.5, 2.2)

Management: (3.75, 5.4)

Delivery / operations: (3.75, 4.4)

Risk support / challenge: (3.75, 3.4)

People and information: (3.75, 2)

Governing body: (11.5, 8.7)

Independent internal audit: (11.5, 3.9)

Resources and mandate: (9.7, 6.6)

Findings and: (13.3, 6.7)

accountability: (13.3, 6)

Access to people and information: (8.4, 1.6)

Concurrent role relationships in the IIA Three Lines Model, September 2024 revision of the 2020 model—not three required departments or sequential approvals. Independence rests on resources, protected access, and direct accountability, not another checking pass.
Different claims call for different examinations. Evaluation decisions and governance assurance develop the methods.
ClaimRelevant examinationWhat remains separate
The task is performed acceptablyRepresentative cases and competent judgmentsBenefit under everyday use
The organization benefitsBaseline comparison including remaining human workPermission for wider exposure
Operations can receive the serviceAcceptance criteria and readiness with operational usersEvery future operating condition
A supplier owes a controlApplicable agreement or attestationActual configuration and enforcement

Separate prompts or models for implementation and testing can add scrutiny, but they do not establish independent errors or organizational independence. Similarly, an AI-generated safety argument identifies claims to inspect rather than proving its own completeness. Leaders must provide reviewers with competence, time, access, and a route to the authority that can resolve disagreement.

The disposition can be conditional acceptance, restricted use, further examination, or withheld authorization. Record the supported scope, remaining uncertainty, responsible decision maker, and changes that trigger reconsideration. NIST's framework treats residual-risk decisions and post-deployment monitoring as continuing responsibilities. A passing test and a funded project do not combine into unrestricted permission to operate.

Organizational learning

Turn field findings into investment

Forward deployed engineering brings engineers close to customer work to learn the domain and implement useful solutions; its home chapter explains those responsibilities. Leadership decides what the findings justify beyond the engagement. Productization makes a capability repeatable and supported. Repeated requests are a reason to investigate shared requirements, not proof that one implementation or operating arrangement fits them all.

Kevin Bai's platform-versus-customer distinction keeps unique behavior local while using deployment discoveries to identify generalizable capabilities. The leadership choices are broader than accepting or rejecting a feature: fund local adaptation, create a supported shared capability, offer explicitly staffed service work, or decline the request. Each choice needs an owner and a maintenance boundary.

In this illustrative contrast, two requests both ask for automated case summaries.
RequirementPotentially sharedStill locally owned
Read the same case systemA supported integrationAccess approval for each use
Produce a summaryCommon interface and reusable implementationDifferent domain definitions of an acceptable summary
Keep the service usefulShared technical supportDistinct review, escalation, and staffing duties

The common integration can deserve investment without centralizing every policy decision. CNCF's platform guidance prioritizes common needs while allowing capabilities outside the platform. Before expanding, preserve the conditions behind the original result: users, domain rules, data access, review effort, support, and beneficiaries. Establish which recur in the receiving setting. A success produced by unusually intensive specialist support does not by itself justify expansion without that support.

Change assumptions and incentives

Revise the governing choice

Single-loop learning corrects action while preserving governing goals and rules. Double-loop learning reconsiders those goals or rules and then changes action. Argyris's 2002 explanation makes the distinction explicit. Reducing review effort to meet an existing savings target is single-loop improvement. Reconsidering whether that target remains justified after discovering necessary review work changes the governing assumption.

More reporting does not automatically produce that reconsideration. In Argyris's 1977 management-information analysis, tighter financial information helped address some deviations while defensive relationships concealed harder organizational problems. A dashboard can make a target's shortfall visible without making it acceptable to challenge the target. Leadership must permit the investment thesis itself—not just execution quality—to be examined.

Escalation of commitment means increasing resources devoted to an existing course despite adverse feedback. In Knee-Deep in the Big Muddy, Barry Staw's 1976 experiment assigned 240 business students hypothetical investment decisions. Commitment was greatest when participants owned the earlier allocation and received negative results. This bounded experiment illustrates a danger of defending one's prior choice, not a universal explanation of executive behavior. Continued investment can be rational when future benefits justify it; prior personal responsibility is not itself such a benefit.

Make truthful reporting workable

Psychological safety is a shared belief that interpersonal risk-taking is safe. Amy Edmondson's 1999 study of 51 manufacturing teams associated it with learning behavior and considered contextual support and leader coaching alongside interpersonal beliefs. The observational study does not establish an AI-specific causal effect. Its relevance is practical: admitting uncertainty and seeking help need conditions beyond a reporting channel. Safety to challenge is compatible with demanding standards for the work.

Incentives are consequences that influence behavior, including recognition, workload, and professional standing. Rewarding licenses used, demonstrations delivered, or output generated can favor activity without establishing benefit. KCS guidance recommends outcome goals while using activity counts diagnostically. Recognize accurate reporting and justified stopping; Enterprise AI explains implementation with affected teams.

Workforce commitments need the same honesty. State changed duties, learning resources, support coverage, and employment expectations instead of assuming saved task time warrants fewer staff. The OECD's workplace survey report found consultation associated with more favorable worker reports in finance and manufacturing, but not a causal productivity estimate. Skills were discussed more often than job loss or wages. Leadership should resource and address the whole transition; change-management practice belongs with the teams doing it.

Renewal and retirement

Reallocate without abandoning duties

Review commitments when material assumptions change, before renewal deadlines, and often enough to address consequential uncertainty. Compare observed benefits with expectations, then examine remaining expenditure, operating duties, scarce capabilities, and feasible alternatives. The available dispositions are different decisions: continue within limits, expand, redirect, narrow, pause, or retire. A project status does not specify the resources or authority each requires.

Sunk costs are past expenditure that cannot be recovered. They should not determine the next choice. Future exit costs, unavoidable commitments, reusable assets, and scarce staff time still matter. Cancelling a project's allocated share does not save money if total expenditure remains unchanged.

Keep a short decision record that makes reallocation explicit.

  • Changed basisObservations that changed the thesis, including unresolved benefits.
  • Remaining commitmentProspective spending, service duties, and constrained capacity.
  • AlternativesFeasible uses of released resources and what each displaces.
  • DispositionDecision, reasons, authorized scope, and accountable owner.
  • ReconsiderationNext deadline or observation that requires another decision.

Retirement is funded transition work. Government Digital Service guidance, first published in 2016, requires attention to continuing user needs, knowledge transfer, communication, API consumers, and protection of retained information. Assign those duties and finance them even when new development stops. Enterprise AI develops the operating practice. Ending a roadmap does not end obligations to people already relying on the service.

A narrower service or staged replacement can preserve demonstrated value while releasing capacity from unsupported expansion. That is neither automatic retreat nor a defense of the original plan. It is the central leadership discipline: commit resources to work that remains worthwhile, make the authority match the consequences, and preserve responsibility for what the organization has already undertaken.

Open questions

  1. Measuring productivity becomes harder when tools change who participates, which tasks are attempted, and how concurrent work is timed. Progress requires credible comparisons that represent changed working practices without mistaking preference for measured benefit.

  2. Practical supplier independence remains difficult to value. Behavioral compatibility, retained expertise, transition time, and common upstream resources all matter. More informative replacement studies would report completed transitions and their full costs, not only candidate performance.

  3. Organizations still need workable ways to scale scrutiny without making approval the limiting resource or weakening challenge. Progress would show that reviewers retain access and influence while consequential decisions remain timely and accountable.

Follow the curated reading path through the speakers and demonstrations behind this entry.

18 min

AI Engineer World's Fair 2026 · 2026

Forward Deployed Engineering 101

Kevin Bai

Cited in this entry

Clarifies the economic boundary between reusable platform capabilities and customer-specific work, including the maintenance obligations that remain after reuse.

Watch talk

Explore more talks

The rest of the library, beyond the curated path. Cited talks support this entry; reviewed transcripts were processed in full. Metadata candidates have not been reviewed as sources or verified as topic members.

20 matching talks

Every catalogued talk on this subject: Leadership

TalkSpeakerEventYear
Eno ReyesAI Engineer World's Fair 20262026
Sandipan BhaumikAI Engineer Europe 20262026
Sonny Merla, Mauro Luchetti, Mattia RedaelliAI Engineer Europe 20262026
Rossella Blatt Vital, Deepsha MenghaniAI Engineer World's Fair 20252025
Dan BjornnAI Engineer World's Fair 20262026
Alex AtallahAI Engineer World's Fair 20252025
Amir HaghighatAI Engineer World's Fair 20252025
Jeremy Silva, Chris HernandezAI Engineer World's Fair 20252025
Fuzzing in the GenAI Era

Transcript reviewed

Leonard TangAI Engineer World's Fair 20252025
Michal CichraAI Engineer Europe 20262026
Martin Harrysson, Natasha ManiarAI Engineer Code 20252025
Sanja GrbicAI Engineer World's Fair 20262026
Brian ScanlanAI Engineer Europe 20262026
Nathaniel Whittemore (NLW)AI Engineer Code 20252025
Shirsha ChaudhuriAI Engineer Summit 20252025
Jia WuAI Engineer World's Fair 20262026
The New Lean Startup

Cited in this entry

Sid BendreAI Engineer World's Fair 20252025
Ben HylakAI Engineer World's Fair 20262026
Vision: Zero Bugs

Cited in this entry

Johann Schleier-SmithAI Engineer Code 20252025
Theo BrowneAI Engineer World's Fair 20262026

References

Coverage and source review
Processed transcripts
24 processed in full · 4 in the curated path
Automated source review
Passed
Metadata candidates
0 unreviewed; not verified topic membership
Corpus version
1bd8e407b26a07b33815594e1b2db5f41827119a2b3cb6fbf240f9fc571fc767

Automated review checks source support; it is not publication approval.

A synthesis of selected conference talks and technical references. Citations link to the source material; they do not imply that every talk on this subject is included.

  1. Government Digital Service: Artificial Intelligence Playbook for the UK Government

    The playbook defines AI team capability beyond model development: understanding users, integrating software, handling data responsibly, testing with users, measuring service performance, operating, improving and retiring the service. Teams need technical and domain expertise plus access to legal, commercial, security and privacy specialists. Capability needs change across the lifecycle; staffing may combine hiring, contractors and internal development. Organizational support includes a sourcing strategy, learning-needs assessment, leadership access, feedback mechanisms and sufficient staff time and tools to adapt to emerging risks.

  2. The engineer of the future is the person who is able to choose what is worth doing — Addy Osmani

    Let agents investigate, implement, test, and report in the inner loop while accountable owners decide, verify, approve, and own production outcomes in the outer loop.

  3. Thomson Reuters: Fourth-Quarter and Full-Year 2025 Results Presentation

    In its February 5, 2026 presentation, Thomson Reuters contrasted its 2023-onward emphasis on broad tool access, learning days, certification, forums and experimentation with additional 2025-onward organizational commitments. The latter included leadership oversight, three-year AI roadmaps across business segments and functions, dedicated change-management and AI expertise, business-specific agent development and process mapping. The case shows an organization adding explicit planning and delivery capacity to employee experimentation.

  4. R. H. Coase: The Nature of the Firm

    In November 1937, Coase examined why firms coordinate some activities through managerial direction instead of separate market transactions. Market exchange incurs contracting costs; internal organization can reduce some of them, but organizing additional activities also creates costs and mistakes. His explanation places the firm’s boundary where further internal coordination ceases to be preferable to market exchange or another firm’s organization. Internal authority remains bounded by the underlying agreement.

  5. Chris Argyris: Organizational Learning and Management Information Systems

    Argyris’s 1977 article investigates why management information systems can disappoint despite technical investment. Systems can help correct deviations from existing targets while leaving the targets and defensive organizational behavior unchallenged. In a newspaper-management example, tighter financial information helped address some errors but also concealed harder problems associated with competitive, defensive relationships among managers. Improving information delivery therefore did not automatically improve the organization’s capacity to question its governing assumptions.

  6. James G. March: Exploration and Exploitation in Organizational Learning

    March’s February 1991 paper distinguishes exploration—searching, experimenting and discovering possibilities—from exploitation—refining and executing existing competence. Both consume scarce resources, but their returns differ in timing, uncertainty and who receives them. Models of organizational learning and competition show how rapid improvement in exploitation can reinforce short-term effectiveness while weakening longer-term adaptation. Exclusively exploring can instead leave an organization paying for experiments without developing usable competence.

  7. D. Sculley and colleagues: Hidden Technical Debt in Machine Learning Systems

    Google’s Sculley and colleagues argued in 2015 that rapidly building a useful prediction system can create substantial continuing maintenance obligations. Their analysis identifies data dependencies, hidden feedback loops, undeclared consumers, configuration problems and integration code that embeds assumptions. These liabilities exist across the system, so improving model code alone does not resolve them. The paper explicitly recognizes that taking on technical debt can be strategically reasonable, provided its continuing obligations are addressed.

  8. Harvard Institute for Strategy and Competitiveness: Creating a Successful Strategy

    The institute explains strategy through a distinctive value proposition, activities configured to deliver it, explicit tradeoffs, and activities that reinforce one another. Choosing what not to do prevents incompatible demands from undermining the chosen position. Strategic differentiation therefore concerns the organization’s activity system, not merely possession of a particular technology.

  9. HM Treasury: The Green Book (2026)

    Appraisal starts with the case for change, a theory connecting intervention to outcomes, business-as-usual conditions, objectives and strategic fit. It compares alternative approaches against that baseline and considers lifetime costs, nonmonetized effects and risks rather than ranking solely by financial ratios. Sensitivity analysis changes important assumptions; switching values identify changes that overturn a decision. Real-options analysis examines flexibility exercised after new information arrives. Sunk expenditure should not determine future choices, while alternative uses of existing resources still matter. Evaluation subsequently checks whether expected outcomes and costs materialized.

  10. ACCA: Relevant costs

    For a decision, relevant financial effects are future cash flows changed by that decision. Reallocating existing fixed overhead does not create savings when total expenditure remains unchanged; additional fixed costs caused by a choice do matter. Opportunity cost includes benefits forgone by choosing one use of a resource over another. The published labor example distinguishes already-paid idle capacity, newly hired capacity and scarce staff diverted from productive work.

  11. HM Treasury and Government Finance Function: The Government Efficiency Framework

    The framework distinguishes direct spending reductions from monetizable benefits that leave spending unchanged, such as handling more queries with the same staff. Efficiency claims must preserve outcomes, subtract delivery costs, avoid consequential costs elsewhere and avoid double counting. Benefits methods require a baseline, documented assumptions and evidence. A benefit owner works with finance and analysts, and monitoring continues into normal operations because benefits may arise after project delivery.

  12. METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

    Joel Becker, Nate Rush, Beth Barnes and David Rein studied 16 experienced open-source developers working on 246 tasks in familiar repositories during February–June 2025. Tasks were randomized to permit or prohibit AI assistance. With the studied tools, AI access increased completion time by 19%, although participants subsequently estimated that it had made them about 20% faster. The experiment separates perceived usefulness from measured task duration under a specified workload.

  13. METR: February 2026 developer-productivity follow-up update

    In its February 24, 2026 update, METR explained why its later developer study did not provide a reliable estimate of current AI productivity effects. Developers’ reluctance to work without AI affected participation and task selection; changed compensation and simultaneous agent use also complicated comparisons and time measurement. The researchers were revising the study design rather than treating the follow-up’s apparent result as a clean update.

  14. How Forward Deployed Engineering is done at Factory

    Define an outcome or ROI story at the beginning that connects changes in engineering behavior to core business goals.

  15. ROI Institute: Introduction to the ROI Methodology

    The methodology calculates ROI percentage as monetary program benefits minus program costs, divided by program costs, multiplied by 100. This differs from the benefit-cost ratio, which divides benefits directly by costs. It calls for isolating the program's contribution to observed improvement and including fully loaded costs such as analysis, development, team time, overhead and evaluation.

  16. ROI Institute: ROI Basics

    The authors require separating a project's contribution from other influences on business measures before monetizing benefits. For increased sales, they use profit margin rather than treating all sales revenue as benefit. Their introductory example compares one year of monetary benefits with the project's full direct and indirect costs, making the benefit period explicit.

  17. FinOps terminology: ownership, depreciation and utilization

    FinOps defines ownership cost broadly, including acquisition, management/support, communications and labor; depreciation distributes asset cost over time, while activity-based costing can allocate staff hours times hourly rates. Illustrative engineering model: period cost=a(H-S)/L+energy+maintenance+staffing+software/network+service charges, where H is hardware purchase cost, S assumed residual value, L useful life in matching periods and a the workload's allocated share. Do not count both the full purchase and its depreciation in the same period model. With sustained busy throughput q, available time T and utilization u, completed volume is approximately q*u*T under a stable workload; fixed cost per unit therefore rises as utilization falls. Energy expense is measured kWh times the applicable tariff.

  18. Erik Brynjolfsson, Daniel Rock and Chad Syverson: The Productivity J-Curve

    The authors’ October 2018 working paper explains how general-purpose technologies require complementary investments in processes, products, business models and human capital. These intangible investments are often poorly measured. Their model shows how measured productivity can initially understate improvement while organizations build complements, then overstate it when benefits arrive without corresponding recognition of the earlier investment. The contribution concerns both organizational complements and the timing of their measurement.

  19. APM: What is portfolio management?

    A portfolio groups projects and programmes to manage investment toward strategic benefits or operational efficiency. Portfolio management selects, prioritizes and controls that work against strategy and delivery capacity. APM explicitly recognizes that individual project priorities may need to give way to the wider portfolio. Portfolio plans identify dependencies, timescales and deliverables; portfolio risks include resource availability, implementation capacity and investment constraints.

  20. Bank of England: Financial Stability in Focus—Artificial Intelligence in the Financial System

    The Bank’s April 2025 analysis explains that externally supplied AI models create operational dependencies, while internally built models can still depend on cloud infrastructure and external data suppliers. Reliance on a small number of providers can transmit disruption across institutions, particularly when rapid migration is infeasible. It also identifies common model components and data libraries as possible shared weaknesses. Internal development and multiple customer-facing products therefore do not establish independence of their underlying resources.

  21. FinOps Foundation: Allocation

    Shared technology costs can be funded centrally or assigned to cost centers. Allocation methods include fixed shares, proportional shares and variable proportions based on proxy measures. FinOps recommends an explicit strategy for each shared-cost category and recognizes that organizations may combine methods and revise them as their information and practices develop.

  22. The Production AI Playbook: Deploying Agents at Enterprise Scale

    Define business success and build a representative evaluation dataset before comparing models; reuse that dataset to assess provider upgrades.

  23. Robert L. McDonald and Daniel Siegel: The Value of Waiting to Invest

    McDonald and Siegel model an irreversible investment whose benefits and costs evolve uncertainly. Investing now gives up the alternative of investing later after conditions change. Consequently, a positive difference between expected benefits and immediate cost need not be sufficient to justify immediate commitment: the option to wait also has value. The analysis compares mutually exclusive investment timings rather than treating deferral as doing nothing without consequence.

  24. Guidelines for Managing Projects: How to organise, plan and control projects

    The guidance calls for agreed scope and exclusions, with changes assessed against that baseline. Stakeholder analysis examines changes to work, responsibility, authority, and maintenance duties, and identifies daily contacts and escalation paths. Changes beyond the project manager's authority go to the responsible owner or board. Benefits planning specifies the benefit, measurement units, timing, method, and responsibility; the business case is updated as costs and expected benefits change.

  25. Most Enterprise Agentic Projects Are Doomed — Here’s Why

    The speakers recommend that finance 'think like a VC': fund a portfolio of AI experiments rather than require every project to promise a fixed solution and predictable payback upfront.

  26. Cabinet Office: The Sourcing Playbook

    A delivery-model assessment compares in-house, purchased and hybrid provision for a service or individual components. Criteria include strategy, transition, skills, capacity, asset ownership, continuing quality, management requirements, risk and whole-life cost. Tests and pilots should establish objectives, scope, resources, timescales and time to assess results before scaling. Exit planning connects outgoing-provider duties with incoming-provider or internal mobilization, specifying activities, resources, accountabilities, dependencies, timelines and acceptance standards.

  27. Oliver E. Williamson: Transaction-Cost Economics—The Governance of Contractual Relations

    Williamson’s October 1979 article distinguishes transactions by uncertainty, recurrence and relationship-specific investment. Asset specificity means that equipment, skills or other investments lose substantial value in alternative uses. Specialized training and accumulated knowledge can make an initially competitive supplier relationship costly to replace. Conversely, standardized inputs permit easier substitution. The paper compares market exchange, internal organization and contracting arrangements that support adaptation; elaborate governance also imposes costs when applied to simple transactions.

  28. UK Research and Innovation: Collaboration Agreements

    UKRI’s collaboration guidance calls for partners to agree how research will be managed and coordinated, who supplies funding, and how responsibilities and liabilities are divided. Agreements should also address intellectual property, reporting, publication, access, confidentiality, default, termination and dispute resolution. A partnership therefore requires arrangements for continuing coordination and disagreement, not merely agreement on an initial deliverable.

  29. How Intuit uses LLMs to explain taxes to millions of taxpayers

    Prompts create behavioral lock-in in addition to contractual vendor lock-in; same-vendor upgrades still require substantial evaluation.

  30. Emma Casey, David Roberts, David Sim and Ian Beaver: When Your LLM Reaches End-of-Life

    Verint’s April 29, 2026 preprint describes selecting replacements for a model in a commercial question-answering service. Candidate models first had to meet internal vetting, licensing, cost and regional-availability constraints. Comparisons then examined answer correctness, refusals, output structure, style and response time against the incumbent. The authors identified Nova 2 Lite and Qwen3-32B as suitable candidates under their tests. Covering all required regions and modalities could require retaining more than one model.

  31. UK Government: Data Ownership Model

    The model separates accountable data owners, who authorize major changes and answer for them, from stewards responsible for everyday management. Owners oversee meaning, quality, use and access; stewards facilitate access processes and investigate, triage and remediate quality problems. Common definitions and authoritative sources support sharing. Its RACI legend distinguishes Responsible, doing the work; Accountable, owning the outcome; Consulted, providing input; and Informed, receiving updates.

  32. The engineer of the future is the person who is able to choose what is worth doing — Addy Osmani

    Distrust does not create review capacity; verification must become cheaper, clearer, and harder to skip as generation scales.

  33. How Intuit uses LLMs to explain taxes to millions of taxpayers

    Intuit uses tax analysts as prompt engineers and initial evaluators, allowing data science and ML staff to concentrate on quality metrics and test datasets.

  34. APM Body of Knowledge, Seventh Edition: Transition into Use

    APM treats business readiness as work throughout delivery, including skill gaps, operating impacts and legitimate dissent. Transition involves agreed acceptance criteria, testing with operational users, documentation and transfer of responsibility. Higher-risk transitions need contingency arrangements. Adoption requires continuing support, and benefits tracking remains accountable after handover. The guidance warns against claiming savings by increasing operating costs elsewhere.

  35. Operating Model Canvas

    The authors describe an operating model through the work needed to deliver a service, the people performing it, locations and assets, supporting information systems, suppliers, and management processes. These arrangements translate strategy into operating choices. Their tools include decision grids and process-owner grids, and support describing both current and intended operations.

  36. AWS: Generative AI operating models in enterprise organizations with Amazon Bedrock

    AWS distinguishes decentralized delivery within business units, centralized delivery through a common team, and federated arrangements combining business-unit applications with centrally managed reusable capabilities. Its federated example keeps business-specific development near domain expertise while a central team curates and maintains shared services. It describes continuing product ownership for evolving those services and the operating model.

  37. Most Enterprise Agentic Projects Are Doomed — Here’s Why

    Treat governance and deployment speed as engineering debt, because faster code creation can shift the constraint to review and release.

  38. One Registry to Rule them All - Sonny Merla, Mauro Luchetti, & Mattia Redaelli, Quantyca

    AmplifAI separates central guideline setting from country and corporate execution, with governance, platform, and factory as distinct program responsibilities.

  39. The New Lean Startup

    Use the harvester-and-cultivator model to distinguish product delivery from cross-product infrastructure.

  40. FinOps Foundation: Invoicing and Chargeback

    Showback makes attributed technology costs visible to responsible groups; chargeback places those expenses into their official budgets or accounting arrangements. FinOps does not treat chargeback as inherently more mature. Its administrative burden may be unnecessary when costs already map clearly to one cost center. Finance and technology stakeholders should choose an approach that fits their organization’s accounting and accountability needs.

  41. FinOps Foundation: Forecasting

    Forecasting establishes an agreed expectation of future technology spending and value using historical spending, planned changes and lifecycle assumptions. Engineering, product, finance and leadership collaborate on the model. Forecasts inform budgets; application budget owners are responsible for meeting the forecast or seeking additional funding for shortfalls. Models document timing, implementation, pricing and total ownership cost.

  42. Forward Deployed Engineering 101

    Keep genuinely customer-specific behavior local, and move generalizable capabilities into the platform over time.

  43. HM Treasury: The Orange Book, October 2004

    A risk owner is responsible for ensuring that an identified risk is managed and monitored and needs sufficient authority to do so; that person need not perform every mitigating action. Residual risk is the exposure remaining after controls are applied. Its tolerability depends on acceptable impact and frequency, and it may require reassessment.

  44. Department for Work and Pensions: Accounting officer system statement, 2025

    DWP describes written delegations of spending authority and responsibilities, alongside separate risk-management and assurance arrangements. Its risk governance assesses the exposure the organization is willing to take to achieve objectives and escalates ineffective mitigation or exposure outside agreed limits. Risks also escalate when they cannot be managed at the current level, affect multiple areas, or require authority beyond the risk owner's control.

  45. NIST AI Risk Management Framework: Core

    NIST assigns executive leadership responsibility for AI development and deployment risk decisions, with documented responsibilities, communication and training for personnel and partners. Proceeding with development or deployment requires determining whether intended purposes and objectives are achieved. Risk responses include mitigation, transfer, avoidance and acceptance, with residual risks documented. Continuing responsibilities include monitoring third-party resources and assigning authority to disengage or deactivate systems whose outcomes conflict with intended use. Post-deployment plans include feedback, appeal, override, incident response, recovery and change management.

  46. Intercom: AI is approving our pull requests

    Intercom's engineering and security authors describe responding to a review bottleneck with a bounded automated-approval pilot. Their system rejects changes considered too large or broad, allows engineers to request human review, and records review comments, approval, tests and merge events. Engineers retain responsibility for observing their changes in production and rolling them back when necessary. The authors acknowledge that review cannot catch all infrastructure, usage-pattern or third-party failures.

  47. Department for Science, Innovation and Technology: Introduction to AI Assurance

    The guidance describes AI assurance as measuring, evaluating and communicating whether an AI system is trustworthy. Appropriate techniques examine different properties and can operate throughout the lifecycle. Assurance requires access to relevant information about the system and its management, and communicates limitations, risks and mitigations to decision makers. Its purpose is to support justified confidence and informed decisions, rather than substitute a generic approval label for examination.

  48. The Institute of Internal Auditors: The Three Lines Model

    The model separates management’s delivery and risk-management responsibilities from independent internal audit. Independence requires accountability to the governing body, access to necessary people, resources and information, and freedom from interference. The governing body resources the audit plan and permits private access without management present. Specialist risk teams may support and challenge delivery while remaining part of management. The lines distinguish concurrent roles, not three mandatory departments or sequential approval gates.

  49. Moving away from Agile: What's Next?

    Use the talk's proposed MECE measurement framework to connect inputs, operational outputs, developer experience, quality, and economic outcomes.

  50. Vision: Zero Bugs

    Ask the LLM for explicit risk analysis and safety cases that connect possible failures to mitigations in the code.

  51. Vision: Zero Bugs

    Use separate prompts for implementation and testing, with different foundation models as an optional further separation.

  52. Forward Deployed Engineering 101

    The speaker defines FDE as an enterprise-scale design partnership built on shared platform primitives, rather than independent software projects for every customer.

  53. CNCF Platforms White Paper

    CNCF describes platforms as capabilities and interfaces serving internal users such as application developers and data scientists. It recommends prioritizing common needs across teams, allowing teams to operate capabilities outside the platform when necessary, and gathering user feedback to guide investment. Platform teams may own the user-facing interface without implementing or operating every underlying service.

  54. How Forward Deployed Engineering is done at Cognition

    Cognition frames forward deployed engineering as expanding product-market fit through both customer delivery and feedback that changes the product.

  55. Chris Argyris: Double-Loop Learning, Teaching, and Research

    Argyris distinguishes correcting errors while retaining governing values from correcting errors by changing those values and then changing actions. The latter is double-loop learning: reconsidering the rules underlying action, rather than only improving execution within them.

  56. Barry M. Staw: Knee-Deep in the Big Muddy—A Study of Escalating Commitment to a Chosen Course of Action

    Staw’s 1976 paper tested whether responsibility for a previous investment changes subsequent funding after adverse results. In a role-playing experiment with 240 business students, participants either made the initial allocation or inherited another executive’s decision, then received favorable or unfavorable performance information. Commitment to the previously selected division was greatest when participants were responsible for the earlier decision and received negative consequences. The result illustrates escalation of commitment: increasing resources devoted to an existing course despite adverse feedback.

  57. Amy Edmondson: Psychological Safety and Learning Behavior in Work Teams

    Edmondson’s 1999 study defines team psychological safety as a shared belief that interpersonal risk-taking is safe. Across 51 teams in a manufacturing company, psychological safety was associated with learning behavior, which was associated with team performance. The study jointly considers interpersonal beliefs and structural conditions such as contextual support and leader coaching, rather than treating either organizational structure or attitudes as a sufficient explanation.

  58. KCS v6 Practices Guide: Summary

    The KCS authors recommend setting goals for outcomes while using activity counts to understand trends. They describe knowledge management as organizational change involving values, interactions and processes, with executive funding, sustained communication and supported coaching. Performance assessment should consider value created by individuals and teams through both qualitative and quantitative evidence.

  59. OECD: The impact of AI on the workplace

    In OECD's 2022 finance and manufacturing surveys, workers at organizations consulting workers or their representatives were more likely to report improved performance and working conditions from AI. Skills and training were the most commonly discussed consultation topic; job loss and wages were discussed less often. Employers reported that consultation could produce changes to AI strategies, guidelines or collective agreements. The report distinguishes task automation from overall employment change.

  60. Everything we knew about software has changed — Theo Browne

    Do not accept an unsuitable implementation merely because substantial effort has already been invested in it.

  61. Government Digital Service: Retiring Your Service

    Retirement requires deciding how the service’s existing user needs will be met, sharing knowledge with replacement teams, explaining changes and giving API consumers time to adapt. The guidance also requires continued support for protecting retained information and managing any transfer to a new owner. Ending active delivery therefore leaves transition and information-management work that must still be supported.

  62. How Forward Deployed Engineering is done at Factory

    Missions is described as a long-running harness for bounded work whose completion can be verified, with human involvement concentrated in planning.

  63. Missing pieces of workflow automation

    Evaluate the whole business process rather than merely inserting agents into its existing tasks.

  64. The Rise of Open Models in the Enterprise

    The speaker reports that switching frontier providers is feasible, but still requires renewed evaluations and prompt tuning.

  65. Your Fine-Tuned Model Is Tech Debt: A 50x ROI House of Cards

    The speaker's 'calcification tax' describes maintenance complexity that made both model switching and architectural change too costly.

  66. The engineer of the future is the person who is able to choose what is worth doing — Addy Osmani

    High agency means owning outcomes with judgment, including deciding that a possible task is not worth pursuing.