Contents
  1. Purpose and development
    1. Assistance beyond one request
    2. Learning habits and sharing initiative
  2. Understanding the person
    1. Preferences depend on the situation
    2. Learn preferences without overgeneralizing
  3. Entrusting work
    1. Define standing delegation
    2. Ask for the right decision
  4. Applications and time
    1. Preserve meaning across accounts
    2. Reconsider work when it resumes
  5. Initiative and attention
    1. Act, ask, notify, or stay quiet
  6. Ongoing control
    1. Make responsibilities inspectable
  7. Privacy and changed authority
    1. Keep information within its purpose
    2. Withdraw authority and leave cleanly
  8. Evidence over time
    1. Test changing circumstances
    2. Measure sustained usefulness
  9. Check understanding
  10. Open questions
  11. Selected talks
  12. References
  13. Talk library
← All topics

Personal Agents

A personal agent should reduce the effort of managing everyday responsibilities: explaining the same preferences, carrying information between applications, remembering unfinished work, and deciding what needs attention. That requires more than successful individual requests. The assistant must keep its understanding revisable, its commitments visible, and its authority bounded as circumstances change.

Purpose and development

Assistance beyond one request

An agent is software that chooses actions using observations. A personal agent organizes those choices around one person's continuing work. It may prepare a meeting, retain an unresolved scheduling decision, and return to it when new information arrives. The underlying observation–action loop is general; the personal-agent problem is maintaining a useful relationship across those separate occasions.

Separate three responsibilities. Personal understanding includes relevant facts and preferences. Unfinished work records commitments that still need attention. Authority determines which operations the assistant may perform. Remembering that someone prefers afternoon meetings does not mean a meeting remains necessary, and neither fact permits sending an invitation.

Earlier assistant research already distinguished managing commitments and workload from planning and executing tasks. Karen Myers and colleagues' PExA paper, published in 2007, joined these responsibilities with monitoring and explanation. This remains a useful design distinction: keeping track of work and performing it are related but different jobs.

Continuity does not require uninterrupted computation. Retained state and a recoverable trigger can let work resume after the worker has stopped; runtime scheduling supplies that machinery. Nor does a friendly personality establish continuity. The useful comparison is between work removed and work introduced: less repeated explanation and coordination, weighed against setup, review, correction, and recovery. A short field study of goal assistance found that duplicated logging and incorrect context could themselves create effort.

Learning habits and sharing initiative

Personal-assistant research predates modern LLM-based assistants. Three recurring problems explain its development: learning an individual's habits, deciding when help is welcome, and coordinating responsibilities that span tasks. Broader language interfaces extend the available work without resolving those problems automatically.

Learning habits, sharing initiative, carrying out work

  1. July 1994Calendar ApprenticeLearned scheduling regularities from ordinary calendar use while allowing people to override suggestions.Sources & context

    Contributors: Tom Mitchell, Rich Caruana, Dayne Freitag, John McDermott, and David Zabowski.

    What changed: Experience With a Learning Personal Assistant described suggestions for appointment fields rather than manually maintained customization rules. Corrections supplied learning examples; the authors reported approximately five user-years of experience among a handful of users.

  2. 1999LookOutCombined email-to-calendar assistance with editable proposals and decisions about when to help, ask, or defer.Sources & context

    Contributors: Eric Horvitz.

    What changed: Principles of Mixed-Initiative User Interfaces considered uncertain intent, attention, intervention benefits, and interruption costs. LookOut demonstrated how automated contributions could coexist with direct user control.

  3. 2003CALOBegan integrating learning, reasoning, and information management across interrelated office responsibilities.Sources & context

    Contributors: SRI-led collaboration within DARPA’s Personalized Assistant that Learns program.

    What changed: The five-year Cognitive Assistant that Learns and Organizes effort covered task management, scheduling commitments, preparing information products, and coordinating resources. This broadened the design problem beyond individual recommendations.

  4. October 2011SiriEntered the consumer phone as an integrated iPhone 4S feature.Sources & context

    Contributors: Apple; Siri, Inc., with research origins attributed by SRI to CALO and joint work with EPFL.

    What changed: Product integration followed separate commercialization events: SRI spun off Siri, Inc. in 2007, and Apple acquired it in April 2010. The October 2011 milestone marks phone integration, not the beginning of the research.

  5. February 25, 2026Gemini on AndroidAnnounced a limited preview of user-requested background tasks with monitoring, takeover, and stop controls.Sources & context

    Contributors: Google.

    What changed: The announcement described multi-step execution in selected applications, initially focused on food, grocery, and rideshare tasks. It presented a supervision design and rollout plan, not measured completion reliability.

Learning habits, sharing initiative and coordinating responsibilities remain relevant as assistance reaches phones and background execution. Chronological spacing is not to scale.

Mixed initiative means that either the person or the system can initiate a contribution. Learning a habit, offering help and carrying out a delegated task each require different decisions about when the assistant should act.

These contributions are complementary. A fixed rule can enforce a limit, a learned preference can rank acceptable choices, and a language interface can express a new task. None removes the need for the others. The broader history of shared-control interfaces develops this coexistence; here, the next step is deciding what the assistant should learn about one person.

Understanding the person

Preferences depend on the situation

Personal context consists of facts and circumstances relevant to the present task. User modeling represents that information so it can influence choices. Facts describe the situation; hard constraints exclude unacceptable options; preferences rank the remaining choices. A preference is therefore a conditional default, not a prohibition or permission to act. PTIME, a personalized calendaring assistant, illustrates why the conditions matter: disliking early meetings can coexist with accepting one requested by a manager.

Consider these example scheduling statements. Their consequences differ even though all belong to the same task.
StatementRoleConsequence
The meeting lasts 30 minutes.Fact about the taskEvaluate slots of sufficient length.
Do not overlap an existing commitment.Hard constraintExclude conflicting slots.
I generally prefer afternoons.Standing preferenceRank suitable afternoon slots higher.
For this meeting, find a morning slot.Current instructionApply the local exception without deleting the general preference.
Prepare options, but do not send invitations.Delegation boundaryPermit preparation, not communication.

Applicability depends on task, role, audience, and time. A concise engineering summary and a detailed personal journal can both suit the same person. Retained preferences and prior decisions are memory: information supplied to later interactions. Personalization through memory explains why a local instruction can override a default without changing the person's enduring profile.

Even apparently factual profiles require care. In a reported travel-memory error, discussion of Thailand and Turkey as alternatives became a claim that both trips occurred, although only Thailand was visited. Considering, choosing, planning, and completing are different statuses. Compressing them into one biographical sentence can make later personalization confidently wrong.

Learn preferences without overgeneralizing

Preference elicitation obtains information that distinguishes the available choices. Start with the pending decision rather than an exhaustive questionnaire. PTIME combined initial elicitation with refinement from subsequent scheduling choices. This permits useful assistance before a detailed profile exists, while leaving room for contextual exceptions.

Explicit feedback states a judgment or instruction. Implicit feedback is behavior interpreted as a signal, such as editing a suggestion. Calendar Apprentice's published example changed a proposed 60-minute meeting to 30 minutes and used the interaction as a learning example. But an observed correction does not, by itself, explain its future scope: this meeting may be unusual, or the person may want a new default.

Make that scope controllable. Correcting this result should be distinct from applying a preference to future work of this kind. Preserve who supplied the instruction and its circumstances; when two instructions genuinely conflict, clarify rather than applying an automatic newest-wins rule. The 2019 Guidelines for Human-AI Interaction recommend granular feedback and explaining its effect on future behavior. Memory's correction mechanisms carry an accepted change into later use.

Acceptance is ambiguous too. Someone may choose the first adequate option because searching is costly. Chaney, Stewart, and Engelhardt's 2018 recommendation-system simulations showed how learning from behavior already influenced by recommendations could make consumption more homogeneous without improving utility. For assistants, the implication is to retain what was offered and under which circumstances—not treat every acceptance as an enduring preference. Silence supplies even less information.

Preference drift is a genuine change in what someone wants over time; it must be distinguished from a temporary exception or an earlier mistaken inference. Retain scoped preferences and accepted decisions when they have a concrete future use, following memory admission. Personalization need not change model parameters: retrieving a preference and placing it in the next request changes the input to a fixed model. Prompting explains that boundary.

Entrusting work

Define standing delegation

Delegation entrusts specified work and bounded discretion to an agent acting on someone's behalf. Standing delegation covers a class of future actions while stated conditions continue to hold. A desired outcome is not a complete delegation: arranging a meeting can involve reading availability, preparing options, changing records, and communicating commitments, each with different consequences.

OAuth lets a resource owner delegate access to a client application without sharing the owner's password. An authorization server issues an access token, a credential the client presents to the resource server that holds the protected information. A scope describes requested or granted access; the grant can be narrower than the request. Logging into an assistant establishes a session, but connecting a calendar or mailbox requires separate permission to access that service. Neither step alone specifies every action the person intends.

An application-level delegation contract should make the following dimensions explicit.
DimensionWhat to specify
Representation and purposeThe person represented and the responsibility being delegated.
Accounts and resourcesThe permitted account, organizational boundary, calendar, mailbox, or other resource.
OperationsSeparate reading, preparing, modifying, communicating, and committing resources.
LimitsAllowed recipients, amounts, action counts, and aggregate limits across tasks.
LifetimeActivation, expiry, renewal, and withdrawal conditions.
ReportingWhat must be reported, where it may be delivered, and when intervention is required.

Classify preparation by its actual effect. Composing text privately differs from saving it into a connected mailbox: a Gmail draft is already a stored resource. Sending then deletes the draft and creates a distinct sent message. Keep the prepared artifact, reviewed content, and observed external result separate.

Limits must hold together at execution. Least privilege grants only necessary powers; complete mediation checks access at the protected operation. Preferences can shape a proposal but cannot bypass this boundary. See action enforcement. Cumulative limits also require shared accounting: concurrent tasks must not each consume the same remaining allowance. Repeated success or past approvals never silently enlarge the grant; accumulated effects remain a separate responsibility.

Ask for the right decision

Human involvement can resolve different missing conditions. Clarification establishes what the person means. Consent is informed agreement to a specified activity. Confirmation approves a concrete proposed action. Asking which calendar to use clarifies intent; agreeing to an ongoing monitoring activity establishes its intended scope; reviewing an invitation confirms particular recipients and content. These decisions should not be compressed into one generic approval button.

Existing delegation can permit routine work without another prompt. Ask when material intent is uncertain, authority is missing, or the agreed review boundary is reached. Consider recipients, consequences, and reversibility. An application confirmation does not establish every permission needed for personal-data processing; Privacy and Data Governance distinguishes those layers. Asking constantly is not a substitute: repeated prompts can encourage reflexive approval.

Bind confirmation to the reviewed action. For example, approval to send a draft from the work account to its named colleague does not cover sending the same text to an external mailing list. Material changes require renewed review; denial, expiry, and silence do not approve execution. The final check must compare the action being executed with the authorized details. Approval binding supplies the execution contract; review design makes those details inspectable.

Duration deserves particular attention. Malkin, Wagner, and Egelman's SOUPS 2022 study involved 23 participant pairs using a simulated proactive voice assistant. Some misunderstood permission scope or assumed continuing rules would expire; only a minority used the review feature during the sessions. This supports making duration and review discoverable, but understandable controls still need independent enforcement.

Applications and time

Preserve meaning across accounts

Cross-application work must preserve both whom the assistant represents and which account it uses. The represented person is the person on whose behalf the work is done; the acting agent is the software performing it; the authenticated account is the identity used to access a particular service. One person can have separate work and personal accounts. A tenant is an organizational or customer boundary within a service, so choosing an account also requires preserving the intended organizational context. Identity foundations develops these distinctions. OAuth token exchange can represent the subject and actor separately, but the application must still select and check the intended resources.

Carry a task's meaning through each handoff: its source references, selected account and destination, intended operation, applicable delegation, and resulting object references. A system of record is the application authoritative for a particular object. A task record can connect an email and calendar event without replacing either application's authoritative state.

For email-to-calendar assistance, preserve the source message reference, selected calendar, attendee addresses, time zone, and notification choice. Google's Events.insert treats calendar selection, attendees, notification behavior, and response status separately. Creating an event does not establish attendee acceptance. A handoff should retain the returned event reference rather than reporting an undifferentiated scheduling success.

One task connects distinct application records

One task connects distinct application recordsA mailbox source message supplies a reference to the retained task. Represented person and acting agent stay distinct from service accounts. Task and selected calendar context enter validation. A permitted successful write creates a separate event; unresolved target or authority means no write. Attendee response is separately observed and connected to that event.Mailbox · source authorityCalendar · destination authoritySource messageSource account + message IDReference retained, not replacedRetained taskRepresented person ≠ acting agentGoal + source message referenceAttendees, time zone, notification choiceService account is a separate identitySelected destinationAccount + tenant where applicableCalendar IDValidate target + current authorityProposed invitation and exact targetCreated eventNew event ID + notification choiceNo calendar writeAttendee responseAccepted, declined, tentative, unknownReferenceProposed invitationTarget contextPermitted write succeedsFailed or unresolved checkObserved separatelyEvent creation ≠ attendee acceptance · a notification choice does not guarantee delivery
An example application design retains the source message and destination event as separate authoritative records. The represented person, acting agent and service account have different roles. Event creation does not establish attendee acceptance or notification delivery.

The outgoing identity can differ from the login too. Gmail's send-as resource distinguishes the primary login address, outgoing From address, default sending address, and reply-to address. Showing the right account name therefore does not establish the right sender. Verify the identity that the operation actually uses, especially when crossing work and personal contexts.

Reconsider work when it resumes

A continuing commitment needs an addressable record: its goal, current disposition, pending dependency, completion condition, and next review point. It should be possible to distinguish waiting for information from waiting for permission. PExA connected task execution with calendar commitments and monitoring; the same separation helps an assistant preserve unfinished work without treating every remembered intention as an active assignment.

A runtime preserves and resumes execution. A checkpoint records a recovery boundary, not renewed permission. At resumption, reconsider whether the goal is still wanted, whether the person already completed it elsewhere, whether the supporting information changed, and whether authority still applies. Retire or redirect the affected task while preserving unrelated commitments. Directing work across long waits explains the underlying control machinery.

Consider an assistant waiting before changing an existing calendar event. The person edits the event directly during the wait. An ETag identifies a resource version; Google's conditional modification can reject an update with HTTP 412 when its If-Match version is stale. The assistant should inspect the current state and reconsider the task, not simply overwrite the human edit. This detects a conflict; it does not interpret changed intent or make an availability check and later event insertion atomic.

A stale conditional update preserves the newer event

A stale conditional update preserves the newer eventTime advances down four lanes. T observes existing event E with ETag v1 and waits. A person edits E to v2. Current authority permits this update attempt; T sends If-Match v1, receiving412 while E remains v2. A separately permitted fresh read returns current E. T reconsiders; no automatic retry.Task TPersonCalendar serviceCurrent authorityObserve existing E · ETag v1Person edits E → ETag v2Current authority permits this attempt; otherwise withholdModify E · If-Match: v1412 Precondition Failed · E unchanged at v2Fresh read of E, under current read permissionCurrent E · ETag v2 → reconsider TT retains goal + old observationSame event E; changed ETagCompare v1 with current v2v1/v2 are example ETags · refresh does not renew permission, infer intent or automatically retry
Task T retains E’s earlier observation while the person edits E. The authorized stale update receives 412 and leaves E at v2. A permitted fresh read supports reconsideration under current authority; the conflict does not reveal the person’s intent.

Longer lifetimes make this distinction increasingly practical. Google's August 26, 2026 Gemini Live announcement described handing scheduled work across applications to Spark, including work lasting days or weeks. Such continuity increases the occasions for changed circumstances. The announcement does not establish a complete changed-intent or revocation contract; those remain requirements of the responsibility being delegated.

Initiative and attention

Act, ask, notify, or stay quiet

Proactive assistance initiates a suggestion or permitted work when an opportunity arises, rather than waiting for a fresh request. An event is a reason to evaluate an intervention, not automatically to interrupt. First establish that observation and the proposed action are permitted. Then consider relevance, uncertainty, urgency, reversibility, attention cost, and the cost of waiting. Horvitz's mixed-initiative principles explicitly included doing nothing and deferring assistance alongside asking and helping.

Adjustable autonomy allocates individual decisions between the system and a person. The 2002 transfer-of-control framework considers decision quality, timely human response, waiting, and intermediate-action costs. It rejects a universal all-or-nothing choice of autonomy. Apply this reasoning only among permitted alternatives: expected usefulness cannot create authority that the assistant lacks.

For a routine scheduling update, the response changes with the conditions. These are example policy choices, not universal thresholds.
ConditionAppropriate responseReason
Relevant update; preparation is authorized; no immediate decision needed.Prepare options.Useful progress does not require interruption.
The destination calendar is ambiguous.Ask a targeted clarification.The answer changes the pending operation.
A permitted change completed; the reporting policy allows a later digest.Queue a notification.Completion and immediate attention are different needs.
Nothing changed since a declined suggestion.Remain quiet.Repetition adds no useful information.
A deadline approaches, but the required permission is absent.Use the agreed escalation path; withhold commitment.Urgency does not supply authorization.

Buy time without expanding authority

A fictional room hold awaits confirmation of a revised invitation. Finalizing would send it. Example policy: wait for scheduled review when it precedes expiry; otherwise use one permitted free 120-minute extension if it preserves that opportunity. If neither works, use the agreed early escalation path. Times are minutes from now; extension execution is assumed immediate and successful.

Free extension available?
Standing delegation permits extension?
Extend once, then wait for scheduled review.
The permitted extension moves expiry from +15 to +135 min, beyond review at +75 min.
Finalization withheld until confirmation of the revised invitation. Silence and urgency supply no approval.
Opportunity preserved by timing, within existing authorityOriginal expiry 15 minutes, review 75 minutes, effective expiry 135 minutes. Selected action: extend. Finalization remains withheld.Now+60+120+180Original expiry+15 minScheduled review+75 minEffective expiry+135 minMinutes from now · review must precede expiry, including after extension

A review exactly at expiry is too late in this example. Changing the timing does not confirm the revised content or authorize sending. This is a stated policy, not a universal optimum.

The extension changes when a decision is needed, not who may make it. Confirmation of the revised invitation remains required whether the assistant waits, extends or escalates.

Nonresponse and repetition are separate hazards. The Electric Elves failure report describes unwanted decisions after a five-minute human-response timeout. Another assistant delayed a meeting almost fifty times in five-minute increments: locally repeated choices ignored cumulative nuisance. Neither incident supplies a modern failure rate, but both explain why a policy needs bounded follow-up rather than treating silence as assent or each reminder as an isolated event.

Specify quiet periods, digest delivery, duplicate suppression, suggestion expiry, and a limit on unresolved escalations. These controls need separate meanings. OpenClaw's heartbeat documentation separates periodic cadence, active hours, context, and delivery destination from tool policy. Disabling periodic turns does not disable every event-driven wake. Cadence controls when the assistant considers work; authorization controls what it can do.

Do not assume fewer notifications always means better assistance. A two-week smartphone experiment found better reported attention and perceived productivity with three daily batches than ordinary delivery, while eliminating notifications increased reported anxiety and fear of missing out. This was not an agent-workflow trial or a universal optimal schedule. Workflow-native suggestions, such as Tegon's optional assistance, offer another pattern: make help available where the person is already working.

Ongoing control

Make responsibilities inspectable

Oversight should expose responsibilities, not require watching every tool call. For each standing responsibility, show the represented account, purpose, limits, expiry, active tasks, waiting decisions, last verified outcome, and next expected intervention. Make the retained preferences behind a decision inspectable too. Alma's editable personal context illustrates user-visible memory controls, although its demonstration does not establish downstream deletion guarantees.

Controls should name their scope. Correction changes an erroneous result or understanding. Steering redirects a task. Pause retains work for possible resumption. Cancellation requests that its remaining work stop. Withdrawal removes a standing delegation that could create future tasks. Cancelling today's preparation should not silently cancel every future responsibility, while withdrawing the responsibility should not leave its future triggers active. The interface chapter develops persistent controls and honest stopping semantics.

Visibility must survive the original interaction. A person who starts work on a laptop should be able to inspect and redirect the same task later from another authorized device. The durable-session account explains why a bidirectional connection alone is insufficient: a second client needs access to shared session state, not merely a new connection unrelated to the running task.

An action receipt records an identified operation and its observed outcome. If cancellation races with an external write, the outcome can remain unknown; cancellation does not roll back completed changes. Preserve the operation reference and reconcile against the receiving application before repeating it. A receipt should distinguish requested stopping, confirmed stopping, completed effects, and unresolved status.

Compensation performs a new action to address a completed effect, such as correcting an already-sent notice. It is not erasure of history, can require fresh authority, and can itself fail. A usable control surface therefore preserves what happened and the remaining remedy instead of replacing everything with “cancelled.” Keep receipts concise and protect sensitive supporting content rather than copying entire messages into routine logs.

Privacy and changed authority

Keep information within its purpose

Privacy concerns appropriate use, not only protection against outsiders. Helen Nissenbaum's Privacy as Contextual Integrity, published in 2004, distinguishes what information belongs in a social setting from how it should flow between people. Applied to assistants, access to work and household sources does not automatically justify combining or redistributing them. A scheduling response can disclose availability without disclosing the private reason someone is unavailable.

Data minimization limits information to what the permitted purpose needs. Google's Freebusy.query accepts selected calendars and a bounded interval, returning busy intervals without event titles, descriptions, or attendee lists. That can support availability-only scheduling without retrieving private explanations. Calendar-specific errors must remain unavailable coverage, not become free time. The query neither reserves a slot nor grants permission to disclose its result.

This conceptual example requests busy intervals directly; private event titles remain source context. An unavailable calendar does not establish free time. Reading availability neither reserves a slot nor grants permission to disclose it.

Recipients include more than an addressed email recipient. A notification may appear on a lock screen or in a shared conversation. Correspondents and household members also have interests in records the account owner can access. Review the output's purpose, audience, and destination, including inferences that reveal more than any individual source. Derived disclosure controls explain that broader boundary.

Collecting less reduces available detail; retaining a scoped summary can reduce repeated access but still preserve sensitive claims. Selected local processing can avoid some external transfers, yet local storage does not imply local inference or appropriate disclosure. Kitze's local-file migration account expresses a preference for ownership and direct access, not a complete privacy guarantee. Trace the actual local and cloud data paths, then specify permitted uses, readers, and retention for each retained representation.

Withdraw authority and leave cleanly

Ending assistance is not one delete operation. Work, credentials, delegation, and personal records have different lifetimes. The application should translate the person's request into explicit changes across those objects, then report what actually completed. Changed authority and memory forgetting provide the underlying distinctions.

Use this mapping as an offboarding contract, not an assumption that one provider performs every change automatically.
RequestAffected scopeWhat does not follow automatically
Pause or cancel this task.Covered pending work and active operations.Completed effects are not reversed; unrelated tasks remain.
Disconnect this account.Credentials and future access through that connection.Retained information and already-started work are not necessarily removed.
Withdraw this responsibility.Future triggers, task creation, pending work, and approvals relying on that delegation.A separate enrollment path must not recreate the same authority.
Stop using this information; correct or delete it.Specified uses or retained records and their derived copies.These requests are not interchangeable and need distinct completion evidence.

Token revocation invalidates an access credential. RFC 7009 acknowledges propagation delay and policy-dependent treatment of related tokens. RFC 8693 does not generally create automatic revocation linkage between an exchanged token and its input. Test downstream credentials separately rather than assuming one successful revocation closes every path.

Replacement identities create another path. In the Agent Auth workshop, revoking one agent did not prevent later reads: the host created a replacement with default read capabilities. That observation explains the limit of individual-agent revocation; it does not prove that old tokens remained valid. Withdrawal must govern new enrollment as well as existing identities. Similarly, inspect event subscriptions as well as periodic schedules.

Withdrawal must govern replacement enrollment

Withdrawal must govern replacement enrollmentReported path: old agent A is revoked; the same host H enrolls distinct B with a fresh read grant, then B reads the mailbox. Required path: H checks the withdrawn responsibility before replacement enrollment or grant. Covered authority is withheld and there is no mailbox read through that path. This is not a demonstrated repair.Reported workshop behaviorHost HOld agent A revokedHost can enroll replacementsDistinct agent BNew identity + fresh read grantNo successful read from old AMailboxRead succeeds using BOld-token validity unprovenRequired application behavior · not a demonstrated repairSame host HNew replacement attemptBefore enrollment or grantCheck current responsibilityResponsibility withdrawnCovered authority withheldSame mailboxNo new grant through this pathNo resulting read through itEnroll + grantRead as BConsultStopped before new authority is issuedExisting grants, queued work and active operations need separate checks; prior effects and records remain.
The workshop described a new identity with a fresh read grant, not a successful read using the revoked identity. The required current-responsibility check belongs before replacement enrollment or grant issuance. It is an application requirement, not a demonstrated repair.

An orderly exit should leave usable task outcomes and appropriate user-controlled records, without exporting credentials or unauthorized third-party material. Report stopped responsibilities, disabled connections, completed external effects, retained records, and unresolved cleanup. Recovered workers and replacement agents must consult current authority rather than restore old grants. A complete offboarding claim requires end-to-end tests across these paths, including pending approvals and already-started operations.

Evidence over time

Test changing circumstances

Longitudinal evaluation assesses behavior across time and repeated interactions. A personal-agent case needs earlier information, an intervening change, a later opportunity, and an observable outcome. Include preserved commitments and appropriate restraint: asking, deferring, or remaining quiet can be the correct behavior. A final task-completion score cannot express all of these requirements.

Example regression cases apply the chapter's control requirements across multiple interactions.
Intervening changeExpected later behaviorObservable check
A morning exception is specified for one meeting.Respect the exception; retain the afternoon default elsewhere.Inspect both the current proposal and a later unrelated proposal.
The person switches the active application account.Preserve or clarify the task's intended account.Inspect the actual account and resource used by the operation.
A suggestion is declined; no material information changes.Do not repeatedly surface the same suggestion.Observe subsequent notification opportunities.
Permission is withdrawn before a queued task resumes.Block covered work, including replacement-identity paths.Check resource access and new task creation after resumption.
One commitment is cancelled.Retain unrelated authorized commitments.Inspect remaining tasks and their later outcomes.

Replay only information available at each decision point. Feeding a later correction into an earlier replay hides the very error being tested. Check authoritative application state rather than accepting the assistant's claim of success, and keep unavailable outcomes unresolved. The methods belong to simulation and replay, incomplete feedback, and memory evaluation.

Haoran Zhang and colleagues' π-Bench, a May 14, 2026 preprint, separates proactive intent resolution from final completeness across 100 multi-turn tasks. Removing preceding sessions reduced proactivity more than completeness in its three-model ablation. But its users were simulated and experiments shared one adapted scaffold. It supports testing continuity, not a claim of months-long human benefit. Extra clarification turns can reduce human burden, so turn count is not that burden.

Measure sustained usefulness

The unit of benefit is the whole person–assistant workflow. Compare continuing assistance with credible alternatives: ordinary application features, fixed reminders, or an on-demand assistant with a small explicit profile. Memory baselines isolate one component; they do not establish whether proactive work and added supervision improve the whole experience.

Keep a measurement ledger whose categories answer different questions.

  • Useful outcomesCompleted commitments, avoided omissions, and results that remain useful after review.
  • Human effortSetup, direct work, repeated explanation, review, correction, duplicated logging, and repair.
  • InitiativeHelpful interventions, unwanted interruptions, and useful opportunities the assistant missed.
  • Control and privacyWrong-account actions, unauthorized commitments, inappropriate disclosures, and failed withdrawal.
  • Continued useChanging needs, reduced use, abandonment, reasons for leaving, and missing follow-up.

Needs can change quickly. Yan Xu and colleagues' From Goals to Actions, reported at CUI 2025, studied 14 participants for two to four weeks around their 2024 resolutions. Initial demand for discovering actions declined after the first week as routines, tracking, and encouragement became more relevant. Incorrect context could reduce relevance without prompting correction. The small, short study does not establish long-term net benefit, but it explains why a successful onboarding experience is an incomplete evaluation.

Similarly, PTIME evaluated preference agreement, reasoning, learning, and perceived usefulness separately. Its 15 participants used invented events in a fake calendar over four weeks. Those results cannot establish real-world time savings. Acceptance and satisfaction are useful observations, but neither substitutes for measured work and consequences.

A within-person comparison observes the same participant under different conditions. It reduces some differences between people, but the periods can still differ in task mix. Novelty can affect early behavior; learned habits can carry into the next condition; benefits and failures may appear only later. People who stop using the assistant must remain in the analysis, with missing follow-up identified rather than counted as success. Live experiment design explains how to choose and interpret the comparison.

Keep privacy and authority failures beside effort outcomes, not inside a score where saved minutes cancel unauthorized actions. Separate unlike outcomes, then use findings to narrow responsibilities, improve context, change notification policy, or remove an unhelpful feature. More autonomous activity is not the objective. Continuing assistance succeeds when the person has less work to manage while retaining meaningful control over what the assistant understands and does.

Open questions

  1. Sustained net benefit remains difficult to establish because needs, task mix, and participation change while setup and repair costs accumulate. Progress would mean months-long comparisons with credible alternatives, complete effort accounting, and follow-up that includes people who leave.

  2. Preference revision must distinguish genuine change from exceptions and choices shaped by the assistant's own suggestions. Progress would combine low-effort correction with tests showing that local edits neither become universal rules nor prevent appropriate adaptation.

  3. End-to-end withdrawal spans credentials, replacement identities, triggers, pending approvals, retained information, and already-started effects. Progress would provide testable cessation contracts and receipts that identify unresolved paths instead of treating one revoked credential as complete offboarding.

  4. Useful initiative must adapt to attention and deadlines without interpreting nonresponse as permission. Progress would measure missed opportunities, repeated nuisance, and human effort together while keeping authorization independent of predicted usefulness.

Follow the curated reading path through the speakers and demonstrations behind this entry.

19 min

AI Engineer World's Fair 2024 · 2024

The Adversarial Path to the Personal Assistant

Sumit Agarwal

Cited in this entry

Concrete personal scheduling and profile examples show why importing data is only the start: uncertainty and follow-up still need management. Distinguish the demonstrated calendar import from the proposed email-monitoring extension.

Watch talk

Explore more talks

The rest of the library, beyond the curated path. Cited talks support this entry; reviewed transcripts were processed in full. Metadata candidates have not been reviewed as sources or verified as topic members.

26 matching talks

TalkSpeakerEventYear
Vinoth GovindarajanAI Engineer World's Fair 20262026
Steve KorshakovAI Engineer World's Fair 20262026
Sarthak AggarwalAI Engineer World's Fair 20262026
Rene BrandelAI Engineer World's Fair 20252025
Damien MurphyAI Engineer World's Fair 20252025
Tobin SouthAI Engineer World's Fair 20252025
Notion's Token Town

Cited in this entry

Sarah SachsAI Engineer World's Fair 20262026
Bennet FennerAI Engineer Europe 20262026
Zhou YuAI Engineer Summit 20252025
Angel Ortmann LeeAI Engineer World's Fair 20262026
Soumith ChintalaAI Engineer Summit 20252025
The End of Apps

Cited in this entry

KitzeAI Engineer Europe 20262026
Rami AlhamadAI Engineer World's Fair 20252025
Shivam VermaAI Engineer World's Fair 20252025
Mehedi HassanAI Engineer Europe 20262026
Shlok KhemaniAI Engineer World's Fair 20262026
Rafal Wilinski, Vitor BaloccoAI Engineer World's Fair 20252025
Michael GrinichAI Engineer World's Fair 20252025
Identity for AI Agents

Cited in this entry

AI Engineer Code 20252025
Jared HansonAI Engineer World's Fair 20252025
Bobby Tiernay, Kam SweenAI Engineer World's Fair 20252025
Codex, Behind the Harness

Transcript reviewed

Dominik KundelAI Engineer World's Fair 20262026
Ravi MadabhushiAI Engineer World's Fair 20262026
Nick Nisi, Lizzie SiegleAI Engineer World's Fair 20252025
Filip KozeraAI Engineer World's Fair 20252025
Sam BhagwatAI Engineer World's Fair 20262026

References

Coverage and source review
Processed transcripts
30 processed in full · 5 in the curated path
Automated source review
Passed
Metadata candidates
1 unreviewed; not verified topic membership
Corpus version
1bd8e407b26a07b33815594e1b2db5f41827119a2b3cb6fbf240f9fc571fc767

Automated review checks source support; it is not publication approval.

A synthesis of selected conference talks and technical references. Citations link to the source material; they do not imply that every talk on this subject is included.

  1. Memory overview — LangChain

    Thread-scoped memory holds an ongoing conversation and other execution state; cross-thread memory stores information such as user preferences or facts in application-defined namespaces. LangGraph uses checkpoints for the former and stores for the latter. Memory may be updated during the request path or asynchronously, with different latency and freshness tradeoffs. A single profile is easy to retrieve but harder to update safely as it grows; a collection of smaller records shifts complexity toward retrieval, consolidation, and deletion. Semantic memory means stored facts, not the similarity-search algorithm used to retrieve them.

  2. OWASP Access Control

    Authentication establishes identity; authorization decides which actions that identity may perform on particular resources. A user allowed to initiate a transfer must still be authorized for the source account. Least privilege limits the authority of running code and service accounts, while centralized checks reduce inconsistent enforcement. In an AI application, tool availability and a model-produced argument are therefore insufficient grounds to execute a business operation; the application must apply resource- and action-level policy.

  3. An Intelligent Personal Assistant for Task and Time Management

    Karen Myers and colleagues' 2007 PExA paper describes a CALO component delivered in 2006 to reduce office task overload. Time management concerns commitments, reminders and workload; task management concerns planning, execution and oversight. Delegation means the user chooses tasks to allocate and bounds the assistant's autonomy; the assistant requests missing information and confirms important decisions. Proactive assistance means initiating communication about problems, commitments or useful feedback. PExA integrates task and calendar management with procedure learning, execution monitoring and explanation.

  4. Persistence — LangGraph

    LangGraph distinguishes checkpoints that persist a thread's graph state from stores containing application-defined data shared across threads. Checkpoints support conversation continuity, interruption, and failure recovery; a store serves facts or preferences outside the current graph state. An in-memory saver loses its checkpoints when the process restarts, so process recovery requires a persistent backend. Checkpoint accumulation also needs retention management. These are runtime persistence decisions, separate from how much conversation text is included in a model call.

  5. From Goals to Actions: Designing Context-aware LLM Chatbots for New Year's Resolutions

    Yan Xu, Brennan Jones, Hannah Nguyen, Qisheng Li and Stefan Scherer studied a context-aware chatbot with 14 participants for two to four weeks around their 2024 resolutions, reporting at CUI 2025. Personalized suggestions helped participants discover concrete actions, but demand for discovery declined after the first week while needs for routines, progress tracking and encouragement emerged. Missing or incorrect context reduced relevance; users sometimes avoided correcting it. Redundant logging across existing tools also imposed effort. The prototype combined conversation with inspectable action recommendations and editable context.

  6. Principles of Mixed-Initiative User Interfaces

    Mixed-initiative assistance combines automated contributions with direct user control. Horvitz recommends considering uncertain intent, attention, intervention benefits, interruption costs, and the possibility of deferring assistance. Clarifying dialogue also has a cost; users need efficient invocation, termination, and correction. LookOut illustrates this through email-to-calendar assistance: it interprets meeting information, displays proposed appointment fields, and lets the person edit and save them. When an exact time cannot be inferred, it can instead show a relevant calendar interval. Its decision alternatives include doing nothing, asking, and providing assistance.

  7. PTIME: Personalized Assistance for Calendaring

    PTIME combines initial preference elicitation with refinement from scheduling choices and presents ranked alternatives under meeting constraints. Preferences concern favored tradeoffs: its example contrasts shortening an afternoon meeting with preserving its duration in the morning. Context matters: disliking early meetings can coexist with accepting them from a manager. Its evaluation separates preference-model agreement, reasoning performance, learning, and perceived usefulness. Fifteen CALO participants scheduled meetings over four weeks, but invented events for a fake calendar. Selected options need not be uniquely preferred; tie feedback was not logged. Interviews reported both perceived usefulness and dissatisfaction with prototype speed and stability.

  8. Feedback Loops are All You Need

    Useful meeting outputs depend on role-specific needs: sales may need deal-focused summaries, while engineering may need action items, blockers, and tickets.

  9. Lessons from Studying Every Memory System

    A synthesized profile can incorrectly turn discussion of alternatives into claims that both events occurred.

  10. Experience With a Learning Personal Assistant

    Tom Mitchell, Rich Caruana, Dayne Freitag, John McDermott and David Zabowski described CAP, the Calendar Apprentice, in July 1994. Rather than requiring users to maintain customization rules, CAP learned scheduling regularities from routine calendar use. It suggested meeting duration, location, time and date while allowing overrides. Its published example shows a user replacing a suggested 60-minute duration with 30 minutes; the interaction becomes a learning example. The authors reported approximately five user-years of accumulated experience among a handful of users.

  11. Guidelines for Human-AI Interaction

    The guidelines recommend making capabilities and likely errors understandable, timing assistance to context, supporting dismissal and correction, and explaining system behavior. For adaptation over time, they recommend limiting disruptive changes, accepting granular preference feedback, explaining how feedback affects future behavior, and providing controls over monitoring and operation. Application implication: a personal assistant should make the future scope of a correction visible instead of silently treating every edit as a universal preference change.

  12. How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility

    Allison Chaney, Brandon Stewart and Barbara Engelhardt studied feedback loops in which a recommender learns from behavior already influenced by its previous recommendations. Their simulations distinguish underlying user preferences from choices affected by what was shown and its ranking. Repeated training on this influenced behavior could make consumption more homogeneous without improving utility. For personal assistants, this supports treating acceptance of a surfaced suggestion as an observation made under particular exposure conditions, not an unqualified statement of enduring preference.

  13. Developing Taste in Coding Agents: Applied Meta Neuro-Symbolic RL — Ahmad Awais, Command Code

    The speaker describes reflective context engineering as a mechanism for updating learned preferences when observed behavior changes.

  14. Training language models to follow instructions with human feedback

    InstructGPT begins with a pretrained language model. Supervised fine-tuning updates it using demonstrations of desired responses to prompts. Preference training then fits a reward model to human comparisons of candidate responses; PPO updates the response policy to increase predicted reward, with a penalty for departing from the supervised policy. These stages optimize learned parameters using datasets and objectives. Supplying an instruction or tool observation during ordinary inference instead changes the current input to that trained policy. A favorable preference score represents the learned comparison objective, not a proof of factual or program correctness.

  15. You Didn't Ship a Bug. You Just Wrote It for a Human.

    Start agents with least privilege, limit permission duration, and require just-in-time authorization for elevated scopes.

  16. RFC 6749 — OAuth 2.0 Authorization Framework

    OAuth separates the resource owner, the client requesting access on that owner's behalf, the resource server accepting access tokens, and the authorization server issuing them after authentication and authorization. Its example delegates access to protected photographs without giving the client the owner's password. Scope describes requested or granted access using authorization-server-defined strings. The authorization server may grant less access than requested and must report a differing granted scope.

  17. Identity for AI Agents

    Authentication to the agent and permission to access upstream resources are separate steps, even when both applications use the same credentials.

  18. IT Admin for the AI Workforce — Sarthak Aggarwal, Decawork

    Represent the agent actor, accountable owner, represented subject, and delegation context separately.

  19. How to Secure Agents using OAuth

    Authorize financial or commercial actions per transaction, potentially with specific amounts or budgets, rather than relying only on broad read/write scopes.

  20. Beyond Conversation: Why Documents Transform Natural Language into Code

    Background agents can respond to explicit events or infer a task from contextual events, rather than require a fresh chat request.

  21. Gmail API: Create and Send Draft Emails

    A Gmail draft is an unsent message held in a resource with a stable draft ID. Replacing its content changes the underlying message ID. Sending deletes the draft and creates a new message with a new ID and the SENT label; drafts.send returns that message. Application implication: the assistant's task record should distinguish the draft container, the reviewed content, and the resulting sent message rather than treating one identifier as proof of all three.

  22. The Protection of Information in Computer Systems

    Saltzer and Schroeder describe complete mediation, fail-safe defaults and least privilege: check authority for accesses, deny absent permission, and limit a component’s granted powers. These principles apply during recovery as well as ordinary execution. In an agent loop, tool proposals and retrieved instructions must therefore pass an independently enforced authorization boundary before they can affect protected resources. Model capability does not establish the caller’s authority.

  23. Beyond Conversation: Why Documents Transform Natural Language into Code

    Introduce human review that can reject or edit outputs and also correct the agent's underlying logic.

  24. Runtime Permissions for Privacy in Proactive Intelligent Assistants

    Nathan Malkin, David Wagner and Serge Egelman's SOUPS 2022 study examined runtime permissions with 23 participant pairs interacting with a simulated proactive voice assistant. Participants differed over control and interruption costs. Some misunderstood permission scope or assumed continuing rules would expire; some hesitated between one-time and lasting permission because future preferences might change. Only a minority used the available review feature during sessions. Participants also requested microphone controls and distinctions between household members, and did not regard application permissions as sufficient protection against the listening device itself.

  25. OWASP: Transaction Authorization

    Approval must concern the significant transaction details the user actually reviewed. Store and verify those details server-side, protect them against substitution and enforce the authorization state sequence. If transaction data changes, invalidate the previous challenge or restart authorization. Credentials should be transaction-specific and time-limited. A final gate tied to execution verifies that this transaction was authorized. Application inference: bind approval to an immutable transaction version containing the action, target and material arguments, and reject execution if that version no longer matches.

  26. CIAM for AI: Authn/Authz for Agents — Michael Grinich, CEO of WorkOS

    Escalation to humans, or human-in-the-loop approval, can fail through consent fatigue even when it satisfies an organization's approval process.

  27. RFC 8693: OAuth 2.0 Token Exchange

    OAuth token exchange can represent the subject on whose behalf access is requested separately from the actor exercising it. Requests can identify target services and service-specific scopes, with issuance subject to authorization-server policy. Exchanging a token does not generally invalidate the input token or create a continuing linkage between input and output tokens. In particular, propagation of revocation is not a general property of this protocol; it depends on the implementation, token type or deployment.

  28. Notion's Token Town

    The talk's progression—AI as thought partner, assistant, teammate, and system—distinguishes drafting, individual task execution, repeated processes, and interacting workflows. Notion hypothesizes that a durable system of record enables the last transition.

  29. Google Calendar API: Events.insert

    Event creation targets a calendarId; the primary keyword selects the currently authenticated user's primary calendar. Attendees require email addresses, while display names are optional. Notification behavior is controlled separately through sendUpdates. Attendee response states distinguish no response, rejection, tentative acceptance, and acceptance. Start and end timestamps require an offset unless an explicit time zone is supplied; recurring events require a named zone for recurrence expansion. Event IDs are unique within a calendar and differ from iCalUIDs. These fields make account, calendar, recipient, timing, notification, and response state separate parts of an invitation operation.

  30. Gmail API: Users.settings.sendAs Resource

    A Gmail send-as identity can be the account's primary login address or a custom From address. sendAsEmail identifies the outgoing From address; displayName and replyToAddress are separate fields. isPrimary identifies the login address, while isDefault identifies the default sending address. Custom aliases also have a verification status. Application implication: identifying the logged-in account alone does not establish which sender identity or reply address a proposed message will use.

  31. The Protection of Information in Computer Systems: Basic Principles

    Least privilege limits each user and program to permissions needed for its job. Complete mediation requires authority checks on every access to every object, including initialization, recovery, shutdown, and maintenance. It requires reliable identification of request sources and care with cached authorization when permissions change. Fail-safe defaults base access on explicit permission. Applied to a harness, these principles imply that protected operations must pass through an enforcement mechanism that the requesting program cannot bypass or modify; a prompt instructing the model to behave is not that mechanism.

  32. Google Calendar API: Get Specific Versions of Resources

    Calendar resource updates and deletions can carry an If-Match header containing the previously retrieved ETag, a resource-version identifier. If the resource changed, the server returns 412 Precondition Failed instead of applying the conditional modification. The client can retrieve the current version before resolving the conflict. Insert operations do not support conditional modification. Application implication: an assistant can detect an intervening human edit to an existing event, but this mechanism does not make an earlier availability check and a later event insertion atomic.

  33. Turn your voice into action with new productivity features in Gemini Live

    Google's August 26, 2026 announcement describes Gemini Live handing long-running and scheduled work to Spark across Docs, Sheets, Drive and the web, including jobs spanning days or weeks while the app is unused. Its examples include maintaining a weekly family meal plan and turning emailed school events into calendar entries. The announcement also describes conversational inbox operations and continuity across past conversations and connected applications. App connections are chosen in Personal Intelligence settings.

  34. Why the Elf Acted Autonomously: Towards a Theory of Adjustable Autonomy

    Adjustable autonomy concerns whether and when an agent makes a decision or transfers decision-making control, commonly to a person. A transfer-of-control strategy can contain several steps, including waiting or taking intermediate actions to obtain more time for input. The paper models decision quality, likelihood of a timely response, waiting costs, and intermediate-action costs. It argues against treating autonomy as one permanent all-or-nothing choice and establishes that no strategy dominates across all domains under its model.

  35. Electric Elves: What Went Wrong and Why

    The Electric Elves authors report failures when personal assistants generalized learned decisions and acted after users failed to respond. An initial policy waited indefinitely for human input; adding a five-minute timeout followed by autonomous action produced unwanted cancellations and an unwanted presentation commitment. Another assistant delayed a meeting almost fifty times in five-minute increments, following a learned rule while ignoring cumulative nuisance to participants. The report identifies uncertainty about user intent, unavailable human responses, and the cost of action sequences as distinct problems.

  36. Heartbeat — OpenClaw

    OpenClaw documents a heartbeat as a periodic agent turn that can inspect context and surface matters needing attention. Its configuration separates cadence, active hours, context inclusion and delivery destination. The default owner route requires a concrete owner identity and avoids group delivery; another setting can instead follow the last conversation, including groups. Disabling recurring cadence does not disable every event-driven wake. The documentation explicitly directs users to tool policy and sandboxing, rather than heartbeat frequency, to control command execution.

  37. Batching Smartphone Notifications Can Improve Well-Being

    This randomized field experiment compared ordinary notification delivery with hourly batching, three daily batches, and no notifications. Three daily batches improved reported attention and perceived productivity relative to ordinary delivery; eliminating notifications increased reported anxiety and fear of missing out. The authors emphasize that their blanket notification policy ignored differences in notification relevance and that participants could not modify its settings. They call for personal control and longer studies.

  38. Don't just slap on a chatbot: building AI that works before you ask

    Proactive AI should supplement agency, keep recommendations optional, and make its changes easy to reverse.

  39. My AI Thinks I'm Eating My Feelings (and Other Nutritional Insights)

    Accumulate new user information across interactions while letting users inspect, add to, and remove that information.

  40. Why Your AI UX Is Broken (and It's Not the Model's Fault)

    A bidirectional transport alone does not make an in-progress task visible or reachable from other devices.

  41. gRPC lifecycle: cancellation is not rollback

    Client and server can disagree about an RPC's success: the server may finish while its response arrives after the client's deadline. gRPC explicitly warns that cancellation does not roll back changes already made. Application implication: cancellation during a mutation may leave an unknown outcome. Retain an operation identifier, query authoritative status or reconcile the resulting state, and use an idempotent retry contract before resubmitting. If an effect must be reversed, that requires a separate supported compensating operation rather than assuming cancellation undid it.

  42. Compensating Transaction pattern

    Compensation performs new, business-specific actions to counter completed steps of an eventually consistent workflow. It differs from transaction rollback: intervening concurrent work must be preserved, the exact original state may be unattainable, and cancellation may incur charges. Record completed steps and the information needed to compensate them. Compensation order need not exactly reverse execution, and some steps can run in parallel. Compensation can itself fail, so persist progress, resume from failure, and make retryable steps idempotent. Where automated recovery is impossible, alert an operator with diagnostic information. For irreversible effects, an application must define an acceptable remedy or escalation rather than claim the action has been undone.

  43. OWASP Logging Cheat Sheet

    OWASP advises against directly logging passwords, access tokens, session identifiers, connection strings, encryption keys, sensitive personal information, and payment data; remove or appropriately protect sensitive fields. Validate and sanitize event data crossing trust boundaries to prevent log injection. Restrict and periodically review read access, record access to logs, protect transfer and storage, and apply retention and disposal rules to debug logs, backups and extracts as well as primary logs. Applied to agents, collect the identifiers, timings, outcomes and selected diagnostic fields needed for investigation; do not assume entire prompts, retrieved documents or tool responses are safe to retain.

  44. Privacy as Contextual Integrity

    Helen Nissenbaum's 2004 account defines contextual integrity through appropriate information handling within social contexts. One set of norms concerns what information belongs in a situation; another concerns how it may flow between people. Medical, employment, financial and friendship relationships have different expectations. Information being available in one setting does not make its transfer into another appropriate, and public settings are not norm-free. Applied to personal assistance, combining work and household information requires attention to recipients and purposes, not merely whether the assistant can access both accounts.

  45. Google Calendar API: Freebusy.query

    Freebusy.query accepts calendar identifiers and a bounded time interval and returns busy intervals grouped by calendar, without event titles, descriptions, or attendee lists in its response schema. Response time zone defaults to UTC unless specified. Busy intervals have inclusive starts and exclusive ends. The response can contain calendar-specific or group-specific errors, including missing resources and internal failures. Application implication: availability-only scheduling can avoid retrieving private event descriptions, and failed calendar queries must remain unavailable coverage rather than being interpreted as free time.

  46. The End of Apps

    The speaker favors local files and self-hosted services to give agents direct access while retaining control over files, memory, and sessions.

  47. NIST Privacy Framework 1.0: lifecycle and minimized audit evidence

    The framework inventories data elements, processing purposes, actions, owners and flows. Policies define permitted uses and retention periods; the data lifecycle aligns with system development and operations. Authorizations must be maintained and revocable, access limited by least privilege, and deletion and destruction performed under policy. Audit records themselves must incorporate data minimization. Engineering application: define the decision evidence needed for review, its purpose, authorized readers, retention trigger and disposal method before logging. Retain the necessary decision, model and policy versions and relevant evidence without indiscriminately copying personal data into logs, prompts or backups. Where review requires sensitive evidence, constrain fields, access and retention rather than treating auditability as permission to keep everything. Assess removal and disclosure across downstream copies and service providers.

  48. gRPC cancellation and handler cooperation

    Cancelling a client RPC does not generally interrupt application code already running in the server handler. Long-running handlers must check cancellation and cease processing cooperatively; cancellation of downstream RPCs is automatic in some language implementations and must be propagated explicitly in others. Applied to agents, stopping the orchestration loop must also stop new dispatches and signal cancellation to active operations where supported. Merely abandoning an await does not establish that remote work stopped.

  49. RFC 7009: OAuth 2.0 Token Revocation

    RFC 7009 specifies token invalidation and acknowledges possible propagation delay between servers. Whether related tokens and the underlying grant are also revoked depends on server policy, with different recommendations for refresh-token and access-token revocation. A successful HTTP 200 response also covers an already-invalid token; a 503 response requires assuming the token still exists. Clients must accommodate unexpected invalidation rather than assuming a connected account remains usable indefinitely.

  50. Full Workshop: Agent Auth Protocol — Paola Estefanía de Campos, Better Auth

    Revoking one agent identity did not stop read access in the demo because the host could create a replacement identity with default read capabilities.

  51. π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

    Haoran Zhang and colleagues' May 14, 2026 preprint evaluates 100 multi-turn tasks across five personas, with persistent workspaces, hidden requirements and dependencies between sessions. It measures proactivity separately from final completeness: a reactive assistant can eventually finish after the simulated user supplies missing requirements. Removing preceding sessions in a three-model ablation reduced proactive intent resolution more than final completeness. The authors also distinguish turn count from human burden because useful clarification can add turns.

  52. Build AI Systems for Discernment, Not Approval - Angel Ortmann Lee, Duolingo

    Define system success, concrete metrics, and required improvement data before designing the interaction.

  53. Identity for AI Agents

    Async Auth is described as an agent-initiated approval flow that returns an access token containing the transaction details approved by the user.

  54. The Adversarial Path to the Personal Assistant

    The calendar example preserves both possible school dates instead of prematurely choosing one, while checking imported events against existing commitments.

  55. The Adversarial Path to the Personal Assistant

    The proposed email feature narrows monitoring to specific information categories rather than claiming general inbox comprehension.

  56. Developing Taste in Coding Agents: Applied Meta Neuro-Symbolic RL — Ahmad Awais, Command Code

    Corrections to generated code can serve as implicit preference feedback, reducing the need to restate the same instructions.

  57. Don't just slap on a chatbot: building AI that works before you ask

    Tegon's suggestion mode observes issue composition and proactively asks contextual questions within the existing workflow.

  58. Full Workshop: Agent Auth Protocol — Paola Estefanía de Campos, Better Auth

    The email demonstration separates initial read authorization from a later request to send, requiring another approval before adding the send capability.

  59. 75 Years of Innovation: CALO (Cognitive Assistant that Learns and Organizes)

    SRI's retrospective dates the five-year CALO collaboration to 2003 and places it within DARPA's Personalized Assistant that Learns program. Its ambition extended beyond individual recommendations to interrelated office responsibilities: organizing information, managing tasks, scheduling commitments, preparing information products and coordinating resources. SRI describes an integrated collection of learning, reasoning and information-management components, with subsequent commercial descendants including Siri, Tempo AI and Desti.

  60. Siri — SRI

    SRI attributes Siri's technology to its CALO research and joint work with EPFL. It spun off Siri, Inc. in 2007; Apple acquired the company in April 2010; Siri was unveiled as an integrated iPhone 4S feature in October 2011. These are distinct commercialization events connecting an institutional assistant-research program to a consumer phone product.

  61. Gemini's multi-step tasks on Android

    Google's February 25, 2026 announcement described an early preview of Gemini performing multi-step tasks in selected Android apps after a user request. Tasks could proceed in the background, with notifications allowing the person to monitor progress, take over or stop execution. Google described a constrained virtual window and an initial focus on food, grocery and rideshare applications. This changes the assistant's responsibility from explaining steps to carrying them out while retaining a supervision surface.