Research & Blog

Approval Should Expire When an AI Agent Materially Changes

A release manifest, configuration comparison, and observed use review can keep enterprise approval tied to the AI agent that is actually running.

Alex Georges, PhD5 min read
Share

A one time approval should expire when an AI agent changes materially. Material drift in observed deployment behavior should also trigger reassessment when the prior review no longer describes actual use. For enterprise AI risk leaders, the stake is whether an approval still applies to the system that can act in production. The decision should move from calendar based or model version only review to release level change control. An approval belongs to a defined configuration, not to an agent name that persists while its operating boundary changes.

Time alone should not make a change material. A scheduled review can still uncover problems, but elapsed time is not the expiration trigger. The trigger is a difference that could make the prior evidence, assumptions, or conditions unreliable.

The approval boundary is larger than the model

Anthropic's guidance says agents plan and execute tasks with minimal human intervention while retaining mechanisms for human control through configurable permissions and approval workflows. [S2] The same guidance identifies model behavior, tool access, and operational environments as layers that need security. [S2] These statements should be treated as qualitative guidance, not quantitative evidence that a particular control produces a particular outcome.

Together, these premises explain why a model version cannot define the complete approval boundary. Tool permissions can change which actions are available. A different operating environment can change exposure and control coverage. Modified prompts, user flows, or approval settings can change how often a person reviews an action. The model may remain unchanged while the system that reaches users becomes meaningfully different.

Observed behavior adds another dimension. Anthropic's autonomy report examines Claude Code and public API data. [S1] In that report, the longest running sessions nearly doubled from under 25 to over 45 minutes within a few months. [S1] The share using full automatic approval rose from about 20% among new users to over 40% among users with greater experience. [S1] Those experienced users also interrupted the agent more frequently. [S1]

These measurements do not prove that greater autonomy is safe or risky. They show why measured autonomy is one decision input rather than complete proof. More automatic approval and more frequent intervention can coexist. A risk leader therefore needs evidence about actual use, control operation, and consequences, not a single autonomy score.

Define materiality by its effect on the prior decision

A universal quantitative threshold for material change remains undefined. A practical standard is to ask whether the difference could alter an approval conclusion, invalidate an assumption, change the required controls, or expand the consequences of an agent action.

The approval manifest should cover at least five categories:

  1. The model, system instructions, prompts, routing, memory, retrieval, and media that shape agent behavior.

  2. The tools, data connections, permission scopes, and action limits available to the agent.

  3. The user flow, including where people review plans, approve actions, provide clarification, or receive results.

  4. The production controls, monitoring, operating environment, and response procedures that constrain the release.

  5. The intended users, use cases, transaction types, data classes, and expected patterns of human involvement.

A difference within any category is not automatically material. A corrected label may have no effect on the approval reasoning. A new tool with authority to create an external transaction is much more likely to require reassessment. The decision should turn on the relationship between the change and the prior risk conclusion, with the rationale recorded for later review.

This standard also prevents a routine calendar from becoming a substitute for change analysis. A scheduled review may be useful for detecting stale assumptions or missed drift. It should not preserve an approval for a materially different configuration, and it should not force a complete reassessment merely because a date passed when nothing relevant changed.

Put the rule into release control

The operating workflow begins with an approval manifest for the release that was actually reviewed. It should identify the model, prompts, tools, media, user flow, production controls, permissions, operating environment, intended use, and material assumptions.

At each proposed release, the operator should compare the candidate configuration with that manifest. A minimum process has five steps:

  1. Preserve the approved baseline and its evidence so that later comparisons do not depend on memory.

  2. Record every configuration difference, including changes outside the model itself.

  3. Classify each difference against documented materiality criteria and identify affected assumptions.

  4. Reassess changed components and any interfaces through which the change could alter the prior conclusion.

  5. Route an accountable decision to reapprove, constrain, defer, or retire the changed release.

An immaterial difference can retain the existing approval when the reviewer records why the evidence still applies. A material difference ends reliance on that approval for the changed configuration until review resolves it. This does not necessarily revoke approval for an older configuration that remains unchanged and within its conditions.

The exact reassessment path remains a design question. Enterprises still need to specify who can classify a difference, which evidence is required, when a wider review is necessary, and which temporary constraints apply while a decision is pending. The process should be explicit even when those thresholds vary by use case.

Observed use can make an unchanged release stale

A release manifest is necessary but not sufficient. Operators can use the same formal configuration differently over time. Session duration can expand, automatic approval can become more common, users can introduce new data classes, or the mix of tools and transactions can move beyond the assumptions used in review.

No single movement should automatically expire approval. Drift becomes material when it makes an approval assumption unreliable. Examples include a human review rate falling outside the assessed operating condition, a new user group applying the agent to a different decision, or production activity expanding into a consequence class that the review did not cover.

This calls for a comparison between approved use and observed use, not continuous reapproval for every fluctuation. Measured autonomy, intervention frequency, tool use, user population, and consequence type can serve as indicators. The accountable reviewer must still decide whether the difference is material.

Keep the economic claim bounded

Release difference triggered reassessment may focus review effort on changed components while reducing reliance on evidence for another configuration. This is an operating inference, not quantitative evidence of savings, avoided loss, or faster launch. A narrow change may permit a narrow review, while a change with broad dependencies may still require assessment of the complete release.

A narrow company position

AetherLab's position is that the complete release is the unit of approval. This is an operating position, not a product description. It does not claim that institutional adoption or evidence acceptance is already proven. The materiality threshold and reassessment path remain open design questions.

The rule for enterprise approval

Enterprise AI risk leaders should adopt a specific change control rule: approval belongs to a defined release configuration. A material configuration difference, or material observed use drift that makes the prior boundary stale, ends reliance on that approval until accountable review resolves the difference. Time can prompt inspection, but time alone is not material change. Measured autonomy can inform the decision, but it cannot decide it.

Sources

  1. Measuring AI agent autonomy in practice \ Anthropic
  2. Trustworthy agents in practice \ Anthropic
  3. SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents \ Anthropic
  4. Sabotage evaluations for frontier models \ Anthropic
ai agentsapproval governancechange control
Share

Related analysis

See what an adversarial assessment finds in your AI system.

AdversarialScan red-teams your text and image surfaces against your break-goals, and the Evidence Pack turns the findings into approval-ready governance evidence.

Ask about the Evidence Pack

Leave your email and we'll walk you through what an Evidence Pack contains for your use case: severity-scored findings, business-impact mapping, and the approval record.