Skip to main content

Agent Reliability Engineering for reliable and governable autonomous AI agents

Agent Reliability Engineering (ARE) is the software engineering discipline for making autonomous AI agents reliable, governable, observable, recoverable, and accountable in production.

Intelligence is not enough.

AI agents are becoming capable of making decisions, using tools, writing code, interacting with systems, and taking consequential actions. Capability is increasing rapidly. Reliability must increase with it. Without reliability, you only have liability.

CAPABILITY→ RELIABILITY→ TRUST→ AUTHORITY→ ADOPTION

What is Agent Reliability Engineering?

Canonical definition

Agent Reliability Engineering (ARE) is a software engineering discipline for making autonomous AI agents reliable, governable, observable, recoverable, and accountable in production.

ARE treats reliability as an architectural property—designed into an agent's identity, authority, contracts, controls, and recovery paths before it receives production authority.

Why do AI agents need reliability engineering?

Traditional software generally executes deterministic or bounded instructions. Autonomous AI agents introduce:

  • Probabilistic behavior
  • Dynamic planning
  • Tool selection
  • Unscoped data access
  • Delegated execution
  • Emergent workflows
  • Model drift
  • Prompt injection
  • Authority expansion
  • Multi-agent interactions

The more capable an agent becomes, the more important it becomes to engineer the boundaries within which that capability operates.

The Reliability Surface

The Reliability Surface encompasses the behavioral dimensions and engineering controls that determine whether an autonomous AI agent can be trusted to operate in production.

Each row is a dimension a team can audit independently — a gap in any one of them can undermine the rest, regardless of how well the others are engineered.

Reliability Debt

Reliability Debt is the accumulated risk created when an autonomous system's capabilities, complexity, or authority grow faster than its reliability engineering.

CAPABILITY GROWTH→ RELIABILITY LAG→ RELIABILITY DEBT→ OPERATIONAL RISK→ LOSS OF TRUST→ RESTRICTED AUTONOMY

Progressive Autonomy

Agents should not receive unlimited authority simply because they are capable of performing an action. They should earn greater authority through demonstrated reliability.

01

Observe

Proposed actions are logged, not executed.

02

Draft

The agent prepares actions for a human to execute.

03

Act with approval

Execution requires explicit sign-off.

04

Act within scope

Autonomous execution inside a narrow, verified boundary.

05

Expanded authority

Scope widens only as evidence accumulates.

Authority should increase when reliability improves and contract when reliability deteriorates.

How ARE fits into the engineering landscape

DisciplinePrimary focus
SREReliability of software services and infrastructure
AI SafetyPreventing harmful AI outcomes
AI SecurityProtecting AI systems, data, models, and infrastructure
AI GovernanceOrganizational accountability, policy, risk, and compliance
AI EvaluationMeasuring model and agent behavior
AREEngineering autonomous systems so these concerns operate together reliably in production

ARE does not replace these disciplines. It connects their relevant engineering practices around the operation of autonomous systems.

Agent Reliability Engineering Concepts

ARE defines a common vocabulary for engineering reliable autonomous AI systems.

AI Agent Reliability

The measurable degree to which an agent behaves as intended, within authority, and recovers safely.

Reliability Surface

The full set of behavioral dimensions and controls that determine agent trustworthiness.

Reliability Debt

Accumulated risk from capability growing faster than reliability engineering.

Progressive Autonomy

Authority that expands with demonstrated reliability and contracts when it deteriorates.

Agent Authority

The explicit, scoped set of actions an agent may take and the data those actions may touch.

Agent Governance

The organizational structures and accountability mechanisms that govern agent behavior.

Agent Reliability Controls

The concrete engineering mechanisms that implement the Reliability Surface.

Reliability Maturity

A five-level model for how systematically reliability engineering has been implemented.

Missing a term?

New concepts are added through community review, not by editorial fiat.

Propose a concept

The Agent Reliability Engineering Manifesto

Agents are intelligent. They are not yet reliable.

The manifesto lays out the founding principles of the discipline — developed in the open, subject to community review before each version is finalized.

Principles of Agent Reliability Engineering

  1. 01

    Governance is architectural, not operational.

    Why

    Governance built as a meeting process, rather than as part of the system itself, will be bypassed under pressure the first time it's inconvenient.

    Implication

    Authority checks and audit logging are implemented in code and infrastructure, not left to a review committee's memory.

  2. 02

    Authority is earned, not assumed.

    Why

    An agent's ability to call a tool is not the same as its right to use it unsupervised.

    Implication

    New capabilities launch at the most restrictive stage of Progressive Autonomy and expand only against evidence.

  3. 03

    Reliability must be observable, or it does not exist.

    Why

    A team cannot manage what it cannot see; unobserved agents accumulate Reliability Debt invisibly.

    Implication

    Every consequential action is logged with enough context to reconstruct why the agent took it. Prefer OpenTelemetry spans over proprietary logs. Enrich each consequential span with identity, authority scope, data scope, and the policy verdict so any OTEL backend can audit the run without a second control plane.

AI Agent Reliability Maturity Model

Five levels for assessing how systematically an organization has implemented the Reliability Surface across its autonomous agents.

LEVEL 1

Unmanaged

No defined identity or authority boundaries; agent actions are effectively unbounded and untracked.

LEVEL 2

Observed

Actions are logged, but authority is still broad and largely ungoverned.

LEVEL 3

Governed

Explicit authority scopes and policy enforcement exist; approval gates are in place.

LEVEL 4

Accountable

Every action is attributable to an identity and policy; incidents produce traceable root cause.

LEVEL 5

Adaptive

Authority expands and contracts automatically based on measured reliability.

Agent Reliability Engineering Glossary

Core terms in the emerging Agent Reliability Engineering vocabulary. Canonical concept pages are being published as the working draft develops.

Agent Authority
The explicit, scoped set of actions an agent is permitted to take.
Reliability Surface
The full set of behavioral dimensions and controls that determine agent trustworthiness.
Progressive Autonomy
Authority that expands with demonstrated reliability and contracts when reliability deteriorates.
Reliability Debt
Accumulated risk from capability growing faster than reliability engineering.

About Agent Reliability Engineering

ARE is intended to be an open discipline rather than a proprietary methodology.

ARE was proposed by Mike Hogan and is currently a working draft, published openly for review and revision. It has not yet been adopted as a formal standard by any industry body.

Mike Hogan's professional work includes Trustabl. ARE is published as an independent, openly licensed discipline rather than as Trustabl product documentation.

Research and further reading

This manifesto was not written in a vacuum. It builds on academic work that measured the capability/reliability gap, enterprise studies that documented agent failures after pilots, identity and runtime work from vendors who now treat agents as first-class principals, and the frustration of developers who can ship intelligent agents faster than they can govern them.

The list starts with the paper that treats reliability as a measurable engineering problem. We encourage these authors, and anyone else measuring failures, designing controls, or operating agents in production, to add their insights to the manifesto.

Academic research

Metrics

Towards a Science of AI Agent Reliability

Rabanser, Kapoor, Kirgis, Liu, Utpala, and Narayanan (Princeton). Twelve metrics across consistency, robustness, predictability, and safety; capability gains have not produced matching reliability gains.

arXiv:2602.16666
Benchmark

ReliabilityBench

Evaluates consistency, robustness to task perturbations, and fault tolerance under injected tool/API failures. Single-run success rates hide how agents behave under production-like stress.

arXiv:2601.06112
Benchmark

τ-bench

Yao, Shinn, Razavi, and Narasimhan. Tool-using agents must follow domain policy across multi-turn user interaction; pass^k shows even strong models are inconsistent across retries.

arXiv:2406.12045
Computer-use

On the Reliability of Computer Use Agents

Succeeding once is not the same as succeeding again. Unreliability comes from execution stochasticity, task ambiguity, and behavioral drift across repeated OSWorld runs.

arXiv:2604.17849
MCP

MCP Tool Descriptions Are Smelly!

Most MCP tool descriptions contain defects that mislead tool choice and arguments. Cleaning them can raise success — and also add steps or regress some tasks — so contracts, not just prompts, matter.

arXiv:2602.14878
Assurance

Engineering Trustworthy Agentic AI for Critical Systems

Treats trustworthiness as an engineering property — safety, robustness, transparency, accountability, security — mapped onto an assurance workflow rather than a benchmark score.

arXiv:2607.18548
Evaluation

Evaluation Scores Are Perishable Knowledge Claims

Averaging eval signals inflates trust. Scores have formality, scope, and a validity window; weakest-link ranking of HELM models does not match mean ranking.

arXiv:2607.26191

Press and analysts

Survey

Boomi / Forrester: Agentic AI Readiness Gap

86% of surveyed enterprises have moved agents beyond pilots; only 34% trust the actions those agents take. "Agentic chaos" correlates with about $2.1M in extra failure cost.

Boomi study
Forecast

Gartner: 40% demote or decommission by 2027

Governance applied as binary — locked down or fully trusted — is predicted to drive rollback after production incidents, not before them.

Coverage
Survey

VentureBeat: the agent evaluation gap

About half of surveyed enterprises shipped an agent that passed internal evaluations and then failed a customer. Few fully trust automated evaluation as a production gate.

VentureBeat
Framework

Forrester AEGIS and Why AI Agents Fail

AEGIS frames enterprise guardrails for agentic systems. Companion analysis covers compounding errors, goal misalignment, orchestration risk, and why genAI review habits do not transfer.

AEGIS overview

Commercial insights

Identity

Microsoft Entra Agent ID

Agents get distinct identities, blueprints, sponsorship, Conditional Access, and lifecycle — not shared service principals or human accounts.

Microsoft Learn
Security

Zero Trust for AI agents

Microsoft's Zero Trust for AI guidance applies verify-explicitly and least-privilege to agents, memory, tools, and runtime — not only to users and devices.

Security blog
Governance

Copilot Studio and Foundry governance

Zoned environments, data policies, Agent 365 inventory, and Foundry task-adherence controls treat production agents as governed workloads.

Copilot Studio
Runtime

NVIDIA OpenShell

Out-of-process sandbox, filesystem, network, and process policy so an agent cannot lift its own limits after prompt injection or drift.

NVIDIA Technical Blog
Identity

Okta for AI Agents

Treats agents as first-class non-human identities: discovery, ownership, scoped tokens, and a kill switch when behavior goes off-policy.

Okta
Platform

Databricks Agent Bricks

Identity, tool and data access, traces, and continuous evaluation in one governed execution path instead of after-the-fact review.

Databricks
Investment

AWS Forward Deployed Engineering

$1B to embed engineers with customers so agentic systems ship against real data, governance, and operating constraints.

AWS
Investment

Google Cloud $750M partner fund

Partner and FDE investment aimed at production agentic deployments, not model access alone.

Google Cloud
Safety

Smarter models make your agents less safe

Trustabl on why more capable models expand the action surface faster than guardrails, contracts, and least-privilege policy keep up.

Trustabl
Contribute

Add a source

Research, incident data, or a control pattern that belongs here should be proposed in the open. The list should grow with the field.

Open an issue

Help define the discipline

Agent Reliability Engineering is being developed in the open. Contributions from engineers, researchers, security professionals, operators, architects, and AI practitioners are encouraged.

Review

Review the Manifesto

Read the current draft and leave comments directly against specific passages.

Comment on the Google Doc
Build

Contribute code & artifacts

Documentation, diagrams, tooling, schemas, examples, and reference implementations.

GitHub
Discuss

Submit a proposal

Propose terminology, controls, metrics, and case studies with the working group.

Community