The Enterprise Agent Stack Reference Architecture is now available.
Vendor-Neutral Category Reference Architecture

The Reference Architecture for Production AI Agents

A 15-layer vendor-neutral framework for building, running, and governing autonomous AI agents in enterprise environments.

"The model proposes. The runtime decides. The evidence proves."
Interactive Enterprise Agent Reference Architecture
Hover layers to inspect Evidence & Control Plane bindings
ENTERPRISE AGENT APPLICATIONSLayers 1
CopilotsWorkflowsOperationsCustomer AgentsEngineeringResearch
↓ Objectives & Intent
AGENT FRAMEWORK AND RUNTIMELayer 2
Objective IDsPlanning LoopsState MachineError RecoveryDelegation
↓ Proposed Actions & Queries
MODELS AND INTELLIGENCELayer 3
Models • Gateways • Dynamic Routing
Hosted, Private VPC & Open-Weight Fallbacks
TOOLS AND CAPABILITIESLayer 4
MCP • APIs • Database Connectors • Systems
Capability Gateways & Schema Validation
↓ RAG & Hydration
CONTEXT AND MEMORYLayers 5 – 6
Permission-Aware Retrieval • Document ACLs • Context Bundles • Validated Memory Store
EVIDENCE AND CONTROL PLANE (Spans All Layers)
Layers 7 – 14

Identity • Ownership • Authority • Policy-as-Code • Human Approval • Pre-Action PEP • Budget Control • Evaluation • Observability • Decision Receipts • Revocation

"The model proposes. The runtime decides. The evidence proves."
DEPLOYMENT AND OPERATIONSLayer 15
Public Cloud • Customer VPC • On-Premises • Air-Gapped • Edge • Kubernetes • Global Kill-Switch
Evidence and Control PlaneActive Control Boundary
Enforced Pre-Action Controls:
  • Pre-action PDP/PEP policy evaluation
  • Runtime budget reservations
  • Human-in-the-loop escalation
Durable Audit Output:
Cryptographic Decision Receipt & Immutable Audit Trail
Why EnterpriseAgentStack.com Exists

The Category Problem

Enterprise agent deployments are currently failing not because the underlying LLM models are weak, but because they are being deployed into production without a complete control and governance architecture.

The Control Plane Solution

A production agent framework must bind model proposals with real-time pre-action policy evaluation, authority delegation, runtime budget limits, human escalation, and tamper-evident Decision Receipts.

The Vendor-Neutral Standard

Keep the models and frameworks that already work (LangChain, CrewAI, OpenAI, Anthropic, AWS). Add the 15-layer control and evidence architecture across them.

The Prototype-to-Production Gap

Building an agent is easy. Operating one responsibly is the stack problem.

A prototype proves that an agent can perform a task. A production architecture must prove that the agent can operate within enterprise boundaries, recover from failure, and leave evidence behind.

A Typical PrototypeUncontrolled State

Needs only intelligence and basic API tool connections to demonstrate a working demo loop.

  • System Prompt in codebase
  • Single LLM API model key
  • Basic Agent Framework loop
  • Un-scoped Tool API functions
  • Local / ephemeral demo interface
  • No identity or owner registration
  • No pre-action policy check
  • No budget caps or retry limits
  • No audit-grade Decision Receipts
Core Risk: Capability grows faster than ownership and control.
A Production SystemGoverned & Assured

Requires full identity, authority, pre-action policy enforcement, evidence, and operations.

Objective & Accountable Owner
Unique SPIFFE Workload Identity
Scoped & Versioned Model Gateways
Governed Read/Write Tool Schemas
Permission-Aware RAG Context
Pre-Action Policy Enforcement (PEP)
Real-time Runtime Budget Caps
Risk-Based Human Escalation
OpenTelemetry Trajectory Traces
Continuous Production Evals
Signed Cryptographic Decision Receipts
Instant Global Kill-Switch
Production Guarantee: Operates within enterprise boundaries, recovers from failure, and leaves immutable evidence behind.
5 Architectural Groups

Organized into 5 Cohesive Architecture Groups

Every production agent system spans these 5 distinct functional groups, bound together by the Evidence and Control Plane.

Group ALayers agent-applications, agent-framework-runtime

Business and behavior

Defines business objectives, permitted autonomy, and runtime orchestration framework.

2 Architecture Layers
Group BLayers models-intelligence, tools-mcp-capabilities, context-knowledge, agent-memory

Intelligence and capabilities

Models, tool capabilities, permission-aware context retrieval, and managed memory.

4 Architecture Layers
Group CLayers agent-identity-ownership, authority-delegation, runtime-policy-enforcement, human-approval, runtime-budget-governance

Identity and control

Workload identity, delegated authority, pre-action runtime policy, human approval, and budgets.

5 Architecture Layers
Group DLayers agent-observability, agent-evaluation, evidence-audit

Reliability and evidence

Episode tracing, continuous evaluation, decision receipts, and audit trails.

3 Architecture Layers
Group ELayers deployment-operations

Infrastructure

Cloud, VPC, on-premises, edge, and physical execution operations with SRE controls.

1 Architecture Layers
15 Architecture Stack Layers

Explore Every Layer of the Enterprise Agent Stack

Click any layer card to inspect minimum production requirements, common failure modes, required interfaces, enterprise ownership, and evidence output.

View Perspective:
Layer 1Group A

Agent Applications and Workflows

"Start With the Business Objective, Not the Model"

Defines business intent, success criteria, and user interaction surfaces for autonomous workflows.

/layers/agent-applications
Layer 2Group A

Agent Framework and Runtime

"Control How the Agent Plans, Retries, Recovers, and Delegates"

Manages step execution, memory hydration, tool execution loops, retry strategies, and sub-agent delegation.

/layers/agent-framework-runtime
Layer 3Group B

Models and Intelligence

"Use the Right Model for the Objective, Data, Risk, Cost, and Environment"

Routes prompts and context to appropriate models based on task complexity, data sensitivity, latency, and cost rules.

/layers/models-intelligence
Layer 4Group B

Tools and Capabilities

"Define What Every Tool Can Do Before an Agent Uses It"

Schema validation, read/write classification, resource scoping, transaction limits, idempotency enforcement, and side-effect controls.

/layers/tools-mcp-capabilities
Layer 5Group B

Context and Knowledge

"Give Agents the Right Context, Under the Right Permissions, at the Right Time"

Enforces permission-correct retrieval, source identity verification, freshness limits, provenance tracking, and effective-time boundaries.

/layers/context-knowledge
Layer 6Group B

Agent Memory

"Control What Agents Remember, Reuse, Share, and Forget"

Manages memory types, origin tagging, validation state, decay, expiration, deletion, and cross-agent memory contamination controls.

/layers/agent-memory
Layer 7Group C

Agent Identity and Ownership

"Give Every Production Agent an Identity, Purpose, and Owner"

Assigns verifiable SPIFFE/OAuth workload identity, logical IDs, environmental scopes, approved purpose definitions, and named business/technical/security owners.

/layers/agent-identity-ownership
Layer 8Group C

Authority and Delegation

"Define What Each Agent May Do—and What It May Delegate"

Enforces authority scoping across 12 dimensions: capability, resource, purpose, environment, time, amount, customer, region, risk, approval, delegation, and budget.

/layers/authority-delegation
Layer 9Group C

Policy and Runtime Enforcement

"Evaluate Policy Before Consequential Actions Occur"

Evaluates policy-as-code against current identity, authority, context, tool, model, budget, and approval state to issue Allow, Deny, Require Approval, Degrade, Throttle, Replan, Escalate, or Terminate decisions.

/layers/runtime-policy-enforcement
Layer 10Group C

Human Escalation and Approval

"Put People in the Decisions That Need Them"

Routes risk-based approval requests with structured evidence objects, expected effects, reversibility options, and budget impacts to human reviewers.

/layers/human-approval
Layer 11Group C

Runtime Budget Governance

"Control Agent Cost While the Work Is Still Running"

Monitors spend across 11 dimensions: objective budget, model cost, tokens, tool cost, retrieval cost, compute, duration, rows scanned, retries, replanning, and delegation depth.

/layers/runtime-budget-governance
Layer 12Group D

Agent Observability

"Observe Complete Agent Episodes, Not Isolated API Calls"

Captures full agent trajectories (objective, requester, agent identity, user identity, context, plans, model calls, tool calls, delegations, policy decisions, budget events, approvals, side-effects, outcomes).

/layers/agent-observability
Layer 13Group D

Evaluation

"Evaluate the Entire Path From Objective to Outcome"

Measures objective interpretation, context selection, plan quality, model behavior, tool arguments, recovery efficiency, policy adherence, cost efficiency, and outcome correctness.

/layers/agent-evaluation
Layer 14Group D

Audit and Decision-Grade Evidence

"Preserve Evidence That Explains What Happened and Why"

Answers 12 material audit questions: What happened? Which identity acted? Which authority applied? Which objective? Which context? Which model? Which tool? Which policy? Which approval? What runtime decision? What side-effect? Was outcome verified?

/layers/evidence-audit
Layer 15Group E

Deployment and Operations

"Run the Same Controls Across Cloud, Private, On-Prem, Edge, and Physical Systems"

Enforces environment isolation, signed deployments, version pinning, shadow mode, canary rollouts, rollback, instant kill-switches, credential rotation, and disaster recovery.

/layers/deployment-operations
Core Runtime Engine

The Evidence and Control Plane in Action

Models, frameworks, tools, and data do not govern themselves. Watch how pre-action policy evaluation, budget reservations, human approvals, and cryptographic evidence bind every single agent tool call.

"The model proposes. The runtime decides. The evidence proves."
10-Step Governed Execution Flow
Step 1 of 10Agent Framework / Model
PROPOSED

Agent Proposes Tool Call

Agent proposes tool invocation: execute_wire_transfer({ amount: "$5,000", recipient: "vendor_881" })

Runtime Control Plane TraceLATENCY: 8.2ms
REQ: agt_9824 -> tool_gateway.call("execute_wire_transfer", { amount: 5000 })
Signed Decision Receipt SnapshotID: #DR-99201
{
  "episode_id": "ep_88192a",
  "step": 1,
  "action": "propose_action",
  "decision": "PROPOSED",
  "timestamp": "2026-07-22T23:38:00.000Z",
  "proof_hash": "sha256:8f9a2b7c..."
}
Auto-advancing step 1 / 10
Architecture Audit Questions

Every production agent architecture should answer these questions

Before granting autonomous agents access to enterprise data or production systems, verify if your platform can definitively answer these 14 control questions.

Question #1Registry

Which agents exist?

Click to inspect requirement
Question #2Ownership

Who owns each agent?

Click to inspect requirement
Question #3Governance

Which models may it use?

Click to inspect requirement
Question #4Capability

Which tools may it call?

Click to inspect requirement
Question #5Context

Which data may it retrieve?

Click to inspect requirement
Question #6Memory

Which memory may it reuse?

Click to inspect requirement
Question #7Authority

Which actions may it take?

Click to inspect requirement
Question #8Delegation

Can it delegate authority?

Click to inspect requirement
Question #9FinOps

What is its budget?

Click to inspect requirement
Question #10Approval

Which actions require approval?

Click to inspect requirement
Question #11Control

Can the runtime stop it?

Click to inspect requirement
Question #12Security

Can the organization revoke it?

Click to inspect requirement
Question #13Observability

Can every episode be reconstructed?

Click to inspect requirement
Question #14Evidence

Can the outcome be verified?

Click to inspect requirement
Can your current agent platform answer all 14 questions?

Take the interactive assessment to calculate your exact control score.

6-Stage Maturity Ladder

Measure how close your agent architecture is to governed production

Move from local prompt experimentation to decision-grade assured operations. Select a stage to view its architectural capabilities and core risks.

Stage 3 BenchmarkGoverned

Identity, ownership, permissions, and approvals.

Added Capabilities & Features:

Agent registry
Named owners
Workload identity
Parameter RBAC
Human approval workflows

Core Architectural Risk:

"Governance may remain static while runtime conditions change."
What stage is your organization currently operating at?

Our 7-minute assessment scores your exact maturity stage across all 15 layers.

18 Production Acceptance Criteria

An agent stack is not production-ready until these controls work

Use this interactive acceptance checklist during architecture reviews to verify control completeness before promoting autonomous agents into live enterprise environments.

Interactive Production Control Verification(0 / 18 Passed)
Filter Area:
Every agent is registered.Identity

All active and historical agents exist in a single verified Enterprise Agent Registry.

Every agent has an accountable owner.Ownership

Named business and technical owners are assigned with verified contact details.

Every production instance has a workload identity.Identity

Unique non-human machine credentials (SPIFFE/mTLS) isolated per environment.

Every agent has an approved purpose.Governance

Documented purpose contract defining permitted business tasks and boundaries.

Model use is approved and versioned.Models

LLMs are routed through governed gateways with pinned version numbers.

Tool access is scoped.Tools

APIs and MCP tools enforce strict schema parameters and read/write classification.

Context access is permission-aware.Context

RAG search preserves original source user ACLs and data sensitivity rules.

Memory has origin, scope, and expiration.Memory

Long-term memory is validated before promotion and has automated TTL decay.

Material actions are evaluated before execution.Policy

Pre-action Policy Enforcement Point (PEP) evaluates Allow/Deny before tool runs.

Runtime budgets are bounded.Budgets

Hard monetary, token, compute, and retry caps enforced active during execution.

High-risk actions require appropriate approval.Approval

Financial or irreversible actions trigger human-in-the-loop review with evidence.

Complete episodes are traceable.Observability

OpenTelemetry spans link context, plan, model call, tool call, and outcome.

Evaluations run continuously.Evaluation

Continuous sampling evaluates trajectory correctness and policy adherence in prod.

Revocation works.Control

Global kill-switch suspends agent and revokes sub-agent delegation in under 5 seconds.

Side effects can be verified.Evidence

Database and external API modifications produce verified side-effect logs.

Incidents can be replayed.Observability

Full episode trajectories can be replayed deterministically for post-mortems.

Rollback is tested.Operations

Automated state rollback recovers systems from partial execution failures.

Production ownership is defined.Operations

SRE runbooks and incident response rotas assigned specifically for agent failures.

Strategic Architecture Decision Guide

You do not need to rebuild the entire stack

Most enterprises already have valuable components. Keep the components that work. Build what differentiates you. Buy what is operationally heavy. Integrate what already exists. Standardize controls across them.

Build when capability is strategically unique and proprietary

Apply This Strategy When:

Capability is strategically unique and competitive differentiator
Domain workflow logic is proprietary
Integration is deeply internal to custom legacy systems
Control requirements are unusual or military-grade
Internal platform team has capacity and expertise to operate 24/7
Layer Contracts & Interfaces

Production architecture depends on contracts between layers

Products do not become a stack merely because they are purchased together. They become a stack when their identities, policies, decisions, events, and evidence can be connected through explicit interfaces.

Layer 1 ContractAgent Applications and Workflows
Incoming Inputs:
Business Objective Request • User Identity • Target Systems Context
↓ Layer Transformation
Required Outputs:
Objective ID • Boundary Constraints • Approved Purpose Record
Layer 2 ContractAgent Framework and Runtime
Incoming Inputs:
Proposed Action Plan • Objective Context • State History
↓ Layer Transformation
Required Outputs:
Step Execution Event • Delegation Request • Termination Receipt
Layer 3 ContractModels and Intelligence
Incoming Inputs:
Model Request Payload • Sensitivity Classification • Cost Boundary
↓ Layer Transformation
Required Outputs:
Model Response • Usage Metrics Record • Selected Model Metadata
Layer 4 ContractTools and Capabilities
Incoming Inputs:
Tool Call Request Payload • Authority Token • Transaction Parameters
↓ Layer Transformation
Required Outputs:
Execution Receipt • Side-Effect Record • Verification Outcome
Layer 5 ContractContext and Knowledge
Incoming Inputs:
Context Query • Requester User Identity • Purpose Scope
↓ Layer Transformation
Required Outputs:
Context Bundle • Permission Decision Receipt • Provenance Map
Layer 6 ContractAgent Memory
Incoming Inputs:
Memory Write Candidate • Validation Proof • Scope Tags
↓ Layer Transformation
Required Outputs:
Memory Store Receipt • Hydrated Memory Context • Deletion Event
Layer 7 ContractAgent Identity and Ownership
Incoming Inputs:
Agent Registration Request • Ownership Attestation
↓ Layer Transformation
Required Outputs:
Agent Passport • Workload Identity Token • Revocation Credential
Layer 8 ContractAuthority and Delegation
Incoming Inputs:
Authority Request • Delegator Passport • Objective Scope
↓ Layer Transformation
Required Outputs:
Signed Authority Manifest • Delegated Lease Token • Revocation Tree
Layer 9 ContractPolicy and Runtime Enforcement
Incoming Inputs:
Proposed Action • Agent Passport • Runtime Context State
↓ Layer Transformation
Required Outputs:
Policy Decision (Allow/Deny/Approval) • Policy Decision Record • Enforcement Instruction
Layer 10 ContractHuman Escalation and Approval
Incoming Inputs:
Approval Request Object • Target Approver Role
↓ Layer Transformation
Required Outputs:
Approval Decision (Approved/Rejected/Modified) • Approval Receipt • Condition Constraints
Layer 11 ContractRuntime Budget Governance
Incoming Inputs:
Budget Allocation • Estimated Action Cost • Current Spend Telemetry
↓ Layer Transformation
Required Outputs:
Budget Authorization Decision • Cost Reservation Receipt • Intervention Event
Layer 12 ContractAgent Observability
Incoming Inputs:
Span Event Telemetry • Agent Passport Header • State Transition Payload
↓ Layer Transformation
Required Outputs:
Episode Trace Record • Trajectory Graph • Anomaly Alert
Layer 13 ContractEvaluation
Incoming Inputs:
Completed Episode Trace • Golden Reference Dataset • Evaluation Criteria
↓ Layer Transformation
Required Outputs:
Evaluation Scorecard • Failure Analysis Report • Release Decision Gate
Layer 14 ContractAudit and Decision-Grade Evidence
Incoming Inputs:
Episode Events • Policy Receipts • Verification Records
↓ Layer Transformation
Required Outputs:
Evidence Manifest • Decision Receipt • Compliance Audit Package
Layer 15 ContractDeployment and Operations
Incoming Inputs:
Deployment Package • Environment Configuration • Kill-Switch Signal
↓ Layer Transformation
Required Outputs:
Deployment Manifest • Health Status Telemetry • Revocation Confirmation
Enterprise Accountability Map

Every layer needs a named enterprise owner

An unowned layer becomes an uncontrolled boundary. Assigning explicit organizational responsibility ensures incident response, policy enforcement, and SLA alignment.

Layer #Architecture LayerLikely Enterprise OwnerRisk of Missing Ownership
Layer 1Agent Applications and WorkflowsBusiness owner and product teamTeams start with a model and tools but cannot clearly state what success or acceptable autonomy means.
Layer 2Agent Framework and RuntimeAI platform teamInfinite retry loops consuming massive API costs on broken tools.
Layer 3Models and IntelligenceModel governance and AI platform teamData leakage by sending PII to public, non-compliant third-party models.
Layer 4Tools and CapabilitiesPlatform engineering and application ownersAgent calling a write endpoint with unvalidated parameters causing database corruption.
Layer 5Context and KnowledgeData, knowledge, and security teamsAgent retrieves salary records because vector search indexed files without ACL metadata.
Layer 6Agent MemoryAI platform and data governance teamAn agent hallucinates a customer discount, stores it in long-term memory, and honors it for months.
Layer 7Agent Identity and OwnershipIAM and security teamRogue agents created by shadow IT running under developer personal credentials.
Layer 8Authority and DelegationSecurity, risk, and business ownersAn agent with database read credentials delegates unrestricted write commands to a child worker.
Layer 9Policy and Runtime EnforcementSecurity, compliance, and architecture teamsPost-action logging alerts security 2 hours AFTER an unauthorized wire transfer occurred.
Layer 10Human Escalation and ApprovalBusiness and risk ownersApproval spam causing human reviewers to blind-approve all requests without reading evidence.
Layer 11Runtime Budget GovernanceFinOps, platform, and business ownerSurprise $50k cloud bill caused by recursive sub-agent planning loop over a weekend.
Layer 12Agent ObservabilitySRE and AI platform teamSecurity incidents cannot be investigated because logs show LLM responses but not tool parameters.
Layer 13EvaluationAI quality and domain ownersUpdating a system prompt breaks 30% of downstream tool call parameters unnoticed.
Layer 14Audit and Decision-Grade EvidenceRisk, compliance, and security teamInability to defend an agent decision during a regulatory inquiry because raw logs were rotated out after 30 days.
Layer 15Deployment and OperationsCloud, platform, and SRE teamsDeploying an updated prompt directly to production without canary testing breaks live customer sessions.
Stack Vulnerabilities & Failure Analysis

What fails when a layer is missing?

Architectural gaps produce compound failure modes. Here is what happens in enterprise production when critical control layers are omitted.

No Registry

"The organization cannot identify or locate every active production agent."

No Ownership

"Incidents have no accountable responder or defined SLA runbook."

No Authority Layer

"Raw API credentials become a substitute for permission, exposing write endpoints."

No Context Controls

"Agents retrieve information outside original requester or purpose boundaries."

No Memory Governance

"Generated hallucinated conclusions silently become trusted future context."

No Runtime Enforcement

"Policies describe what should happen but cannot stop what is happening."

No Budget Controls

"Agents enter expensive retry, planning, or delegation loops."

No Evaluation

"A successful LLM HTTP 200 run is mistaken for a correct business outcome."

No Decision Evidence

"The organization cannot reconstruct why an action occurred during an audit."

No Instant Revocation

"Access persists through sessions, delegations, tools, and credentials after compromise."

The Governed Execution Layer

Enterprise Hypervisor

Enterprise Hypervisor connects agent identity, authority, context, models, tools, policies, budgets, approvals, evidence, and outcomes at runtime.

Commercial Bridge: Keep the platforms that already work (OpenAI, Anthropic, AWS, LangChain, CrewAI, Okta). Add one control and evidence layer across them.
Runtime Workflow:
Agent Proposes Action
Enterprise Hypervisor Evaluates: Identity • Authority • Context • Policy • Budget • Approval
Allow • Deny • Require Approval • Degrade • Terminate
Decision Receipt & Evidence Record
Specialized Systems for the Governed Stack

Capability-First Unifying Systems

Agent population and authorityAvailable

EnterpriseACP

Register agents, assign ownership, define authority, manage lifecycle, and revoke access.

Unifying Capability
Governed executionAvailable

Enterprise Hypervisor

Evaluate whether a specific action may occur now under current identity, policy, context, budget, and approval conditions.

Unifying Capability
Decision authorizationAvailable

Decision Hypervisor

Control high-impact decisions, decision objects, review rules, evidence requirements, and outcome conditions.

Unifying Capability
Context and evidenceAvailable

EnterpriseContextGraph

Preserve permission-correct, time-correct context, lineage, decisions, policies, and evidence.

Unifying Capability
Context deliveryAvailable

EnterpriseContextServer

Deliver the authorized context required for the approved objective.

Unifying Capability
Tools and capabilitiesAvailable

AgentPlugin

Define, register, mediate, and verify agent capabilities and tool operations.

Unifying Capability
ModelsAvailable

LocalModels

Discover, evaluate, deploy, route, and govern models under enterprise control.

Unifying Capability
Model packagingIn development

OpenModels

Package model artifacts, manifests, provenance, and reproducible distributions.

Unifying Capability
Analytics and observabilityAvailable

HeadlessAnalytics

Analyze complete agent episodes, evaluations, costs, interventions, failures, and evidence.

Unifying Capability
VerificationAvailable

EnterpriseVerification

Verify execution, controls, side effects, and required outcomes.

Unifying Capability
Physical executionDesign partner

EON Robotics

Apply governed execution to robots, machines, edge systems, and physical environments.

Unifying Capability
Enterprise Advisory & Architecture Review

Ready to govern your production agent fleet?

Book a 1-on-1 architecture review session with a Principal Enterprise AI Architect, or evaluate your stack controls with our 20-question diagnostic.