The Agentic Architecture Framework: A TOGAF-Style Method for AI & Data

Enterprise architecture has a trusted method for designing technology change: TOGAF and its Architecture Development Method (ADM). But TOGAF was built for a world of deterministic systems — applications with known inputs, predictable outputs, and release cycles you could plan on a Gantt chart.

AI systems don’t behave that way. They are probabilistic. They learn, drift, hallucinate, and compound cost silently in multi-step loops. The data they depend on is rarely ready. And the governance models built for traditional software either block AI entirely or wave it through unexamined.

This article proposes the Agentic Architecture Framework — a practitioner’s method for architecting AI systems on a real data foundation. If you know TOGAF, you already know its shape: principles, a development method, a content framework, reference architectures, and governance. What’s different is what it confronts: data readiness, evaluation, model economics, and human judgment as first-class architectural concerns.

This isn’t a standard. It’s a field method — drawn from shipping production AI and data platforms at scale — for anyone who has to answer the question: where does AI actually earn its keep here, and how do we build it so it keeps earning it?

Why TOGAF alone isn’t enough for AI

TOGAF’s ADM is excellent at what it was designed for: aligning business, data, application, and technology architecture through a disciplined, iterative cycle. But apply it directly to an agentic AI initiative and four gaps appear:

  1. Data readiness is assumed, not assessed. TOGAF’s Phase C covers data architecture, but AI needs something stricter: a verdict on whether the data is fit for machine consumption — quality, lineage, freshness, access, and bias.
  2. Evaluation has no home. Traditional architecture assumes you can specify behavior. AI behavior is emergent. Without a dedicated evaluation discipline — eval harnesses, red-teaming, offline-to-online validation — you’re shipping hope.
  3. Cost is not an architecture concern. TOGAF plans migration cost; it doesn’t account for token spend compounding across agent loops at 3 AM. In A
  4. Governance is a gate, not a loop. Phase G governs implementation, but AI systems need continuous governance — drift detection, feedback loops, and human-in-the-loop placement that evolves as the system learns.

Most AI failures are data
failures wearing an AI costume.

In AI, cost is architecture.

The Agentic Architecture Framework keeps TOGAF’s skeleton and rebuilds the muscle for AI.

The framework at a glance

Six components, mirroring the structure architects already know:

  1. Architecture Principles — the 12 rules every AI initiative is judged against.
  2. The AI Delivery Cycle — the method itself: Preliminary + six phases (A–F), iterative, not waterfall.
  3. The Content Framework — what you produce in each phase (artifacts and deliverables).
  4. Reference Architectures — reusable blueprints: the RAG blueprint, the agent blueprint, the eval harness blueprint.
  5. Governance — mapped directly to the NIST AI Risk Management Framework (Govern, Map, Measure, Manage).
  6. The Maturity Model — five levels, from Exploring to Leading, so you can locate any organization honestly.

Component 1: The 12 principles

Principles are the framework’s conscience. Every design decision should be traceable to one of these:

  1. Business value first, models second. No AI initiative exists without a named business outcome and an ROI thesis.
  2. Data quality is an AI prerequisite, not a data problem. If the data isn’t fit, the AI isn’t ready. Say so early.
  3. Evaluation before enthusiasm. No model, agent, or pipeline reaches production without an eval harness that can catch regressions.
  4. Humans in the loop by design. Place human judgment where the cost of error is highest — deliberately, not as an afterthought.
  5. Cost as an architecture concern. Token spend, compute, and data pipeline cost are designed in, with budgets per workflow.
  6. Deterministic where possible, probabilistic where valuable. Don’t pay for an agent to do what a rule, a lookup, or a smaller model does better.
  7. Lineage is non-negotiable. Every AI output should be traceable to the data and model version that produced it.
  8. Governance that accelerates. Guardrails ship with the system — as code, as evals, as contracts — never as a committee that meets after launch.
  9. Start narrow, earn the right to expand. One workflow, done well, beats five pilots that never reach production.
  10. Security and privacy by construction. Data access, PII handling, and model exposure are architectural inputs, not audit findings.
  11. Observability from day one. If you can’t see what the system did and why, you can’t operate it.
  12. Learn in production. Feedback loops, drift detection, and retraining aren’t maintenance — they’re the system working as designed.

Component 2: The AI Delivery Cycle

The method. Iterative, tailorable, and familiar to anyone who’s run an ADM cycle. Requirements management runs continuously through the center — same as TOGAF.

Preliminary — Principles & Posture. Establish the AI principles (above), define risk appetite, stand up the governance scaffolding, and tailor the framework to the organization. Maps to NIST AI RMF Govern.

Phase A — Vision & Value. Define the business outcome in plain language. Where does AI earn its keep here? Produce the AI Vision: the outcome, the ROI thesis, the explicitly-named non-goals. Get stakeholder sign-off before a single model is evaluated.

Phase B — Data Foundation. Assess data readiness brutally: quality, completeness, freshness, lineage, access rights, and bias. Remediate or descope. Output: the Data Readiness Assessment — a go/no-go verdict per data source. This phase kills more bad AI projects than any other, and that’s its job.

Phase C — Intelligence Architecture. Design the AI itself: model selection and routing, retrieval architecture, agent decomposition, tool use, and the deterministic-vs-probabilistic split. Build-vs-buy decisions live here. Output: the Intelligence Architecture with cost projections per workflow.

Phase D — Evaluation. Build the eval harness before production: offline benchmarks, golden datasets, red-teaming, and the promotion criteria for online testing. Define what “good” means numerically — and what triggers a rollback. Output: the Evaluation Plan with pass/fail thresholds.

Phase E — Governance & Deployment. Map risks (NIST AI RMF Map), measure them (Measure), and put controls in place (Manage). Place human-in-the-loop checkpoints where error cost is highest. Deploy in stages behind the eval harness. Output: the Risk Register, the HITL design, and the staged rollout plan.

Phase F — Operate & Evolve. Run it like a production system: cost monitoring against the budgets set in Phase C, drift detection, feedback loops feeding back into evals, and a defined change process for model updates. Output: the Operations Runbook — and the input to the next cycle.

Then cycle again. AI systems are never “done”; each cycle deepens the data foundation, sharpens the evals, and expands scope only where value is proven.


Component 3: The Content Framework

Every phase produces artifacts (working products) and deliverables (signed-off outputs). The core set:

PhaseKey deliverableSupporting artifacts
PreliminaryAI Principles sign-offRisk appetite statement, governance charter
A — Vision & ValueAI VisionROI thesis, non-goals list, stakeholder map
B — Data FoundationData Readiness AssessmentData quality scorecard, lineage map, access matrix
C — Intelligence ArchitectureIntelligence ArchitectureModel routing table, cost projections, build-vs-buy analysis
D — EvaluationEvaluation PlanGolden datasets, red-team findings, promotion criteria
E — Governance & DeploymentRisk Register + rollout planHITL design, control mappings, rollback procedures
F — Operate & EvolveOperations RunbookCost dashboards, drift reports, feedback-loop specs

A deliverable is reviewed and signed off. An artifact is lived in. Confuse the two and you get shelfware.


Component 4: Reference architectures

Don’t redesign what repeats. The framework ships with three starting blueprints, tailored per engagement:

  • The RAG blueprint. Ingestion → chunking → embeddings → vector store → retrieval → generation → eval. With the data-readiness gates from Phase B baked into ingestion, and the eval harness from Phase D wrapped around retrieval quality.
  • The agent blueprint. Planner → tools → memory → human checkpoints → output validation. With loop budgets (max steps, max cost per run), deterministic fallbacks, and HITL placement per Principle 4.
  • The eval harness blueprint. Golden dataset → offline eval → shadow mode → canary → production monitoring. The promotion ladder every AI system climbs, with rollback criteria at each rung.

Component 5: Governance — the NIST AI RMF mapping

Governance in this framework isn’t a separate track; it’s woven through the cycle, mapped to the NIST AI Risk Management Framework’s four functions:

  • Govern → Preliminary phase. Culture, accountability, risk appetite, and the principles themselves.
  • Map → Phases A–C. Context, categorization, and risk identification as the vision, data, and architecture take shape.
  • Measure → Phase D. Evaluation, red-teaming, and quantified risk — the numbers behind the verdicts.
  • Manage → Phases E–F. Controls deployed, risks tracked, feedback loops running in production.

If your organization already runs a NIST AI RMF program, this framework slots into it. If it doesn’t, the framework gives you the program.


Component 6: The Maturity Model

Five levels. Be honest about where you are — the framework meets you there:

  1. Exploring. AI curiosity, no production systems. Focus: Preliminary + Phase A. Pick one outcome.
  2. Experimenting. Pilots and demos exist. Focus: Phase B + D. Kill what the data can’t support; eval what survives.
  3. Piloting. First production AI workloads. Focus: Phase E. Governance and HITL are the bottleneck — fix them.
  4. Scaling. Multiple AI systems in production. Focus: Phase F + reference architectures. Cost control and reuse separate winners here.
  5. Leading. AI is a managed capability with feedback loops, mature evals, and governed expansion. Focus: cycle speed — how fast can you run the Delivery Cycle well?

Most enterprises sit at 2. Most AI strategies assume they’re at 4. The maturity assessment exists to close that gap before it closes your budget.

How to use it

  1. Start with the maturity assessment. One workshop, honest scoring, no vanity.
  2. Run Preliminary once — principles and posture are enterprise-level, not per-project.
  3. Run the cycle per initiative — Phases A through F, tailored. A narrow internal tool gets a light pass; a customer-facing agent gets the full treatment.
  4. Build the reference architectures as you go. Your second RAG system should be 70% reuse.
  5. Let Phase F feed Phase A. The learnings from operations are the most valuable input to the next vision.

Closing

TOGAF gave enterprise architecture a shared language for the deterministic era. The agentic era needs its own — one where data readiness is a verdict, evaluation is a discipline, cost is designed in, and governance ships with the system instead of chasing it.

That’s what this framework is for. It’s a starting point, deliberately: tailor it, argue with it, and let production teach you the rest.

The best framework is the one your teams actually run — cycle after cycle, until “where does AI earn its keep?” is a question you can answer with evidence.


Jugal Shah writes about production AI, agentic systems, and enterprise data architecture at aideeva.com, where he publishes The Agentic Enterprise newsletter.

Comments

Thanks for the comment, will get back to you soon… Jugal Shah