Skip to content

Research behind the standard

The AI Workflow Standard did not start from a blank page. Two research reports established what already existed, where it fell short, and which of the draft’s own assumptions did not survive contact with primary sources. Both are preserved in the repository under docs/research/ and are downloadable from this site. They are historical context: later accepted decisions supersede their recommendations, and the current requirements live in the standard, the implementation profile and the capability matrix.

Do vendor-neutral workflow standards for AI and agentic systems already exist, and if not, what exactly is missing? Its answer was that workflow standards certainly exist, but there is no single, broadly accepted, vendor-neutral standard for an end-to-end AI workflow. The report surveyed process notations, executable workflow languages, agent definition formats, interoperability protocols and governance bodies, mapped them against the needs of agentic work, and proposed a reference model and roadmap.

Its working definition became the basis of clause 1 of the standard: an AI workflow is a governed, stateful execution in which human, agentic and deterministic participants cooperate toward a declared intent, using delegated capabilities within explicit authority and policy boundaries, while producing durable artifacts, execution records and evidence sufficient to verify the outcome. Its shortest statement of purpose is that the standard defines what remains invariant while an autonomous execution is allowed to vary.

Family What it contributes What it lacks for agentic work
BPMN 2.0, DMN 1.5, CMMN 1.1 Mature vocabulary for tasks, events, gateways and lanes; deterministic, inspectable decisions; adaptive case management with discretionary work and milestones Topology fixed in advance; no notion of model-mediated decisions, plan mutation or cross-agent delegation
WfMC, XPDL, Wf-XML The historical reference model and interchange format Teaches that notation, executable semantics and interchange are three different standardization problems
Open Workflow Specification 1.0.3 The strongest candidate for a deterministic execution substrate: switch, fork, wait, retry, subworkflow, with a conformance test kit Defines no autonomy, intent or authority semantics
OpenAPI Arazzo 1.1.0, AsyncAPI Standard descriptions of API-operation sequences Not an agent or workflow language
Agent Spec 26.3.0 The closest thing to portable agent and flow interchange Vendor-originated governance; adapter equivalence unproven
MCP 2026-07-28 The capability, tool and context boundary Should be a binding, not something to replace
A2A 1.0 The delegated-task boundary: agent cards, tasks, messages, artifacts and a task state model Does not define parent intent, topology or success criteria
AG-UI draft Human and application interaction events, including human-in-the-loop Draft, not ratified, and not the workflow definition
OpenTelemetry GenAI conventions Observability vocabulary for invoking workflows, agents and tools Terminology still contested; align rather than fork
OASF, AGNTCY Participant and capability descriptions for discovery and matching No execution semantics
AAIF Workflows and Process Integration working group The most directly relevant venue; its charter names handoffs, persistence, retries, idempotency, recovery, human approval and portability Chartered in May 2026; no published specification at the time of the survey
NIST AI Agent Standards Initiative, W3C community groups, ISO/IEC SC 42, IEEE P3777 Agent identity and authority as formal policy; the governance, lifecycle and evaluation envelope Policy and incubation, not executable specifications; workflow standards should reference them rather than duplicate them

The landscape report ranked the unresolved problems. The high-priority ones map directly onto clauses of AIWS-001.

  1. No common metamodel. Workflow, run, task, step, agent, session, handoff, plan and artifact meant different things in every framework. Clause 4 fixes the vocabulary.
  2. Intent had no success contract. Existing standards encode actions and conditions far more strongly than the desired outcome and acceptable evidence. Intent is not the same thing as an initial prompt. Clause 6 requires criteria, acceptance methods and evidence to be declared up front.
  3. Authority is not authentication. Knowing who an actor is says nothing about what an autonomous participant may decide, delegate or cause. Clause 8 defines grants, attenuation and enforcement at a trusted boundary.
  4. Plans mutate. An agentic execution constructs part of its own future topology, and nothing governed when an agent could create tasks or revise a plan. Clause 6 governs plan revisions inside a declared envelope.
  5. Durable state was not portable. Checkpoint, suspend, resume and recovery were entirely per-runtime. Clause 11 defines what must be durable and how recovery behaves; migration stays experimental.
  6. Side effects, retries and compensation. Retrying “reason about this document” is different from retrying “pay invoice.” Clause 10 requires effect classification, attempt records, retry permission and separate compensation.
  7. Completion is not verification. COMPLETED cannot safely mean “the agent said it is finished,” and an artifact is not evidence. Clauses 7 and 13 separate execution, verification and acceptance.
  8. Conformance semantics. A portable YAML file is worthless if two runtimes interpret delegation, failure, retry and completion differently. Clause 16 and Annex A define assessment.

Medium-priority gaps, approval ownership and expiry, memory versus state versus conversation, capability matching, resource envelopes across tokens, money, time, tool calls and delegation depth, and evaluation semantics, became clauses 9, 12 and 13 and parts of 14.

The second report asked a harder question of the edition 0.1 draft: which requirements are supported by existing standards or implementation evidence, which need correction, and which remain design choices? It read primary sources, from IETF RFCs and the Cedar authorization algorithm to Temporal, LangGraph, Step Functions, Stripe, WS-HumanTask, PROV-O, SACM, RATS and SCXML, and found several places where the draft had assumed more than the sources provide. Each finding produced a decision recorded in Part A of the standard.

Finding Source Decision in edition 0.2
A token-exchange actor chain is informational; it is not a permission chain, so carrying identities does not restrict delegation RFC 8693 Attenuation is computed and enforced, not inferred from identity
Rich authorization requests define no generic way to compare two permission objects; an intersection described in prose is insufficient RFC 9396 A versioned authority profile with a three-result containment check; an unsupported comparison is indeterminate, never permission
Cached introspection results go stale RFC 7009, RFC 7662 Declared, measurable revocation freshness, clock tolerance and in-flight boundaries
An authorization engine that skips a policy on error can return Allow alongside a policy error Cedar Mandatory policy checks block the affected dispatch when they cannot answer
Runtime continuation chains show that “run” was too tightly coupled to one engine invocation; a linked new run must not reset an exhausted budget Temporal Mission, ContractRevision and ExecutionSegment; cumulative accounting across runs
Resuming an interrupt reruns the node containing it; resume does not always mean continuing after the interrupted line LangGraph Recovery revalidates and replays recorded decisions rather than assuming a resume point
Engine delivery guarantees are scoped, and an idempotency key is a contract with namespace, retention, fingerprint and concurrency rules, not a field Step Functions, Stripe, transactional outbox Delivery claims are stated separately from business-effect claims; retry needs proof or a target guarantee
A durable pre-dispatch record does not make a remote write atomic Outbox pattern Outcome unknown is distinguished from effect not applied; UNKNOWN is a first-class state
A review that rejects is a successful review, not a failed execution WS-HumanTask Acceptance is its own record and state, separate from execution
A single wait reason loses information when branches wait in parallel SCXML Independent WaitCondition records with explicit ALL and ANY joins
Conformance tests, quality benchmarks and security evaluations answer different questions Open Workflow test kit, agent benchmarks Four separate assessment tracks: control conformance, outcome quality, adversarial robustness, interoperability

The report also identified reuse candidates rather than inventions: PROV-O for provenance, with artifacts as entities, attempts as activities and participants as agents; SACM for assurance rationale and counter-evidence; the RATS separation of evidence producer, technical assessor and acceptance authority; WS-HumanTask for approval ownership, claim and escalation; and the OWASP transaction-authorization comparison of what was approved with what is about to be dispatched. None became a mandatory core dependency.

  • Do not invent another notation, transport or telemetry protocol. BPMN, MCP, A2A and OpenTelemetry exist; replacements would fragment rather than standardize. The standard defines bindings to them instead.
  • Do not standardize prompts, model internals, or the choice of model, framework, database, vendor or interface.
  • Do not equate portability with identical execution. Conformance means different valid trajectories preserve the same intent, policy invariants, authority boundaries, lifecycle semantics and verification obligations.
  • Do not start from a large serialization schema. Establish semantic interoperability first and a wire format second. The SDK’s closed JSON schema came after the semantics, as a profile.
  • Do not let the planning model’s own judgment be the authority-containment mechanism.
  • Keep experimental alignments experimental. Mission-bound authorization and delegation receipts are individual Internet-Drafts with no IETF endorsement. Cross-runtime migration is deferred because the reviewed runtimes establish local recovery, not universal checkpoint interchange.
  • Make no prompt-injection-proof claim. The research treated adaptive-attack results as disconfirming evidence that defenses can degrade under pressure, and assigned adversarial robustness its own assessment track.

Both reports state their limits. No runtime, policy engine, benchmark or adapter was installed or tested during the research; standards inspection was selective; some full texts were unavailable. The comparative matrices are analytical mappings, not official ratings. The supplemental report closes by noting that its findings do not turn the draft into a ratified or tested standard.

The landscape report proposed a modular architecture so that protocol churn never forces a major version of the standard: a core of terminology, metamodel, lifecycle and conformance; a governance profile of intent, authority, autonomy, budgets and evidence; bindings for MCP, A2A, AG-UI, API workflows and asynchronous events; execution profiles for Open Workflow, Agent Spec and BPMN or CMMN mapping; and an observability profile aligned with OpenTelemetry. It recommended testing the standard against deliberately diverse use cases, such as incident escalation, vulnerability remediation, corrective action, invoice payment with irreversible effects, multi-agent research and long-running regulatory approval, because a standard tested only against a research agent that summarizes documents misses most of the hard semantics.

The supplemental report recommended a design-decision review before adding more requirements, then a small reference contract and simulator with four fixtures: a bounded document review with human acceptance, a delegated task whose authority is revoked while it waits, an external write that commits before its acknowledgement is lost, and a continuing segment that must preserve consumed budget and unresolved work. Those became edition 0.2, the shared scenario corpus in Annex A, and eventually the three native SDKs, whose conformance scenarios still exercise exactly those cases.