# Workflow Standards for AI and Agentic Systems

## Executive summary

**Yes, workflow standards exist today. No, there is not yet a single, broadly accepted, vendor-neutral standard for an end-to-end AI/agentic workflow.** As of **September 4, 2026**, the standards landscape is better understood as a stack of complementary specifications that solve different portions of the problem.

At the mature process-modeling layer, **BPMN, DMN, and CMMN** provide standardized models for processes, decisions, and adaptive case work. BPMN is standardized as ISO/IEC 19510 as well as by OMG; DMN's current formal OMG release is 1.5; and CMMN's formal release remains 1.1. These are mature conceptual standards, but none was designed around model-driven participants that may dynamically plan, delegate, discover tools, or modify their own execution path. citeturn12search0turn4view2turn3search2

At the executable-workflow layer, the CNCF **Open Workflow Specification**, renamed in June 2026 from **Serverless Workflow**, is particularly relevant. Version **1.0.3** provides a vendor-neutral YAML/JSON DSL with control flow, events, waits, retries/error handling, forks, loops, subworkflows, scripts, containers, HTTP/OpenAPI, gRPC and AsyncAPI operations. It is arguably the strongest existing candidate for the **deterministic execution substrate** of an AI workflow standard, but it does not itself define the semantics of autonomy, intent, authority delegation, AI planning, or probabilistic decision making. citeturn9view0turn0search15

At the agent-definition layer, Oracle's **Agent Spec/Open Agent Specification** is the closest existing effort to a portable declaration of both **Agents** and **Flows**. Its current repository identifies release **26.3.0**; it supports JSON/YAML descriptions and adapters/runtimes for multiple agent frameworks. Its significance is high, but its governance remains project/vendor-originated rather than a neutral international or foundation standard, and implementation-independent conformance semantics are still evolving. citeturn6search1turn6search0turn6search4

At the interoperability boundaries, an increasingly coherent stack has emerged:

| Boundary | Leading specification |
|---|---|
| Agent/application ↔ tools, resources, context | **MCP** |
| Agent ↔ agent / delegated task | **A2A** |
| Agent ↔ human-facing application | **AG-UI** |
| Agent/workflow ↔ telemetry systems | **OpenTelemetry GenAI semantic conventions** |
| Agent ↔ capability/discovery metadata | **OASF / AGNTCY** |
| Workflow ↔ APIs/events | **OpenAPI, Arazzo, AsyncAPI** |
| Enterprise process semantics | **BPMN / DMN / CMMN** |
| Executable workflow | **Open Workflow Specification** |

MCP's current normative specification is dated **2026-07-28**. A2A reached **1.0** in March 2026 and became an AAIF-hosted project in August 2026. AG-UI is promising but its intended 1.0 behavioral specification is explicitly still **draft and not ratified**. OpenTelemetry now defines `invoke_agent`, `invoke_workflow`, `execute_tool` and related GenAI telemetry concepts, but active issues in 2026 show that workflow/agent terminology and lifecycle semantics are still being worked out. citeturn2search0turn1search1turn13search3turn16view1turn21search1turn21search11

Most importantly, the **Agentic AI Foundation Workflows & Process Integration Working Group** was approved on **May 13, 2026** specifically to establish a “shared workflow model and common terminology for agentic AI systems,” including tasks, tool calls, branches, handoffs, long-running execution, persistence, retries, idempotency, failure recovery, human approval, workflow interchange, and portability across runtimes. Its charter explicitly identifies the difficulty of translating deterministic workflow models into the nondeterministic models used by autonomous agents. This is currently the strongest standards-community evidence that the missing abstraction is real rather than merely a product-framework problem. citeturn16view3

The central conclusion of this research is therefore:

> **The industry does not need another general-purpose transport protocol or another framework-specific workflow DSL. It needs a common semantic model above the emerging protocol stack.**

That missing model should define at minimum:

**Intent → Workflow Definition → Run → Plan → Task → Step Attempt → Participant → Capability → Authority → State → Artifact → Evidence**

and it should establish normative semantics for **dynamic task creation, delegation, authority attenuation, waits, retries, idempotency, checkpointing, replay, compensation, human intervention, success verification and nondeterministic decision recording**.

The most consequential gaps are not YAML syntax. They are **meaning**:

1. What exactly is a workflow versus an agent, plan, task, step, conversation, session, or run?
2. How does a goal or **intent** become a verifiable execution contract?
3. How much autonomy does an agent possess, and how is that authority delegated?
4. What happens when the workflow topology is discovered during execution rather than known beforehand?
5. How is durable state transferred across runtimes?
6. How do retries and compensation work when actions have real-world side effects?
7. What evidence demonstrates that the workflow succeeded rather than merely stopped?
8. How can the same definition run on two implementations with predictably equivalent semantics?

Those should be the core of an **AI Workflow Standard**, with MCP, A2A, AG-UI, OpenTelemetry, Open Workflow, Arazzo, BPMN/DMN/CMMN and related standards reused rather than replaced. The AAIF W&PI charter explicitly favors standards that layer on top of existing execution environments rather than replacing them. citeturn16view3

## Standards landscape and detailed profiles

The following profiles distinguish **standards**, **specifications/protocols**, **community incubation efforts**, and **implementation projects**. This matters: an open-source project can be extremely influential without defining portable semantics, while an ISO/OMG specification can define semantics without supplying a runtime.

The example snippets below are deliberately minimal illustrations of each model rather than complete conforming documents.

**BPMN — Business Process Model and Notation**

**Purpose and scope.** BPMN standardizes the representation of business processes and collaborations. Its core abstraction is a process containing flow elements such as activities, events and gateways, connected through sequence and message flows; pools and lanes represent participants or organizational responsibility. BPMN was deliberately designed to bridge human-readable business process modeling and implementable process semantics. OMG's formal specification is **BPMN 2.0.2**, and ISO/IEC 19510:2013 is based on BPMN 2.0.1 and remains current following ISO's review. Governance is through OMG and, for ISO/IEC 19510, ISO/IEC. citeturn4view0turn12search0

Its strengths for AI workflows are its mature vocabulary for **tasks, events, gateways, subprocesses, parallelism, participant lanes and messages**. Its limitation is equally important: the process model normally identifies the allowable control topology. An autonomous participant dynamically inventing a new subtask, selecting an unknown agent at runtime, changing its plan, or deciding when sufficient evidence has been gathered is not BPMN's native center of gravity. That makes BPMN highly reusable as a business-facing notation, but insufficient by itself as the semantic model for agentic execution. citeturn12search0

```text
Start
  ↓
Review Request [Human Task]
  ↓
Exclusive Gateway
  ├─ approved → Execute Change [Service Task]
  └─ rejected → Notify Requester
  ↓
End
```

Primary sources: [OMG BPMN](https://www.omg.org/spec/BPMN/2.0.2/About-BPMN), [ISO/IEC 19510](https://www.iso.org/standard/62652.html). citeturn4view0turn12search0

**DMN — Decision Model and Notation**

**Purpose and scope.** DMN standardizes the representation of decisions and business logic that can complement process models such as BPMN. Its principal concepts include decisions, input data, business knowledge, decision requirements and the FEEL expression language/decision-table model. The current formal OMG release is **DMN 1.5**, adopted in 2024; OMG also has later beta work in progress, demonstrating continuing evolution rather than a frozen specification family. citeturn4view2turn11search0turn11search1

DMN is especially valuable in an AI workflow standard because it lets a workflow explicitly distinguish a **deterministic, inspectable policy decision** from a probabilistic model decision. A high-risk workflow should often use DMN-like logic for mandatory approval or policy gates even if AI determines the underlying plan. citeturn3search1

```text
decision RequireHumanApproval
inputs:
  risk_class
  external_side_effect

rule:
  if risk_class = "high"
     or external_side_effect = true
  then approval_required = true
```

Primary source: [OMG DMN 1.5](https://www.omg.org/spec/DMN/1.5/About-DMN). citeturn4view2

**CMMN — Case Management Model and Notation**

**Purpose and scope.** CMMN addresses adaptive and case-oriented work in which not every activity can be completely sequenced in advance. Its formal OMG release is **1.1**. Concepts include the **Case**, Case Plan Model, stages, tasks, milestones, event listeners, sentries and case information. CMMN therefore has more conceptual affinity with agentic work than BPMN does where execution depends heavily on circumstances encountered during the case. citeturn3search2

For an AI workflow standard, CMMN is particularly valuable prior art for **optional/discretionary work, events that enable future work, milestones rather than strictly prescriptive sequences, and long-lived stateful cases**. It still does not define AI autonomy, model decisions or cross-agent delegation protocols, but its adaptive-work model should be studied before inventing new concepts. citeturn3search2

```text
Case: Investigate Incident

Available work:
  - Review evidence
  - Request additional evidence
  - Escalate to specialist   [discretionary]
  - Interview employee      [discretionary]

Milestone:
  "Root cause established"
```

Primary source: [OMG CMMN 1.1](https://www.omg.org/spec/CMMN/1.1/About-CMMN). citeturn3search2

**Workflow Management Coalition artifacts**

The Workflow Management Coalition's historical work is critical prior art even though it is not a current AI-agent standards program. Its **Workflow Reference Model** decomposed workflow infrastructure into interfaces involving process definitions, client applications, invoked applications, workflow-engine interoperability, and administration/monitoring. WfMC also produced **XPDL**, an interchange format for process definitions, and **Wf-XML** for workflow interoperability. The official WfMC document archive continues to expose the reference model, audit/interface material and conformance artifacts. citeturn10search3turn10search16

XPDL eventually represented BPMN constructs and was explicitly positioned around process-model interchange. This is a useful historical lesson: **graphical notation, executable semantics and interchange are different standardization problems** and should not be conflated in an AI workflow specification. citeturn10search1turn10search18

```xml
<WorkflowProcess Id="incident-review">
    <Activities>
        <Activity Id="triage"/>
        <Activity Id="investigate"/>
    </Activities>
</WorkflowProcess>
```

Primary sources: WfMC's [official site](https://www.wfmc.org/) and official public-document/XPDL materials. citeturn10search16turn10search3turn10search1

**Open Workflow Specification — formerly Serverless Workflow**

This is one of the most directly reusable technologies in the landscape. The project provides a **vendor-neutral, open workflow DSL** and currently publishes specification version **1.0.3**. In June 2026 the Serverless Workflow project was renamed **Open Workflow Specification** while retaining technical continuity. The project remains under CNCF/Linux Foundation governance; its CNCF status is Sandbox, meaning that although it has reached a 1.x specification, the project itself is still in CNCF's early-stage project category. citeturn9view0turn0search15turn8search10

Its DSL already encompasses many capabilities an AI workflow requires: **HTTP/OpenAPI, gRPC and AsyncAPI calls; event emission/listening; loops; forks and competition; switches; waits; try/catch; subworkflows; scripts; containers; and fault-handling constructs**. Several runtimes and SDKs implement the model. citeturn9view0turn0search7

Illustrative Open Workflow-style YAML:

```yaml
document:
  dsl: "1.0.3"

do:
  - gather:
      call: http
      with:
        method: GET
        endpoint: "${ .source }"

  - decide:
      switch:
        - continue:
            when: "${ .confidence < 0.8 }"
            then: investigate

  - investigate:
      run:
        workflow:
          name: deeper-review
```

The important standards-design conclusion is that Open Workflow already solves a large fraction of **durable deterministic orchestration syntax**. A new AI standard should strongly consider defining an agentic semantic profile or mapping around it rather than reinventing `switch`, `fork`, `wait`, `retry` and subworkflow primitives. citeturn9view0turn16view3

Primary source: [Open Workflow / Serverless Workflow specification site](https://serverlessworkflow.io/). citeturn9view0turn0search15

**OpenAPI Arazzo**

The OpenAPI Initiative's **Arazzo Specification 1.1.0**, published in May 2026, standardizes machine-readable descriptions of **sequences of API operations and the dependencies among them in order to achieve an outcome**. Version 1.1 added AsyncAPI integration, allowing described workflows to span synchronous API operations and asynchronous/event-driven interactions. Arazzo is governed by the OpenAPI Initiative under the Linux Foundation. citeturn14search1turn14search4turn14search20

Its core concepts include source descriptions, workflows, steps, parameters, outputs, dependencies, runtime expressions and success/failure actions. It is not an autonomous-agent workflow language, but it is very relevant whenever the “workflow” consists of interoperable API calls. citeturn14search1

```yaml
arazzo: 1.1.0
workflows:
  - workflowId: resolveIncident
    steps:
      - stepId: getIncident
        operationId: getIncident
      - stepId: updateIncident
        operationId: updateIncident
```

Primary source: [OpenAPI Arazzo Specification](https://spec.openapis.org/arazzo/latest.html). citeturn14search1turn14search20

**Open Agent Specification / Agent Spec**

Oracle's **Agent Spec** describes itself as a framework-agnostic, declarative configuration language for agentic systems. The repository currently identifies release **26.3.0**. It distinguishes two important runnable constructs: **Agents** and **Flows**, supports JSON/YAML serialization, and uses runtimes/adapters to execute definitions against frameworks such as LangGraph, AutoGen and CrewAI. Oracle's WayFlow is a reference runtime. citeturn6search1turn6search0

Its flow model includes multiple node types, including agent, flow, LLM, API, tool, map, start, end and branching nodes. The project also has tracing concepts for agents, flows, nodes, tools and LLM operations and supports OpenTelemetry export. citeturn6search2turn7view0turn6search4

```yaml
name: incident-investigation

agents:
  - name: investigator
    model: enterprise-model
    tools:
      - search_incidents
      - inspect_logs

flow:
  start: triage
  nodes:
    - name: triage
      type: agent
      agent: investigator
    - name: approval
      type: branch
```

Agent Spec is arguably the **closest current implementation-oriented project to an AI workflow interchange format** because it explicitly attempts to make both agents and flows portable. Its weakness as a universal standard today is governance and semantic authority: it is an open Oracle-led project, not yet a neutral standards body's ratified workflow standard, and framework adapters necessarily raise questions about precisely equivalent behavior across runtimes. citeturn6search1turn6search3

Primary source: [Oracle Agent Spec repository](https://github.com/oracle/agent-spec). citeturn6search1

**MCP — Model Context Protocol**

MCP's current normative specification is dated **2026-07-28**. It defines a protocol through which AI applications interact with capabilities and context exposed by MCP servers. Core abstractions include **Tools**, **Resources**, **Prompts**, and client/server capabilities such as sampling and related interaction mechanisms. The project is now hosted by the Agentic AI Foundation. citeturn2search0turn2search6turn2search10turn2search30turn13search14

MCP is important precisely because of what it **does not** attempt to be: it does not define an end-to-end workflow graph. It standardizes the capability/context boundary. The 2026 work also reflects movement toward long-running and agentic scenarios, while maintaining a deliberately composable protocol architecture. citeturn2search2turn2search7turn2search19

Illustrative JSON-RPC-style interaction:

```json
{
  "jsonrpc": "2.0",
  "id": 17,
  "method": "tools/call",
  "params": {
    "name": "inspect_incident",
    "arguments": {
      "incident_id": "INC-1042"
    }
  }
}
```

For an AI workflow standard, MCP should be treated as the standard **Capability Invocation** binding, not replaced with a competing tool protocol. citeturn2search6turn16view3

Primary source: [MCP Specification](https://modelcontextprotocol.io/specification/2026-07-28). citeturn2search0

**A2A — Agent2Agent Protocol**

A2A reached **1.0** in March 2026 and is now production-stable at the 1.0 protocol level. In August 2026 it became an AAIF-hosted project. Its governance lineage includes donation into Linux Foundation governance and a technical steering structure involving multiple major technology companies. citeturn1search1turn1search9turn13search3

A2A's canonical data model centers on **Agent Cards, Tasks, Messages and Artifacts**. A client sends messages that may initiate a Task, while results should be represented as Artifacts. The protocol supports streaming and long-running task interaction and specifies version negotiation independently of particular JSON-RPC, gRPC or HTTP+JSON bindings. citeturn16view0

A2A 1.0 defines task states including:

`SUBMITTED → WORKING → COMPLETED`

with terminal alternatives `FAILED`, `CANCELED` and `REJECTED`, plus interrupt states `INPUT_REQUIRED` and `AUTH_REQUIRED`. The protocol explicitly distinguishes terminal and interrupted states. citeturn17view1turn17view2

```json
{
  "task": {
    "id": "task-42",
    "contextId": "ctx-9",
    "status": {
      "state": "TASK_STATE_WORKING"
    }
  }
}
```

A2A is therefore an excellent candidate for the **delegated-task boundary** in an AI workflow standard. It still does not define the parent workflow's complete intent, topology, success criteria, or rules for deciding when delegation should occur. citeturn16view0turn17view3

Primary source: [A2A 1.0 Specification](https://a2a-protocol.org/latest/specification/). citeturn16view0

**AG-UI — Agent User Interaction Protocol**

AG-UI standardizes the event-based connection between an agent and a user-facing application. Its current behavioral specification is explicitly labeled **draft, not yet ratified**, describing intended 1.0 behavior without a compatibility promise until frozen. The protocol uses a request plus an ordered stream of typed events carrying text, tool activity, reasoning-related presentation data, shared state and progress. citeturn16view1

Its event model includes run-level and optional step-level lifecycle signals such as run start/finish/error and step start/finish, as well as tool-call and state-management concepts. It is particularly important for **human-in-the-loop** workflows in which a user may provide information, approve an operation or otherwise interact while the execution is in progress. citeturn1search22turn1search38turn1search26

```json
{
  "type": "RUN_STARTED",
  "threadId": "thread-12",
  "runId": "run-93"
}
```

AG-UI is therefore the strongest candidate in this landscape for the workflow's **human/application interaction binding**, but it should not be mistaken for the workflow definition itself. citeturn16view1

Primary source: [AG-UI draft specification](https://docs.ag-ui.com/spec/draft). citeturn16view1

**OpenTelemetry GenAI semantic conventions**

OpenTelemetry's GenAI work is attempting to standardize **how AI execution is observed**, not how it is orchestrated. OpenTelemetry itself became a CNCF Graduated project in May 2026, while the GenAI semantic conventions are still moving rapidly and have been moved to their own repository. citeturn21search20turn16view2

Current agent-related conventions include concepts such as:

```text
invoke_workflow
  ├─ invoke_agent
  │    ├─ chat
  │    └─ execute_tool
  └─ invoke_agent
```

The current `invoke_workflow` convention describes a coordinated process composed of multiple agents or GenAI operations and differentiates it from a standalone agent invocation. Open issues during 2026 continue to discuss nodes, tasks, workflow grouping, long-running lifecycle events and even the glossary for workflow/agent/session concepts. The dedicated GenAI repository has not yet published formal GitHub releases, another indication that the conventions are actively evolving. citeturn21search1turn21search5turn21search9turn21search14turn21search11turn21search3

This work is unusually valuable to an AI Workflow Standard because observability is forcing the ecosystem to answer the same ontological questions: **what constitutes an agent, workflow, node, task, tool call, conversation and execution?** The workflow standard should coordinate vocabulary with OpenTelemetry rather than creating incompatible terminology. citeturn21search11turn16view3

Primary source: [OpenTelemetry GenAI agent/workflow semantic conventions](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-agent-spans.md). citeturn21search1

**OASF / AGNTCY — Open Agentic Schema Framework**

OASF addresses another adjacent gap: how an agent describes **what it is and what it can do**. AGNTCY's Open Agentic Schema Framework defines extensible records describing agent attributes, skills, capabilities and relationships and is designed to work alongside technologies such as A2A and MCP. citeturn15search1turn15search10

```json
{
  "name": "incident-investigator",
  "skills": [
    "incident.triage",
    "log.analysis"
  ],
  "capabilities": [
    "mcp",
    "a2a"
  ]
}
```

It is not a workflow execution standard. Its likely role in a workflow standard is the **Participant/Capability Description** layer, especially for runtime discovery and matching a Task's requirements to an agent's declared capabilities. citeturn15search4turn15search22

Primary source: [AGNTCY OASF documentation](https://docs.agntcy.org/oasf/). citeturn15search1

**AAIF Workflows & Process Integration Working Group**

This is the most directly relevant current standards initiative.

The AAIF W&PI Working Group charter was approved **May 13, 2026** and updated **July 23, 2026**. Its mission is explicitly to establish a shared workflow model and terminology for agentic systems and produce portable, interoperable workflow patterns across frameworks, tools and execution environments. citeturn16view3

Its scope already includes almost every major unresolved topic identified in this report:

```text
Workflow primitives
Tasks / tool calls / branches / sub-agents
Ordering / parallelism / conditions
Agent handoffs
Long-running execution
State persistence
Partial completion
Resume
Retries
Idempotency
Failure recovery
Audit histories
Human approvals
Interruptions
Interchange formats
Cross-runtime portability
Deterministic and probabilistic workflows
```

The charter is especially significant because it explicitly states that industry frameworks are misaligned on definitions and that there is a challenge translating deterministic workflow models into the nondeterministic models used by autonomous agents. It also states that the WG should layer standards over existing execution environments rather than replace them. citeturn16view3

The WG therefore appears to be the most natural neutral venue in which an AI Workflow reference model could be advanced, although its work is itself still early: its initial goals include establishing taxonomy, prioritizing use cases and producing a workflow reference architecture. citeturn16view3

Primary source: [AAIF Workflows & Process Integration WG](https://github.com/aaif/wg-workflows-and-process-integration). citeturn16view3

**Agentic AI Foundation project portfolio**

As of September 4, 2026, AAIF's hosted-project portfolio includes **MCP, goose, AGENTS.md, agentgateway and A2A**. MCP, goose and AGENTS.md were part of the foundation's original project set; agentgateway joined in June 2026 and A2A in August 2026. citeturn13search14turn18search6turn13search8turn13search3

These should not all be categorized as workflow standards:

| AAIF project | Relevance to an AI workflow standard |
|---|---|
| MCP | Standard tool/context interface |
| A2A | Standard inter-agent task/communication interface |
| AGENTS.md | Convention for supplying repository-specific instructions to coding agents |
| agentgateway | Runtime gateway/control plane for MCP, A2A, LLM and service traffic |
| goose | Agent framework/reference implementation rather than workflow interchange standard |

AAIF therefore increasingly resembles a **standards and implementation ecosystem around agent interoperability**, while the W&PI Working Group addresses the missing workflow semantics that connect these pieces. citeturn13search14turn13search8turn16view3

One correction to an easy misconception is worth making explicit: **AG-UI is important and adjacent, but it is not listed among AAIF's hosted projects on AAIF's current project page as of September 4, 2026.** citeturn13search14

**NIST AI Agent Standards Initiative**

NIST launched its **AI Agent Standards Initiative** on **February 17, 2026** to foster industry-led technical standards and open protocols supporting secure, interoperable AI agents. NIST has associated work on agent security, identity and authorization, with additional research, guidelines and deliverables planned through its public-input process. citeturn19search0turn19search1turn19search8

This is not currently an executable workflow specification and therefore has no workflow YAML or lifecycle schema to implement. Its importance is instead institutional: it means **agent interoperability, identity, authority and security are now formal U.S. standards-policy topics**, and any serious AI Workflow Standard should maintain liaison with this initiative. citeturn19search1turn19search4

Primary source: [NIST AI Agent Standards Initiative](https://www.nist.gov/artificial-intelligence/ai-agent-standards-initiative). citeturn19search1

**W3C agent-protocol incubation**

The **W3C AI Agent Protocol Community Group** is active in 2026 and holds regular meetings. Its stated mission is to develop open interoperable mechanisms for agents to **discover, identify and collaborate across the Web**. This is a Community Group incubation effort rather than a W3C Recommendation. citeturn20search2turn20search1

Several related W3C Community Groups appeared during 2026, including work on **agent memory interoperability, agent trust, agent integrity verification and agent conformance/benchmarking**. Their existence is important evidence that memory, identity, trust and conformance are becoming separate standardization domains, but these efforts are very early and should not be treated as mature W3C standards. citeturn20search7turn20search24turn20search6turn20search36

Primary source: [W3C AI Agent Protocol Community Group](https://www.w3.org/community/agentprotocol/). citeturn20search2

**ISO/IEC and IEEE adjacent standards**

ISO/IEC JTC 1/SC 42 remains the central ISO/IEC committee for AI standardization. It has already produced standards relevant to the governance envelope around an agent workflow, including **ISO/IEC 42001:2023** for AI management systems, **ISO/IEC 5338:2023** for AI system life-cycle processes and **ISO/IEC TS 8200:2024** for controllability of automated AI systems. Work under development in 2026 includes guidance for human oversight and AI conformity assessment. None currently constitutes a portable agentic-workflow execution language. citeturn22search0turn22search13turn22search19

IEEE also has active agent-oriented standards work such as **P3777**, a proposed framework for benchmarking AI agents and their capabilities/performance. Again, this addresses evaluation rather than workflow interchange, but its evaluation model may eventually matter to standardized workflow evidence. citeturn22search3

These standards reinforce an important architectural point: workflow execution standards should **reference organizational governance, human oversight and evaluation standards rather than duplicate them**.

## Comparative mapping

The matrices below are analytical mappings of each specification's primary scope, not official ratings assigned by the standards organizations.

Legend: **● = first-class/direct coverage**, **◐ = meaningful but partial/adjacent coverage**, **○ = largely outside scope**.

| Standard/specification | Conceptual model | Execution semantics | Lifecycle | Humans | Agents | Services | Nondeterminism / autonomy |
|---|---:|---:|---:|---:|---:|---:|---:|
| BPMN citeturn12search0 | ● | ● | ◐ | ● | ◐ | ● | ○ |
| DMN citeturn4view2 | ● | ◐ | ○ | ○ | ○ | ◐ | ○ |
| CMMN citeturn3search2 | ● | ● | ● | ● | ◐ | ● | ◐ |
| WfMC/XPDL citeturn10search3turn10search1 | ● | ◐ | ◐ | ● | ○ | ● | ○ |
| Open Workflow 1.0.3 citeturn9view0 | ● | ● | ● | ◐ | ◐ | ● | ◐ |
| Arazzo 1.1 citeturn14search1 | ● | ● | ◐ | ○ | ○ | ● | ○ |
| Agent Spec citeturn6search1turn6search0 | ● | ● | ◐ | ◐ | ● | ● | ● |
| MCP 2026-07-28 citeturn2search0turn2search6 | ◐ | ◐ | ◐ | ◐ | ● | ● | ◐ |
| A2A 1.0 citeturn16view0turn17view2 | ● | ◐ | ● | ◐ | ● | ◐ | ● |
| AG-UI draft citeturn16view1 | ◐ | ◐ | ● | ● | ● | ○ | ◐ |
| OTel GenAI citeturn21search1turn21search14 | ● | ○ | ◐ | ◐ | ● | ● | ◐ |
| OASF/AGNTCY citeturn15search1 | ● | ○ | ○ | ○ | ● | ◐ | ◐ |
| AAIF W&PI citeturn16view3 | ● | ● | ● | ● | ● | ● | ● |
| NIST Agent Initiative citeturn19search1 | ◐ | ○ | ○ | ◐ | ● | ◐ | ● |

The second matrix shows the integration and governance dimensions.

| Standard/specification | Authority / governance semantics | Tool integration | Portability / interchange | Observability | Intent / outcome | Durable state / recovery |
|---|---:|---:|---:|---:|---:|---:|
| BPMN | ◐ | ◐ | ● | ◐ | ◐ | ◐ |
| DMN | ● for decision rules | ○ | ● | ○ | ◐ | ○ |
| CMMN | ◐ | ◐ | ● | ◐ | ◐ | ● |
| WfMC/XPDL | ○ | ◐ | ● | ◐ | ○ | ◐ |
| Open Workflow | ◐ | ● | ● | ◐ | ◐ | ● |
| Arazzo | ○ | ● | ● | ○ | ● | ◐ |
| Agent Spec | ◐ | ● | ● | ● | ◐ | ◐ |
| MCP | ● at capability boundary | ● | ● | ◐ | ○ | ◐ |
| A2A | ● for auth/task boundary | ◐ | ● | ◐ | ◐ | ● |
| AG-UI | ◐ | ● client-side tools | ● | ◐ | ○ | ◐ |
| OpenTelemetry GenAI | ○ | ● telemetry for calls | ● telemetry | ● | ○ | ◐ |
| OASF | ◐ | ◐ | ● | ○ | ○ | ○ |
| AAIF W&PI | ● target scope | ● | ● | coordinated with OTel/AAIF WG | ● potential | ● |
| NIST initiative | ● major concern | ○ | ● objective | ◐ | ○ | ○ |

The ratings are based on the respective official scopes: Open Workflow's DSL is execution-centric; MCP is capability-centric; A2A is delegated-task-centric; AG-UI is user-interaction-centric; OpenTelemetry is telemetry-centric; and AAIF W&PI explicitly spans the semantic integration problem. citeturn9view0turn2search2turn16view0turn16view1turn21search1turn16view3

A useful architectural picture emerges:

```text
                         BUSINESS SEMANTICS
                  BPMN / DMN / CMMN / ISO governance
                              │
                              ▼
                    AI WORKFLOW SEMANTICS
                ┌────────────────────────────┐
                │ Intent                     │
                │ Tasks / plans / authority  │
                │ lifecycle / state / proof  │
                │ nondeterministic semantics │
                └────────────────────────────┘
                     │        │        │
             ┌───────┘        │        └───────────┐
             ▼                ▼                    ▼
        Open Workflow      Agent Spec           Arazzo
        orchestration      agent/flow        API sequences
             │
       ┌─────┼───────────────┬───────────────────┐
       ▼     ▼               ▼                   ▼
      MCP   A2A            AG-UI            OpenAPI /
     tools  agents       humans/apps      AsyncAPI/events
       │     │               │                   │
       └─────┴───────────────┴───────────────────┘
                              │
                              ▼
                    OpenTelemetry GenAI
```

That picture leads to a critical distinction:

**Protocol interoperability is already advancing faster than workflow-semantic interoperability.**

MCP can let an agent invoke a tool. A2A can let an agent delegate a task. AG-UI can let an application display a request for intervention. OpenTelemetry can record those operations. What none of those protocols independently answers is:

> **Under what declared intent, authority, policy and execution contract were those operations allowed, how are they related as one workflow, and what constitutes verified completion?**

AAIF's workflow charter is effectively an acknowledgement of that gap. citeturn16view3

There is also a fundamental semantic discontinuity between classical and agentic orchestration.

A conventional workflow often looks like:

```text
Definition
   ↓
Known set of allowable transitions
   ↓
Execution chooses among those transitions
```

A sufficiently autonomous workflow can instead look like:

```text
Intent
   ↓
Agent creates Plan v1
   ↓
Discovers capability
   ↓
Creates Task not enumerated in original plan
   ↓
Delegates Task
   ↓
Receives new evidence
   ↓
Revises Plan to v2
   ↓
Requests human authority
   ↓
Continues
```

That is not simply “a BPMN graph with an LLM in one box.” It is a workflow in which **the execution can construct part of its own future topology**. AAIF explicitly identifies deterministic-versus-nondeterministic workflow semantics as a standardization problem, while OpenTelemetry's ongoing workflow/node discussions demonstrate the same issue from the telemetry side. citeturn16view3turn21search5turn21search9

I would call this:

> **governed dynamic orchestration**

rather than merely “nondeterministic workflow,” because pure nondeterminism misses the defining requirement: autonomous variation must remain inside an explicit control envelope.

## What remains unresolved

The following priorities are a synthesis of gaps visible across the standards above. “High” means the gap prevents reliable semantic interoperability across agent runtimes; “Medium” means partial standards already exist but require composition; “Low” means a solution already exists or the gap should deliberately not become the center of a new standard.

| Priority | Gap | Why it matters |
|---|---|---|
| **High** | **Common workflow metamodel and vocabulary** | `workflow`, `run`, `task`, `step`, `agent`, `session`, `conversation`, `handoff`, `delegation`, `plan` and `artifact` are not consistently defined across frameworks. AAIF and OpenTelemetry are actively working on this exact vocabulary problem. citeturn16view3turn21search11 |
| **High** | **Intent and success-contract semantics** | Existing orchestration standards usually encode actions and conditions more strongly than desired outcome, acceptable evidence and completion criteria. Agentic execution needs a stable contract even when the plan changes. |
| **High** | **Authority and autonomy model** | Authentication answers who an actor is; it does not fully answer what an autonomous participant may decide, delegate or cause on behalf of another actor. NIST is separately studying agent identity/authority, while A2A and MCP cover pieces of authorization. citeturn19search8turn16view0turn2search0 |
| **High** | **Dynamic topology and plan mutation** | An agent may create/decompose tasks and revise plans at runtime. Classical graph standards do not normatively describe when such mutations are permitted or how they map to the original definition. AAIF explicitly flags deterministic-to-nondeterministic translation. citeturn16view3 |
| **High** | **Portable durable execution state** | Long-running workflows need checkpoint, suspend, resume and cross-runtime recovery semantics. A2A provides task continuity and Open Workflow provides durable-workflow concepts, but there is no universal portable agent-workflow checkpoint model. citeturn16view0turn9view0turn16view3 |
| **High** | **Side-effect, retry, idempotency and compensation semantics** | Retrying “reason about this document” is different from retrying “pay invoice,” “delete record,” or “open valve.” AAIF explicitly puts retries, idempotency and consistency with external stateful systems in scope. citeturn16view3 |
| **High** | **Evidence, provenance and verifiable completion** | `COMPLETED` cannot safely mean “the agent said it is finished.” Enterprise workflows need artifacts and evidence mapped to objective success criteria and audit history. A2A already distinguishes communication from output artifacts; OTel records execution but is not itself a system-of-record semantics. citeturn16view0turn16view3 |
| **High** | **Conformance semantics** | A portable YAML document has little value if Runtime A and Runtime B interpret delegation, failure, retry or completion differently. Emerging W3C conformance work underscores that protocol definition and conformance testing are separate concerns. citeturn20search36 |
| **Medium** | **Human intervention/approval semantics** | AG-UI and A2A already provide useful interruption mechanisms, but cross-runtime approval ownership, expiry, escalation and resumption require common workflow semantics. citeturn16view1turn17view2 |
| **Medium** | **Portable memory/context boundaries** | MCP, A2A and emerging W3C memory work touch context, but there is no universally accepted rule distinguishing durable workflow state, conversational history, agent memory and evidence. citeturn20search7turn2search10 |
| **Medium** | **Capability discovery and matching** | A2A Agent Cards and OASF address capability description, but standards are still converging on how workflow requirements map to dynamically discovered participant capabilities. citeturn17view3turn15search1 |
| **Medium** | **Budgets, deadlines and resource envelopes** | Autonomous execution needs limits on tokens, money, time, tool calls, delegation depth and other resources. Existing protocols expose pieces but do not provide one portable workflow-level budget contract. |
| **Medium** | **Evaluation semantics** | NIST and IEEE are developing agent-evaluation approaches, while OpenTelemetry records operational data. A workflow standard will need a portable means to attach evaluation requirements and evidence without reinventing evaluation science. citeturn19search21turn22search3 |
| **Low** | **A new graphical notation** | BPMN already provides a mature graphical process language. The initial AI standard should prioritize semantics and mappings rather than creating a competing diagram vocabulary. citeturn12search0 |
| **Low** | **Another tool or agent transport protocol** | MCP and A2A already occupy these interfaces and are now both under AAIF. Building replacements would fragment rather than standardize the ecosystem. citeturn13search14turn13search3 |
| **Low** | **Another telemetry protocol** | OpenTelemetry is CNCF Graduated and is actively defining agent/workflow semantic conventions. The correct strategy is liaison and semantic alignment. citeturn21search20turn21search1 |
| **Low** | **Standardizing prompts or model reasoning internals** | AAIF W&PI explicitly places prompt engineering outside its workflow scope. Portability should concern observable contracts, decisions, artifacts and effects rather than proprietary internal reasoning. citeturn16view3 |

Three of these gaps deserve special emphasis.

**Intent is not the same thing as an initial prompt.**

For interoperability, an intent should be a durable machine-readable contract containing at least:

```yaml
intent:
  outcome: "Resolve the identified safety deficiency"

  successCriteria:
    - "corrective action implemented"
    - "verification evidence accepted"
    - "required approvals recorded"

  constraints:
    - "do not alter production controls without approval"

  completionPolicy:
    evidenceRequired: true
    verifier: independent
```

A prompt can change. A plan can change. An agent can change. **The intent contract is what allows the system to decide whether those changes still pursue the same authorized objective.**

**Authority is not the same as authentication.**

A workflow needs an authority representation closer to:

```yaml
authority:
  principal: "agent:investigator"
  capability: "system:change-record"
  actions: ["read", "draft"]
  prohibitedActions: ["approve", "deploy"]

  delegation:
    permitted: true
    maxDepth: 2
    mayExpandAuthority: false

  sideEffects:
    approvalRequired: true

  expiresAt: "2026-09-05T18:00:00Z"
```

The critical invariant should be:

> **Delegated authority may be equal to or narrower than parent authority; it must never silently become broader.**

A2A and NIST's identity/authority work provide pieces of the foundation, but this workflow-level attenuation rule needs to be explicit in the reference model. citeturn16view0turn19search8

**Completion is not verification.**

A useful standard should distinguish:

```text
Execution completed
        ≠
Desired outcome verified
```

A2A already gives us an important separation between **Messages** and output **Artifacts**. An AI workflow model should go one step further and distinguish the Artifact from the **Evidence** establishing that a success criterion was met. citeturn16view0

For example:

```text
Task:
  "Patch vulnerability"

Artifact:
  pull-request-318

Evidence:
  regression-test result
  security scan result
  independent review approval

Success Criterion:
  vulnerability cannot be reproduced

Verification:
  PASSED
```

That distinction is especially important for safety, compliance, cyber, financial and operational workflows.

## Proposed reference model

The proposed model below is intentionally small. It does not attempt to replace every concept in BPMN, Open Workflow, MCP, A2A or OpenTelemetry. It defines the **semantic glue** those specifications currently lack when combined into one governed agentic workflow.

| Object | Normative meaning |
|---|---|
| **WorkflowDefinition** | Versioned declaration of the workflow's invariant contract, permitted structure and execution requirements. |
| **Intent** | Outcome, success criteria, constraints and completion/verification requirements. |
| **Run** | One execution instance of a WorkflowDefinition. |
| **Plan** | Mutable, versioned strategy for satisfying the Intent during a Run. Unlike the WorkflowDefinition, a Plan may evolve. |
| **Task** | Goal-oriented unit of responsibility assignable or delegable to a Participant. |
| **StepAttempt** | One concrete, observable attempt to perform an operation within a Task. Retries create additional attempts rather than rewriting history. |
| **Participant** | Actor performing work: Human, Agent, Service or deterministic Workflow. |
| **Capability** | Operation a Participant can provide, potentially bound through MCP, A2A, OpenAPI or another protocol. |
| **AuthorityGrant** | Scoped permission under which a Participant may use a Capability or delegate work. |
| **State / Checkpoint** | Durable information necessary to resume execution. |
| **Artifact** | Persistent output produced or consumed by Tasks/Steps. |
| **Evidence** | Verifiable information demonstrating satisfaction or violation of an Intent criterion or policy. |
| **Event** | Immutable record of lifecycle transition, decision, delegation, invocation or other significant execution occurrence. |

The distinction between **WorkflowDefinition, Plan, Task and StepAttempt** is particularly important.

```text
WorkflowDefinition = what is authorized and invariant

Plan               = current proposed strategy

Task               = what someone/something is responsible for accomplishing

StepAttempt        = what actually happened
```

That lets the standard accommodate both ends of the spectrum:

```text
Fully deterministic

Definition == almost the complete Plan
```

and:

```text
Highly agentic

Definition = intent + boundaries + mandatory gates
Plan       = substantially discovered during execution
```

The reference model can then map cleanly onto existing standards:

- BPMN/CMMN represent business-facing workflow/case structure. citeturn12search0turn3search2
- DMN represents deterministic decisions and policy logic. citeturn4view2
- Open Workflow represents executable deterministic control constructs. citeturn9view0
- Agent Spec can serialize/configure agents and flows. citeturn6search1
- MCP binds Capabilities to tools/resources. citeturn2search6turn2search10
- A2A binds delegated Tasks to remote Agents. citeturn16view0
- AG-UI binds human/application interaction. citeturn16view1
- Arazzo binds API-operation sequences. citeturn14search1
- OASF can describe participant skills/capabilities. citeturn15search1
- OpenTelemetry provides observational spans and events. citeturn21search1

The proposed architecture is:

```mermaid
flowchart TB
    WD["WorkflowDefinition"]
    I["Intent<br/>Outcome • Success Criteria • Constraints"]
    POL["Policy & Authority<br/>AuthorityGrant • Budgets • Approvals"]
    RUN["Run"]
    PLAN["Plan<br/>mutable + versioned"]
    TASK["Task"]
    ATT["StepAttempt"]
    PART["Participant"]
    CAP["Capability"]
    STATE["State / Checkpoint"]
    ART["Artifact"]
    EVID["Evidence"]
    EVENT["Event / Audit Record"]

    WD --> I
    WD --> POL
    RUN -->|"instantiates"| WD
    RUN --> PLAN
    PLAN -->|"decomposes into"| TASK
    TASK -->|"assigned / delegated to"| PART
    PART -->|"offers"| CAP
    TASK --> ATT
    ATT -->|"invokes"| CAP
    RUN --> STATE
    ATT --> ART
    ART --> EVID
    EVID -->|"verifies"| I
    ATT --> EVENT
    TASK --> EVENT
    PLAN --> EVENT
    POL -->|"constrains"| PLAN
    POL -->|"constrains"| TASK
    POL -->|"constrains"| ATT

    MCP["MCP"]
    A2A["A2A"]
    AGUI["AG-UI"]
    API["OpenAPI / Arazzo / AsyncAPI"]
    OTEL["OpenTelemetry"]

    CAP -. "tool binding" .-> MCP
    TASK -. "agent delegation" .-> A2A
    PART -. "human interaction" .-> AGUI
    CAP -. "service binding" .-> API
    EVENT -. "telemetry projection" .-> OTEL
```

The standard's lifecycle should derive as much as possible from existing task protocols while adding workflow-level durable states.

A concise Run/Task state model would be:

```text
CREATED
   ↓
READY
   ↓
RUNNING
   ├────────→ WAITING_INPUT ─────┐
   ├────────→ WAITING_AUTH ──────┤
   ├────────→ WAITING_EXTERNAL ──┤
   ├────────→ PAUSED ────────────┤
   │                             │
   └←────────────────────────────┘
   │
   ├────→ COMPLETED ───→ VERIFYING ───→ VERIFIED
   │                                  └→ INVALIDATED
   ├────→ FAILED
   ├────→ CANCELED
   └────→ REJECTED
```

This intentionally aligns its core task states with A2A's `SUBMITTED`, `WORKING`, `INPUT_REQUIRED`, `AUTH_REQUIRED`, `COMPLETED`, `FAILED`, `CANCELED` and `REJECTED`, rather than inventing incompatible terms. The additional distinction between execution completion and verification belongs at the workflow/outcome layer rather than inside A2A. citeturn17view1turn17view2

The execution semantics should include the following normative rules.

**Deterministic versus agentic operations must be explicit.** A branch based on a DMN/FEEL expression and a branch chosen by an LLM are semantically different. A conforming execution history should indicate whether a decision was deterministic, human, model-mediated or delegated. DMN provides a natural deterministic-decision reuse point. citeturn4view2

**Plans are mutable; definitions are versioned and stable during a Run.** An agent may modify the Plan only within the WorkflowDefinition's policy envelope. Material changes outside that envelope require a new definition version or explicit authorization.

**Task creation is governed.** A Participant may dynamically create a Task only when the applicable AuthorityGrant permits planning/decomposition.

**Delegation attenuates authority.** A child Task receives an explicit subset of the parent's authority. The protocol binding may be A2A, but the authority semantics belong to the workflow standard.

**Retries create attempts.** A failed StepAttempt remains part of history; retry creates `attempt=2`, preserving auditability.

```text
Task T-17
  StepAttempt 1 → failed
  StepAttempt 2 → failed
  StepAttempt 3 → succeeded
```

**Side effects require idempotency or declared compensation semantics.** AAIF already identifies retries, idempotency, failure recovery and consistency with external stateful systems as workflow concerns. citeturn16view3

**Waits are durable.** `WAITING_INPUT`, `WAITING_AUTH`, timers and external waits must survive process/runtime restarts without relying on an in-memory LLM session.

**Checkpoint and conversation are different.** A checkpoint contains the minimum durable state required to continue the workflow. A conversation is contextual interaction state. A memory may span workflows. Conflating the three makes portability extremely difficult; current OpenTelemetry and W3C work confirms that session/memory terminology remains unsettled. citeturn21search11turn20search7

**Decisions with material effects are recordable.** The standard need not expose private chain-of-thought. It should record observable decision metadata: input references, decision mechanism/class, applicable policy, chosen action, authorization and resulting effect.

**Success requires a criterion, not merely an exit condition.** A workflow can terminate successfully at the execution layer yet fail verification at the outcome layer.

The governance model should similarly be machine-readable.

```yaml
policy:
  autonomy:
    planning: permitted
    dynamicTaskCreation: permitted
    delegation:
      permitted: true
      maximumDepth: 3

  authority:
    externalWrite:
      default: deny
    productionChange:
      approvalRequired: true

  resourceEnvelope:
    deadline: "PT8H"
    maximumDelegations: 12
    maximumToolCalls: 250

  verification:
    independentVerifierRequired: true
```

This is where the proposed standard can add something that neither classical workflow languages nor current agent protocols fully provide:

> **A workflow should define not only the sequence of work, but the bounded space of work the system is authorized to discover.**

That is, in my assessment, the defining semantic innovation required for an agentic workflow standard.

## Roadmap to a vendor-neutral standard

The recommended strategy is to create a **semantic interoperability standard first and a serialization second**. Beginning with a large YAML schema would prematurely encode unresolved terminology.

A nine-step program is appropriate.

| Step | Deliverable | Key decision |
|---|---|---|
| **Charter and representative use cases** | 10–15 canonical workflows spanning deterministic, agentic, human-in-loop and long-running execution | What must interoperable implementations actually be able to exchange? |
| **Normative vocabulary** | Precise definitions of WorkflowDefinition, Run, Intent, Plan, Task, StepAttempt, Participant, Capability, AuthorityGrant, State, Artifact, Evidence and Event | Eliminate framework-specific ambiguity before syntax is designed. |
| **Core metamodel** | UML/JSON-schema-neutral conceptual model plus relationships and cardinalities | What is invariant across runtimes? |
| **Lifecycle and execution semantics** | State machines, waits, retries, delegation, failure, checkpointing, cancellation, compensation and verification | What behavior must two conforming runtimes share? |
| **Authority and autonomy profile** | Machine-readable grants, delegation attenuation, approval gates, budgets and side-effect classifications | How is autonomous freedom bounded? |
| **Existing-standard mappings** | Normative mappings to MCP, A2A, AG-UI, Open Workflow, Arazzo, OpenTelemetry, BPMN/DMN/CMMN and OASF | Prevent unnecessary duplication. |
| **Interchange representation** | Minimal JSON/YAML representation with extension mechanism | Encode the model only after semantics stabilize. |
| **Conformance and interoperability suite** | Golden workflow cases, event traces, expected transitions and cross-runtime tests | Demonstrate semantic rather than merely syntactic portability. |
| **Neutral governance and version-one process** | AAIF proposal/liaison structure, public change control, compatibility rules and reference implementations | Ensure no vendor owns the meaning of a workflow. |

An approximately 18-month standards-development path beginning after the current September 2026 research point could look like this:

```mermaid
gantt
    title Proposed AI Workflow Standard Roadmap
    dateFormat  YYYY-MM-DD
    axisFormat  %b %Y

    section Foundation
    Charter and canonical use cases       :a1, 2026-10-01, 60d
    Vocabulary and terminology            :a2, 2026-11-01, 90d
    Core conceptual metamodel             :a3, 2026-12-15, 120d

    section Semantics
    Lifecycle and execution semantics     :b1, 2027-02-01, 150d
    Authority and autonomy profile        :b2, 2027-02-15, 150d
    Existing-standard mappings            :b3, 2027-03-15, 150d

    section Interchange
    JSON/YAML interchange draft           :c1, 2027-06-01, 120d
    Reference adapters and bindings       :c2, 2027-07-01, 150d

    section Conformance
    Interoperability/conformance suite    :d1, 2027-09-01, 150d
    Cross-runtime pilot implementations   :d2, 2027-10-01, 150d

    section Standardization
    Candidate specification v0.9          :milestone, e1, 2028-01-15, 0d
    AAIF / neutral governance review      :e2, 2027-12-01, 100d
    Candidate v1.0                        :milestone, e3, 2028-03-15, 0d
```

The first use cases should be deliberately different from one another. A standard tested only against “research agent summarizes documents” will miss most of the hard semantics. Suitable canonical cases would include an IT incident with human escalation, a software vulnerability remediation workflow, an EH&S corrective-action workflow, an invoice/payment process with irreversible effects, a multi-agent research process, and a long-running regulatory approval case. That diversity forces the model to confront durable state, side effects, independent verification, human authority and adaptive planning rather than optimizing exclusively for conversational agents.

For conformance, each reference scenario should contain:

```text
Input definition
Expected invariant policies
Permitted plan variations
Required lifecycle transitions
Authority constraints
Required artifacts
Required evidence
Expected telemetry projection
Failure/retry scenarios
Interoperability checkpoints
```

Two runtimes would **not** need to produce identical model outputs. That would defeat the purpose of probabilistic agents. They would need to preserve the same **semantic invariants**.

For example:

```text
Runtime A chooses:
  inspect logs → inspect configuration → ask specialist

Runtime B chooses:
  inspect configuration → inspect logs → ask specialist

Both conform if:
  authority boundaries are preserved
  mandatory approval occurs
  required evidence is collected
  no prohibited side effect occurs
  success criteria are verified
  execution history is interoperable
```

This distinction between **behavioral equivalence** and **trajectory identity** is essential. Agentic workflow conformance cannot mean “produce exactly the same sequence of tokens or tool calls.” It should mean:

> **Different valid execution trajectories preserve the same declared intent, policy invariants, authority boundaries, lifecycle semantics and verification obligations.**

The standard's version-one profile should consequently be intentionally conservative. It should standardize:

```text
WHAT the workflow is trying to accomplish
WHO may participate
WHAT participants are authorized to do
HOW work may be delegated
HOW execution state changes
HOW external capabilities are invoked
HOW side effects are governed
WHAT artifacts are produced
WHAT evidence proves completion
HOW execution is observed
HOW the definition is exchanged
```

It should **not** standardize:

```text
Which model must be used
How an agent internally reasons
How prompts must be engineered
Which vector database stores memory
Which framework executes the workflow
Which orchestration vendor hosts it
Which user interface renders it
```

That division aligns well with AAIF's explicit decision to layer workflow standards over independently evolving execution frameworks and to leave prompt engineering, model training, telemetry details, security-specialist questions and evaluation-quality questions to related groups where appropriate. citeturn16view3

The governance recommendation is likewise straightforward: **do not create a new foundation if AAIF's Workflows & Process Integration group can serve as the neutral home or primary liaison.** Its approved charter already covers the target problem and its parent foundation hosts both MCP and A2A. citeturn16view3turn13search14

A sensible standards architecture would be:

```text
AI Workflow Standard
│
├── Core
│   ├── terminology
│   ├── metamodel
│   ├── lifecycle
│   ├── execution semantics
│   └── conformance
│
├── Governance Profile
│   ├── intent
│   ├── authority
│   ├── autonomy
│   ├── budgets
│   └── evidence
│
├── Bindings
│   ├── MCP Binding
│   ├── A2A Binding
│   ├── AG-UI Binding
│   ├── OpenAPI/Arazzo Binding
│   └── Async Event Binding
│
├── Execution Profiles
│   ├── Open Workflow Profile
│   ├── Agent Spec Profile
│   └── BPMN/CMMN Mapping
│
└── Observability Profile
    └── OpenTelemetry GenAI Mapping
```

This is preferable to a monolithic standard because protocol innovation can continue independently. MCP 2027 could evolve without requiring “AI Workflow 2.0”; only the MCP binding would need updating.

## Conclusions and primary-source map

The research supports five firm conclusions.

**First, workflow standardization is not starting from zero.** BPMN, DMN, CMMN, WfMC, Open Workflow and Arazzo supply decades of process, decision, case, execution and interchange prior art. citeturn12search0turn4view2turn3search2turn10search3turn9view0turn14search1

**Second, agent interoperability is rapidly standardizing at the boundaries.** MCP now provides a strong capability/tool boundary, A2A provides a stable 1.0 agent/task boundary, AG-UI is developing the human/application boundary, OASF addresses capability description, and OpenTelemetry is converging on observability semantics. citeturn2search0turn16view0turn16view1turn15search1turn21search1

**Third, the missing problem is semantic composition.** None of those boundary protocols alone specifies a portable end-to-end contract for intent, plans, dynamic task creation, authority delegation, workflow-level state, side-effect semantics, evidence and verified completion. AAIF W&PI's charter targets almost exactly this missing layer. citeturn16view3

**Fourth, a viable AI Workflow Standard should not equate portability with identical execution.** Probabilistic participants make identical trajectories neither realistic nor desirable. The standard should instead require **semantic invariants**: equivalent authority, lifecycle, policy, evidence and outcome semantics.

**Fifth, the best strategic position is therefore not “replace all workflow standards.” It is “standardize the semantic contract that composes them.”**

A concise proposed definition is:

> **An AI workflow is a governed, stateful execution in which human, agentic and deterministic participants cooperate toward a declared intent, using delegated capabilities within explicit authority and policy boundaries, while producing durable artifacts, execution records and evidence sufficient to verify the outcome.**

And an even shorter formulation of the standard itself would be:

> **The AI Workflow Standard defines what remains invariant while an autonomous execution is allowed to vary.**

That is the design center I would recommend.

For convenient follow-through, the principal official sources used in this analysis are:

| Area | Primary/official source |
|---|---|
| BPMN | [OMG BPMN 2.0.2](https://www.omg.org/spec/BPMN/2.0.2/About-BPMN) and [ISO/IEC 19510](https://www.iso.org/standard/62652.html) citeturn4view0turn12search0 |
| DMN | [OMG DMN 1.5](https://www.omg.org/spec/DMN/1.5/About-DMN) citeturn4view2 |
| CMMN | [OMG CMMN 1.1](https://www.omg.org/spec/CMMN/1.1/About-CMMN) citeturn3search2 |
| Open Workflow | [Open Workflow / Serverless Workflow](https://serverlessworkflow.io/) citeturn9view0turn0search15 |
| Agent Spec | [Oracle Agent Spec](https://github.com/oracle/agent-spec) citeturn6search1 |
| MCP | [MCP 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28) citeturn2search0 |
| A2A | [A2A Protocol Specification](https://a2a-protocol.org/latest/specification/) citeturn16view0 |
| AG-UI | [AG-UI draft specification](https://docs.ag-ui.com/spec/draft) citeturn16view1 |
| OpenTelemetry GenAI | [Agent/workflow semantic conventions](https://github.com/open-telemetry/semantic-conventions-genai/blob/main/docs/gen-ai/gen-ai-agent-spans.md) citeturn21search1 |
| AAIF workflows | [Workflows & Process Integration WG](https://github.com/aaif/wg-workflows-and-process-integration) citeturn16view3 |
| AAIF project ecosystem | AAIF official project materials and Linux Foundation announcements citeturn13search14turn13search3turn13search8 |
| Arazzo | [OpenAPI Arazzo](https://spec.openapis.org/arazzo/latest.html) citeturn14search1 |
| OASF | [AGNTCY OASF](https://docs.agntcy.org/oasf/) citeturn15search1 |
| NIST | [AI Agent Standards Initiative](https://www.nist.gov/artificial-intelligence/ai-agent-standards-initiative) citeturn19search1 |
| W3C | [AI Agent Protocol Community Group](https://www.w3.org/community/agentprotocol/) citeturn20search2 |
| ISO AI governance | [ISO/IEC JTC 1/SC 42](https://www.iso.org/committee/6794475.html) and [ISO/IEC 42001](https://www.iso.org/standard/42001) citeturn22search0turn22search19 |
| IEEE agent evaluation | [IEEE P3777](https://standards.ieee.org/ieee/3777/12350) citeturn22search3 |

As of **September 4, 2026**, the standards window is unusually favorable: MCP and A2A have reached meaningful protocol maturity, Open Workflow has reached a 1.x DSL, Agent Spec is testing cross-framework portability, OpenTelemetry is defining the observable vocabulary, NIST has formally launched an agent standards initiative, and—most consequentially—the AAIF has created a working group whose charter expressly calls for a shared, vendor-neutral agent workflow model. citeturn2search0turn13search3turn9view0turn6search1turn21search1turn19search1turn16view3

The opportunity is therefore not hypothetical. **The pieces of an AI Workflow Standard already exist; what is missing is the specification that makes those pieces mean the same thing when assembled into one autonomous, governed workflow.**