# AI Workflow Standards — Supplemental Research

**Research date:** 7 September 2026  
**Purpose:** Supplement the original landscape report and test the foundations of AIWS-001, edition 0.1, before further standard drafting.  
**Audience:** A prospective standards author and implementers of governed AI workflows.  
**Status:** Research and design assessment; not a revised standard or a conformance assessment.

## Executive finding

**The case for a shared AI workflow contract remains persuasive, but more of its foundations already exist than our earlier work recognized.** The next step should be selective reuse and explicit integration rules. We should make fewer claims of conceptual novelty and be more precise about what existing technologies actually guarantee.

This research expanded beyond MCP, A2A, and orchestration formats into authorization standards, human-task specifications, distributed execution, provenance, assurance cases, and experimental agent evaluation. It identified three kinds of result:

1. **Established foundations to reuse:** OAuth token exchange and rich authorization, W3C provenance and state-machine semantics, OMG assurance-case modeling, and classical human-task management.
2. **Implementation behavior that exposes gaps:** policy engines can handle errors differently; resumption can repeat code; workflow execution guarantees differ; idempotency guarantees have conditions and expiry windows.
3. **Still-unsettled proposals:** mission-level agent authorization, cryptographic delegation receipts, generalized cross-runtime migration, and the precise rules connecting intent to acceptable autonomous behavior.

The most consequential changes recommended for AIWS-001 are:

- Separate an **authorized mission or contract instance** from individual runtime execution segments.
- Define how permission scopes are compared; a mathematical “intersection” in prose is insufficient.
- Explicitly handle errors in mandatory policy checks.
- Distinguish audit playback, crash recovery, and new execution.
- Add evidence-support reasoning and a separate acceptance decision where the use case needs them.
- Specify multiple simultaneous waits and the aggregation of parallel-task status.
- Separate control conformance, task quality, adversarial robustness, and interoperability testing.

These are research-driven recommendations. They have not yet been incorporated into edition 0.1 or validated in running implementations. The evidence and limitations behind each are described below.

## 1. Scope and evidence method

The research question was: **Which requirements in our proposed standard are supported by existing standards or implementation evidence, which need correction, and which remain design choices?**

The baseline was the supplied *Workflow Standards for AI and Agentic Systems* and the subsequently created *AI-Workflow-Standard-Review-and-Draft.md*. This supplement does not repeat their complete protocol inventory or re-audit every historical release date.

The additional research used primary standards publications, official product documentation, original research papers, and first-party research repositories. Discovery covered authority, durable execution, side effects, human decisions, provenance and assurance, lifecycle semantics, and testing. Follow-up focused on disconfirming evidence: optional guarantees, failure behavior, scope exclusions, and the actual status of emerging proposals.

The source register distinguishes published specifications, product behavior, guidance, and experimental proposals. Dates refer to a named edition where established; living documentation is identified as such. Access date for all sources is 7 September 2026. Current documentation and individual drafts can change, so future implementation bindings must pin revisions.

This is a focused technical research supplement, not a systematic literature review, legal-compliance analysis, security audit, or empirical runtime comparison. It maps findings to selected requirement families rather than claiming a complete clause-by-clause external validation of all 70 requirements.

## 2. Delegated authority has standards foundations, with important limits

### 2.1 Delegation identity does not itself enforce delegation restrictions

**Source finding.** RFC 8693 defines OAuth token exchange and distinguishes delegation from impersonation. Its actor claim can represent the current actor and historical actors. However, nested prior actors are informational for access-control purposes; they are not a permission chain that a consumer must recursively evaluate. The RFC also describes exchange as a one-time event without a tight ongoing linkage between input and output tokens. [OAuth Token Exchange, RFC 8693](https://www.rfc-editor.org/rfc/rfc8693.html), §§2.1 and 4.1.

**Implication for our draft.** AU-03 and AU-05 remain useful obligations, but an A2A or OAuth binding cannot satisfy them simply by carrying actor identities. It needs a trusted mechanism that issues appropriately constrained child authority and propagates revocation according to a declared policy.

The distinction is practical: knowing that Agent B is acting for Agent A does not establish that B inherited A's spending limit, document restrictions, or grant expiry. Those properties need independently enforceable semantics.

### 2.2 Permission containment needs a defined language

**Source finding.** RFC 9396 supplies structured `authorization_details` for fine-grained permissions. It explicitly does not provide a universal comparison procedure for arbitrary authorization-detail objects; their semantics depend on the authorization type. Simple object comparison can therefore produce the wrong interpretation of increased or reduced access. [Rich Authorization Requests, RFC 9396](https://www.rfc-editor.org/rfc/rfc9396.html), §7.

**Recommended change.** Require each authority profile to define the meaning of its fields, resource matching, quantity units, time boundaries, and a procedure for testing whether a proposed child grant is contained within a parent grant. Unsupported comparisons should return **indeterminate**, not permission.

For example, “may edit files under `/policies/`” cannot be safely compared with “may edit documents tagged `safety`” unless the system can resolve the actual resource sets under a trusted rule. A natural-language judgment by the planning model should not be the authority-containment mechanism.

### 2.3 Revocation has a freshness problem

**Source finding.** RFC 7009 defines token revocation and recognizes related-token invalidation where applicable. RFC 7662 permits introspection-response caching and requires attention to the security/performance trade-off: cached active status can become stale. [Token Revocation, RFC 7009](https://www.rfc-editor.org/rfc/rfc7009.html); [Token Introspection, RFC 7662](https://www.rfc-editor.org/rfc/rfc7662.html), §4.

**Recommended change.** Extend AU-05 with a measurable freshness contract: maximum accepted age of authorization information, propagation delay, behavior when the authority service is unavailable, and the boundary after which an already-committed action cannot be prevented.

“Check permission on every call” is insufficient if every check uses the same stale cache. Conversely, “instant revocation everywhere” is not a credible guarantee without an identified enforcement architecture and failure assumptions.

### 2.4 Deny-by-default and deny-on-error are different

**Source finding.** Cedar combines default denial and explicit prohibition precedence with **skip-on-error**: a policy that errors is excluded from the authorization result. Its documentation explains that applications may inspect diagnostics and choose a stricter response. An Allow decision can therefore coexist with a policy error. [Cedar authorization algorithm](https://docs.cedarpolicy.com/auth/authorization.html).

**Recommended change.** AU-02 and AU-04 should require explicit identification of mandatory policy checks. If one is unavailable, malformed, or indeterminate, the affected consequential dispatch should be blocked regardless of an Allow returned by another check. This is a proposed integration requirement, not a claim that Cedar's behavior is defective.

**New test:** A permit rule succeeds while a mandatory prohibition rule errors because a required attribute is absent. Verify that the workflow's enforcement boundary blocks dispatch and records the missing attribute.

## 3. There is direct emerging overlap with our intent-and-authority model

### 3.1 Mission-Bound Authorization

**Source finding.** Karl McGuinness's *Mission-Bound Authorization for OAuth 2.0*, revision 00 dated 6 July 2026, proposes a durable approved mission artifact connecting intent, derived authority, and consent. Its scope distinguishes mission-related token issuance from a separate runtime-enforcement layer. The inspected Datatracker record classifies it as an **individual Internet-Draft**, not an adopted IETF standard. [Mission-Bound Authorization](https://datatracker.ietf.org/doc/draft-mcguinness-oauth-mission/).

**Assessment.** This is directly relevant prior work for our core proposal. We should compare terminology and authorization boundaries before inventing another mission representation. The existence of a draft establishes overlap, not correctness, adoption, or deployment maturity.

**Recommendation.** Keep it in an experimental alignment track. Do not make this individual draft a mandatory core dependency. Compare its approval artifact with our Intent, AuthorityGrant, and Approval records, then document which workflow obligations remain outside its scope.

### 3.2 Delegation receipts

**Source finding.** Ryan Nelson's *Delegation Receipt Protocol for AI Agent Authorization*, revision 10 dated 13 June 2026, proposes user-signed authorization objects and an append-only record before runtime control. The inspected Datatracker record is also an individual Internet-Draft with no IETF endorsement. [Delegation Receipt Protocol](https://datatracker.ietf.org/doc/draft-nelson-agent-delegation-receipts/).

**Assessment.** It provides a concrete comparison point for approval evidence and provenance. It does not establish that every workflow needs end-user cryptographic signing, or that a signed instruction proves correct interpretation or execution.

**Recommendation.** Treat signatures as one possible integrity and attribution mechanism in a profile. Do not require users to manage signing keys for ordinary workflows merely because one proposed design uses that mechanism.

## 4. Durable execution is well developed; its semantics are not interchangeable

### 4.1 Deterministic control can coordinate nondeterministic agents

**Source finding.** Temporal requires workflow code to reproduce the appropriate command sequence during replay. Its activity documentation places responsibility on application designers to make business activities idempotent because activity attempts can repeat. [Temporal Workflow Definition](https://docs.temporal.io/workflow-definition); [Temporal Activity Definition](https://docs.temporal.io/activity-definition).

**Inference.** A model's output can vary on first execution while subsequent recovery uses its recorded result. “The agent is nondeterministic” is therefore not, by itself, a reason to abandon deterministic orchestration. The important boundary is where model calls, human input, and external effects enter the durable history.

**Recommended change.** Replace a simple deterministic-versus-agentic division with two dimensions: the mechanism that selects an action, and the mechanism that commits and recovers that choice. Dynamic planning may run inside a durable, deterministic control framework.

### 4.2 Our use of “run” is too tightly coupled to an engine invocation

**Source finding.** Temporal documents a workflow execution chain whose executions share a workflow ID while individual runs have different run IDs. Continue-As-New and retries can connect those runs, and chain-level and run-level timeouts differ. [Temporal Workflow Execution](https://docs.temporal.io/workflow-execution), “Workflow Execution Chain.”

**Recommended change.** Introduce an authorized contract instance above engine-specific execution segments. Preserve its budget, accepted intent revision, authority lineage, and unresolved effects when a runtime continues or replaces a segment. Define explicit relationships between AIWS identities and engine identities.

This also addresses a weakness in I-02: creating a linked new run must not provide a way to reset an exhausted budget or escape outstanding effects. Whether an intent change creates a new mission, a new contract revision, or both remains a standards-design choice. The audit link and authority for the change are mandatory design needs; the exact identity scheme is not yet settled.

### 4.3 Resume does not always mean continuing after the interrupted line

**Source finding.** LangGraph's interrupt documentation says resumption reruns the node containing an interrupt. Code before the interrupt can execute again; the documentation accordingly calls for idempotent side effects. Persistent checkpointing and a stable thread identifier are also required to recover the intended state. [LangGraph Interrupts](https://docs.langchain.com/oss/python/langgraph/interrupts).

**Recommended change.** Each execution binding must document its restart boundary. A checkpoint, node, task, and external operation are not necessarily the same unit. Adapter tests should put an observable write before an approval interrupt, then resume twice and inspect whether the write repeats.

### 4.4 “Replay” needs three definitions

**Design synthesis.** D-04 currently distinguishes audit replay and re-execution, but runtime recovery needs its own explicit category:

| Mode | Purpose | New effects allowed? |
|---|---|---|
| Audit playback | Reconstruct and inspect recorded events | No |
| Recovery replay and continuation | Rebuild state, then continue unfinished authorized work | Historical effects must not be repeated unsafely; new pending work may proceed after checks |
| New execution or experimental rerun | Try the workflow again with current inputs, code, or model | Yes, under an explicitly identified new execution and current authorization |

This distinction reconciles the runtime behavior above without requiring every product to rename its APIs.

## 5. Side-effect safety depends on both the orchestrator and the target

### 5.1 Execution guarantees are scoped

**Source finding.** AWS documents exactly-once workflow execution for Standard Step Functions, subject to explicit Retry behavior. It distinguishes at-least-once Asynchronous Express execution from at-most-once Synchronous Express execution. [Choosing workflow type in Step Functions](https://docs.aws.amazon.com/step-functions/latest/dg/choosing-workflow-type.html).

**Implication.** “All workflow engines are at-least-once” would be an inaccurate generalization. Equally, an engine's execution claim must not be silently extended to every downstream business transaction. AIWS should require the claim's exact boundary, configured retry behavior, target guarantees, and failure assumptions.

### 5.2 Idempotency is a contract, not merely a key field

**Source finding.** Stripe documents returning the stored result for repeated requests using the same key, including stored server-error responses. It permits key removal after at least 24 hours, checks parameter consistency, and treats validation failures and certain concurrent conflicts differently because endpoint execution has not begun. [Stripe Idempotent Requests](https://docs.stripe.com/api/idempotent_requests).

**Recommended change.** Keep E-04–E-06 and add a capability guarantee record: key namespace, retention window, fingerprint comparison, eligible operations, result caching behavior, concurrent-request behavior, and reconciliation method. The 24-hour detail is a Stripe example, not an AIWS default.

### 5.3 Recording intent does not make a remote write atomic

**Source finding.** AWS's transactional-outbox guidance addresses the dual-write problem by coordinating business data and an outgoing-event record. It also identifies duplicate delivery as a concern that requires idempotent consumers. [Transactional Outbox Pattern](https://docs.aws.amazon.com/prescriptive-guidance/latest/cloud-design-patterns/transactional-outbox.html).

**Inference.** E-03's durable pre-dispatch record is necessary for accountability but does not by itself close the gap between an external commit and a lost acknowledgment. Outbox-style coordination helps where local state and dispatch records share a transaction; external targets still need appropriate deduplication, conditional operations, or reconciliation.

**New test:** Commit the external effect, drop its acknowledgment, crash the executor, and restore it. Verify that the system distinguishes “outcome unknown” from “effect did not occur.”

### 5.4 Compensation is a business operation

**Source finding.** Microsoft's compensating-transaction guidance explains that compensation can fail, can require application-specific ordering, and need not restore the original state. Manual intervention can be necessary. [Compensating Transaction Pattern](https://learn.microsoft.com/en-us/azure/architecture/patterns/compensating-transaction).

**Recommendation.** Retain E-07. Add explicit tracking of compensation preconditions and dependencies. For example, an equipment change may not be safely reversible after people or other systems have acted on it. The workflow standard should require a declared disposition and responsible authority, while leaving domain-specific recovery procedures to the relevant operating policy.

## 6. Human approval has richer prior art than “pause and resume”

### 6.1 Human-task ownership and business outcome are separate concepts

**Source finding.** The inspected OASIS WS-HumanTask 1.1 Committee Specification 01 includes potential and actual owners, claiming, delegation, deadlines, and escalation. Its approval example permits a task to complete with an approval or a rejection business result. Its delegation default, when omitted, is permissive. [WS-HumanTask 1.1, CS01](https://docs.oasis-open.org/bpel4people/ws-humantask-1.1-spec-cs-01.html), §§3.5, 4.1, 4.9–4.10.

**Implication.** We can reuse ownership and escalation concepts without importing SOAP infrastructure or permissive defaults. A successfully performed review that rejects a proposal is not a failed execution of the review. This supports keeping task execution, authorization decision, and business acceptance separate.

**Recommended change.** Extend the human-decision model with claimed ownership, reassignment, withdrawal, and explicitly defined multi-approver rules where needed. Preserve deny-by-default delegation in AIWS rather than inheriting a legacy specification's default.

### 6.2 Approval must remain tied to what is actually executed

**Source finding.** OWASP's transaction-authorization guidance calls for presenting significant transaction data, preventing modification and skipped authorization steps, and checking authorization at the final execution gate. It identifies time-of-check/time-of-use risks. [OWASP Transaction Authorization](https://cheatsheetseries.owasp.org/cheatsheets/Transaction_Authorization_Cheat_Sheet.html), §§1.1 and 2.5–2.8.

**Recommended change.** Strengthen H-01/H-02 with an explicit approved-versus-dispatched comparison. Define which fields are material in each action type. An authorized retry of the same logical operation should not automatically consume another business approval, but a materially changed operation should require a new decision. Exact approval-reuse semantics should be specified, not inferred from a UI response.

## 7. Evidence should use provenance and assurance concepts already available

### 7.1 Provenance provides the history of an assertion

**Source finding.** W3C PROV-O defines entities, activities, and agents, with relations covering generation, use, derivation, and responsibility. Its “agent” is a responsibility-bearing actor, not exclusively an AI agent. [PROV-O: The PROV Ontology](https://www.w3.org/TR/prov-o/).

**Proposed mapping.** An AIWS Artifact can be represented as an Entity; a StepAttempt as an Activity; and a Participant as a PROV Agent. Evidence lineage can use generation and derivation relationships. This is a proposed alignment, not an already-standardized AIWS binding. Preserve the broader PROV meaning of agent.

### 7.2 Assurance asks why evidence supports a claim

**Source finding.** OMG SACM 2.3 models structured claims, argumentation, artifact references, and asserted evidence relationships. It also represents counter-evidence and assertions needing further support. [Structured Assurance Case Metamodel 2.3](https://www.omg.org/spec/SACM/2.3/PDF), §§7 and 11.

**Recommended change.** V-01/V-02 should support an assessment rationale that connects a criterion to the evidence and states important assumptions or rebuttals. For complex outcomes, an optional assurance-case profile can reuse SACM instead of inventing another general argument graph.

For an EH&S example, a signed inspection record may support a claim that a specific inspection occurred. It does not automatically establish that every operating condition is safe. The acceptance method must identify what was inspected, which conditions were covered, what remains untested, and who accepts the residual uncertainty. This is an illustrative evidence-design distinction, not a regulatory determination.

### 7.3 Verification and acceptance may require different roles

**Source finding.** The IETF's RATS architecture distinguishes Evidence, a Verifier's appraisal, and the Relying Party's policy for using the resulting attestation. A positive appraisal is not automatically acceptable to every relying party. This is an Informational RFC about remote attestation. [RATS Architecture, RFC 9334](https://www.rfc-editor.org/rfc/rfc9334.html), §§5 and 8.

**Design inference.** For workflow outcomes, distinguish the evidence producer, technical assessor, and business acceptance authority where appropriate. Borrow the separation of responsibilities; do not require hardware attestation for every workflow or claim that RATS proves business outcomes.

The next draft should allow these roles to be combined in low-risk cases and require independence only under an explicit policy. “Independent verifier” should specify independence of identity, responsibility, evidence source, and method where relevant, rather than merely demanding a second model call.

## 8. Lifecycle and semantic composition need more exactness

### 8.1 Multiple waits need an aggregation rule

**Source finding.** W3C SCXML defines hierarchical and parallel state configurations, history, and event-processing semantics. It provides mature prior art for specifying which states are simultaneously active and how transitions are selected. [SCXML Recommendation](https://www.w3.org/TR/scxml/), §§3 and Appendix D.

**Draft weakness identified.** L-02 assigns one wait reason to a run. A parallel run can simultaneously await a human review, an external system, and a timer while another branch remains active. A single reason loses operational information.

**Recommended change.** Record wait conditions per task or obligation; derive run status under explicit aggregation rules. Define precedence for cancellation, active work, unresolved effects, and all-branches-waiting. Adopt the useful state semantics without implying that SCXML's history feature supplies crash durability or that XML is mandatory.

### 8.2 Semantic service composition predates LLM agents

**Source finding.** OWL-S 1.1 describes service profiles, process models, and protocol groundings, including inputs, outputs, preconditions, and effects used for discovery and composition toward an objective. It is a 2004 W3C Member Submission, not a W3C Recommendation. [OWL-S: Semantic Markup for Web Services](https://www.w3.org/submissions/OWL-S/).

**Assessment.** Our distinction between capability syntax and capability meaning has substantial prior art. Add semantic preconditions and expected effects to capability contracts. Avoid claiming that objective-driven composition or an abstraction above transports is uniquely new to LLM workflows.

### 8.3 Event identity is reusable; event truth and ordering need more

**Source finding.** CloudEvents 1.0.2 defines event identity through the combination of `source` and `id`, including duplicate-event identification. It also supplies an event envelope. [CloudEvents 1.0.2 specification](https://github.com/cloudevents/spec/blob/v1.0.2/cloudevents/spec.md).

**Recommended change.** Evaluate CloudEvents as an event-export binding, while retaining explicit AIWS causal references, committed state revisions, integrity controls, and ledger completeness rules. An event identifier should not double as a business-operation idempotency key unless a separate contract makes that relationship valid.

## 9. Conformance tests, quality benchmarks, and security evaluations answer different questions

### 9.1 Improve the specification before declaring implementations conformant

**Source finding.** W3C's QA Specification Guidelines address conformance models, implementation classes, optional features, extensions, and implementation conformance statements. Open Workflow's CTK provides Gherkin scenarios with implementation-specific steps and runners. [W3C QA Framework: Specification Guidelines](https://www.w3.org/TR/qaframe-spec/); [Open Workflow CTK](https://github.com/open-workflow-specification/specification/blob/main/ctk/README.md).

**Recommended change.** Keep the 30 AIWS scenarios, but turn compound requirements into individually assessable assertions. Publish an applicability table assigning each assertion to definition producers, runtimes, enforcement services, exporters, consumers, and adapters. Supply fixtures and observable assertions, not just scenario descriptions. Passing the Open Workflow CTK would demonstrate the relevant DSL behavior; it would not establish compliance with AIWS authority or evidence requirements.

### 9.2 Repeated task success needs empirical evaluation

**Source finding.** The original τ-bench evaluates tool-using agents with domain policies and simulated users, compares resulting database state against intended goals, and introduces `pass^k` for reliability across multiple trials. [τ-bench paper](https://arxiv.org/abs/2406.12045), Yao et al., 2024.

**Recommended change.** Add repeated-trial evaluation where the acceptance method relies on probabilistic behavior. Keep outcome quality separate from control conformance: a runtime may correctly block an unauthorized action even when its agent fails to complete the task. A good average outcome score cannot excuse a mandatory authorization bypass.

### 9.3 One malicious instruction is not a security assessment

**Source finding.** AgentDojo provides an extensible environment for measuring tool-using agents under malicious external content. It separates user-task success from attacker objectives. [AgentDojo paper](https://arxiv.org/abs/2406.13352), Debenedetti et al., 2024.

**Recommended change.** Expand AT-18 into an adversarial test family covering tool results, retrieved documents, delegated messages, memory, and attempted disclosure through otherwise allowed tools. Define the attacker-controlled surface and permitted target action for each case.

**Additional evidence and limit.** CaMeL explores enforcement of control and information flows outside the model. Its repository explicitly describes the implementation as a research artifact that may contain security defects. [CaMeL paper](https://arxiv.org/abs/2503.18813); [CaMeL research repository](https://github.com/google-research/camel-prompt-injection).

This supports examining concrete enforcement mechanisms rather than relying solely on an instruction to ignore malicious content. It does not justify a universal “prompt-injection-proof” claim or prescribe CaMeL as a production dependency.

**Disconfirming evidence.** The 2026 AutoDojo preprint evaluates adaptive attacks and reports that defenses can fare worse when attacks are optimized against them, especially when the user's task leaves the choice of action to untrusted content. [AutoDojo, version 2](https://arxiv.org/abs/2606.15057v2).

The apparent disagreement with AgentDojo's description as an extensible environment is bounded: a framework can be extensible while a particular evaluation uses a fixed attack set. Our test specification should explicitly distinguish fixed test vectors from adaptive attack campaigns. Neither benchmark is an AIWS certification scheme.

### 9.4 Proposed four-part assessment model

The following division is our recommended synthesis:

| Assessment | Question | Appropriate evidence |
|---|---|---|
| Control conformance | Are mandatory workflow rules enforced? | Deterministic tests, denied-action cases, fault injection, state/effect records |
| Outcome quality | Does the agent accomplish the intended task reliably? | Domain evaluation, repeated trials, independently assessed artifacts |
| Adversarial robustness | Can an attacker cause prohibited behavior? | Explicit threat model, fixed and adaptive attacks, disclosure/effect checks |
| Interoperability | Do implementations preserve the declared contract across a boundary? | Shared fixtures, loss reports, adapter mappings, paired-runtime trials |

The reports should disclose their own scope and limitations rather than collapse these results into one “compliant AI” badge.

## 10. Research-driven change register for AIWS-001

“Strong basis” means the supporting source semantics are explicit. It does not mean the recommended AIWS wording has itself been standardized. “Design choice” identifies a decision the cited work informs but does not settle.

| Priority | Affected draft requirement | Recommended disposition | Evidence basis / remaining decision |
|---|---|---|---|
| Critical | M-01, I-02, B-02 | Add an authorized contract/mission identity and runtime-segment mapping. Carry consumed budget and unresolved effects across continuation and linked replacement. | Temporal execution chains; emerging mission draft. Identity model remains a design choice. |
| Critical | AU-03 | Define type-specific authority-containment rules with an indeterminate result. | Strong basis: RFC 9396 does not define generic comparison. |
| Critical | AU-02, AU-04 | Block affected dispatch on failed mandatory checks; inspect policy diagnostics. | Strong basis: Cedar skip-on-error behavior. Mandatory-check classification remains a policy choice. |
| High | AU-05, AU-06 | Specify revocation freshness and trusted remote enforcement, independently from actor history. | RFCs 8693, 7009, 7662; deployment guarantees need testing. |
| High | D-04 | Define audit playback, recovery replay/continuation, and new execution separately. | Temporal and LangGraph provide different replay/resume boundaries. |
| High | E-03–E-06 | Add target guarantee metadata and crash-boundary tests; retain UNKNOWN outcomes. | Stripe and outbox documentation; exact target behavior must be verified per binding. |
| High | H-01, H-02 | Define material approval fields, final comparison, retry reuse, and withdrawal semantics. | OWASP transaction authorization; reuse policy is an AIWS design choice. |
| High | L-01, L-02, P-03 | Represent simultaneous wait obligations and define run-status aggregation. | SCXML and the draft's own parallel-task requirements. |
| High | V-01, V-02 | Add provenance alignment, support rationale, assumptions, and counter-evidence. | PROV-O and SACM; an AIWS mapping must still be specified. |
| High | V-03, V-04 | Separate technical appraisal from accountable acceptance where required. | RATS supplies a useful architectural analogy; workflow acceptance remains domain-specific. |
| High | X-01–X-04, AT-18 | Specify the trust-boundary mechanism and expand adversarial tests beyond one injected instruction. | AgentDojo, CaMeL, AutoDojo; no universal security claim is established. |
| High | T-01–T-04 | Split compound assertions and publish a role/applicability matrix and executable fixtures. | W3C QA guidance and Open Workflow CTK. |
| Medium | O-01, EX-01, BI-01 | Evaluate a versioned CloudEvents export mapping without claiming ordered or trusted delivery from the envelope. | CloudEvents identity/envelope semantics. |
| Medium | E-01, P-02 | Strengthen semantic preconditions/effects and traceability of dynamic capability selection. | OWL-S prior art; capability validation remains implementation work. |
| Medium | C-01, C-02, CH-01 | Distinguish core technical claims from an organizational adoption profile. | W3C conformance guidance; avoids making a library component responsible for every operator obligation. |
| Defer | MG-01–MG-03 | Keep live cross-runtime migration experimental until a specific runtime pair passes transfer and split-ownership tests. | Reviewed runtime docs establish local recovery, not universal checkpoint interchange. No general migration guarantee verified. |

## 11. What this research supports—and what remains unresolved

### Well-supported directions

The additional evidence strengthens the case for explicit authority enforcement, safe handling of retries and unknown effects, durable recovery, approval-to-action binding, separate verification, and scoped conformance claims. It also provides concrete specifications to reuse for provenance, assurance, human-task management, and event exchange.

The strongest contribution we can aim for is a **precise composition contract**: how these pieces preserve authority, effects, state, and acceptance obligations when used together in an AI workflow. This is a design assessment, not a finding that no competing comprehensive proposal exists.

### Decisions that sources do not settle for us

- **Mandatory core size:** How much persistence, independent review, and evidence is appropriate for a small local assistant versus a consequential enterprise workflow?
- **Mission and revision identity:** How should an approved objective evolve without losing continuity or creating an authority reset?
- **Independence:** When is a deterministic validator enough, and when does the decision require a separately accountable human or organization?
- **Resources:** Which units and reservation rules can be guaranteed when provider usage or prices become known only after execution?
- **Cancellation:** Which visible status should summarize stopped dispatch with unresolved external effects?
- **Migration:** What minimum portable checkpoint can two unrelated runtimes actually honor?
- **Assurance scope:** Which threat models justify signatures, external attestation, or stronger evidence protection?

These questions need explicit decisions, representative use cases, and implementation evidence. More source collection alone will not resolve them.

### Research limitations and stopping point

No runtime, policy engine, benchmark, or migration adapter was installed or tested. Product documentation establishes documented behavior, not independent confirmation of an implementation. Papers establish the reported experimental results under their own methods, not general guarantees. This report deliberately avoids current-model rankings and benchmark performance percentages that would distract from the semantic questions.

Detailed standards inspection was selective: relevant clauses and definitions were read rather than every line of every referenced standard. The FIPA Contract Net specification was discovered through its official indexed material, but direct retrieval failed; it remains a follow-up lead and is not used as a substantive foundation here. The full text of commercial ISO/IEC standards was not available and no clause-level conformity mapping is claimed.

Discovery stopped after each selected gap had primary evidence or an explicit unresolved boundary, and follow-up had checked the high-impact exceptions. Further broad protocol cataloging was unlikely to change the immediate draft corrections. The next useful evidence would come from a narrowly scoped semantic model and failure-injection pilot.

## 12. Recommended next step

Prepare a **design-decision review before edition 0.2**. Resolve the mission/run distinction, authority-comparison model, policy-error behavior, lifecycle aggregation, and assessment roles first. These decisions change the metamodel; adding more SHALL statements before settling them would make the draft longer without making it clearer.

Then build a small reference contract and simulator covering:

1. A bounded document-review workflow with human acceptance.
2. A delegated task whose authority is revoked while waiting.
3. An external record update that commits before its acknowledgment is lost.
4. A continuing execution segment that must preserve consumed budget and unresolved work.

Each should have explicit legal state transitions in the technical sense of permitted transitions, failure fixtures, and an observable acceptance result. The project can then evaluate whether an existing runtime and one adapter preserve the contract before expanding to migration or many protocol bindings.

AAIF remains a relevant place for coordination: its public workflow repository identifies related workstreams spanning workflow integration, identity/trust, security, and observability. That supports liaison across specialties, not an assumption that the foundation has accepted our proposal. [AAIF workflow working group](https://github.com/aaif/wg-workflows-and-process-integration).

## Source register

This register records the material sources used in the supplement. Unless a date is shown, the source is living documentation with no fixed publication date established in this review. All were accessed on 7 September 2026. Section-level links above identify the relevant use; the register preserves publication type and scope.

| ID | Source and publisher | Edition/date and evidence type |
|---|---|---|
| R01 | [OAuth 2.0 Token Exchange — RFC Editor / IETF](https://www.rfc-editor.org/rfc/rfc8693.html) | RFC 8693, January 2020; Standards Track |
| R02 | [OAuth 2.0 Rich Authorization Requests — RFC Editor / IETF](https://www.rfc-editor.org/rfc/rfc9396.html) | RFC 9396, May 2023; Standards Track |
| R03 | [OAuth 2.0 Token Revocation — RFC Editor / IETF](https://www.rfc-editor.org/rfc/rfc7009.html) | RFC 7009, August 2013; Standards Track |
| R04 | [OAuth 2.0 Token Introspection — RFC Editor / IETF](https://www.rfc-editor.org/rfc/rfc7662.html) | RFC 7662, October 2015; Standards Track |
| R05 | [Authorization — Cedar Policy](https://docs.cedarpolicy.com/auth/authorization.html) | Official implementation semantics; living documentation |
| R06 | [Mission-Bound Authorization for OAuth 2.0 — Karl McGuinness](https://datatracker.ietf.org/doc/draft-mcguinness-oauth-mission/) | Revision 00, 6 July 2026; individual Internet-Draft |
| R07 | [Delegation Receipt Protocol — Ryan Nelson](https://datatracker.ietf.org/doc/draft-nelson-agent-delegation-receipts/) | Revision 10, 13 June 2026; individual Internet-Draft |
| R08 | [Workflow Execution — Temporal](https://docs.temporal.io/workflow-execution) | Product documentation; replay and execution chains |
| R09 | [Workflow Definition — Temporal](https://docs.temporal.io/workflow-definition) | Product documentation; deterministic workflow behavior |
| R10 | [Activity Definition — Temporal](https://docs.temporal.io/activity-definition) | Product documentation; retries and idempotency |
| R11 | [Interrupts — LangChain / LangGraph](https://docs.langchain.com/oss/python/langgraph/interrupts) | Product documentation; checkpointing and node re-execution |
| R12 | [Choosing workflow type — AWS Step Functions](https://docs.aws.amazon.com/step-functions/latest/dg/choosing-workflow-type.html) | Product documentation; scoped execution guarantees |
| R13 | [Idempotent Requests — Stripe](https://docs.stripe.com/api/idempotent_requests) | API reference; target-specific guarantee |
| R14 | [Transactional Outbox Pattern — AWS](https://docs.aws.amazon.com/prescriptive-guidance/latest/cloud-design-patterns/transactional-outbox.html) | Official architecture guidance |
| R15 | [Compensating Transaction Pattern — Microsoft](https://learn.microsoft.com/en-us/azure/architecture/patterns/compensating-transaction) | Official architecture guidance |
| R16 | [WS-HumanTask — OASIS](https://docs.oasis-open.org/bpel4people/ws-humantask-1.1-spec-cs-01.html) | Version 1.1, Committee Specification 01, 2009; historical specification, not a latest-status claim |
| R17 | [Transaction Authorization — OWASP](https://cheatsheetseries.owasp.org/cheatsheets/Transaction_Authorization_Cheat_Sheet.html) | First-party security guidance, not a workflow standard |
| R18 | [PROV-O — W3C](https://www.w3.org/TR/prov-o/) | Recommendation, 30 April 2013 |
| R19 | [Structured Assurance Case Metamodel — OMG](https://www.omg.org/spec/SACM/2.3/PDF) | SACM 2.3, October 2023; formal specification |
| R20 | [RATS Architecture — RFC Editor / IETF](https://www.rfc-editor.org/rfc/rfc9334.html) | RFC 9334, January 2023; Informational |
| R21 | [SCXML — W3C](https://www.w3.org/TR/scxml/) | Recommendation, 1 September 2015 |
| R22 | [OWL-S — W3C-hosted submission by its authors](https://www.w3.org/submissions/OWL-S/) | Version 1.1, 22 November 2004; Member Submission, not a Recommendation |
| R23 | [CloudEvents — CNCF project](https://github.com/cloudevents/spec/blob/v1.0.2/cloudevents/spec.md) | Fixed specification tag v1.0.2; event envelope/identity |
| R24 | [QA Framework: Specification Guidelines — W3C](https://www.w3.org/TR/qaframe-spec/) | Recommendation, 17 August 2005 |
| R25 | [Conformance Test Kit — Open Workflow Specification](https://github.com/open-workflow-specification/specification/blob/main/ctk/README.md) | Project conformance assets; mutable main branch |
| R26 | [τ-bench — Yao, Shinn, Razavi, Narasimhan](https://arxiv.org/abs/2406.12045) | Original research, 2024; benchmark methodology |
| R27 | [AgentDojo — Debenedetti et al.](https://arxiv.org/abs/2406.13352) | Original research, 2024; adversarial evaluation environment |
| R28 | [Defeating Prompt Injections by Design — Debenedetti et al.](https://arxiv.org/abs/2503.18813) | Original research, first submitted March 2025; CaMeL design |
| R29 | [CaMeL research artifact — Google Research](https://github.com/google-research/camel-prompt-injection) | First-party code and explicit implementation limitations |
| R30 | [AutoDojo — Ma et al.](https://arxiv.org/abs/2606.15057v2) | Preprint, version 2, 19 June 2026; adaptive attacks and task specification |
| R31 | [Workflows & Process Integration — AAIF](https://github.com/aaif/wg-workflows-and-process-integration) | Working-group repository; coordination context |

**Relationship to prior work:** This supplement adds evidence, challenges several assumptions, and proposes specific changes. AIWS-001 edition 0.1 remains a working draft; this research does not turn it into a ratified or tested standard.
