Skip to content

M6: bounded scheduler recovery

Praxis naming: This page documents Praxis, the AIWS workflow engine. Existing engine/* source paths, /engine/... routes, aiws-engine/* protocol identifiers, and existing script names remain unchanged for compatibility.

Milestone status: M6 is complete for the D-M6-01 supervised Windows 11 x64 / Node 24 / local SQLite developer profile. M8 is now active: its plan is complete; qualification and expanded release/support claims remain open.

M6 is in progress. This first increment hardens the scheduler in the completed M5 local engine profile. It adds bounded abandoned-claim recovery, stricter lease checks and real process-termination tests. Local backup/restore was added in slice 3; history retention remains open. Recurring schedule catch-up is implemented in the later fixed-UTC scheduling slice. Persistent clock tracking and trust recovery were added in the following slice.

These are TypeScript engine source APIs for Node.js 24 and SQLite. They are not new APIs in the released TypeScript, Python or Rust SDK 0.3.0 packages.

Process a backlog without one unlimited transaction

Section titled “Process a backlog without one unlimited transaction”

The coordinator processes a configurable number of abandoned claims and pending one-time wake-ups per reconciliation pass. Each transaction commits the records it processed. Remaining work stays durable for the next pass, including after restart.

import { EngineCoordinator } from './engine/src/coordinator.ts';
import { BoundedStorageExecutor } from './engine/src/storage-executor.ts';
const storage = new BoundedStorageExecutor('./engine.sqlite');
const coordinator = new EngineCoordinator(storage, {
automatic: true,
recoveryBatchSize: 100,
wakeupBatchSize: 100,
});
await coordinator.start();
// Register authenticated workers through your host's integration.
// Later, during controlled shutdown:
await coordinator.stop();
await storage.close();

Both batch sizes default to 100 and must be positive safe integers. An explicit value above 1,000 is accepted; there is no hard-coded product ceiling on workflow or task counts. These settings bound rows processed in each scheduler write transaction, not total database size, every query, or total memory use. Large values can occupy SQLite’s writer for longer; measure before increasing them. The coordinator still reads full snapshots elsewhere, so sustained-scale qualification remains open.

For direct repository use, recoverAbandonedClaims(epoch, now, limit = 100) and drainDueWakeups(epoch, now, limit = 100) return the identities processed by that call. The storage executor forwards the same optional limits. Do not repeatedly drain in a tight loop that starves human controls. Automatic reconciliation already revisits pending work; callers using automatic: false must call reconcileOnce() themselves.

Boundary Behavior
Wake-up becomes due now >= dueAt makes a waiting task ready and marks its occurrence fired in the same transaction
Process killed before commit Neither occurrence firing nor task readiness survives
Process killed after commit, before response Both survive; retry does not fire that occurrence again
Abandoned pre-dispatch claim Claim expiry, scheduler reservation release and task requeue commit together
Large recovery backlog Only the configured batch is processed; later passes retain the remaining responsibility
Old coordinator epoch Writes are rejected, including duplicate drain calls
Expired worker claim Renewal, completion and release reject at expiry equality, even before cleanup has run
Clock before original claim time Claim renewal/completion/release reject; renewal also cannot shorten an existing lease
Clock moves backward after firing Fired occurrences stay fired; pending occurrences wait until their stored UTC deadline is reached

A scheduler claim is pre-dispatch responsibility. Once dispatch is authorized, the separate dispatch ledger owns the attempt and its resource exposure. Recovering a scheduler claim does not reconcile an external effect, refund that exposure, clear an UNKNOWN outcome, or grant permission to execute. Readiness still passes through current approval, worker, policy, hold and resource gates.

The scheduling rollback test proves occurrence deduplication only. The subsequent clock recovery slice adds persisted high-water time, wall/monotonic anomaly detection, dispatch holds and verified restoration. Consult that guide for its trust boundary and platform limits.

Scheduler timestamp arguments must use canonical UTC with exactly three fractional digits, for example 2026-09-10T00:00:00.000Z. Offsets, omitted milliseconds, invalid dates and extended years are rejected. This preserves the ordering used by SQLite’s text deadline indexes. Convert a known valid instant with new Date(value).toISOString() before calling the API.

Lease durations must be positive safe integer milliseconds and produce a deadline no later than year 9999. Batch sizes of zero, fractions, negative values, NaN and infinity now fail instead of being silently clamped. Claim renewal cannot revive an expired claim or move its deadline backward. Existing engine callers using SystemClock or ManualClock already produce the supported timestamp form.

No database schema changes or automatic migrations are introduced. Existing rows are not rewritten or scanned for timestamp compatibility. If a custom host previously persisted noncanonical timestamps, inspect and migrate that installation offline before depending on lexical deadline ordering. Automatic claim recovery now handles at most 100 claims per call by default; custom recovery loops must accommodate partial results.

Terminal window
npm run test:engine:m6-scheduling
npm run engine:typecheck
npm run test:engine

The focused suite covers a simulated year of downtime, partial drains across reopen, backward wall-clock movement, expiry equality, invalid arguments, batch settings through the storage worker, SQL write-failure rollback, and four real process kills before/after wake-up and claim-recovery commits. It checks durable state and retries after acquiring a new epoch. A configured batch above 1,000 is tested as well.

The repository’s optional third constructor argument is a test boundary callback. Production hosts should omit it; no network or worker input configures it. Tests pause a child process at an exact boundary and kill it from the parent. An after-commit failure means the operation may already be committed: inspect or retry using durable identities.

A process kill is not a power failure. The SQL trigger fault is an injected write failure, not a full-disk or failed-device test. Accelerated time proves these transition rules, not months of production reliability. The broader M6 checklist remains open.