Durable workflow engines
Durable workflow engines and the Mesh workflow adapter interface
Independent fact-check: review of this document.
Date: 2026-10-04. Status: research plus a proposal. The interface in section 6 is not agreed.
0. Who this is for and what it answers
Mesh is a planned TypeScript framework in which one resource file declares data, operations
and rules, and Mesh derives the rest. Its architecture proposal is in
notes/research/00-synthesis.md. That document says Mesh should “design the seams (stable step
identity, events committed in the same transaction as state) and stop there” for durable workflows
(section “Do not build yet”), and lists “Durable workflows: which engine, if any” as open question 7.
On 2026-10-04 the operator ruled (notes/decisions-2026-10-04.md, row 7) to compare durable engines
now, in order to define Mesh’s workflow adapter interface, with the in-process runner as the
first adapter. Row 8 of the same file makes the first transport a CLI, so nothing here may
assume an HTTP request.
A durable workflow engine runs a multi-step function so that it survives a crash, a deploy or a
months-long sleep: completed steps are recorded and not run again. A job queue only runs
independent jobs with retries; it has no notion of a multi-step function. This document compares
six durable engines and three job queues, adds two cloud-platform workflow products, and proposes
one TypeScript interface that the in-process runner implements completely and that maps onto
Temporal, Inngest, Restate and DBOS.
How facts were gathered. Facts come from official docs, SDK source on GitHub, the npm registry
and the GitHub API, fetched on 2026-10-04. Three read-only research agents did the first pass; I
re-fetched and re-read the sources for the claims the interface depends on most (marked
“re-checked” below). Anything not confirmed from a fetched page is marked UNVERIFIED.
Contributor counts come from the last page number of the GitHub contributors API requested with
anon=1, which also counts contributors who have no GitHub account, so they run slightly higher
than the named-contributor list (pg-boss shows 106 with and 96 without; Trigger.dev 145 and 140).
Contributor count matters because it estimates how many people could fix a bug in the SDK Mesh
would depend on. Release dates matter because they show whether the project is alive: every engine
below had a release within the preceding five weeks, but Graphile Worker’s cadence is bursty: five npm
releases since June 2026 after a 26-month gap, and only one GitHub Release object (see section
2.7), so a recent date there is not a sign of steady activity.
1. Summary
- Every durable engine exposes the same seven operations: start a named workflow with an
input and a dedupe key, run a named step with retries, durable sleep, wait for an external
signal with a timeout, call a child workflow, cancel, and read status/result. Section 4 lists
them and the cases where an engine’s version of an operation is weaker. - They disagree on what identifies a step. Inngest and Cloudflare key on a string name.
Temporal, Restate and DBOS match recorded steps by position and also check the recorded
name, so a rename is a replay failure on all three (Inngest instead silently re-runs a renamed
step). This is the most important divergence for Mesh’s “stable step identity” seam. Because a
Mesh workflow is declared in a resource file, Mesh can keep names stable and declared; replay
safety also needs the control flow to be deterministic given recorded results, which names
alone do not guarantee (section 6.3, 6.6). - Only DBOS commits a durable-workflow step together with application database writes
(pg-boss does the same for a job’s completion), and only DBOS, pg-boss and Graphile Worker can
enqueue work in the caller’s own Postgres transaction, with conditions (section 2.4). For every
other engine, “events committed in the same transaction as state” needs an outbox: a row
written in the action’s transaction and relayed to the engine afterwards. The relay delivers
at least once, so a workflow is started once only while the relay lag is shorter than the
engine’s dedupe memory (section 6.1). Steps also run at least once on every engine, so a step
that writes to the app database needs an idempotency shim (section 6.1, rule R3). The
proposed interface makes both a Mesh-core responsibility. - Running with no HTTP request narrows the field. Temporal and Hatchet workers connect out
over gRPC and DBOS is a library on Postgres. Restate needs a server plus an HTTP service
endpoint (an outbound tunnel exists but only for Restate Cloud/BYOC). Inngest’s default
serve()is an HTTP endpoint; itsconnect()worker avoids that but needs WebSocket support. - Bun is unproven for most workers. The pg-boss, BullMQ, Restate and Inngest docs claim Bun;
Trigger.dev documents an experimental Bun runtime but its CLI must run under Node; Hatchet
documents a client-only mode on Bun; Temporal explicitly warns against non-Node workers.
Mesh targets Bun and Node (CLAUDE.md), so any adapter other than in-process or pg-boss
carries a runtime caveat. - Recommendation. Define the interface as two small contracts,
JobQueueAdapterand
WorkflowAdapter, with a declared capability list in the style of ruling 4 (an unsupported
capability is a build-time error, never a silent fallback). Build the in-process runner first,
then DBOS as the first external adapter (Postgres only, no server, in-process, native
transactions), then Temporal. Restate and Inngest are mappable but need an HTTP endpoint or
an extra server. Trigger.dev, Hatchet and the cloud products are poor fits (section 7).
2. Per-engine findings
Each subsection is ordered: programming model, step identity, determinism and replay, storage and
infrastructure, transactions, retries/timers/signals/children/schedules/dedupe, versioning,
TypeScript typing, runtime, licence and maintenance, client API. A table in section 3 repeats the
key rows side by side.
2.1 Temporal (@temporalio/*, TypeScript SDK)
- Model. A workflow is an exported async function. Activities (the side-effecting steps) are
plain functions in a separate file, called throughproxyActivities<typeof activities>({ startToCloseTimeout })
(workflows,
activities). - Step identity. The workflow type is the function name, with no way to customise it in
TypeScript (“there isn’t a mechanism to customize the Workflow Type”, workflows page). An activity type name can be customised at worker registration
(activities page). Replay matches commands by sequence and name: the docs warn that “adding, removing, or
reorderingawaitcalls on Command-producing APIs” breaks replay
(versioning),
and the core library raises a nondeterminism error when “Activity type of scheduled event … does
not match activity type of activity command” (it also checks the activity id)
(activity_state_machine.rs).
So a rename breaks an in-flight run. (Closed from reviewer evidence.) - Determinism and replay. Event-history replay. Workflow code runs in a V8 sandbox where
Math.random,DateandsetTimeoutare replaced by deterministic versions; side effects go in
activities (workflows page;
definition). - Storage and infrastructure. A Temporal Service (server plus persistence plus visibility
stores). Persistence: Cassandra, PostgreSQL (specific versions 13.18, 14.15, 15.10, 16.6), MySQL 5.7/8.0;
SQLite for dev only
(persistence;
architecture in service docs).
Workers are separate processes started withWorker.create. A server is always required. - Transactions. No atomic step-plus-application-DB commit. The docs recommend idempotent
activities and “idempotency keys for critical side effects”
(activity definition).
No outbox feature found (absence of evidence; UNVERIFIED). - Retries, timers, signals, children, schedules, dedupe. Activities retry by default with
unlimited attempts and exponential backoff (coefficient 2.0, 1 s initial interval, capped at a
maximum interval of 100 seconds, i.e. 100 times the initial interval; the 10-minute figure on the
same page applies to Workflow Task retries, which retry policies do not govern)
(retry policies);
sleep('30 days')andcondition(fn, timeout)
(timers);
defineSignal/defineQuery/defineUpdate
(message passing);
startChild/executeChild
(children);
schedules and cron
(schedules).
Dedupe: the Workflow Id is unique only among open workflows in a namespace. The default
reuse policy, Allow Duplicate, lets a closed run’s id start again; Reject Duplicate checks only
closed runs still within the namespace’s retention period; the default conflict policy for an
open run returns a “Workflow execution already started” error rather than the existing run
(USE_EXISTINGchanges that)
(workflow id). - Versioning.
patched('id')/deprecatePatch('id')markers in history, or Worker Versioning
with Build IDs, which pin a run to the worker build that started it, so old runs finish on old
code (versioning page above). - Typing.
proxyActivities<A>,client.workflow.start(fn, { args })typed from the workflow
function, genericdefineSignal/defineQuery
(workflow.ts). - Runtime. Workers: Node 20/22/24 only; the README says they “strongly discourage running
Temporal Workers in anything except authentic Node.js”. The client is “believed to work” on Bun,
Deno and Workers but is untested (re-checked,
README). Open Bun
tracking issue: #2273. No inbound HTTP
is needed; the server is reached over gRPC. - Licence and maintenance. MIT. 1.24.0 released 2026-09-15; last commit 2026-10-02; 96 contributors with
anon=1; the
figures come from contributors,
release and
commits. Temporal
Cloud terms: UNVERIFIED. - Client.
client.workflow.start/execute/getHandle, handleresult/describe/cancel/terminate/signal/query/executeUpdate
(workflow-client.ts).
2.2 Inngest (inngest)
- Model.
inngest.createFunction({ id, triggers }, async ({ event, step }) => { await step.run("load-task", ...) })
(step.run). - Step identity. The string id of each
step.run. The SDK hashes the id as the state key and
records the step index; reordering steps only logs a warning, because results are matched by id
(durable workflows,
versioning). - Determinism and replay. Memoisation: “Inngest resumes a run by replaying your handler and
returning saved step results.” Code outside a step may run again on every replay. SDK v4
checkpointing is an optimisation that runs steps eagerly in one request
(checkpointing). - Storage and infrastructure. Inngest Cloud, or self-host with
inngest start: single node
with embedded Redis-compatible store plus SQLite (labelled beta), or external Redis and Postgres
for larger installs
(self-hosting). - Transactions. No documented atomic step plus application DB commit; the docs tell you to
use idempotency keys or upserts because “a request can succeed before Inngest records the step
result” (idempotency).
step.sendEventis atomic only with respect to Inngest’s own state. - Retries and the rest. Default 4 retries (0–20),
timeouts.start/finish,step.sleep,
step.sleepUntil,step.waitForEvent,step.waitForSignal,step.invoke, cron triggers,
event-id dedupe for 24 hours, plus concurrency, throttle, rate-limit, debounce, singleton,
priority and batching
(retries,
create,
invoke).
The 24-hour dedupe window matters: a start request retried after 24 hours is not deduplicated. - Versioning. New code applies to in-flight runs; completed steps are reused by matching id;
a changed id is a new step; for rewrites, create a new function id and let the old one drain
(versioning page above). - Typing.
eventType(name, { schema })accepts any library implementing Standard Schema (a shared
validator interface that Zod, Valibot and ArkType implement), giving typed
event.dataand typedsend
(triggers). - Runtime. Docs claim Node, Bun and Deno; npm
enginesisnode>=20.serve()needs an
HTTP endpoint (serve).
connect()is an outbound WebSocket worker with no inbound HTTP, needing Node 22.4+, Bun 1.1+ or
Deno, and limited, on Inngest Cloud, to 3 connections on Free and 20 on paid plans (a self-hosted
server exposes its own connect gateway)
(connect).
Related open issue: #1624, a memory-retention
performance fix (“Lazily instantiate StreamTools TransformStream to avoid … Bun GC
retention”), not evidence that Bun is unsupported. - Licence (resolved). npm
latestsays Apache-2.0 (re-checked,
registry), the published tarball contains an
Apache-2.0LICENSE.md
(package LICENSE),
and the monorepo has no root LICENSE (the GitHub licence endpoint returns 404 while the repository
metadata saysgpl-3.0), so the GPL-3.0 label is probably stale detector metadata; the source
of the detection stays UNVERIFIED, and the SDK that Mesh would import is Apache-2.0 in
practice. The server (inngest/inngest) is under the Server Side Public
License v1 (SSPL: a source-available licence that requires anyone offering the software as a
service to publish their whole service stack) plus a future Apache-2.0 grant
(LICENSE), which matters
for anyone self-hosting. - Maintenance. SDK 4.21.1 released 2026-10-01
(release); server v1.45.1
released 2026-09-17 (release);
58 (SDK) and 51 (server) contributors withanon=1
(SDK,
server). - Client.
inngest.send({ name, data })returns ids. Status and results come from the REST API
(GET /v1/events/{id}/runs, example).
send()returns event ids; a run id exists only once the run is created and is found through
that endpoint, so a run handle needs a lookup. A TypeScript run-handle with an await-result
helper: UNVERIFIED. Single-run cancel and status exist on the server API
(DELETE /v1/runs/{runID},GET /v1/runs/{runID},
apiv1.go); the
docs say cancellation “prevents queued and sleeping runs from continuing and stops active runs
between steps”
(cancellation);
cancelOnevents and
bulk cancel
also exist.
2.3 Restate (@restatedev/restate-sdk)
- Model. Services, virtual objects and workflows with handlers that take a
ctx; steps are
ctx.run("write", () => ..., retryPolicy)
(services,
durable steps). - Step identity. Position in the journal, with the recorded name also compared: the protocol’s
RunCommandMessageholds the completion id and thename, and message equality is full-struct
equality, so a renamedctx.runis a journal mismatch (error 570)
(messages.rs,
key concepts). The name is optional in
the API. (Closed from reviewer evidence.) - Determinism and replay. Journal replay: “Restate replays the journal, skipping completed
steps”. Non-deterministic work goes inctx.run; usectx.randandctx.date.now(). - Storage and infrastructure. The Restate Server, a single binary with an embedded store, plus
your service as an HTTP endpoint (default port 9080); no external Postgres or Redis; needs a
persistent volume (server overview). The server
is under the Business Source License 1.1 (BSL: source-available; free to use except for a
stated competing use, here offering a “Public Restate Platform Service”, and converting to an open
licence after a set date) (re-checked: LICENSE
begins “Business Source License 1.1” and contains the “Public Restate Platform Service” grant).
Restate Cloud is the hosted service; BYOC (“bring your own cloud”) runs Restate’s managed server
inside the customer’s cloud account. - Transactions. None built in. The docs show user-built patterns: an idempotency-token table
written “in the same transaction as the main operation”, and Postgres two-phase commit
(databases guide). - Retries and the rest. Per-
ctx.runretry policy;ctx.sleep; delayed sends; signals,
awakeables and workflow promises; typed service clients; idempotency keys with 24-hour retention;
no native cron
(timers,
external events,
communication,
clients). - Versioning. Immutable deployments: each invocation is pinned to the deployment where it
started, so no patching API exists; a new endpoint is registered per version and old ones are
drained (versioning). Long sleeps keep old
deployments alive. - Typing. Generic clients (
serviceClient<MyService>); known gap: workflow client generics
(issue 662). - Runtime. Docs prerequisite line: “NodeJS >= v22 or Bun or Deno”
(services); Node HTTP/2, Lambda and fetch
handlers (serving). A server plus an HTTP
endpoint is always required;@restatedev/restate-sdk-tunnelavoids an inbound endpoint but is
for Restate Cloud/BYOC. - Licence and maintenance. SDK MIT; server BSL 1.1. SDK 1.17.2 released 2026-09-21; about 19
contributors (19 withanon=1,
contributors;
release), the smallest of
the engines compared. That matters because very few people can fix a bug in an SDK Mesh would
depend on. - Client.
clients.connect({ url }),workflowClient(...).workflowSubmit(...), result
attach. Cancel by CLI or admin API
(managing invocations);
the client library has no cancel method (nocancelin the clients package’s
api.ts
or ingress.ts),
so cancel is admin API or CLI only.
2.4 DBOS (@dbos-inc/dbos-sdk)
- Model. A library.
DBOS.registerWorkflow(fn)andDBOS.runStep(() => stepOne(), { name: "stepOne" }),
or@DBOS.workflow()/@DBOS.step()decorators
(workflows,
steps). - Step identity. Execution order:
function_idis “the monotonically increasing ID of the step
… based on the order in which steps execute”
(system tables); the recorded name is also
checked:DBOSUnexpectedStepErroris the “step has an unexpected recorded name” error
(error.ts).
(DBOSStepNondeterminismErroris a different error: a step recorded twice with different
results.) - Determinism and replay. Checkpoint to Postgres and resume from the last completed step;
steps must be invoked “with the same inputs in the same order”. Cost: one DB write per step plus
two per workflow (architecture). - Storage and infrastructure. Postgres only (the “system database”), no orchestrator; an
optional Conductor, DBOS’s control plane for monitoring and managing workflows, hosted or
self-hosted. - Transactions (re-checked). “DBOS Transactions are a special kind of step intended for
database access. They execute as a single database transaction, atomically committing both
user-defined changes and a DBOS checkpoint”
(transactions); data
sources exist for Knex, Kysely, Drizzle, TypeORM, Prisma, node-postgres and Postgres.js. A
client callenqueueInTransaction(pgClient, options, ...args)enqueues a workflow in your own
transaction; the caller owns commit/rollback and the workflow is not enqueued until commit
(client reference). Conditions (same
page): the client “must be connected to your DBOS system database”, so the app tables must live
in that database for the write to be atomic with the enqueue;duplicationPolicy: 'return-existing'
is not supported inside a transaction and the default'reject'throws
DBOSQueueDuplicatedError, which aborts the caller’s whole transaction; and no existing-run
handle comes back on a duplicate. The same page documentssendInTransaction(client, destinationID, message, topic, idempotencyKey), which sends a message (a signal) to a running
workflow atomically with your writes. DBOS is the only engine here with transactional steps,
transactional start and transactional signal. - Retries and the rest. Step
retriesAllowed(default false),maxAttempts,backoffRate,
timeoutMS; durable workflow timeouts;DBOS.sleep;DBOS.send/recvwith topics (recvtakes seconds and, by default, returnsnullafter 60
seconds) andsetEvent/getEvent; queues with concurrency, rate limits and priority; DB-stored cron
schedules;workflowIDas an idempotency key
(communication,
schedules). Child workflows: “start child workflows withstartWorkflowand await the results from their
WorkflowHandles”; a timeout cancels the workflow and all its children, and
cancelWorkflow(id, { cancelChildren })exists (workflow tutorial, client reference). - Versioning.
DBOS.patch('id')/deprecatePatch(needsenablePatching) and
applicationVersion(default a hash of workflow source); only same-version workflows are
recovered, so deploys are blue-green
(upgrading). - Typing.
DBOSClient.enqueue<typeof Class.method>is type-safe given the function type;
otherwise args are unchecked. Inputs and outputs must be JSON-serialisable (SuperJSON default). - Runtime. npm
enginesnode >= 20. No official Bun statement; a Bun crash report exists
(#1126, closed: “Production crash:
‘Duplicate debug point name’ on Bun 1.3.1+”). Needs no HTTP
server:DBOS.launch()connects to Postgres, andDBOSClient.create({ systemDatabaseUrl })
works from a CLI. Worker-side Bun/Deno support: UNVERIFIED. - Licence and maintenance. MIT. 5.2.11 released 2026-09-29; 26 contributors with
anon=1
(release,
contributors). - Client (client reference).
DBOS.startWorkflow,handle.getResult(),retrieveWorkflow, plus
getWorkflow,listWorkflows,cancelWorkflow,resumeWorkflow,forkWorkflow(id, startStep).
2.5 Trigger.dev (@trigger.dev/sdk v4)
- Model.
task({ id, retry, run })with ordinary async code and no step primitive
(how it works). - Step identity. None below the task: the task
idonly. Durability across attempts comes
fromidempotencyKeyson childtriggerAndWaitcalls. - Replay vs checkpoint. Checkpoint/restore, not replay (CRIU, “Checkpoint/Restore In Userspace”, is a Linux tool that
freezes a running process’s memory and state to disk so it can be resumed later): while a run waits (triggerAndWait,
wait.forof 60 seconds or more) the platform snapshots the process with CRIU and restores it
later. Cloud only; self-hosted has no checkpoints
(self-hosting). A retry is a fresh attempt
ofrun. - Infrastructure. Cloud, or self-host a webapp (Postgres, Redis) plus workers; documented
minimums are 3+ vCPU/6 GB and 4+ vCPU/8 GB
(docker). - Transactions. No documented atomic commit with app DB writes.
- Retries and the rest.
retry.maxAttempts(default 3),maxDuration,wait.for/until,
wait tokens (wait.createToken/forToken/completeToken) for external events,triggerAndWait,
schedules.task,idempotencyKey, queues withconcurrencyLimit
(wait, triggering,
idempotency). No event-bus wait. - Versioning. A run is locked to the code version it started on; children inherit the
parent’s version (versioning); no in-flight migration. - Typing.
tasks.trigger<typeof task>;schemaTaskvalidates payloads with Zod. - Runtime. Node officially; Bun experimental (
runtime: "bun"), but the CLI must run under
Node (bun guide). No HTTP endpoint in your
app; code is built and deployed to the platform’s workers. - Licence and maintenance. The SDK package is MIT (npm; its
LICENSE
reads “MIT License, Copyright © 2023 Trigger.dev”) while the repository root is Apache-2.0
(repo): different licences for
different packages. 4.7.2 released 2026-10-02
(release); 145
contributors withanon=1
(contributors).
2.6 Hatchet (@hatchet-dev/typescript-sdk)
- Model.
hatchet.task({ name, retries, fn }); DAGs viahatchet.workflowplus
dag.task({ name, parents, fn }); durable tasks viadurableTaskwithctx.sleepForand
ctx.waitForEvent(tasks,
DAGs,
durable sleep). - Step identity. Task
name; DAG edges are explicit. How waits are keyed inside a durable task
is UNVERIFIED. - Determinism and replay. Hybrid: durable tasks replay a durable event log from the latest
checkpoint, and “must be deterministic given the event history”; ordinary tasks are at-least-once
and must be idempotent
(durable tasks,
guarantees). - Infrastructure. Postgres is the source of truth; RabbitMQ is optional; an engine plus API
server run as separate processes; Hatchet Lite is a single image for low throughput
(self-hosting). - Transactions. No documented atomic commit with app DB; guidance is idempotent child tasks.
- Retries and the rest.
retries(retry policies),
scheduleTimeout(5 min default) andexecutionTimeout(60 s default)
(timeouts),ctx.waitForEvent(key, celExpr)
(event waits),child.run()
(children), cron, and idempotency and
concurrency strategies keyed by CEL (Common Expression Language, a small expression language for
evaluating conditions on event data)
(idempotency,
concurrency). - Versioning. None in-flight: “deploy a new durable task definition and let the old work
drain” (Temporal migration page). - Typing. Generic input/output, Zod support; no event-schema registry.
- Runtime. Worker is a long-running process over gRPC (
@grpc/grpc-js), no HTTP server; npm
enginesnode>=20. A fetch-based client-only mode is said to work on Bun/Deno/Workers
(changelog); Bun as a worker
runtime is UNVERIFIED. - Licence and maintenance. MIT (SDK and repo). SDK 1.34.0 released 2026-10-02; engine v0.107.0
released 2026-09-15 (a 0.x engine,
release); SDK 1.34.0 from
npm; 89 contributors with
anon=1
(contributors).
2.7 Job queues
These have no workflow, replay or step model. They are listed because the Mesh interface also
needs a jobs contract (a one-off background action) and because two of them can enqueue in the
caller’s Postgres transaction.
| BullMQ | pg-boss | Graphile Worker | |
|---|---|---|---|
| Version, last release | 6.3.11, 2026-10-01 (npm) | 12.36.0, 2026-10-02 (npm) | 0.18.0, 2026-09-08 (npm); 0.x; five npm releases since June 2026 after a 26-month gap (0.16.6 on 2024-04-26), but only one GitHub Release object, which is what 07-ts-foundation-candidates.md counted. Bursty, not dead |
| Storage | Redis; v6 adds an optional Postgres backend (guide) | Postgres 13+, claiming jobs with SKIP LOCKED (a Postgres clause that lets many workers take different rows without waiting on each other) |
Postgres 12+ |
| Composition | FlowProducer parent/child tree (flows) |
boss.flow([...dependsOn]) (jobs) |
manual: a task calls addJob |
| Same-transaction enqueue and completion | Not possible with Redis (inference); Postgres backend: no documented way to join your transaction (UNVERIFIED) | Yes: send(), insert(), fetch(), complete() accept a db option; “If the transaction rolls back, so does the job” (re-checked, adapters). A transactional worker also commits a worker’s writes together with the job’s completion (same page), the only job-level atomic completion here |
Yes: graphile_worker.add_job is a SQL function callable inside your transaction (docs, re-checked); the JS addJob “simply defers to” the SQL function, so a caller can run graphile_worker.add_job on its own transaction (add-job) |
| Retries | attempts, backoff |
retryLimit, retryBackoff |
max_attempts default 25, exponential backoff |
| Dedupe | jobId, deduplication: {id, ttl} |
singletonKey, queue policies |
job_key (modes replace, preserve_run_at, unsafe_dedupe); completed jobs are deleted so keys are not permanent |
| Delay and cron | delay; job schedulers |
startAfter; cron or RRULE (the iCalendar recurrence-rule format, e.g. “every second Tuesday”) schedule() |
run_at; crontab |
| Timeouts | none built in | expireInSeconds default 15 min, heartbeats |
none documented (UNVERIFIED) |
| In-flight versioning | none | none | none; paid “Worker Pro” adds live migration |
| Bun | README says “Node.js / Bun”; Bun test config in repo; ioredis on Bun UNVERIFIED | README: Node 22.12+ or Bun; fromBunSql adapter for Bun 1.4.0+ |
engines Node >= 22.18; Bun UNVERIFIED |
Licence, contributors (anon=1) |
MIT, 212 (contributors) | MIT, 106 (contributors) | MIT, 60 (contributors) |
Sources: BullMQ README,
BullMQ connections,
pg-boss README,
pg-boss queues,
pg-boss database backends,
Graphile tasks,
Graphile job keys,
Graphile requirements,
Graphile project status.
2.8 Cloud-platform workflows (included because their TypeScript APIs are public and documented)
- Vercel Workflow SDK / Workflow DevKit (npm
workflow5.0.1, Apache-2.0, npm; 125
contributors withanon=1,
contributors). Directives"use workflow"and"use step"compiled
by a framework or bundler plugin; deterministic replay against an event log; steps carry a stable
stepIdusable as an idempotency key
(workflows and steps).
Storage is pluggable (“Worlds”): local, Vercel, Postgres, or custom. The Postgres world runs
outside Vercel but is described by its own docs as a “reference implementation … not optimized
for scale, speed, or security” with no authentication on its HTTP routes, and needs a
long-running HTTP server plus graphile-worker
(postgres world). Whether npmlatest5.0.1 is
“generally available” is UNVERIFIED (no GA statement found). Because it requires a bundler
integration and an HTTP server, it does not fit a CLI-first runtime. - Cloudflare Workflows (generally available since April 2025,
changelog).WorkflowEntrypoint.run(event, step)
withstep.do(name, config, cb), where the name (up to 256 characters) is the replay cache key
(rules).
Closed-source, Workers-only, no self-hosting, so there is no way to run it from a Bun/Node CLI.
Limits include 1 MiB per step result and 10,000 steps by default
(limits). First-party
documentation on in-flight versioning: UNVERIFIED. Both are excluded from the adapter
mapping in section 6.5 because they cannot run without their platform.
3. Comparison table
| Engine | Step identity | Durability model | Infrastructure | Native tx with app DB | No inbound HTTP | Bun worker | Licence |
|---|---|---|---|---|---|---|---|
| Temporal | call order, name checked | event-history replay | Temporal Service + Postgres/MySQL/Cassandra | no | yes (gRPC out) | discouraged; client untested | MIT |
| Inngest | string id | memoised step replay | Cloud, or self-host (SQLite or Redis+Postgres) | no | connect() only |
claimed in docs | SDK Apache-2.0; server SSPL+future Apache |
| Restate | journal order, name checked | journal replay | Restate Server + HTTP endpoint | no (user patterns) | no (tunnel on Cloud) | claimed | SDK MIT; server BSL 1.1 |
| DBOS | order, name checked | checkpoint to Postgres | Postgres only | yes (transaction steps, enqueue and send in tx, with conditions) | yes (library) | unverified | MIT |
| Trigger.dev | task id only | CRIU checkpoint (Cloud only) | Cloud, or self-host (Postgres, Redis, no checkpoints) | no | yes | experimental | npm MIT; repo Apache-2.0 |
| Hatchet | task name; DAG edges | DAG plus durable-log replay | Postgres (+ RabbitMQ optional), engine | no | yes (gRPC out) | unverified | MIT |
| BullMQ | job id | none (queue) | Redis or Postgres | no | yes | partial | MIT |
| pg-boss | job id | none (queue) | Postgres | yes (enqueue in tx; transactional worker commits with completion) | yes | yes (docs) | MIT |
| Graphile Worker | task id | none (queue) | Postgres | yes (SQL add_job) |
yes | unverified | MIT |
| Vercel Workflow | stable stepId |
deterministic replay | Vercel, or Postgres world + HTTP | n/a | no | unverified | Apache-2.0 |
| Cloudflare Workflows | step name | memoised replay | Cloudflare only | n/a | n/a | n/a | closed |
4. The common denominator, and what diverges
4.1 What every durable engine supports
All of Temporal, Inngest, Restate, DBOS, Trigger.dev, Hatchet, Vercel and Cloudflare support:
- Start a named workflow with a typed input and a caller-supplied dedupe key (Temporal
Workflow Id; Inngest event id oridempotencyexpression; RestateidempotencyKey; DBOS
workflowID; Trigger.dev and HatchetidempotencyKey; Cloudflare instanceid). - Run a named unit of work durably, with a retry policy.
- Sleep durably for an arbitrary duration.
- Wait for an external event/signal, with a timeout (Temporal signals, Inngest
waitForEvent
andwaitForSignal, Restate awakeables, DBOSrecv, Trigger.dev wait tokens, Hatchet
waitForEvent, Vercel hooks, CloudflarewaitForEvent). - Call a child workflow and wait for its result.
- Cancel a run (Inngest cancels queued and sleeping runs and stops active runs between steps).
- Read status and result by run id.
Job queues give only 1 (with a dedupe key), 2, a start-time delay (not mid-run sleep), cancel,
and some child composition.
4.2 What diverges
| Axis | Divergence | Why it matters for Mesh |
|---|---|---|
| Step identity | name-keyed (Inngest, Cloudflare, Vercel) vs order-keyed with the name also checked (Temporal, Restate, DBOS) vs none (Trigger.dev) | Mesh’s seam is “stable, declared step names”; order-keyed engines need the order to be stable too |
| Determinism | replay engines forbid non-deterministic code in the workflow body; Trigger.dev (checkpoint) does not | The in-process runner must apply the strictest rule so a workflow that runs in-process also runs on any replay engine |
| Transactions | DBOS: transactional steps, enqueue and sendInTransaction (with conditions); pg-boss: enqueue and a transactional worker that commits with job completion; Graphile Worker: enqueue via SQL add_job; none of the others |
decides whether the outbox and the step-effects shim (6.1.5) are needed |
| Infrastructure | none (DBOS: Postgres you already have) to a server cluster | decides what a Mesh app must provision |
| Transport | HTTP endpoint required (Restate, Inngest serve, Vercel) vs outbound worker |
decides CLI fit |
| Versioning | patch markers (Temporal, DBOS), drain-the-old-version (Restate, Trigger.dev, Hatchet, Inngest new id), none for queues | Mesh declares steps statically, so Mesh can diff versions at build time (section 6.6) |
| Dedupe window | Inngest 24 h; Restate 24 h by default; Temporal only among open runs by default (closed runs reusable unless reject policy, then bounded by namespace retention); DBOS and Hatchet per their policies | a relayed start retried after the window creates a duplicate (section 6.1) |
| Cron | native on all compared engines except Restate; Vercel and Cloudflare cron not checked | a Mesh schedule must degrade to a Mesh-owned timer on Restate |
5. Mesh seams the interface must preserve
From 00-synthesis.md: (a) “Step identity by content hash → Stable, declared step names;
persisted sagas break on redeploy”; (b) “events committed in the same transaction as state”;
© “Explicit context value on every call” (actor and tenant are never ambient); (d) “Silent
fallbacks → hard errors”; and from decisions-2026-10-04.md row 4, adapters declare capabilities
and using one an adapter lacks is a build-time error. Row 8 adds (e): a workflow may be started
from a CLI with no HTTP request.
6. Proposed adapter interface (proposal, not agreed)
6.1 Design decisions
- Two contracts, not one.
JobQueueAdapter(fire-and-forget background action) and
WorkflowAdapter(multi-step durable function). A durable engine implements both (a job is a
one-step workflow); pg-boss and Graphile implement only the first. - Steps are generated, top-level, named functions. A step is not a closure written inside
the workflow body. Temporal activities, for example, are functions registered on the Worker
and called by name from a sandboxed workflow (a closure created inside the workflow cannot be
shipped to an activity worker; run inline it would execute inside the deterministic sandbox,
without I/O). So the generator emits each declared step as a named function with an explicit,
serialisable input, and the workflow body callsctx.step("reserve-stock", input). This also
fits the architecture decision that “generated code carries the behaviour; the shared engine
stays thin” (decisions-2026-10-04.md, row 2). For a loop, a step family is declared
(charge[]): its instances are namedcharge[<key>]where the key is a pure function of the
workflow input and earlier recorded step results (an index, a record id). The verifier knows
families, so the declared list stays finite while loops stay possible. What the verifier
checks: a key expression may only read the workflow input and recorded step results, and may
call no non-deterministic source; it does not attempt general data-flow analysis of therun
body. On engines that match a step by activity type or registered function (Temporal), the
family name is the registered step and the key travels in the step input (6.5). - Capabilities are static data read at build time. Ruling 4 requires a build-time error, so
each adapter publishes a manifest of plain data (thecapabilitiesobject below) that the
build’s Verify stage (stage 5,00-synthesis.mdsection 16) reads without starting the
adapter. A definition that useswaitForSignalagainst an adapter withsignals: false
fails the build.register()repeats the check at boot only as a backstop. - Mesh owns an outbox, and the outbox is at-least-once. The data layer writes a
mesh_outbox
row in the action’s transaction (run-time phases 5–7,00-synthesis.mdsection 17). A row has
akind(start,signal,cancel, or a job), the target, the payload, and a dedupe key
equal to the row id. A relay delivers rows after commit and retries until the engine
acknowledges. Because delivery is at least once, a workflow starts exactly once only while the
relay lag is shorter than the engine’s dedupe memory (dedupeWindowin the manifest). The
relay must refuse loudly (seam d) to deliver a row older than that window and raise an
operator error instead of risking a second run. Per-engine consequence: Temporal’s default
reuse policy lets a closed run’s id start again and its default conflict policy errors instead
of returning the open run, so the Temporal adapter must set reuse policyREJECT_DUPLICATE
and conflict policyUSE_EXISTING, and its window is the namespace’s closed-run retention;
Inngest and Restate windows are 24 hours. Signals have no engine-level dedupe on most engines,
so a signal row carries its id and the generated workflow code ignores a signal id it has already
consumed. - Steps are at-least-once; effects on the app database need an idempotency shim. Every engine
can run a step twice (the step commits, the process dies before the checkpoint is recorded, the
step re-runs). So seam (b) does not hold for a step that writes to the app database unless
Mesh adds a rule: R3. Each call a step makes to a Mesh action passes an effect key that
is unique per call,${runId}:${stepName}:${callIndex}(StepScope.effectKey(callIndex)), so
two actions inside one step do not share a key. The action inserts that key into a
mesh_step_effectstable with a unique constraint on the key, as the first statement of its
own transaction, and commits its writes, its events and the stored result in that same
transaction. A repeat finds the row and returns the stored result. The unique constraint matters
because two attempts of one step can run at the same time (on Temporal, Inngest and Restate a
timed-out attempt is not stopped when the retry starts): the losing transaction fails on the
conflict, rolls back, and then returns the stored result, so the effect applies once. Rows must
be kept longer than the longest retry span plus the replay horizon of the engine (otherwise a
late retry finds no row and re-applies the effect); the retention is a per-adapter setting
with a safe default of the engine’s maximum run retention. ThestepNamein an effect key is always the instance name:reserve-stockfor a plain
step andcharge[3]for a family instance, never the family name alone. If the key used the
family name, every instance would share one key and R3 would return the first instance’s stored
result for all the others, silently skipping the effect. The generator assignscallIndex
per call site (a static index in the generated code); a call inside a loop uses
${index}:${loopKey}where the loop key is deterministic, so the key does not depend on a
run-time counter (which would be unstable underPromise.allor data-dependent loops). If a
reset re-runs steps under the same run id, R3 returns the old recorded results, which is the
intended behaviour (history is replayed, not re-applied). DBOSrewindWorkflowkeeps the run id
and behaves that way; DBOSforkWorkflow(workflowID, startStep, { newWorkflowID })starts a
new execution with a new id, so its R3 keys differ and steps fromstartSteponward re-apply
their effects, while earlier steps are copied
(client reference). Mesh maps fork to
“new run id” and rewind to “same run id”. This is the idempotency-table pattern Restate
documents (databases guide), made safe against
concurrent attempts. It gives exactly-once effect withoutstepInTransaction. Where the adapter declaresstepInTransaction(the
in-process runner, DBOS) the step’s writes and its checkpoint commit together and the shim is
unnecessary. - The same rules cover jobs. A job enqueued from an action goes through the outbox like a
workflow start, unless the job adapter declaresenqueueInTransaction(pg-boss, Graphile
Worker), in which case the enqueue joins the action’s transaction. Job dedupe memory differs:
BullMQdeduplicationhas a ttl, and a Graphilejob_keydisappears once the job is deleted,
so the same lag bound applies.completeInTransaction(pg-boss transactional worker) lets a job
handler’s writes commit with the job’s completion. - No transport in the contract.
start()takes a plain value and returns a handle; whether
the engine is reached over gRPC, Postgres or memory is internal. A CLI can callstart,
attach,signalandcanceldirectly; an HTTP server appears in no signature. An adapter
whose worker needs an inbound HTTP endpoint declaresrunsWithoutHttp: false. - Context is explicit (seam c):
StartOptions.contextis persisted with the run and passed to
every step throughStepScope.context. - A run id is the dedupe key.
RunHandle.runIdequalsStartOptions.dedupeKeyfor every
adapter. The adapter maps it to the engine’s own id (Temporal workflow id, DBOS workflow id,
Restate key) or resolves it lazily (Inngest, whosesend()returns event ids). A child
run started byinvokegets the deterministic id${parentRunId}:${invokeName}(inside a
family,${parentRunId}:${invokeName}[<key>]), mapped to the engine’s child id (Temporal
executeChildworkflow id, DBOSstartWorkflowworkflow id, Restate workflow key), so a
replayed parent re-attaches to the same child instead of starting a new one, and the child’s own
effect keys stay stable.
6.2 Type sketch
/** Opaque to Mesh core; the adapter and the data-layer adapter agree on it. */
type Tx = unknown;
type Duration = number | `${number}${"ms" | "s" | "m" | "h" | "d"}`;
interface RetryPolicy {
maxAttempts: number; // >= 1
initialInterval?: Duration;
backoffFactor?: number;
maxInterval?: Duration;
nonRetryable?: readonly string[]; // error class names
}
interface MeshContext { // seam (c): never ambient
actor: unknown;
tenant?: unknown;
[extra: string]: unknown;
}
/** Static data. Published by the adapter package; read by the build's Verify stage. */
interface WorkflowCapabilities {
startInTransaction: boolean; // start can join the caller's DB tx (see 6.5 DBOS conditions)
signalInTransaction: boolean; // signal can join the caller's DB tx
stepInTransaction: boolean; // a step's app writes commit atomically with its checkpoint
signals: boolean;
childWorkflows: boolean;
schedules: boolean; // native cron
durableSleep: boolean;
cancel: "running" | "pending-only" | "none"; // "running" may take effect between steps
dedupeWindow: Duration | "namespace-retention" | "forever"; // bound for outbox relay lag
runsWithoutHttp: boolean; // worker AND client need no inbound HTTP endpoint
runtimes: readonly ("node" | "bun")[]; // verified runtimes only
requires: readonly ("postgres" | "server" | "saas" | "http-endpoint")[];
}
interface StepScope {
readonly runId: string;
/** The INSTANCE name: `reserve-stock`, or `charge[3]` for a family instance. Adapters rebuild it
* as `family[key]`; they never use the family name (or an engine activity type) alone. */
readonly stepName: string;
/** Per-call key for rule R3: `${runId}:${stepName}:${callIndex}`, where `stepName` is the instance
* name and `callIndex` is a static per-call-site index assigned by the generator. */
effectKey(callIndex: number | string): string; // static call-site index, or `${index}:${loopKey}` in a loop
/** Present only inside `stepInTransaction`: the transaction that commits with the checkpoint. */
readonly tx?: Tx;
readonly attempt: number;
readonly context: MeshContext;
readonly signal: AbortSignal;
}
/** A generated, top-level, registered step. Input and output are JSON-serialisable. */
type StepFn<In, Out> = (input: In, scope: StepScope) => Promise<Out>;
/** What Mesh generates from the resource file. */
interface WorkflowDefinition<I, O> {
name: string; // stable, unique; part of the persisted identity
version: string; // see 6.6
steps: Record<string, StepFn<any, any>>; // declared step functions, by name
stepFamilies?: Record<string, StepFn<any, any>>; // `name[<key>]` instances
run(ctx: WorkflowContext, input: I): Promise<O>; // deterministic body; no I/O
}
interface WorkflowContext {
readonly runId: string;
readonly context: MeshContext;
/** Run a declared step once; its result is recorded. `name` MUST be in `steps`, or `family[key]`. */
step<T>(name: string, input: unknown, opts?: { retry?: RetryPolicy; timeout?: Duration }): Promise<T>;
/** Same, but the step's app-DB writes commit atomically with its checkpoint. Needs stepInTransaction.
* The adapter opens the transaction and hands it to the step as `scope.tx`; the step must use
* it for every app-DB write. Without the capability this is a build error. */
stepInTransaction<T>(name: string, input: unknown, opts?: { retry?: RetryPolicy }): Promise<T>;
sleep(name: string, d: Duration): Promise<void>;
sleepUntil(name: string, at: Date): Promise<void>;
/** Needs signals. Resolves with the payload, or `undefined` when `timeout` elapses.
* `timeout` is REQUIRED; "no timeout" must be written as an explicit very large value, because
* some engines (DBOS `recv`) would otherwise apply a default and return null silently. */
waitForSignal<T>(name: string, opts: { timeout: Duration }): Promise<T | undefined>;
/** Needs childWorkflows. */
invoke<I2, O2>(name: string, child: WorkflowRef<I2, O2>, input: I2): Promise<O2>;
/** Deterministic replacements for non-deterministic sources. */
now(): Date;
random(): number;
uuid(): string;
}
interface WorkflowRef<I, O> { readonly name: string; readonly __io?: [I, O] }
interface StartOptions {
dedupeKey: string; // REQUIRED. The outbox row id. Becomes the run id.
context: MeshContext;
/** Only honoured when startInTransaction. Then `dedupeKey` MUST be fresh (an outbox row id):
* the adapter need not return an existing run for a duplicate and may throw. The handle is
* usable only after the caller commits. */
tx?: Tx;
delay?: Duration;
}
type RunStatus =
| { state: "pending" | "running" | "sleeping" | "waiting" }
| { state: "succeeded" }
| { state: "failed"; error: { name: string; message: string } }
| { state: "cancelled" };
interface RunHandle<O> {
readonly runId: string; // === dedupeKey
status(): Promise<RunStatus>;
result(opts?: { timeout?: Duration }): Promise<O>; // rejects on failed/cancelled
/** `tx` honoured only when signalInTransaction. `signalId` lets the receiver ignore repeats. */
signal<T>(name: string, payload: T, opts: { signalId: string; tx?: Tx }): Promise<void>;
cancel(reason?: string): Promise<void>;
}
interface WorkflowAdapter {
readonly id: string; // "in-process", "dbos", "temporal", ...
readonly capabilities: WorkflowCapabilities; // MUST equal the build-time manifest
/** Boot-time backstop: rejects if a definition needs a missing capability. */
register(defs: readonly WorkflowDefinition<any, any>[]): Promise<void>;
/** Idempotent on `dedupeKey` within `dedupeWindow` when no `tx` is passed. */
start<I, O>(wf: WorkflowRef<I, O>, input: I, opts: StartOptions): Promise<RunHandle<O>>;
attach<O>(runId: string): RunHandle<O>;
/** Process runs in this process until `stop()`. A CLI calls this for `mesh worker`. */
work(): Promise<void>;
stop(): Promise<void>;
}
interface JobQueueAdapter {
readonly id: string;
readonly capabilities: {
enqueueInTransaction: boolean; // pg-boss, Graphile Worker
completeInTransaction: boolean; // pg-boss transactional worker
delay: boolean; cron: boolean; dedupe: boolean;
dedupeWindow: Duration | "until-completion" | "forever";
};
enqueue<P>(job: string, payload: P, opts: { dedupeKey: string; context: MeshContext; tx?: Tx; delay?: Duration; retry?: RetryPolicy }): Promise<{ jobId: string }>;
work(handlers: Record<string, (payload: any, ctx: { context: MeshContext; signal: AbortSignal; tx?: Tx }) => Promise<void>>): Promise<void>;
stop(): Promise<void>;
}
6.3 Rules the adapter contract imposes on generated code
- Step names are declared in the definition (
steps, or a declared family with a key that is a
pure function of recorded results); the verifier fails the build on a duplicate name, an
undeclared name or a family key that reads anything but the input and recorded results (seam a,
stage 5 “Verify”). The uniqueness rule also coversinvoke,sleepandwaitForSignalnames
(child ids derive frominvokeName, and sleep and wait names are step ids on Inngest): two
invoke("x", ...)calls would map to one child id, which fails on Temporal while the first child
is open and returns the first child’s result on DBOS. Inside a family the rule applies per key. - Code in
runoutsidectx.stepmay only callctx.now/random/uuid,sleep,waitForSignal,
invoke, and pure functions of recorded results. This is the strictest replay rule
(Temporal’s), applied everywhere, so a workflow valid in-process is valid on every replay engine. - Step inputs and outputs are JSON-serialisable (DBOS requires it; the others serialise too).
A step cannot close overctx; whatever it needs arrives in its input andStepScope. - Steps are at-least-once. Effects outside the app database need an idempotency key derived from
scope.effectKey(callIndex); effects inside the app database follow rule R3 (6.1.5). - Every
signal,cancelorstartissued from inside an action goes through the outbox (6.1.4)
unless the adapter declares the matching...InTransactioncapability.
6.4 The in-process runner implements everything
The first adapter keeps a journal (mesh_runs, mesh_steps(run_id, step_name, result, attempt),
mesh_step_effects, mesh_outbox) in the app’s own database (SQLite, Postgres, or memory for
tests). Because it shares the transaction with the data layer, startInTransaction,
signalInTransaction and stepInTransaction are all true; signals, childWorkflows,
durableSleep and schedules use a polling loop; dedupeWindow is "forever";
runsWithoutHttp is true; requires is empty. It replays the run body and returns recorded
results by step name. Cancel is "running": cancel() writes a persisted cancel flag on the
run row, and the worker polls it and fires the step’s AbortSignal. A bare AbortSignal could not
work when the CLI that cancels and the process running work() are different processes. This is a
design claim; no code exists yet.
6.5 Mapping tables
Legend: native = maps directly; shim = Mesh or the adapter builds it; cannot = the
engine cannot honour it and the capability is declared false.
Temporal
| Mesh operation | Temporal |
|---|---|
register |
Worker.create({ workflowsPath, activities, taskQueue }); the generator emits the declared steps as the registered activities and a workflow wrapper whose ctx.step(name, input) is proxyActivities<…>()[name](input), where name is the declared step name, or for a step family the family name (charge) with the instance key passed in the input (consistent with the step row) (native, because steps are top-level functions, 6.1.2) |
step(name, input) |
an activity whose type is the declared step name; Temporal matches by position and checks the activity type, so a rename breaks in-flight runs (6.6). For a step family the activity type is the family name (charge) and the instance key travels in the input and the activity id, because charge[3] is not a registered activity type; the adapter rebuilds the instance name charge[<key>] from that key (not from the activity type) when it builds StepScope.stepName, so effect keys stay per instance |
stepInTransaction, signalInTransaction, startInTransaction |
cannot (no atomic commit with the app DB); use rule R3 and the outbox |
sleep |
sleep() (native) |
waitForSignal |
defineSignal + condition(fn, timeout) (native) |
invoke |
executeChild with workflowId = ${parentRunId}:${invokeName} (native) |
start with dedupeKey |
workflowId = dedupeKey with reuse policy REJECT_DUPLICATE and conflict policy USE_EXISTING (native with non-default policies; with defaults a closed run’s id can start again). dedupeWindow: "namespace-retention": Reject Duplicate checks only closed runs still retained |
cancel |
handle.cancel() (native) |
schedules |
client.schedule.create (native) |
| Constraints | Node workers only (Bun discouraged); server required; version maps to Worker Versioning (Build IDs); patched() needs both code branches kept by hand |
Inngest
| Mesh operation | Inngest |
|---|---|
register |
createFunction({ id: def.name }) per definition; each declared step function is called from step.run(name, ...) inside the generated handler; plus serve() or connect() |
step(name, input) |
step.run(name, () => fn(input, scope)) where fn is steps[name], or stepFamilies[family] for an instance family[key] (native, name-keyed; a rename silently re-runs the step) |
stepInTransaction, signalInTransaction, startInTransaction |
cannot; use R3 and the outbox |
sleep |
step.sleep(name, d) (native) |
waitForSignal |
shim. Inngest does not buffer signals: the signal API answers 404 “No signal found” when no run is waiting (signals.go); the signal string is global, not per run; and by default duplicate signals fail the function (onConflict) (reference). Mesh signals are buffered (Temporal, DBOS and Restate buffer), so the adapter composes the signal string as ${runId}:${name}, and on a 404 the relay checks GET /v1/runs/{runID}: while the run is live it treats the 404 as “not yet waiting, retry later” (the row stays undelivered without tripping the stale-row refusal, 8.4); if the run is terminal (finished, failed or cancelled) the relay dead-letters the row with an operator error (seam d), so a signal for a run that will never wait cannot retry forever or block later signals to the same run. Alternative: step.waitForEvent with an if match on the run id; no evidence found that it matches events sent before the wait starts, so it does not avoid the problem and needs the same terminal-run rule |
invoke |
step.invoke (native); step.invoke does not let the caller set the child’s Inngest run id, so the adapter puts the deterministic Mesh child id ${parentRunId}:${invokeName} into the invoke payload and the generated child handler uses it as its runId (which its StepScope and effect keys depend on) |
start with dedupeKey |
shim: event id = dedupeKey, which dedupes for 24 h only (dedupeWindow: "24h"); send() returns event ids, so the adapter resolves the run through GET /v1/events/{id}/runs and caches the mapping |
attach / status / result |
shim: GET /v1/runs/{runID} (or the events endpoint); SDK await helper UNVERIFIED |
cancel |
DELETE /v1/runs/{runID} or cancelOn: declare cancel: "running" (queued and sleeping runs do not continue; active runs stop between steps) |
runsWithoutHttp |
only with connect() (WebSocket; on Inngest Cloud 3 or 20 connections by plan) |
| Constraints | server is SSPL; SDK is Apache-2.0; replay reruns the body, so rule 6.3.2 is mandatory |
Restate
| Mesh operation | Restate |
|---|---|
register |
a restate.workflow({ name, handlers: { run } }) per definition, served over HTTP; ctx.step(name, input) is ctx.run(name, () => fn(input, scope)) with fn resolved from steps[name], or stepFamilies[family] for an instance family[key]; the run handler first stores ctx.request().id in workflow state (see waitForSignal) |
step(name, input) |
ctx.run(name, fn, retryPolicy) (native; journal position and recorded name are compared, so a rename is a journal mismatch) |
stepInTransaction, signalInTransaction, startInTransaction |
cannot natively; the idempotency-table pattern is exactly rule R3, which is not atomic with the journal |
sleep |
ctx.sleep (native) |
waitForSignal |
shim. A Restate signal is identified by the target invocation id and a name and is resolved from inside another handler (ctx.invocation(id).signal(...).resolve(...), external events); Mesh addresses runs by workflow key and sends signals from the outbox relay, which runs outside Restate; an awakeable’s generated id would have to be published first. So the generator adds a shared handler signal(name, payload) to every workflow; the relay calls it through the workflow client with the workflow key and the handler resolves the run’s signal. To address the run handler’s invocation, the run handler stores ctx.request().id in workflow state as its first action and the shared handler reads it; a signal that arrives before that id is stored fails with a retryable error and the relay retries. A workflow promise resolves only once, so repeated signals use signals, not promises |
invoke |
ctx.workflowClient(...) call with workflow key ${parentRunId}:${invokeName} (native) |
start with dedupeKey |
idempotencyKey or the workflow key (native, 24 h default retention) |
cancel |
admin API PATCH :9070/invocations/{id}/cancel (the client library has none); the adapter calls the admin API |
schedules |
cannot (no native cron); Mesh runs its own timer and calls start |
| Constraints | server plus HTTP endpoint always. The worker (work()) needs an inbound HTTP endpoint (runsWithoutHttp: false); a CLI that only calls start, signal or cancel makes outbound HTTP calls to Restate’s ingress and admin API, which ruling 8 (Mesh’s own transport) does not forbid. Server is BSL 1.1 |
DBOS
| Mesh operation | DBOS |
|---|---|
register |
DBOS.registerWorkflow(fn) per definition; declared steps become DBOS.registerStep; DBOS.launch() |
step(name, input) |
DBOS.runStep(() => fn(input, scope), { name, retriesAllowed, maxAttempts }) with fn resolved from steps[name] or stepFamilies[family] (native; order and recorded name checked via DBOSUnexpectedStepError) |
stepInTransaction |
a DBOS transaction step on the Drizzle/Kysely/Knex/Prisma/pg data source (native, atomic with the checkpoint). Needs Mesh’s data-layer transaction to be that data source (open question 8.1) |
sleep |
DBOS.sleep(ms) (native) |
waitForSignal |
DBOS.recv(topic, timeoutSeconds); the adapter always passes an explicit timeout, converts Duration to seconds, and uses a very large value for “no timeout”, because the default of 60 s returns null |
invoke |
startWorkflow inside a workflow with workflowID = ${parentRunId}:${invokeName}, awaiting the child handle (native) |
start with dedupeKey |
workflowID = dedupeKey (native); on a duplicate the default returns the existing run |
startInTransaction |
client.enqueueInTransaction(pgClient, ...), with conditions: the client must be connected to the DBOS system database (app tables must live there for atomicity); 'return-existing' is unsupported and the default 'reject' throws DBOSQueueDuplicatedError, aborting the caller’s transaction. So the start contract’s “a second call returns the same run” does not hold with tx; the contract (6.2) therefore requires a fresh dedupeKey when tx is passed, which an outbox row id always is. The workflow is not enqueued until the caller commits |
signalInTransaction |
client.sendInTransaction(pgClient, workflowId, message, topic, idempotencyKey) (native, same system-database condition) |
cancel |
cancelWorkflow(id, { cancelChildren }) (native) |
schedules |
DBOS.createSchedule (native) |
| Constraints | Postgres only (no SQLite adapter); Bun worker support UNVERIFIED; version maps to applicationVersion (default a hash of workflow source) or DBOS.patch |
Operations each engine cannot honour (capability declared false; using it is a build error):
| Engine | Cannot honour |
|---|---|
| Temporal | stepInTransaction, signalInTransaction, startInTransaction; dedupe beyond namespace retention; Bun worker |
| Inngest | stepInTransaction, signalInTransaction, startInTransaction; dedupe beyond 24 h; runsWithoutHttp unless connect() |
| Restate | stepInTransaction, signalInTransaction, startInTransaction, schedules; runsWithoutHttp for the worker |
| DBOS | non-Postgres databases; startInTransaction for a duplicate key; the transactional capabilities when app tables are outside the DBOS system database |
6.6 Versioning running workflows
Replay engines break when a persisted run meets changed code. A hash of the declared step names is
a necessary check, not a sufficient one: replay safety depends on the body being deterministic
given recorded results (“the same steps with the same inputs in the same order (given the same
return values)”,
DBOS workflow tutorial), so a
changed branch condition alters the run-time order without touching the declared list. Because
Mesh commits its generated code, the guard (stage 8, 00-synthesis.md section 16) can hash the
generated run body as well as the step names and the step-function inputs, and fail CI when that
hash changes for a released workflow without a migration entry. Renaming a step is breaking on
Temporal, Restate and DBOS (name checked) and silently re-runs the step on Inngest. Adapters map the
policy: Temporal Worker Versioning (Build IDs) or hand-written patched() branches, DBOS
applicationVersion or DBOS.patch, Restate, Trigger.dev and Hatchet drain old deployments,
Inngest adds a new function id. Whether Mesh should automate any of this is a decision for the
operator.
7. Fit assessment
| Engine | Adapter fit | Reason |
|---|---|---|
| In-process | first adapter | implements every operation, no infrastructure |
| DBOS | best external fit | Postgres only, library, native transactions for steps, starts and signals, no HTTP; unknowns are Bun worker support (must be tested before DBOS is chosen as the first external adapter), the shared-transaction integration, and the system-database condition |
| pg-boss | JobQueueAdapter |
native enqueue and completion in a transaction, Bun documented |
| Temporal | good second adapter | mature, strict model matches the contract (generated steps become registered activities); needs a server, Node workers, non-default dedupe policies |
| Inngest | possible | name-keyed identity fits; run-id shim, 24 h dedupe, HTTP or connect(), SSPL server are costs |
| Restate | possible but awkward | worker needs an HTTP endpoint, no cron, BSL server, 19 contributors |
| Graphile Worker | JobQueueAdapter |
0.x, Node-only, but SQL-level enqueue is the cleanest transactional job primitive |
| BullMQ | JobQueueAdapter for Redis shops |
no same-transaction enqueue |
| Hatchet | poor | task-name model, no transactions, engine process to run |
| Trigger.dev | poor | no step primitive; checkpointing is Cloud only; a Mesh step would have to become a child task |
| Vercel, Cloudflare | excluded | bound to their platforms |
8. Open questions
- Can Mesh’s data-layer transaction be exposed as a DBOS data source, and can Mesh’s app tables
live in (or share a connection with) the DBOS system database? Without both, DBOS cannot give
stepInTransaction,startInTransactionorsignalInTransactionatomically, and DBOS
degrades to the outbox plus rule R3 like the other engines. Not tested. - Are step families (6.1.2) enough for dynamic control flow, or does the resource-file syntax need
first-classfor-eachandparallelsteps? Not decided; MX tag design is out of scope here. - Should the outbox relay be core or an extension? The synthesis puts “events committed in the
same transaction as state” in core; this document keeps it there. - What should the relay do when a row is older than the adapter’s
dedupeWindow: stop and alert,
or start anyway after a human check? This document proposes “stop and alert” (seam d). - Still UNVERIFIED (none changes the interface; DBOS worker Bun support changes only the
adapter order): Inngest’s TypeScript run-handle helper and the source of its GPL-3.0
detection (Restate’s client-library cancel is closed: there is none); DBOS worker-side Bun and
Deno support, which must be tested before DBOS is picked as the first external adapter;
the Trigger.dev idempotency-key TTL; that Temporal
has no outbox-style feature (an absence, not a claim needing a source);
Hatchet’s durable-wait keying and Bun worker support; BullMQ ioredis on Bun and enqueue in a
caller’s transaction on its Postgres backend; Graphile timeouts and Bun; Vercel GA status,
run-level dedupe and cron; Cloudflare in-flight versioning and duplicate-id behaviour;
Temporal Cloud terms.
9. Not changed
No other research document was edited. 07-ts-foundation-candidates.md lists Inngest SDK 4.21.0,
Trigger.dev 4.7.0, Hatchet SDK 1.33.2 and pg-boss 12.35.1; the engines have since released
4.21.1, 4.7.2, 1.34.0 and 12.36.0 (one to three weeks later). Its open item on the Inngest
licence (line ~1413) can now be closed: the SDK is Apache-2.0. That document is left as it is.
10. Review corrections (round 1)
The independent review (notes/team-lead/reviews/08-durable-engines-review.md) found 30 issues;
all were applied. I re-fetched the sources for the Temporal retry cap, the Temporal reuse and
conflict policies, the DBOS enqueue conditions and recv default, the Inngest run endpoints, the
Graphile npm release history, the Trigger.dev SDK licence and the pg-boss transactional worker; each
matched the reviewer. No finding was rejected. One nuance: the Temporal retry-policy page states
both “max interval of 10 minutes” (intro) and “100 seconds” (default section); the document uses
the specific default (100 s = 100 × the 1 s initial interval) and says the 10 minutes belongs to
Workflow Task retries, as the reviewer reads it.
- H1 outbox is at-least-once; dedupe is bounded by
dedupeWindow; relay refuses old rows;
Temporal needsREJECT_DUPLICATE+USE_EXISTING(6.1.4, 6.5). - H2 rule R3 and
mesh_step_effectsfor app-DB writes in steps; signals and cancels go through
the outbox; DBOSsendInTransactionadded; seam (b) claim rewritten (1.3, 6.1.4–6.1.5, 6.3.4–6.3.5). - H3 DBOS enqueue-in-transaction conditions;
startcontract requires a fresh key withtx
(2.4, 6.2, 6.5, 8.1). - H4 steps are generated top-level registered functions; Temporal mapping fixed (6.1.2, 6.2, 6.5).
- M1 “order-keyed, name checked” for Temporal, Restate, DBOS; UNVERIFIED items closed. M2
Temporal retry cap is 100 s. M3 Inngest cancel is"running", status viaGET /v1/runs.
M4 Inngest start/attach is a shim;runIdis the dedupe key. M5 step families resolve
the naming contradiction. M6 hash is necessary, not sufficient; guard hashes therunbody.
M7 pg-boss transactional worker andcompleteInTransactionadded. M8 DBOSrecv
timeout is explicit and in seconds;waitForSignalrequirestimeout. M9 Graphile release
cadence corrected. M10 capabilities are build-time static data. M11 job capability
names and outbox rule for jobs. M12 glosses added for CRIU, BSL, SSPL, BYOC, SKIP LOCKED,
CEL, RRULE, Conductor, Standard Schema, Build IDs; maintenance numbers linked;anon=1explained. - L1–L14 quote fixed (L1), Bun claim qualified (L2), Inngest #1624 reworded (L3), Trigger.dev
licence split resolved (L4), DBOS children closed (L5), Inngest GPL left UNVERIFIED with the
stale-metadata observation (L6), cron claim narrowed (L7), Inngestconnect()limits scoped to
Cloud (L8), Restate HTTP cost restated (L9), Temporalpatched()vs Build IDs (L10),
StepScope.contextadded (L11), persisted cancel flag (L12), persistence link and versions
corrected (L13), DBOS #1126 titled (L14).
11. Review corrections (round 2)
The round-2 review (notes/team-lead/reviews/08-durable-engines-review.md, “## Round 2”) confirmed
the 30 round-1 fixes and raised 10 findings on the revised design. All were applied. I re-fetched
the Restate clients package (no cancel) and the Inngest signals handler (404 “No signal found”);
both matched.
- N1 R3 key is per call (
effectKey(callIndex)), the effects table has a unique key inserted
first in the action’s transaction (so concurrent attempts cannot both apply), rows have a stated
retention rule, and reset/fork behaviour is stated (6.1.5, 6.2). - N2 Inngest
waitForSignalis a shim (404 on early signals, global names, duplicate failure);
relay retries on 404 without tripping the stale-row refusal (6.5). - N3 Restate
waitForSignalis a shim with a generated sharedsignalhandler per workflow (6.5). - N4 Temporal step families register the family name as the activity type, key in the input;
the verifier’s family-key check is stated precisely (6.1.2, 6.5). - N5 Child runs get the deterministic id
${parentRunId}:${invokeName}on every engine
(6.1.9, 6.5). - N6
StepScope.txdefined forstepInTransaction(6.2). N7 section 4.2 transactions row
refreshed. N8 garbled sentence in 2.1 fixed. N9 the two missing UNVERIFIED items added to
8.5. N10 Restate client cancel closed: admin API or CLI only (2.3, 6.5). - Added the reviewer’s note that DBOS worker Bun support must be tested before DBOS is picked as
the first external adapter (section 7, 8.5).
12. Review corrections (round 3)
The round-3 review (same file, “## Round 3”) confirmed the 10 round-2 fixes and raised 7 findings.
All were applied as written; the DBOS forkWorkflow signature (P7) is taken from the reviewer’s
fetch of the DBOS client reference and was not re-fetched by me.
- P1 (medium) effect keys use the instance name (
charge[3]);StepScope.stepNameadded;
Temporal rebuilds the instance name from the key in the input; Inngest, Restate and DBOS look
upstepFamiliesas well assteps(6.1.5, 6.2, 6.5). - P2 Inngest signal relay checks
GET /v1/runs/{runID}on a 404 and dead-letters the row when
the run is terminal; thewaitForEventalternative needs the same rule (6.5). - P3 Restate: the
runhandler stores its invocation id in workflow state first; an early
signal fails retryably (6.5). - P4 Inngest child: the adapter passes the Mesh child id in the invoke payload (6.5).
- P5
callIndexis a static per-call-site index from the generator (6.1.5). - P6 verifier uniqueness covers
invoke,sleepandwaitForSignalnames (6.3.1). - P7 DBOS
forkWorkflowstarts a new execution with a new id (effects re-apply from
startStep);rewindWorkflowkeeps the id; the UNVERIFIED item is closed (6.1.5, 8.5). - Nits (final acceptance): the Temporal
registerrow now says the proxied activity is the family name with the instance key in the input;effectKeytakesnumber | stringto match the${index}:${loopKey}form.