elsa-core/doc/wiki/workflow-runtime.md
2026-06-16 23:27:59 +02:00

14 KiB

Workflow Runtime

Workflow runtime owns starting, dispatching, resuming, canceling, logging, and recovering workflow executions. It is the layer that turns definitions into running instances and responds to triggers, bookmarks, background work, and admin operations.

Start in src/modules/Elsa.Workflows.Runtime.

Feature Wiring

WorkflowRuntimeFeature registers and configures:

  • IWorkflowRuntime
  • IWorkflowDispatcher
  • IStimulusDispatcher
  • IWorkflowCancellationDispatcher
  • runtime stores:
    • bookmark, bookmark queue, bookmark queue dead-letter, trigger, workflow execution log, and activity execution stores
  • workflow matcher, starter, invoker, resumer, canceler, restarter
  • trigger indexer and bookmark manager
  • background workflow, stimulus, task, and activity dispatch
  • bookmark queue worker and queue purger
  • distributed lock provider
  • execution cycle registry
  • graceful shutdown machinery
  • runtime startup and recurring tasks

It also configures WorkflowsFeature to use the runtime commit state handler.

Runtime Stores

Important runtime entities:

The default runtime feature uses memory stores. EF Core runtime persistence is wired by EFCoreWorkflowRuntimePersistenceFeature, which replaces runtime store factories on WorkflowRuntimeFeature.

Dispatch Paths

flowchart TB
    Start["Start workflow request"] --> Starter["DefaultWorkflowStarter"]
    Trigger["Trigger/stimulus"] --> Stimulus["StimulusSender / TriggerInvoker"]
    Bookmark["Bookmark resume"] --> Resumer["BookmarkResumer / WorkflowResumer"]
    Instance["Dispatch existing instance"] --> Dispatcher["WorkflowDispatcher"]
    Starter --> Invoker["WorkflowInvoker"]
    Stimulus --> Matcher["WorkflowMatcher"]
    Matcher --> Dispatcher
    Resumer --> Dispatcher
    Dispatcher --> Runtime["LocalWorkflowRuntime"]
    Runtime --> Runner["IWorkflowRunner"]

Key files:

Transactional Dispatch Outbox

Hosts can opt into at-least-once workflow dispatch for dispatch calls made from inside a running workflow:

services.Configure<WorkflowDispatcherOptions>(options =>
{
    options.UseTransactionalOutbox = true;
});

When enabled, TransactionalWorkflowDispatcher writes the command to IWorkflowDispatchOutboxStore before the parent workflow state commits, and stores the outbox item ID in the parent WorkflowState.Properties. WorkflowDispatchOutboxProcessor delivers only records whose owner workflow state contains that committed marker. This prevents a crash between workflow-state commit and mediator enqueue from silently losing the dispatch: the durable outbox record is already present, and the committed marker authorizes delivery after restart.

Operational notes:

  • The default outbox store uses IKeyValueStore; production hosts should pair this option with durable workflow instance persistence and durable key-value persistence.
  • Delivery is at-least-once. If the process crashes after sending a command but before deleting the outbox record, the processor may send it again.
  • Workflow definition dispatches generated by DispatchWorkflow/BulkDispatchWorkflows include a child workflow instance ID. DispatchWorkflowRequestHandler treats that ID as the idempotency key for outbox-routed commands and skips duplicate create-and-run attempts when the instance already exists.
  • Outbox processing is serialized with the configured distributed lock provider. Poison items are abandoned after WorkflowDispatcherOptions.MaxOutboxDeliveryAttempts, and missing-owner items are removed after WorkflowDispatcherOptions.OrphanedOutboxItemRetention.
  • Dispatch calls outside a workflow execution context continue to use the regular background dispatcher.

Triggers And Bookmarks

Triggers start workflows. Bookmarks resume suspended workflow instances. Runtime indexes and queries them through:

Bookmark queue processing is handled by:

Expired bookmark queue items are moved to the dead-letter store before they are removed from the active queue. Processing failures increment DeliveryAttempts; when BookmarkQueuePurgeOptions.MaxDeliveryAttempts is reached, the queue item is dead-lettered with the last exception type and message. BookmarkQueuePurgeOptions.Ttl controls active queue expiry, and BookmarkQueuePurgeOptions.DeadLetterTtl controls how long dead-letter records are retained before the purger deletes them.

Operators can inspect and manage dead-lettered bookmark queue items through the workflow API:

  • GET|POST /elsa/api/bookmark-queue/dead-letters: requires read:bookmark-queue:dead-letters.
  • GET /elsa/api/bookmark-queue/dead-letters/{id}: requires read:bookmark-queue:dead-letters.
  • POST /elsa/api/bookmark-queue/dead-letters/{id}/replay: requires replay:bookmark-queue:dead-letters; replay creates a new active queue item and marks the dead-letter item as no longer replayable.
  • DELETE /elsa/api/bookmark-queue/dead-letters/{id}: requires delete:bookmark-queue:dead-letters.

Read responses return a dead-letter view model for audit and replay status. Resume options are omitted from these responses because they can contain workflow input and property values.

Execution Logs

Workflow and activity execution logs flow through sinks and stores:

API endpoints under WorkflowInstances/Journal, ActivityExecutions, and ActivityExecutionSummaries expose this data.

Background Work

Runtime has several background paths:

  • BackgroundWorkflowDispatcher for workflow dispatch.
  • BackgroundStimulusDispatcher for stimulus dispatch.
  • BackgroundTaskDispatcher for RunTask.
  • LocalBackgroundActivityScheduler for background activity execution.
  • BackgroundActivityInvoker for executing background activity work.

These paths matter for tests: a workflow may return before background activity or bookmark work has completed.

Graceful Shutdown And Recovery

Recent graceful shutdown work added node-local quiescence and drain concepts. Source landmarks:

The design intent is captured in specs/002-graceful-shutdown/plan.md.

Ingress source adapters are currently registered by modules such as HTTP and Scheduling so the runtime can pause external event intake during drain.

Runtime Admin

The workflow API includes runtime admin endpoints:

  • GET /elsa/api/admin/workflow-runtime/status: requires read:workflow-runtime; ManageWorkflowRuntime is also accepted for backward compatibility.
  • POST /elsa/api/admin/workflow-runtime/pause: requires ManageWorkflowRuntime.
  • POST /elsa/api/admin/workflow-runtime/resume: requires ManageWorkflowRuntime.
  • POST /elsa/api/admin/workflow-runtime/force-drain: requires ManageWorkflowRuntime.

Endpoint code lives under Elsa.Workflows.Api/Endpoints/RuntimeAdmin. The service behind these endpoints is WorkflowRuntimeAdminService.

Distributed Runtime

Distributed runtime support lives in Elsa.Workflows.Runtime.Distributed. It layers distributed coordination and resilience support on top of the base runtime. When making runtime changes, check whether the distributed project has a parallel worker or dispatcher that must honor the same semantics.

Distributed Lock Provider Safety

The default workflow runtime lock provider is file-system based and writes under App_Data/locks. That provider is useful for single-host development and tests, but it is not safe for clustered deployments where nodes have separate file systems. When UseDistributedRuntime() is enabled, Elsa logs a startup warning if it detects the default file-system provider or the no-op provider unless the host explicitly acknowledges local-only lock semantics:

elsa.UseWorkflowRuntime(runtime =>
{
    runtime.UseDistributedRuntime();

    // Single-host development/test only. Suppresses the startup warning.
    // Do not use this for clustered production deployments.
    runtime.DistributedLockingOptions = options => options.AllowLocalLockProviderInDistributedRuntime = true;
});

Production clustered deployments must configure an IDistributedLockProvider backed by infrastructure shared by all nodes. Common Medallion providers include:

  • Redis: DistributedLock.Redis with Medallion.Threading.Redis.RedisDistributedSynchronizationProvider.
  • SQL Server: DistributedLock.SqlServer with Medallion.Threading.SqlServer.SqlDistributedSynchronizationProvider.
  • PostgreSQL: DistributedLock.Postgres with Medallion.Threading.Postgres.PostgresDistributedSynchronizationProvider.

Example SQL Server setup:

using Medallion.Threading.SqlServer;

elsa.UseWorkflowRuntime(runtime =>
{
    runtime.UseDistributedRuntime();
    runtime.DistributedLockProvider = _ =>
        new SqlDistributedSynchronizationProvider(configuration.GetConnectionString("SqlServer"));
});

Example PostgreSQL setup:

using Medallion.Threading.Postgres;

elsa.UseWorkflowRuntime(runtime =>
{
    runtime.UseDistributedRuntime();
    runtime.DistributedLockProvider = _ =>
        new PostgresDistributedSynchronizationProvider(configuration.GetConnectionString("PostgreSql"));
});

Example Redis setup:

using Medallion.Threading.Redis;
using Microsoft.Extensions.DependencyInjection;
using StackExchange.Redis;

builder.Services.AddSingleton<IConnectionMultiplexer>(_ =>
    ConnectionMultiplexer.Connect(configuration.GetConnectionString("Redis")));

elsa.UseWorkflowRuntime(runtime =>
{
    runtime.UseDistributedRuntime();
    runtime.DistributedLockProvider = sp =>
        new RedisDistributedSynchronizationProvider(sp.GetRequiredService<IConnectionMultiplexer>().GetDatabase());
});

When To Change This Layer

Change runtime for dispatch semantics, trigger/bookmark indexing, background work, execution logs, recovery, cancellation, graceful shutdown, or runtime stores. If a change only affects how definitions are saved or described, it belongs in management. If it only changes HTTP endpoint activity behavior, start in Elsa.Http.