* refactor(auth): remove the vestigial per-author script permission plumbing
#7975 is closed won't-do: authoring a workflow is a trusted act, and a
per-author gate would not change what a script can do once it runs. The host
switch stays the control, and it is per language, so an untrusted author gets
a host with the switch off rather than a permission.
That settles what the code was still half-carrying. WorkflowDefinitionScriptAuthorizationService
took a ClaimsPrincipal it never read, and could return a MissingPermission
reason nothing produced; two call sites branched on that reason to send a 403
that could not happen. The expression-descriptor endpoint kept a map from
expression type to per-author permission whose values went unused even before
the permissions were retired -- it only ever tested membership, and the
decision was always IsBrowsable. Each of these reads as an authorization gate
to anyone scanning the file, and none of them is one.
The principal, the unreachable reason, and both dead branches are gone. The
map becomes a set of the expression types the host can switch off, which is
what it was actually being used as. Behaviour is unchanged: the only failure
is a language the host disabled, which is a property of the deployment and
so a 400 naming the switch, never a 403.
PermissionNames loses ExecuteCSharpExpressions and ExecutePythonExpressions,
which existed only for that map and the test mirroring it. Five other legacy
constants there are also unreferenced but belong to other modules; they are
left alone rather than swept up here.
Two tests asserting the host-and-user case were exact duplicates of the
host-only case once the principal stopped mattering, so they go with it.
The migration guide said deployments lose per-author granularity "until
#7975 lands" and advised disabling host code until then. That promise is
withdrawn and replaced with the actual guidance.
Closes#7975
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(wiki): drop the retired exec:* permissions from the scripting guide
Review found doc/wiki/expressions-and-scripting.md still telling operators
that API callers "must have the exec:csharp-expressions permission" to
author, publish, dispatch or execute workflows containing C#, and the same
for Python. Those permissions no longer exist, so the instruction cannot be
followed and describes a gate that is not there.
Both sections now say what is actually true: the host switch is the whole
control, there is no per-caller permission because a workflow runs under the
server's authority rather than the caller's, and an untrusted author gets a
host with the switch off. The switches are noted as independent, since
enabling Python while leaving C# off is a real posture.
My earlier sweep searched for the issue number rather than the permission
strings, which is why this file was missed.
Refs #7975
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Adds durable, identity-neutral, workflow-bound human tasks, and reconciles the
REST surface with the approved Studio contract.
- Flat summary/detail DTOs, a global capability descriptor, and workflow context
captured at activation.
- Scope is part of the list authorization predicate; manager decisions require
manage:user-tasks; a denied command answers 404 so it cannot prove a task exists.
- Guest sessions are task-scoped, action-allowlisted, and revoked when the task
closes. Invitations resolve by token hash through the repository, wait in a
Data Protection encrypted outbox, and are rate limited per caller.
- Masked form values are disclosed only through an audited reveal command.
- Store-specific concurrency failures are translated into a single
UserTaskRevisionConflictException, so a concurrent edit returns the documented
revision-conflict result behind any provider instead of a 500.
EF Core (SQLite, SQL Server, PostgreSQL, MySQL, Oracle) and VNext persistence,
hosted due/reconciliation/delivery workers, docs, and 49 tests.
Note: this branch also carries two commits inherited from its branch point that
are not part of User Tasks and are squashed in here — the revert-version
allocation change from #7917 (WorkflowDefinitionPublisher.RevertVersionAsync now
allocates from the last version rather than the latest) and an NU1903 package
pin. Merged deliberately rather than rebased out.
* feat(auth): add the permission model and evaluator (Phase 1)
Additive only. Nothing changes behavior: no endpoint declares against this
yet, and no existing enforcement path routes through it.
A permission is {resource}:{verb}, both axes open and string-keyed. A
trailing wildcard on the resource axis matches the named node and every
descendant at any depth, so workflows/definitions/* covers
workflows/definitions itself; * on the verb axis matches any verb.
Wildcards are the only construct with forward reach.
A bare * parses to *:* at parse time rather than being special-cased in
the evaluator, so superuser stays an ordinary grant and a stored or seeded
* keeps authorizing across the vocabulary migration without a lock-out
window.
Adds:
- Permission, with parsing that rejects a value containing a comma, since
the persistence converter joins collections with one
- CoreVerbs, the recommended set modules should reuse; a convention rather
than a closed vocabulary
- PermissionMatcher, one matching rule shape on both axes
- IPermissionEvaluator, the single place permission decisions are made,
skipping malformed claims so one bad stored grant cannot deny a principal
- PermissionRequirement and PermissionAuthorizationHandler
- The descriptor catalog in core: PermissionDescriptor now carries the
verbs a resource supports and marks non-core ones, and the registry can
report what a wildcard covers today
External Authentication keeps its own descriptor types for now; it moves to
the core catalog with the other modules in Phase 2, which keeps this change
purely additive.
55 unit tests cover the matcher table, wildcard forward reach, the
counterpart that concrete grants stay frozen, absence-is-denial, and the
seeded * case.
Refs #7974
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(auth): contribute the permission catalog from every module (Phase 2)
Still additive. Existing endpoints keep their legacy declarations; nothing
changes behavior for them.
Every module exposing protected endpoints now declares its resources and
the verbs each accepts, following the pattern already proven in External
Authentication -- constants and descriptors colocated -- refined to one
constant per resource, with the verb supplied separately. 47 resources
across 15 modules, matching the settled vocabulary.
Descriptors are discovered from the same assemblies as a module's
endpoints, in AddFastEndpointsFromModule. Registering them per module
would let the catalog and the endpoints drift, which is the failure this
model exists to remove; tying them to one registration makes the catalog
necessarily describe the endpoints that exist.
Adds:
- GET /identity/permissions, the catalog a role editor renders from, so
no client hard-codes permission strings
- GET /identity/permissions/reach, reporting what a wildcard covers today.
This is the mitigation for forward reach on the resource axis: a
wildcard is useful precisely because it covers things that do not exist
yet, so an author needs to see what it reaches now
- GET /identity/me/permissions, resolving wildcards to concrete verbs so a
client needs no matching logic, and listing denied resources with an
empty verb list so "denied" is distinguishable from "unknown"
- IPermissionGrantValidator, wired into Roles/Create and Roles/Update,
which previously persisted request.Permissions after only the
caller-subset check. Concrete segments validate against the catalog;
wildcards validate structurally and are accepted even when they match
nothing today, since installing a module later is what gives such a
grant meaning
- RequirePermission(resource, verb) and RequireAuthenticatedOnly() on the
endpoint base classes, with the six copy-pasted ConfigurePermissions
bodies collapsed into one implementation
New endpoints require new-format grants, so during the transition they
authorize only for holders of *, which parses to *:*. Phase 3 migrates the
rest and closes that gap.
70 unit tests, including the wildcard-accepting validator cases.
Refs #7974
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(auth)!: cut every endpoint over to the permission model (Phase 3)
BREAKING: legacy permission strings no longer authorize. A permanent alias
layer would keep two vocabularies valid forever, so the break is
deliberate and reported rather than absorbed. `*` survives unchanged --
it parses to `*:*` -- so an administrator cannot be locked out while
roles are re-authored.
All 168 declaration call sites across 151 files now use
RequirePermission(resource, verb) with the constants their module
declares, so a typo is a compile error rather than an unreachable
endpoint.
Enforcement consolidated onto IPermissionEvaluator:
- RoleAuthorizationService evaluates containment through the evaluator
rather than by set membership. This matters: a caller holding
workflows/*:view can now delegate workflows/definitions:view, which set
membership got wrong and which would otherwise force administrators to
hold every concrete grant they wish to delegate.
- The two Broker/Logout.cs endpoints declare explicitly. Logout is
authenticated-only; ContinueLogout is anonymous, matching every other
broker callback -- the route handle carries the authority and a
top-level browser navigation sends no Authorization header.
Removes the C#/Python expression permissions (#7975). They conflated an
incoherent execution-side gate -- a workflow runs under the server's
authority, not the caller's, so the check never constrained what a script
could do -- with a meaningful authoring-side one. The host switch
(AllowHostCodeExecution) becomes the single control. This is a deliberate
reduction in control: where host code is enabled, any author who may write
definitions may use C# and Python.
Adds the fail-closed gate. Omitting a declaration previously inherited the
FastEndpoints default with no Elsa-level fallback, so an endpoint could
ship ungated unnoticed. EndpointCoverage asserts every endpoint declares
exactly one of RequirePermission, RequireAuthenticatedOnly or
AllowAnonymous, with no exemption list. Its canary assertion earned its
keep immediately by catching that the gate was scanning an assembly
containing no endpoints.
EndpointPermissionRegistry records what each endpoint declares. The
requirement is attached as an inline policy and is not readable back from
the definition, so this keeps the declaration introspectable -- and lets
tests assert a specific requirement rather than merely that one exists.
Two behavior notes worth calling out:
- The runtime status endpoint previously accepted either the read or the
manage permission. It now requires workflows/runtime:view alone, which
is least privilege; a role holding only control must also be granted
view to read status.
- BPMN interchange repeats the workflow-definitions path locally rather
than taking a dependency on Elsa.Workflows.Api for one constant. It
contributes no descriptor: the resource is owned and described by
Workflows.Api, and the registry keeps one entry per resource.
Also adds a startup validator that logs every stored role permission that
no longer resolves, identified by role, so an upgrade is loud.
188 unit tests pass across Api.Common, Workflows.Api and Identity.
Refs #7974, #7975, #7976
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(auth): revocation bound and role audit notifications (Phase 4)
Default access-token lifetime drops from 1 hour to 15 minutes. This is
the revocation bound: permission claims are issued at sign-in and refresh
re-reads the user's roles, so removing a role takes effect at most one
access-token lifetime later. Refresh already rotates both tokens, so no
client change is required and the refresh lifetime is unchanged.
Adds an optional permission stamp for deployments needing a tighter
bound. The stamp is derived from the user's roles and their permissions
rather than stored as a counter on the user. That avoids changing the
Identity schema, which would have required migrations across all five EF
providers and made this milestone depend on the tenancy work. It also
means every node computes the same value from the same store with no
cross-node cache invalidation, which matters because Elsa has none.
The stamp is issued unconditionally and only validated when enabled, so
turning it on does not invalidate tokens already in flight; an absent
stamp is not treated as a mismatch for the same reason. It changes when a
role is added to or removed from the user and when a held role's
permissions change, but not when an unrelated role changes.
Role create and update now publish typed security notifications per ADR
0007, carrying the resulting grants so a reviewer can reconstruct what a
role conferred at a point in time without replaying every prior event.
This module owns no audit store: a future audit module subscribes and
sets its own retention.
Refs #7974
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(auth): tenancy hardening for identity (Phase 5)
Closes the gaps that made "roles are configurable per tenant" untrue
however the rest of the stack behaved.
Uniqueness becomes per tenant. User.Name, Role.Name, Application.Name and
Application.ClientId carried globally unique indexes, so two tenants could
not both hold a role named Admin. Migrations for all five EF providers
drop the global indexes and create composite ones on (TenantId, Name).
The in-memory user and role stores now scope to the ambient tenant.
Isolation previously existed only on the Entity Framework path, and only
when multitenancy was enabled, so a deployment running the default stores
had none at all. The tenant-agnostic sentinel is honored, matching the EF
query filter, so a shared platform role stays visible from every tenant.
RoleFilter gains TenantId, matching UserFilter, and the role and user list
endpoints pass it explicitly rather than relying on an ambient filter that
only exists on one persistence path.
UserManager.CreateUserAsync sets TenantId explicitly instead of relying on
the EF saving handler, which does not run in memory and left users
unassigned there.
Refs #7974
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs: migration guide, ADR, and security wiki for the authorization model
Adds docs/migrations/authorization-model.md, following the shape of the
external-authentication persistence guide. It leads with the three things
that are not a simple rename, because each silently produces a wrong
result if treated as one:
- The migration expands where new sub-resources are finer-grained than
what they replace, so a one-for-one substitution narrows roles.
- read:* and exec:* become materially more powerful. They are literal
claim values today, authorizing twelve of roughly forty read endpoints;
their replacements work as the names always implied. Any role holding
them needs review by hand, not an automated rewrite.
- The C#/Python expression permissions are removed rather than
translated, which is a deliberate reduction in control where host code
is enabled.
It also states plainly that `*` keeps working, and says to do that first,
since it is what stops an instance locking itself out mid-migration.
ADR 0012 records the model and, more usefully, why a closed verb
enumeration was drafted and rejected: it was justified on implication, but
aggregates were already excluded and no verb implies another, so the
bitwise check was expressing set containment all along.
The security wiki's API Authorization section replaces its Secrets-only
route table with the catalog endpoint as the authoritative source, and
states why read-only mode is a separate axis rather than a permission.
Refs #7974
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(auth): restore the suites after the tenancy and evaluator changes
The whole solution builds with zero errors and every affected suite
passes: 70 Api.Common, 23 Workflows.Api, 95 Identity, 154 External
Authentication unit, 133 External Authentication integration.
Most breakage was test call sites constructing the tenant-aware stores
and the evaluator-backed RoleAuthorizationService directly. Adds
TestTenantAccessor to Elsa.Testing.Shared rather than giving the
production constructors an optional accessor, which would have let a
missing registration silently disable isolation.
Several External Authentication tests created fixtures in tenant-a while
running under the default tenant, so the newly isolating store correctly
stopped finding them. They are now scoped to the tenant their own
fixtures use; JustInTimeProvisioningTests, which genuinely spans two
tenants, is scoped per case.
One production fix came out of it: IdentityFeature now ensures an
ITenantAccessor with TryAdd. The identity stores are tenant-scoped, so a
host that never enables multitenancy would otherwise fail to construct
them -- which is what the DI registration tests were reporting. TryAdd
leaves MultitenancyFeature's own registration untouched.
Refs #7974
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(auth): register the identity services on the classic feature path
Found by running Elsa.Server.Web, not by the test suites: the app failed
at startup with "Unable to resolve service for type RoleSecurityNotifier
while attempting to activate Roles.Update".
RoleSecurityNotifier, the permission stamp services, the memory cache and
the stored-permission validator were registered only in the CShells shell
feature. Elsa.Server.Web uses the classic UseIdentity() path, whose
IdentityFeature registered none of them, so every host on that path
crashed while mapping endpoints. Unit tests did not catch it because they
construct services directly rather than through either feature.
Verified end to end against the running server:
- The seeded admin role stores "*". It parsed to *:* and resolved to
concrete verbs across all 27 registered resources, which is the
bare-wildcard parse rule working on real data rather than in a test.
- GET /identity/permissions returns the catalog for the modules this app
installs -- 27 resources, 0 unverified, categories Dashboard, Identity,
Resilience and Workflows -- rather than all 47, which is correct: the
catalog describes what is installed.
- GET /identity/permissions/reach?resource=workflows/* reports 19 covered
resources.
- A role holding only dashboard:view gets 200 on /dashboard/overview and
403 on /identity/roles, /identity/users, /workflow-definitions and
/identity/permissions, while /identity/me/permissions returns 200
because it declares RequireAuthenticatedOnly -- confirming FR-019's
third declaration state behaves as designed.
- That same principal's /me/permissions lists all 27 resources with 26
carrying an empty verb list, so "denied" stays distinguishable from
"unknown to this server".
- The startup validator logged no unresolvable permissions, as expected
for a seed holding only "*".
Refs #7974
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(auth): discover permission descriptors on the shell host path
Found by running Elsa.ModularServer.Web. The shell host started cleanly
and authorized correctly, but GET /identity/permissions returned zero
resources and /identity/me/permissions returned no grants.
Descriptor discovery was wired into AddFastEndpointsFromModule, which only
the classic module path calls. CShells discovers endpoints from features
implementing its own marker interface, so on a shell host no provider was
ever registered. Authorization still worked, because the evaluator reads
claims and needs no descriptors -- which is exactly why nothing failed
loudly. What silently broke was everything built on the catalog: role
authoring would have rejected every concrete grant as an unknown
resource, introspection returned nothing for clients to render, and the
stored-permission validator would have reported every concrete stored
permission as unresolvable.
ElsaFastEndpointsFeature now contributes descriptors from the loaded Elsa
assemblies, bounded to those and run once per shell.
Verified on the modular host, which installs far more modules than
Elsa.Server.Web:
- 47 resources registered, 0 unverified, across all 12 categories, with
all 17 module-specific verbs present. That is the entire published
vocabulary confirmed against a running server rather than a document.
- Reach reports workflows/* covering 20, external-authentication/*
covering 8, and * covering 47.
- Creating a role with dashboard:view and workflows/*:view succeeds,
confirming a wildcard grant survives authoring validation.
- Creating one with invented/resource:view and secrets:publish is
rejected with 400.
Also makes those rejections actionable. The permission was reported
without the reason, so an operator learned which entry was wrong but not
why; both parts are now in the message, including the supported verbs for
the resource.
Refs #7974
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* wip: bpmn test vocabulary
* fix(auth): make enforcement DI-independent and finish the hub cutover
CI on #7980 was red. Running the full suite locally rather than the
subset I had been checking surfaced 24 failures across four projects,
in three distinct classes.
Enforcement no longer depends on a DI registration. RequirePermission
attached a PermissionRequirement evaluated by a registered handler, so a
host that had not called AddElsaAuthorization got 403 on every endpoint
with nothing to indicate why. Several test hosts wire FastEndpoints
directly and did exactly that. The requirement is now evaluated inline
against a shared stateless evaluator, with a host-registered
IPermissionEvaluator still taking precedence. Registration remains
worthwhile for the catalog and the validator; authorization can no longer
silently fail closed because of a missing one.
Registration also moved from AddFastEndpointsFromModule to
AddFastEndpointsAssembly. Registering an endpoint assembly is what should
guarantee its permissions work, and a host may never call the former.
Finishes T039. The four SignalR hubs still matched hard-coded legacy
permission strings, which no longer exist, so every hub denied access.
They now route through the evaluator like every other enforcement path.
Test fixtures granting legacy strings were updated to the new vocabulary.
Two categories were deliberately left alone: naming tests asserting the
legacy constants still hold their old values, which is true and worth
keeping, and the workflow script authorization tests, which asserted a
MissingPermission outcome that D21 removed -- those now assert the host
switch is the only control.
One test previously pinned that the hub honors a FastEndpoints-configured
permissions claim type. It now asserts the opposite, and says why: Elsa is
the only authority that expands roles into permission claims (ADR 0009),
and this model no longer uses the FastEndpoints permission mechanism, so
its separately configurable claim type is not consulted. That property is
also unreadable outside reflection.
Whole solution builds with 0 errors and every test project passes.
Refs #7974
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(auth): scope the permission-stamp cache to the tenant
Greptile found and reproduced a cross-tenant authorization bug, and it was
mine: Phase 5 made user names unique per tenant rather than globally, but
PermissionStampValidator kept caching by user name alone. Tenant A's
lookup could therefore populate the cache with its own stamp and satisfy a
revoked token belonging to a same-named user in tenant B, without ever
resolving tenant B's user.
Both the cache key and the user lookup are now tenant-scoped. Added
PermissionStampValidatorTests, including the cross-tenant case; verified it
fails without the fix and passes with it.
Also from review:
- Removed the legacy permission constants left unused in the three hubs
after they moved to the evaluator, so no stale vocabulary lingers.
- Narrowed two generic catch clauses. The IL scanner now catches only the
exceptions an unresolvable metadata token actually throws, and the
startup validator rethrows cancellation while still refusing to stop the
host for anything else -- an unreachable or half-migrated store is
exactly when an operator most needs the host up.
Whole solution builds with 0 errors and every test project passes.
Refs #7974
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(docs): correct path in log message for authorization model migration link
Aligns the log message path to the correct documentation directory, changing `docs` to `doc` to avoid confusion and incorrect linking during log output.
* fix(auth): update permissions method to use new syntax
* docs: consolidate docs/ into doc/
The repository had two documentation roots. Merge docs/ into doc/ and
remove the empty docs/ tree.
The two adr/ folders both numbered from 0001, so the identity and
authorization series is renumbered to continue the core series rather
than collide with it:
docs/adr/0001-0012 -> doc/adr/0014-0025
Every reference is updated to match: the Status cross-links between the
renumbered ADRs, the ADR and path links in specs/012-external-authentication
and specs/013-rbac-authorization-model, and doc/wiki/identity-tenancy-security.md.
doc/adr/toc.md gains entries 14-25. doc/adr/graph.dot is regenerated out
to 25; it had been stale since ADR 10 and now also carries the partial
supersession edges declared by the ADRs themselves.
docs/codebase/ and docs/migrations/ move across unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(auth): simplify syntax in PermissionEvaluator and related classes
Streamlined syntax for method definitions by using expression-bodied members and simplified object instantiations across the Authorization module. This includes adjustments in `PermissionEvaluator`, `LocalHostRequirement`, and `WebApplicationExtensions` for better readability and maintainability.
* ci(bounty): point the footer step at the file's real path
The bounty workflow read docs/bounty-footer.md, the path the file had
when the workflow was added in b421b00e1. The file later moved to
doc/bounty/bounty-footer.md and the workflow was never updated, so the
read step has been resolving nothing and the appended comment was empty.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix: stop two silent serialization and test-isolation traps
Two follow-ups from #7957.
ExternalAuthentication tests: the same process-global
EndpointSecurityOptions.SecurityIsEnabled race the shells API tests had,
across the six classes in that assembly that build an endpoint host —
five setting it to false and IdentityLinkAuthorizationTests to true.
Unlike the shells case these all call UseAuthorization(), so it does not
surface as a missing-middleware error: anonymous endpoints answer
401/403, and the authorization test's endpoints come back AllowAnonymous
and stop enforcing what it asserts. A module initializer cannot fix it
since the assembly genuinely needs both values, so the six now share one
collection with DisableParallelization. They are also the only six that
build a host, so nothing else can observe a leaked value.
Unaliased payloads: a payload whose type has no registered serialization
alias is written without a _type discriminator and read back as an
ExpandoObject whose keys carry the state serializer's camel-case naming
policy, so a consumer that published Status finds status. The
degradation is deliberate — the alias registry is an allow-list that
keeps arbitrary CLR type names out of deserialization — but it was
silent. It is now reported once per type, naming the type and both
lossless alternatives, and PublishEvent.Payload documents them. Measured
across the integration suite, only genuine user payload types reach this
path, so the warning does not fire for Elsa's own types.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix: check the log level before claiming the once-per-type warning slot
WarnAboutUnaliasedType claimed a type's single report via TryAdd before
LogWarning applied its level filter, so a type first serialized while
Warning was disabled spent its slot on a call that logged nothing and
then stayed silent forever, including after the level was raised at
runtime. Check IsEnabled first, so the slot is only consumed by a report
that is actually emitted.
The regression test needs the capture to be the only logging provider:
IsEnabled on the composite logger is an OR across providers, so the test
builder's own xunit provider would otherwise keep Warning enabled
regardless of what the test asked for.
Reported by Greptile on #7969.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
A container that schedules a child and then decides the child must not run
had no way to withdraw it. `IActivityScheduler` exposed no removal operation,
and `CancelActivityAsync` no-opped on a context whose status was `Pending`,
so a container could tear a branch down and still have an activity from that
branch execute afterwards, side effects and all. Fixes#7943.
- `IActivityScheduler.RemoveWhere` removes work items and keeps the order the
survivors would have been taken in; implemented in both the FIFO and LIFO
schedulers.
- `CancelActivityAsync` (both the public extension and the internal one used
when a container completes) cancels `Pending` contexts as well as running
ones, and withdraws the work item that would have started the cancelled
activity plus the items it had scheduled for children with no context yet.
Withdrawal is a real removal rather than a terminal status honoured at dequeue
time, because the scheduler is also read: `Flowchart.HasPendingWork` inspects
it to decide whether it may complete, and the work item list is extracted into
the persisted workflow state — a withdrawn-but-queued item would be persisted
and rehydrated with a fresh context after a suspend/resume.
`StateMachine` had hand-rolled the same operation to drop competing triggers by
clearing the scheduler and re-scheduling everything else; it now calls
`RemoveWhere`. `Elsa.Bpmn` no longer needs to refuse a teardown whose subtree
still has queued work, so `BpmnWorkTeardown` drops the `NotSupportedException`
and records the teardown reason on the torn-down activity's journal instead.
BREAKING: `IActivityScheduler` gains a member; external implementations must
add `RemoveWhere`.
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(runtime): let a trigger index payloads under per-payload stimulus names
TriggerIndexingContext.TriggerName is a single field read once after all
payloads have been collected, so one ITrigger could only ever register its
payloads under one stimulus name. An implementation that assigned the name
more than once - which the stimulus extension methods do as a side effect -
had the last write applied to every row, and since Hash derives from the same
name, the earlier payloads were stored under a hash no publisher computes.
Adds an additive, opt-in path: a payload returned from GetTriggerPayloadsAsync
may be wrapped in NamedTriggerPayload, which carries the stimulus name for that
payload alone. The indexer takes name and payload from the same source, so
Hash always matches the Name stored beside it, and the wrapper is unwrapped
before storage so payload consumers (validators, the trigger diff comparer,
the scheduler) see the payload the trigger produced.
TriggerName keeps its existing meaning as the default for payloads that do not
carry their own, so every existing ITrigger indexes identically: same Name,
same Hash, same Payload, same row count. The empty-payload placeholder row is
left alone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(runtime): refuse a nested trigger payload wrapper
NamedTriggerPayload documented that its Payload is never itself a
wrapper, but nothing enforced it. Reject a NamedTriggerPayload whose
payload is another NamedTriggerPayload at construction time, matching
the existing guard against a blank name.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* chore(javascript): update Jint to 4.15.3 and stop blocking on promises
`Engine.Evaluate(...).UnwrapIfPromise()` blocks the calling thread while the
engine's event loop drains, which is exactly the wrong thing to do inside an
`async` method — an expression that awaits a .NET `Task`, such as one calling
`getSecret()`, held a thread pool thread for the duration of the I/O.
`Engine.EvaluateAsync` awaits the returned promise instead, and takes the
cancellation token while it is at it.
The Jint version is moved from 4.4.2 to 4.15.3. `EvaluateAsync` arrived in
4.14.0, but the pin lands past 4.15.2 deliberately: once expressions genuinely
suspend and resume instead of draining the event loop on the calling thread,
they exercise the async suspension machinery 4.15.2 corrected — an `await` on a
right-hand side no longer stores the suspension sentinel, async generators and
`for await...of` preserve loop iteration state across a suspension, and a
suspension node is unwrapped correctly. Shipping the non-blocking change on an
earlier 4.14/4.15 would enable exactly the code paths those releases fixed.
One default changed along the way: since 4.14 `Interop.ArrayConversion` defaults
to `LiveView`, so a CLR array reaches script as a live view over the original
array rather than as a copy. That is observable — a script that sorts an array
would now reorder the workflow's own array, and the value round-trips back as
its original element type rather than as `object[]`. The evaluator therefore
pins the previous `Copy` behaviour so the upgrade is not a behavioural change;
hosts that prefer the live view can opt in through
`JintOptions.ConfigureEngineOptions`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik
* test(javascript): pin the array copy lane and adopt JsString.Create
The array audit. The parent commit fixes `ArrayConversion` to `Copy`, because
4.14 changed the default to `LiveView` and the two differ in behaviour a script
can observe. The existing tests assert what a copy *produces*, which a live view
also satisfies for a value nothing mutates; asking the engine how many
conversions of each kind it performed (4.15.1's interop conversion counters)
pins the lane itself. The second assertion is the more interesting one: an
ordinary evaluation converts no CLR array at all, because Elsa converts
collection-valued variables itself in `ObjectConverterHelper` long before Jint's
array lane could see them. That makes the `ArrayConversion` setting a narrow
compatibility pin rather than something every evaluation depends on.
`JsString.Create` (public since Jint 4.15.3) is adopted in
`JsonElementConverter`, where the string case was the only one still routed
through `JsValue.FromObject` — re-entering the whole conversion pipeline, the
registered object converters and this one included, to arrive at the same call
the number and boolean cases beside it already make directly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f
* perf(javascript): register .NET type globals lazily
A fresh Jint engine is built for every expression evaluation, and roughly twenty
.NET types are registered on it before the expression runs: the common types
(`DateTime`, `Guid`, `TimeSpan`, …) plus every non-primitive workflow variable
descriptor type. Each registration builds a `TypeReference`, which describes the
type through reflection. The overwhelming majority of expressions reference none
of them.
`Options.AddLazyGlobal` installs a global whose value is produced by a factory
the first time a script reads the name. Both type registration handlers now run
on `CreatingJavaScriptEngine` — which already carries the `Jint.Options` being
built — and register through it, so a type is only described if an expression
actually mentions it. A script that uses `Guid.NewGuid()` still sees `Guid`; a
script that uses none of them pays for none of them.
Types whose name cannot be written as a JavaScript identifier are skipped while
we are here, since no script can reach them. That covers constructed generic
types (``IDictionary`2``, which two different variable descriptors both claimed)
and array types (`Byte[]`).
Moving the registrations to engine construction has a consequence worth pinning
beyond the ordering: a global installed by the host through `configureEngine` is
no longer overwritten by the built-in registration of the same name. The lazy
global for `Guid` is already installed when that callback runs, and
`Engine.SetValue` goes through `[[Set]]` on the global object, which reads the
current value — running the factory once and discarding the result — before
replacing the descriptor. The end state is the host value, but only the test
says so.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f
* perf(javascript): register the common functions lazily
The same argument as the type globals, applied to the other bulk registration a
fresh engine pays for. `ConfigureEngineWithCommonFunctions` installs
twenty-seven functions on every engine — `getVariable`, `toJson`, `newGuid`, the
base64 helpers, the deprecated GUID pair — and each `engine.SetValue(name,
Delegate)` builds an interop function wrapper around the delegate: a JavaScript
function object, plus the signature metadata and invoker lookups Jint resolves
per delegate type. An expression such as `variables.Foo` or a comparison calls
none of them.
This was recorded in the PR as deferred, because a lazy version had to keep the
`NonEnumerable` flag `SetValue(string, Delegate)` applies — otherwise the
functions would start appearing in `Object.keys(globalThis)`, which is
observable. Jint 4.15.3's `Engine.Advanced.AddLazyGlobal` takes a `PropertyFlag`
and is the post-construction counterpart of the options-time API used for the
type globals, so the flag is passed explicitly and the port is a line per
function. The CLR delegate is now created inside the factory as well, so a
function nothing reads costs one closure rather than a closure, a delegate and a
wrapper.
The laziness is invisible, and the tests say so rather than leaving it implied.
`AddLazyGlobal` installs the property itself eagerly and defers only its value,
so `in`, `hasOwnProperty` and `Object.getOwnPropertyNames` answer immediately
without materialising anything; `Object.keys(globalThis)` still omits them; the
descriptor still reports writable, non-enumerable and configurable; two reads of
a function are the same value, which is what says the factory ran once and the
result was stored rather than recomputed per read; and a script can still
overwrite one.
Not separately measured. The measured table in the PR predates this commit and
was not re-run for it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f
* refactor(javascript): let Jint expose enums as names and type its converters
`EnumToStringConverter` turned every CLR enum crossing into JavaScript into
`Enum.ToString()`. Jint 4.15 does that with `Interop.EnumConversion`, so the
converter goes away.
It is not only fewer moving parts. The converter could only see values that
crossed the interop boundary, and a constant read off a registered enum type
does not: `LogPersistenceMode.Include` came back as the underlying number while
the same value held in a workflow variable came back as `"Include"`, so
`mode === LogPersistenceMode.Include` was always false. The built-in switch
covers both directions and they now agree. Values written back to the CLR keep
accepting the member name and the number, as they did before.
That fix changes what an expression using a constant *numerically* produces, and
the hazard worth calling out is persisted state: a workflow variable holding a
number written from a constant before the upgrade no longer compares equal to
that constant after it, which reaches in-flight and resumed instances rather
than only new ones. Typed conversions hold in both directions, so activity
inputs and typed variable reads are unaffected. Recorded in the 3.8.0 changelog.
The two remaining converters register through the overload that declares the CLR
types they handle. A converter that does not declare them has to be offered
every value crossing the boundary, which costs the engine its compiled
member-read and method-invoker lanes for every wrapped .NET object; declaring
`byte[]` and `JsonElement` keeps those lanes for everything that cannot produce
one.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f
* perf(javascript): register only the accessors an expression names
Every evaluation registers a getter and a setter for each variable in scope, a
getter for each workflow input, and a getter for each (activity, output) pair in
the enclosing container. The last of those is the expensive one: naming those
accessors walks every node of the workflow and resolves each against the
activity registry, so the cost of setting up an engine grows with the size of
the workflow rather than the size of the expression. Almost no expression uses
any of them.
Jint reports the free identifiers of a prepared program through
`Prepared<T>.ReferencedGlobals`, which is exactly the question being asked here:
an identifier the expression never mentions cannot be read from it. The
evaluator now prepares the script before configuring the engine — the parse was
already cached, so this only reorders work — and passes the set on
`EvaluatingJavaScript`. The accessor handler registers a name only if it is in
the set, and skips the workflow walk entirely when no identifier of the
`get{Output}From{Activity}` shape appears.
The filter is sound for an expression that names an accessor the way one is
meant to be named, and not for one that builds or reaches a name at run time.
Four such forms are detectable and each turns the filter off. A direct `eval`
call is reported as `HasDirectEvalCall`. An *indirect* `eval` call and the
`Function` constructor are deliberately not flagged as direct calls — Jint
reports the identifiers `eval` and `Function` in the set instead, and says so,
because that is the signal a host is meant to act on. Missing that signal is not
theoretical: `new Function('return getMyVariable()')()` would regress from
working to a `ReferenceError`, since Function-constructed code resolves only
against the global scope, and `var e = eval; e('getMyVariable()')` would fail the
same way. The fourth is a reference to `globalThis`, which reaches a global
without naming it. All four are pinned.
What stays undetectable is reaching the global object without naming it at all —
a sloppy-mode top-level `this`, or `[].constructor.constructor(…)`. An
expression written that way loses the generated accessor but not the data:
`getVariable(name)`, `getInput(name)` and `getOutputFrom(activityId, outputName)`
are always registered and reach the same values.
Note this does not replace the regular expression in
`ConfigureEngineWithVariables`, which extracts the member names in
`variables.Foo`. Those are not free identifiers and Jint deliberately does not
report them; the set only says whether `variables` itself is referenced.
The filtering is invisible from inside a script, which also makes it unprovable
from there, so two of the tests hold on to the engine and assert on the globals
directly.
While here, the cancellation token is registered as an engine constraint.
Passing it to `EvaluateAsync` only covers the awaiting part; Jint's own remarks
say the parameter cannot preempt the synchronous evaluation loop and point at
this constraint, so the token read as more coverage than it delivered. A
cancellation constraint is amortizable, so the interpreter keeps its tight-loop
fast path, and a default token registers nothing. #7891 supersedes this with the
full constraint set.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f
* perf(javascript): build marshalled variables in Jint's shaped representation
`ConvertToJsObject` built the JavaScript object a workflow variable is copied
into one `DefineOwnProperty` call at a time, with an explicit descriptor. That
lands the object in Jint's per-object property dictionary, so every marshalled
variable carries its own descriptors even though sibling variables and the
repeated nested payloads of one document present the same key set.
`JsObject.CreateFromEntries` defines the same writable, enumerable and
configurable properties but builds through the shared-layout path. Both wins
land inside a single evaluation — half the descriptor allocations, since the
explicit descriptor caused a second one inside `ValidateAndApplyPropertyDescriptor`,
and one shared layout across objects presenting the same keys. Nothing carries
across evaluations: layouts are interned per engine and per-node inline caches
live in per-engine handler trees that only engage on a second evaluation on the
same engine, which a fresh-engine-per-evaluation host never reaches.
Reaching the shared layout is silent: `CreateFromEntries` falls back to the
ordinary property dictionary whenever a key or a growth guard says the layout
cannot continue, and the object behaves identically either way. A test now asks
the engine whether the object actually got one. `Engine.Advanced.HasSharedShape`,
added in 4.15.3, is the part of that answer Jint documents as a contract — the
finer-grained `GetObjectRepresentation` names an internal representation that may
be renamed or subdivided in any release — and `CreateFromEntries` is one of its
three documented success cases. The assertion is not vacuous: against the
previous property-by-property build it is false, including for the
`CreateDataProperty` variant in #7892, because that object is created through
`Intrinsics.Object.Construct` rather than built as entries.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f
* test(javascript): run the scripting suites with Jint's host-contract verifiers on
Jint has a set of verifiers for the extension points it *trusts* — the ones it
cannot afford to re-check on a hot path, so a violation is otherwise silent.
They used to be compiled out of Release, which made "run your suite against a
Debug Jint" the only way to reach them, and Jint ships Release-only. 4.15.3
moves them behind an AppContext switch read once at type initialization, so a
host that never sets it pays nothing (the JIT folds the guards away) and a test
host can turn them on against the shipped package.
The one Elsa is subject to today is the object-converter type declaration this
PR introduced. Registering a converter with `AddObjectConverter(converter,
handledTypes)` promises the engine the converter produces values only for those
types, and in exchange the compiled interop lanes are kept for every member that
cannot produce one. Nothing links that promise to the converter's own
`TryConvert` switch: a case added there and not added to the registration is
silently skipped on exactly the members the declaration excluded, and honoured
everywhere else. `ByteArrayConverter` and `JsonElementConverter` are consistent
today; this is what would report it if they drifted.
Elsa defines no `ObjectInstance` subclass, so the rest of the verifiers have
nothing to check here yet. They are a standing guard for the day a host handler
or a satellite module adds one.
Wired as a module initializer, because the switch has to be set before the first
use of any Jint type. Duplicated across the three suites that reference
`Elsa.Expressions.JavaScript` rather than shared: `Elsa.Testing.Shared.Integration`
would be the obvious home, but it is a published package, and a module
initializer there would flip a process-wide switch for every external consumer
of it as well. A one-line test pins that the initializer ran, since Elsa
satisfies the contracts it is subject to and nothing else would notice the
checks going away.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f
---------
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Sipke Schoorstra <sipkeschoorstra@outlook.com>
Brings the 3.8.0 release line into main, including the package-manifest
runtime-kind mechanism (src/PackageManifest.props + src/PackageManifestHints.cs)
that main did not have. All 72 manifest-producing packages now declare
compatibility.runtimeKinds = ["elsa.server"]; the two Bpmn modules added on
main pick this up automatically via their ShellFeatures directory.
Conflict resolutions:
- .specify/feature.json, CONTEXT.md, ROADMAP.md, build/_build.csproj: took
main's, which is newer in every case. Verified byte-identical to main
afterwards, so nothing from the release branch was dropped.
- NuGet.Config: union of package sources, minus valence-consolelogstream-feedz.
main removed that feed deliberately in 7389e0a67 and now consumes
ConsoleLogStreaming 1.1.0 from nuget.org.
- Elsa.sln: union of main's Bpmn projects and the release branch's
ExternalAuthentication projects; the two sets are disjoint.
Reverted an unintended revert:
release/3.8.0 had lost commit 33181b2c9 ("test: cover Oracle bulk upsert SQL
generation") through an evil merge in c557c455a. That commit is present at the
merge base, so git resolved the release branch's older content as an
intentional change and would have silently undone it on main. It is three
coupled pieces:
- test/unit/Elsa.Persistence.EFCore.UnitTests (deleted, plus its Elsa.sln
project declaration and NestedProjects entry)
- InternalsVisibleTo("Elsa.Persistence.EFCore.UnitTests")
- the fix itself in BulkUpsertExtensions.GenerateOracleUpsert: internal
visibility, ISqlGenerationHelper.DelimitIdentifier quoting, and explicit
CAST(... AS NVARCHAR2(...)) on string columns
Dropping the third would have been an Oracle runtime regression: unquoted
identifiers lose case, and ODP.NET binds .NET strings as VARCHAR2 while Elsa's
Oracle migrations declare NVARCHAR2, causing a datatype mismatch. Merge base
and main are identical for that file and every hunk on the release side is a
revert plus cosmetics, so main's version was kept in full.
Accepted deliberate release-branch changes, verified as real refactors rather
than losses: AI EF Core migrations moved into the provider projects
(5c0d8b0f4), and AIPersistenceFeature.cs renamed to
EFCoreAIPersistenceShellFeatureBase.cs (ShellFeatures/ still present, so
manifest generation is unaffected).
Verified: dotnet build Elsa.sln succeeds with 0 errors and 2 pre-existing
NU1903 warnings; Elsa.Persistence.EFCore.UnitTests passes 1/1; all 72 emitted
manifests declare elsa.server and Elsa.Api.Common emits none.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(core): a throwing FaultSignal handler must not escape the middleware (#7911)
The signal is sent from inside the catch whose whole job is to stop exceptions
escaping the activity pipeline. A handler that threw went straight through it:
no incident, no strategy, and the original fault lost along with it.
The send is now guarded. A handler that throws is treated as not having handled
the fault, so the incident strategy runs exactly as it would with no handler
present. That is the conservative direction: a handler that failed part way
through may have left the faulted activity in any state, and an incident is a
better answer than silence. Its exception is logged at error level, because a
broken fault handler is a defect in its own right rather than a workflow
outcome.
Covered both ways, since a handler that already claimed the fault before
throwing is the case that could plausibly have been mistaken for success.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(core): let cancellation from a fault handler propagate
The handler guard caught OperationCanceledException along with everything else,
so a cancellation raised while an ancestor was being offered a fault was logged
as a broken handler and handed to the incident strategy. A deliberately
cancelled run reported itself faulted.
Cancellation is excluded now, matching how this repository already keeps the two
apart: the workflow-level exception middleware cancels and rethrows before its
general catch, and WorkflowRunner declines to record cancellation as the
workflow's exception.
Caught by review on #7924.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(core): pin that a handled fault stays in the execution log
The justification for dropping the incident is that the journal keeps the
evidence. That was asserted in several places and guarded nowhere.
It holds because ExecutionLogMiddleware writes the Faulted entry from a catch
that rethrows, and it is registered inside ExceptionHandlingMiddleware, so the
entry lands before the fault is ever offered to an ancestor. Swapping those two
registrations would make a handled failure disappear from the record with
nothing failing, which is what this test now prevents.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* fix(core): a fault a container claimed is not an incident (#7911)
RecoverFromFault reset the counts and the status but left behind the two other
things Fault recorded: the ActivityIncident and the exception. So a container
that successfully handled a child's fault still left the workflow carrying an
incident.
That is not cosmetic. Code reads a non-empty WorkflowExecutionContext.Incidents
as "this workflow failed" without looking further; HttpWorkflowsMiddleware is
one, and it hands the caller a fault response. A workflow whose container caught
the error and finished normally was reported to its caller as failed.
RecoverFromFault is now the inverse of Fault: it removes the incident Fault
appended, matched on this activity's node id and most recent first so an
activity that faults, recovers and faults again keeps the incident that was
never recovered, and it clears the recorded exception so the activity does not
sit in Running carrying one.
The execution log still records the failure, so nothing is hidden from anyone
reading the journal. Two integration assertions that encoded the old behaviour
are updated; they were written from the reasoning this change corrects.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(core): tie an incident to the execution that raised it, not its node
Recovery matched the incident to remove on ActivityNodeId, which identifies the
static workflow node rather than an execution of it. A node inside a loop,
retried, or run concurrently raises one incident per execution, all under the
same node id, so recovering one execution could remove another's incident and
leave its own behind.
ActivityIncident now carries the ActivityInstanceId of the execution that raised
it, and recovery matches on that. Within a single execution the most recent is
still taken, so fault, recover, fault again keeps the incident that was never
recovered. The property is optional: an incident recorded against the workflow
itself has no execution, and so do incidents persisted before this existed.
Caught by review on #7923.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(api-client): mirror ActivityInstanceId on the client incident model
The server model gained the property in the previous commit and the API client
carries a hand-maintained copy of it. Left alone, a client deserializing an
incident would silently drop the only field that says which execution raised it.
Also records two consequences of recovery that were implicit: it relies on the
incident collection preserving insertion order to pick an execution's newest
incident, which holds only because the collection is list-backed; and clearing
the exception also clears it from the activity's execution record, which is
intended for the same reason the incident goes, with the journal keeping the
evidence either way.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Review caught that the contract advertised something that cannot work. It
offered handlers three ways to terminalize the faulted activity - cancel,
complete, or reschedule - but CompleteActivityAsync returns immediately unless
the activity is Running, and throughout the handler it is still Faulted, since
recovery runs only after the handler returns. Completing inline did nothing at
all, silently, leaving the child Running.
Measured, same container, handler completing the faulted child:
inline complete -> child Running, no output, Running/Suspended
TransitionTo(Running), complete -> child Completed, "after", Finished/Finished
So a supported path exists; it just needed writing down. Document it on
FaultSignal, note that it is not licence to call RecoverFromFault (which also
rewrites the fault counts), and note that cancelling and rescheduling need no
equivalent step. Cover it with an integration test asserting that completing the
child with a substitute result fires the container's completion callback and
resumes its sequencing.
Refs #7911
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Review raised that TrySendSignalAsync delivers to the faulting activity before
walking ancestors, so an activity that throws and also handles FaultSignal can
claim its own fault and suppress the incident strategy.
That is real, but it is the channel's existing dispatch, which #7911 chose
deliberately over a variant of it, and SignalContext.IsSelf exists so handlers
can discriminate. It also grants no capability: an activity that catches its own
exception never faults at all, ending Finished/Finished with zero incidents,
which is a cleaner suppression than self-handling (incident still recorded,
activity left Running, workflow suspended).
So dispatch is unchanged. What was missing is that none of this was written
down: the contract describes the handler as an enclosing container and never
mentioned self-receipt. Document it on FaultSignal, including how a handler
that wants ancestors-only semantics opts out, and add a test so the behavior is
pinned rather than incidental.
Refs #7911
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A container activity had no way to learn that one of its children faulted.
ExceptionHandlingMiddleware caught the exception, called context.Fault(e) and
handed off to the workflow-global IIncidentStrategy; the container's completion
callback never fired, because the child never completed.
Add a seam on the ancestor-bubbling signal channel that already exists:
- FaultSignal(Exception, ActivityExecutionContext), beside CancelSignal. Its XML
doc carries the contract, including why a handler must not call
RecoverFromFault and why the CompleteActivityAsync sweep is a backstop rather
than the mechanism.
- An internal bool-returning TrySendSignalAsync, since SignalContext
.StopPropagationRequested is internal and SendSignalAsync reported nothing.
SendSignalAsync keeps its public signature and delegates to it.
- ExceptionHandlingMiddleware sends the signal after faulting and, when an
ancestor stops propagation, calls RecoverFromFault once and returns instead of
raising an incident.
RecoverFromFault now transitions to Running only when the activity is still
Faulted. It is called after the handler runs, so the unconditional transition
would otherwise undo a handler that cancelled or completed the faulted child.
The counts are still reset unconditionally, and the one pre-existing caller is
unaffected.
Behavior is unchanged when nobody handles the signal: verified by running the
new unhandled-fault theory against the pre-change middleware, and by
IncidentStrategyTests and Primitives/FaultTests passing unmodified.
Refs #7911
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This forward-ports the FailOnValidationErrors publish toggle to the 3.8
mainline. It is purely additive and changes no default behavior: publishing
a workflow definition still fails when validation errors are present
(FailOnValidationErrors defaults to true, i.e. strict = opt-out).
Setting FailOnValidationErrors to false — via
WorkflowManagementFeature.UseFailOnValidationErrors(false) or the
ManagementOptions.FailOnValidationErrors option — allows publication to
succeed while surfacing the validation errors as warnings on the publish
result. This enables publishing workflows that intentionally leave required
properties blank (for example, an empty Cron expression used to disable a
trigger).
This is the 3.8 counterpart to #7740, which introduces the same toggle as
opt-in (default false) on the 3.6.x line. References #7738.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
* Avoid null endpoint DTO metadata in tests
* Enforce console logs hub read permission
* Remove unused console logs hub import
* Support mapped endpoint metadata in auth tests
* Reduce console log capture throughput impact
* Address Copilot console logs review
* Refactor task scheduling to support tenant-level background work and enhance logging functionality.
* Introduce ConsoleStreamHook for stdout/stderr tee and enhance logging validation. Adjust test cases and startup warnings for distributed lock provider usage.
* Refactor console logging pipeline with capture optimization and new ConsoleLogsHost; update tests accordingly.
* Add Ansi SGR parser for console logs and associated unit tests
* Remove ANSI color renderings and parsers; integrate ConsoleLogScopeAccessor for improved logging context with workflow instance ID support.
* Address console logs code quality feedback
* Address PR review feedback
* Preserve console logs extension points
* Stabilize console logs host lifecycle
* Address final automated review comments
* Tighten console log capture shutdown
* Address console log review feedback
* Address follow-up review feedback
* Cover final review feedback
* Avoid recursive console provider initialization
* Guard console host lease shutdown
* Preserve console log scope and provider lifetime
* Correlate console log scope fallback
* Tighten console scope correlation
* Expose host services during provider construction
* Redact ANSI-normalized console lines
* Add OpenTelemetry diagnostics backend foundation
* Add OTLP HTTP ingestion parsing
* Document OpenTelemetry diagnostics setup
* Enforce OpenTelemetry hub permissions
* Remove `ConsoleCaptureTee` and related services and tests
* Add OpenTelemetry HTTP ingestion integration test
* Use pipeline contributors for console log context
* Update CShells package versions to 0.0.24-preview.132
* Add OpenTelemetry ingestion security tests
* Add OpenTelemetry API authorization tests
* Filter live console logs by workflow instance
* Add OpenTelemetry hub tests
* Add OpenTelemetry gRPC metadata hook
* Assert OpenTelemetry workflow tags survive ingestion
* Mark OpenTelemetry core build verified
* Enhance console logging with activity execution metadata and extend test coverage.
* Address console logs stream consumption comment
* Wire OpenTelemetry diagnostics into core sample
* Address Core diagnostics review feedback
* Address Core Copilot follow-up feedback
* Add OpenTelemetry metric instrument names
* Address Core Copilot provider feedback
* Address Core Copilot diagnostics follow-up
* Address Core Copilot live feed feedback
* Address Core Copilot store feedback
* Integrate OpenTelemetry for logging, tracing, and metrics in ModularServer and update launch settings and docker-compose configuration.
* Refactor to replace `ConsoleLogStream.Core` with `ConsoleLogStreaming.Core` across codebase and update `ConsoleStreamHook` installation.
* Add diagnostics OpenTelemetry backend
* Fix OpenTelemetry live hub subscription
* Fix modular OpenTelemetry exporter endpoints
* Add CShells logging configuration in appsettings.json
* Remove obsolete unit tests and helper classes
* Restore default activity exception handling
* Simplify type serialization and alias management
This commit refactors the internal type serialization and alias management system to reduce boilerplate, improve robustness, and simplify the developer experience:
- Removed numerous explicit `ExpressionOptions` type alias registrations across various modules.
- Updated `TypeJsonConverter` and polymorphic serialization to reliably handle types using assembly-qualified names when a short alias is not explicitly registered.
- Streamlined `ExcludeFromHashConverter` to strictly adhere to `ExcludeFromHashAttribute` for hash calculations, removing complex `JsonIgnoreCondition` logic.
- Eliminated several helper classes (`WorkflowJsonTypeResolver`, `WorkflowTypeValidator`, `IWorkflowTypeRegistry`, `WorkflowFactoryDictionary`, `JavaScriptExceptionTypeAliasRegistrar`, `WorkflowRuntimeTypeAliasRegistrar`) and their associated unit tests, simplifying the codebase.
Additionally, this commit introduces a comprehensive markdown document (`product-website-feature-source.md`) outlining Elsa's core features, Studio capabilities, extension ecosystem, and architectural selling points, intended as source material for the product website.
* Refine type serialization for improved robustness and alias handling
This commit further enhances the type serialization and deserialization mechanisms:
* Centralizes type resolution and alias management through `IWellKnownTypeRegistry` and `WorkflowJsonTypeResolver`.
* Prioritizes registered type aliases when serializing type metadata in `PolymorphicObjectConverter`, resulting in more concise JSON output.
* Enhances deserialization in `PolymorphicObjectConverter` and `VariableMapper` to gracefully handle unknown or non-instantiable types, providing fallbacks and logging warnings.
* Simplifies `TypeJsonConverter` by delegating complex type resolution logic to the `WorkflowJsonTypeResolver`.
* Adds `JsonArray` to the well-known type aliases for direct recognition.
* Fix console logs packaging and workflow type resolution
* Fix console log metadata and type resolution
* Address Copilot review feedback
* Enhance type resolution, improve console log handling, and update tests
- Streamlined `WorkflowDictionaryExtensions` for better workflow registration validation.
- Refined `ConsoleLogsAuthorizationTests` with the new `SetJsonRequest` helper to improve test requests handling.
- Updated `OrderDefinition` to ignore JSON serialization for `KeySelector`.
- Enhanced `WorkflowRuntimeFeature` for improved workflow registration and type alias configuration.
- Added tests to ensure `ConsoleLogProvider` metadata filtration in various scenarios.
- Improved type serialization logic in `WorkflowJsonTypeResolver`.
- Updated README to fix references related to diagnostics.
- Optimized `ExcludeFromHashConverter` for property serialization conditions.
- Modified `TriggerIndexer` for streamlined trigger management.
- Tested payload checks in `PublishEventTests`.
- Adjusted `Endpoint` in `ConsoleLogs` for automatic JSON request handling.
- Ensured registration of workflow type aliases in `WorkflowsFeature`.
* Restore CLR workflow registration compatibility
* Align JSON island serialization fixtures
* Add Console Logs Services and Enhance Endpoint Handling
- Introduced `ActivityExecutionsEndpointTests` to validate route exposure.
- Added `ConsoleLogCaptureHostedService` for console log streaming.
- Implemented `ConsoleStreamJsonConverter` for JSON conversion of console streams.
- Developed `ElsaConsoleLogRecentBuffer` to handle recent log buffering.
- Updated `ConsoleLogsAuthorizationTests` with new test cases for stream filter mapping.
- Consolidated console log provider dependencies and registration, including recent buffering.
- Enhanced `ElsaConsoleLogProvider` to use recent buffer for filtering.
- Adjusted `Program.cs` for streamlined logging service setup.
* Enhance type resolution and test coverage; streamline console log integration
- Added `ConsoleStreamHook` for streamlined log streaming.
- Updated `WorkflowJsonTypeResolverTests` to improve type resolution and test new scenarios.
- Simplified type resolution by removing trusted assembly checks.
* Fix CI smoke and package restore failures
* Fix Docker smoke image project paths
* Fix Docker Python runtime packages
* Fix Docker CA smoke teardown
* Refresh Elsa roadmap
* Implement background processors and mediation coordination
- Added `BackgroundCommandProcessor`, `BackgroundJobProcessor`, and `BackgroundNotificationProcessor` classes for handling commands, jobs, and notifications, respectively.
- Introduced `MediatorBackgroundProcessingCoordinator` to coordinate the execution of all background processors.
- Implemented `MediatorBackgroundTask` for wrapping `MediatorBackgroundProcessingCoordinator` in `BackgroundTask`.
- Added unit tests for `MediatorBackgroundTask` to ensure proper start and stop behavior.
- Refactored `BackgroundCommandSenderHostedService` to utilize `BackgroundCommandProcessor`.
- Introduced 'elsa-roadmap-refresh' skill configuration for roadmap updates.
* Address workflow type resolution review feedback
* Address follow-up review feedback
* Restore recent console logs execute path
* Address Copilot follow-up review
* Decouple workflow JSON aliases from expressions
* Fix workflow management unit test setup
* Fix console logs recent endpoint handler shape
* Respect workflow JSON strict type aliases
* Remove unused console log contracts reference
* Address Copilot review feedback
* Address Copilot follow-up comments
* Synchronize ring buffer dropped count
* Address background processor strategy replay
* Fix: do not resume interrupted workflows that are already finished
* Avoid fixed timestamps in restart workflow test
---------
Co-authored-by: Sipke Schoorstra <sipkeschoorstra@outlook.com>
* feat(workflows-runtime): add quiescence machinery foundation for graceful shutdown
Introduces the container-scoped quiescence signal, ingress-source contract,
burst registry, and the Interrupted workflow sub-status — the foundational
primitives the drain orchestrator and admin endpoints will build on. No
behaviour change yet: workflows continue to run and shut down exactly as
before. The new types are registered but no host-stop or pause path drives
them.
Highlights:
* IQuiescenceSignal — composable Drain + AdministrativePause flags;
forward-only drain, reversible pause, idempotent transitions, optional
persistence via IKeyValueStore.
* IIngressSource + IForceStoppable — uniform contract for components that
inject external events (HTTP, schedulers, message consumers, internal
workers, third-party modules); IIngressSourceRegistry collects and
surfaces their states.
* IBurstRegistry — atomic counter for in-flight workflow execution
bursts, with per-burst ingress attribution and FR-018 inconsistency
detection (a source claiming Paused but starting bursts is flipped to
PauseFailed).
* WorkflowSubStatus.Interrupted — new value distinct from Suspended,
Cancelled, Faulted; semantics: "last burst force-cancelled by graceful
drain; resumable on next runtime generation". Mirrored on the API client
enum.
* GracefulShutdownOptions — drain deadline, per-source pause timeout,
stimulus-queue back-pressure policy, pause-persistence policy.
Configurable via UseWorkflowRuntime(...).ConfigureGracefulShutdown(...).
* PermissionNames.ManageWorkflowRuntime — single permission for the
forthcoming admin pause/resume/status/force endpoints.
Implements 31 of 77 tasks for the graceful-shutdown feature
(specs/002-graceful-shutdown). Subsequent commits add the drain
orchestrator (US1 / MVP), Interrupted recovery scan (US3), admin
endpoints (US2), and first-party ingress adapters.
Tests: 25 new xUnit unit tests; 100/100 runtime unit tests pass; all
existing tests continue to pass on net8.0/net9.0/net10.0.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(workflows-runtime): add drain orchestrator + host-stop integration (US1, MVP)
When the host receives a stop signal (SIGTERM, Ctrl+C, orchestrator
rollout), the runtime now drains gracefully: ingress sources are paused
in parallel, in-flight workflow bursts run to their next natural
persistence boundary within a configurable deadline, and any burst that
breaches the deadline is force-cancelled and persisted with the
Interrupted sub-status plus a forensic WorkflowInterrupted log entry.
This is the MVP — without the activation-time recovery scan (PR 3) the
existing timeout-based RestartInterruptedWorkflowsTask still picks up
Interrupted instances, just on its periodic cadence. No regression in
that recovery path (SC-008).
Highlights:
* IDrainOrchestrator + DrainOrchestrator — protocol per the contract:
BeginDrainAsync → parallel ingress pause with per-source timeouts +
IForceStoppable escalation → poll BurstRegistry.ActiveCount until zero
or deadline → on breach iterate live handles, cancel, persist
Interrupted, write log entry. All exceptions are captured into the
returned DrainOutcome; only second-invocation throws.
* Deadline clamping: effective deadline is min(GracefulShutdownOptions.
DrainDeadline, HostOptions.ShutdownTimeout - 500ms safety epsilon),
so the runtime never outlives its host process.
* DrainOrchestratorHostedService — IHostedService.StopAsync wakes the
orchestrator on host stop. Registered AFTER the heartbeat
(Elsa.Hosting.Management) so reverse-order shutdown keeps the
heartbeat alive throughout drain. Prevents sibling-node crash recovery
from false-positive-recovering instances we are gracefully handling
here (FR-029).
* BurstTrackingMiddleware — workflow-execution-pipeline middleware that
registers a BurstHandle for the lifetime of every burst. All nine
IWorkflowRunner.RunAsync overloads ultimately funnel into
pipeline.ExecuteAsync(context), so this single middleware covers the
three "burst choke points" the spec references without nine separate
decorators. Added to UseDefaultPipeline().
* Ingress attribution: optional IngressSourceName property on
DispatchStimulusRequest, DispatchWorkflowDefinitionRequest, and
DispatchWorkflowInstanceRequest. Adapters set it; the middleware reads
it via WorkflowExecutionContext.TransientProperties (helpers in
IngressAttributionExtensions). The BurstRegistry uses the name to
detect the FR-018 invariant violation — a source that reports Paused
but starts a burst is flipped to PauseFailed.
* InterruptedLogExtensions — the LogWorkflowInterruptedAsync helper
that the orchestrator calls when persisting the forensic record.
Tests: 11 new unit tests (DrainOrchestrator parallel-pause +
wait-for-bursts + idempotency + persistence-failure path); 5 new
integration tests (full DI graph resolves, burst-tracking middleware
registers handles end-to-end, no-op drain returns
CompletedWithinDeadline). 100/100 runtime unit tests pass; all
existing tests continue to pass on net8.0/net9.0/net10.0.
Implements 14 of 77 tasks (T032–T045). Subsequent commits add
Interrupted recovery scan (US3), admin endpoints (US2), and
first-party ingress adapters.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(admin-endpoints): add admin endpoints for workflow runtime control
Introduced admin endpoints to manage workflow runtime: `/pause`, `/resume`, `/status`, and `/force` with full authentication and audit logging. Integrated idempotency checks and error handling to ensure reliable runtime control. Added corresponding integration tests for verification.
* fix(workflows-runtime): propagate drain cancellation into running workflows + initialize persisted pause + fix log-comment
Addresses three findings from the PR #7424 code review:
1. **HIGH — Cancellation now propagates into the running workflow.**
`BurstHandle.Cancel()` previously cancelled only its own linked CTS,
which the workflow runner never observes (the runner reads from
`WorkflowExecutionContext.CancellationToken`, captured at context
construction and not part of the linked chain). On deadline breach
the orchestrator would persist `Interrupted`, but the workflow
continued executing and could overwrite the sub-status with whatever
terminal state it eventually reached.
Fix: `BurstHandle` accepts an optional cancel callback at construction.
`BurstTrackingMiddleware` wires it to `context.Cancel()` so the burst's
cancellation triggers the workflow's own cancellation chain — the
workflow transitions to `Cancelled` and stops scheduling new
activities. The orchestrator then awaits `BurstHandle.Disposed` (with
a 2 s settle timeout) before persisting `Interrupted`, ensuring the
runner's terminal commit completes BEFORE the orchestrator overwrites
the sub-status. Race resolved.
The settle timeout is bounded so a non-cancellable activity (genuinely
pathological case) does not block drain — on timeout the orchestrator
logs and proceeds, accepting the runner-clobber for that one
instance, which the existing timeout-based RestartInterruptedWorkflows
recovery picks up afterwards.
2. **MEDIUM — Pause persistence is now actually wired.**
`QuiescenceSignal.InitializePersistedStateAsync` was implemented but
nothing called it on host startup. A host configured with
`PausePersistence = AcrossReactivations` would write the persisted
key on pause, but on subsequent activation the new
`QuiescenceSignal` instance would never read it, so the runtime would
resume dispatching despite the operator having paused.
Fix: `InitializePauseStateStartupTask : IStartupTask` reads the policy
and calls `InitializePersistedStateAsync` once per activation when the
policy demands it. Registered in both `WorkflowRuntimeFeature`
flavours alongside the other graceful-shutdown services.
3. **MEDIUM — Comment in `DrainOrchestrator.PersistInterruptedAsync` no
longer lies.** The previous comment promised a "synthetic log entry"
that the next statement (`return`) prevented from being written. The
comment is now honest about what actually happens: when no instance
row exists, no log entry is emitted, but the burst metadata is still
captured in the drain outcome's logged warning so operators have a
forensic trail.
Tests:
* New unit tests on `BurstRegistry` (now 9, was 6): cancel-callback is
invoked, callback exceptions are swallowed (drain remains best-effort),
`BurstHandle.Disposed` completes on dispose.
* Full suites continue to pass: 103/103 runtime unit tests; 247/247
workflow integration tests.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(workflows-runtime): address PR review feedback + close runner-clobber race + add e2e drain test
Six issues raised in PR #7424 review (one e2e gap, five comments inline):
1. **Runner-clobber race closed via ICommitStateHandler decorator.**
The previous fix wired `BurstHandle.Cancel()` to `WorkflowExecutionContext.Cancel()`
so the workflow's cancellation chain fires on deadline breach, but the e2e
test exposed that `BurstHandle` disposed at the END of the pipeline middleware
(i.e., BEFORE `WorkflowRunner` calls `commitStateHandler.CommitAsync`). The
orchestrator's `await handle.Disposed` therefore returned too early, the
instance row didn't yet exist, and the orchestrator's Interrupted write was
either a no-op (no row) or got clobbered by the runner's subsequent Cancelled
commit.
Fix: `BurstAwareCommitStateHandler` decorates `ICommitStateHandler`. The
middleware no longer disposes the handle in the success path — it stores
the handle in `WorkflowExecutionContext.TransientProperties`, and the
decorator disposes it AFTER `inner.CommitAsync` completes. The exception
path in the middleware still disposes for safety. Result: the orchestrator's
await-disposed sequencing now correctly lands the Interrupted write last.
2. **C1: Null-instance log entry.** `DrainOrchestrator.PersistInterruptedAsync`
now writes a synthetic `WorkflowInterrupted` log entry directly when no
instance row exists, populating only the fields it knows. Previously the
forensic trail was lost.
3. **C2: Force endpoint cached-outcome audit.** Added `WasCached` flag to
`DrainOutcome` (default false). The orchestrator sets it on the cached
return path (`_previousOutcome with { WasCached = true }`). The force
endpoint now skips the audit notification when the flag is true, so
repeated force calls no longer emit spurious `RuntimeForceRequested` events
(SC-007 idempotency restored).
4. **C3: `StateChanged` raised under lock — deadlock risk closed.**
`QuiescenceSignal.BeginDrainAsync`/`PauseAsync`/`ResumeAsync` now do their
transitions under the lock, capture whether a transition occurred, release
the lock, and only then invoke `RaiseStateChanged`. Subscribers that
synchronously call back into the signal can no longer deadlock.
5. **C4: Scheduling source name.** Renamed `scheduling.cron` → `scheduling.triggers`
to honestly reflect the four trigger types the adapter covers (Cron, Timer,
StartAt, Delay). The name is surfaced verbatim in admin status responses.
6. **C5: Hardcoded Retry-After.** `HttpWorkflowsMiddleware`'s 503 response now
sets a reason-aware `Retry-After`: 5 s during drain (host is exiting and
will be replaced shortly), 60 s during administrative pause (indefinite,
so a longer back-off avoids tight retry loops).
Tests:
* New e2e `DeadlineBreachEndToEndTests` (2 tests): verifies that drain
against a real running workflow detects the in-flight burst, force-cancels
it, persists the instance as `Interrupted`, and writes a `WorkflowInterrupted`
log entry — closing the test gap that hid the cancellation-propagation
issue identified in the previous review pass.
* Updated `OperatorForceAfterPreviousReturnsCachedOutcome` to assert
value-equality + the `WasCached` flag instead of reference-equality
(records use `with` for the cached return path).
Full suites pass: 103/103 runtime unit, 249/249 workflow integration.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(workflows-api): move runtime admin endpoints into Elsa.Workflows.Api
Per PR review feedback: rather than introducing a new sub-module
(Elsa.Workflows.Runtime.Admin) for the four pause/resume/status/force
endpoints, fold them into the existing Elsa.Workflows.Api project. That
project already references both Elsa.Workflows.Runtime and
Elsa.Api.Common (FastEndpoints) and is the established home for
client-facing workflow APIs — so the admin endpoints belong there.
Changes:
* New folder src/modules/Elsa.Workflows.Api/Endpoints/RuntimeAdmin/ with
Models.cs and Pause/Resume/Status/Force/Endpoint.cs. Namespaces moved
from `Elsa.Workflows.Runtime.Admin` → `Elsa.Workflows.Api.Endpoints.RuntimeAdmin`.
* Deleted src/modules/Elsa.Workflows.Runtime.Admin/ entirely and removed
it from Elsa.sln. The ShellFeature marker class
(WorkflowRuntimeAdminFeature) is no longer needed — the existing
WorkflowsApiFeature already discovers FastEndpoints in the Workflows.Api
assembly.
* No consumer changes: the endpoints sit in the same routes
(/admin/workflow-runtime/*) and behave identically.
Note on the second architectural point ("update IShellFeature if cleaner"):
the IShellFeature contract is defined in the external CShells NuGet
package, not in this repo, so we cannot add a DeactivateAsync hook
without an upstream CShells change. The current IHostedService.StopAsync
hook continues to work correctly for the host-stop path; per-shell
deactivation would require either a CShells upstream addition or a
separate Elsa-owned shell-feature variant — neither lighter than what we
have today.
Tests: 103/103 runtime unit + 249/249 workflow integration pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(common): introduce IsFatal exception extension + apply to drain best-effort catches
Per PR review feedback on the static-analyzer "Generic catch clause"
comments: rather than catching ALL exceptions in best-effort drain code
paths, narrow the swallow to non-fatal exceptions. Process-fatal
conditions (StackOverflowException, AccessViolationException,
SEHException, ThreadAbortException, OutOfMemoryException) propagate so
the host's failure-fast policy can act on them, while normal failures
(InvalidOperationException, IOException, etc.) continue to be logged
and allowed through so a single misbehaving ingress source / activity
cannot abort the overall drain.
Highlights:
* New `Elsa.Common.Extensions.ExceptionExtensions.IsFatal` utility:
classifies fatal conditions, unwraps reflection-style wrappers
(TypeInitializationException, TargetInvocationException) before
classification, and treats InsufficientMemoryException (the
recoverable OOM subclass) as non-fatal.
* Applied as a `when (!ex.IsFatal())` filter to:
- BurstHandle.Cancel (cancel callback try/catch)
- DrainOrchestrator.PauseOneSourceAsync (per-source exception path)
- DrainOrchestrator.TryForceStopAsync
- DrainOrchestrator.ForceCancelActiveBurstsAsync (per-burst loop)
- DrainOrchestrator.PersistInterruptedAsync (orphan log write,
instance save, log write)
- DrainOrchestrator.DrainAsync outer catch (existing
`not InvalidOperationException` filter extended)
- InterruptedRecoveryScan (per-instance restart loop)
Tests: 7 new unit tests for IsFatal classification (fatal types,
recoverable types, wrapped causes, null tolerance). Full suites:
103/103 runtime unit (incl. 14/14 in Common.UnitTests including new
tests) + 249/249 workflow integration pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(workflows-runtime): integrate CShells 0.0.15 lifecycle hooks (IDrainHandler + IShellInitializer)
CShells 0.0.15 ships the lifecycle framework needed for first-class per-shell
graceful shutdown — IDrainHandler / IShellInitializer / IShellLifecycleSubscriber.
This commit bumps the package, migrates Elsa's existing usage of the removed
0.0.14 API, and registers the runtime drain orchestrator + pause-state initializer
through the new primitives.
Highlights:
* `ElsaShellDrainHandler : IDrainHandler` — bridges per-shell drain into
`IDrainOrchestrator.DrainAsync(DrainTrigger.ShellDeactivation, ct)`. Invoked
by CShells when a shell enters `ShellLifecycleState.Draining`; the drain
handler's CancellationToken is signalled when the per-shell deadline elapses,
so the orchestrator's own deadline-bounded protocol nests cleanly.
Coexists with the host-stop `DrainOrchestratorHostedService`; the
orchestrator's `DrainAsync` is idempotent — second invocations log and skip.
* `InitializePauseStateShellInitializer : IShellInitializer` — replaces the
IStartupTask variant in shell-aware deployments. IShellInitializer fires on
EVERY shell (re)activation, including reactivations after a reload — exactly
what FR-028 requires. The IStartupTask remains for IModule consumers where
there is no shell platform.
Migrations (CShells 0.0.14 → 0.0.15 breaking changes):
* `ActivateShellTenants`: was `IShellActivatedHandler` + `IShellDeactivatingHandler`,
now `IShellInitializer` + `IDrainHandler`.
* `MultitenancyFeature`: registrations updated to the new transient interface,
`using CShells.Hosting` → `using CShells.Lifecycle`.
* `Reload/Endpoint`, `ReloadAll/Endpoint`: `IShellManager` → `IShellRegistry`,
`ReloadShellAsync` → `ReloadAsync` (returns `ReloadResult` with `Error`),
`ReloadAllShellsAsync` → `ReloadActiveAsync` (returns
`IReadOnlyList<ReloadResult>` with per-shell errors aggregated into 503).
Build + restore:
* `Directory.Packages.props`: all CShells.* packages bumped to 0.0.15.
* `NuGet.Config`: added `cshells-feedz` source
(https://f.feedz.io/sfmskywalker/cshells/nuget/index.json) and split the
package-source-mapping pattern into `CShells` (exact) + `CShells.*`
(prefix). Single-pattern `CShells*` does NOT match correctly under
PackageSourceMapping.
Tests: 103/103 runtime unit + 249/249 workflow integration pass on the new
package version.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(workflows-runtime): apply PR #7424 review feedback
Consolidates the architectural fixes asked for during /review:
- Extract IWorkflowRuntimeAdminService to back the four /admin/workflow-runtime endpoints with a single domain service; thin Pause/Resume/Status/Force endpoints to delegating shells.
- Remove StateChanged C# event from IQuiescenceSignal (Constitution VII: no external subscribers existed; mediator was suggested as the alternative if/when it's needed).
- Promote InitializePersistedStateAsync to IQuiescenceSignal, dropping the concrete-cast in both InitializePauseStateStartupTask and InitializePauseStateShellInitializer.
- Invert ingress-source DI to Lazy<IEnumerable<IIngressSource>> to break the cycle through IQuiescenceSignal; ingress adapters take the signal directly via primary constructor.
- Replace Guid.NewGuid().ToString("N") with IIdentityGenerator in InterruptedLogExtensions and DrainOrchestrator.
- Switch admin-audit timestamps to ISystemClock in WorkflowRuntimeAdminService.
- Make GracefulShutdownOptions.StimulusQueueMaxDepthWhilePaused nullable (null = unlimited).
- Rename RuntimeForceRequested → RuntimeForceDrainRequested.
- Apply IsFatal exception filter to drain best-effort catches.
- Rename IBurstRegistry.EnumerateActive → ListActiveBursts.
- Refresh "Phase X" comments to user-story (USx) references.
- Delete unused IngressAttributionExtensions, IngressSourceServiceCollectionExtensions, IngressSourceRegistrationOptions.
- Migrate Elsa.Shells.Api.Tests to CShells 0.0.15 IShellRegistry / ReloadResult surface.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(graceful-shutdown): apply IsFatal filter to deadline-breach test catch
Aligns the test scaffolding's swallow-everything catch with the project standard introduced in c00eee80c so the analyzer no longer flags the bare `catch` clause. The semantics are unchanged — non-fatal exceptions (OCE, TimeoutException, workflow exceptions) are still acceptable test outcomes.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(workflows-runtime): close IngressSourceRegistry first-access race + spelling
Replaces the non-atomic `_entries.Count > 0` early-return guard in
EnsureMaterialized with a double-checked lock against a volatile
`_materialized` flag, so concurrent first callers can no longer both
iterate the source factory and crash one of them with a "Duplicate ingress
source registration" InvalidOperationException. Adds a regression test that
launches 16 readers behind a TaskCompletionSource gate and asserts every
reader observes the full source set without throwing.
Also flips the British spellings introduced in this PR's scope to American
English (materialize/behavior) — project convention going forward.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(constitution): require American English for new code (v1.0.1)
Adds a "Spelling & language" bullet under principle III (Convention-Driven Design): every newly-introduced symbol, comment, identifier, error message, XML doc, commit message, and Speckit artifact uses American English. Established public API symbols (e.g. WorkflowSubStatus.Cancelled) are not renamed retroactively. PATCH bump because this is a clarification of an existing principle, not a new principle.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Add Greploop skill and workflow for GitLab, GitHub, and Perforce integration
- Introduced Greploop, an iterative optimization and review workflow for GitLab MRs, GitHub PRs, and Perforce changelists.
- Added API and GraphQL references for fetching and resolving review skill.
* Remove GenerateWorkflowVariableAccessorsTests; redundant ExpandoObject type check in handlers
* Potential fix for pull request finding 'CodeQL / Untrusted Checkout TOCTOU'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* fix(graceful-shutdown): apply PR #7424 review feedback round 2
Two P1 findings from Greptile:
1. DrainOrchestrator.ForceCancelActiveBurstsAsync was sequential — each
burst was cancelled, awaited up to ForceCancelSettleTimeout (2 s), and
persisted before the next burst's Cancel() ran. Total wall time was
O(N × 2 s) and bursts 2..N kept executing at full speed during prior
bursts' settle waits, defeating the intent of force-cancel under
concurrency.
Refactored to three phases:
- Phase A — cancel every handle synchronously (cheap CTS.Cancel calls)
so all runners observe cancellation simultaneously.
- Phase B — await every Disposed signal in parallel under a single
shared ForceCancelSettleTimeout. Total wall time bounded regardless
of N.
- Phase C — persist Interrupted for each handle sequentially (keeps
DbContext usage single-threaded; per-handle work is small).
Per-phase failures are caught with !ex.IsFatal() and logged so a single
misbehaving handle doesn't abort the rest of the batch.
2. ShellFeatures/WorkflowRuntimeFeature.ConfigureServices was missing the
IWorkflowRuntimeAdminService registration that Features/WorkflowRuntimeFeature
already had. Any CShells deployment that includes the Pause / Resume /
Status / Force admin endpoints (in Elsa.Workflows.Api) would throw
InvalidOperationException at endpoint construction. Added the singleton
alongside the other graceful-shutdown registrations with a comment
pointing out the symmetry with the IModule path.
17/17 graceful-shutdown integration tests pass; 103/103 runtime unit tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Potential fix for pull request finding 'CodeQL / Untrusted Checkout TOCTOU'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* fix(ci): harden greploop.yml against CodeQL Actions findings
CodeQL flagged 12 findings on .github/workflows/greploop.yml after the
prior commit (c08183a3c) addressed an earlier round. Two distinct issue
classes remain:
1. Code injection (× ~10): step-output values
(steps.pr_head.outputs.head_sha / head_repo_owner / head_repo_name /
head_ref) and inputs.pr_number were interpolated directly into shell
`run:` blocks via `${{ ... }}`. Because PR author controls the branch
name and the manual-dispatch input, those values can carry shell
metacharacters. Standard fix: route every such interpolation through
an `env:` block on the step, then reference $VAR inside the script.
Applied to the Resolve, Resolve PR head metadata, and Checkout PR
branch steps.
2. Untrusted Checkout TOCTOU + Checkout of untrusted code in trusted
context: the workflow runs on `issue_comment` (a privileged trigger)
and checks out PR-author code. Mitigations stacked here:
- Author-association gate already restricts the trigger to OWNER /
MEMBER / COLLABORATOR (existing).
- Step-output values now travel via env vars (above).
- Resolve step rejects pr_number that isn't ^[0-9]{1,10}$ — so
downstream `gh pr view` and the prompt argument can't be hijacked.
- Checkout step now validates HEAD_SHA matches ^[0-9a-f]{40}$ and the
repo owner/name match ^[A-Za-z0-9_.-]+$ before either reaches a
URL or a git command.
- Existing TOCTOU guard preserved: re-fetch head SHA at checkout
time and abort if it changed since the initial resolve.
These match the canonical "Securing your GitHub Actions workflows"
patterns recommended by CodeQL.
No functional change to greploop's runtime behaviour.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(graceful-shutdown): drop [SingleNodeTask] from InitializePauseStateStartupTask
Greptile P1: [SingleNodeTask] gates the task to a single cluster winner via
distributed lock, but IQuiescenceSignal is a singleton scoped to each node's
DI container — each node holds its own in-memory QuiescenceState. With the
attribute, only the winning node restored the persisted pause; every other
node started with QuiescenceReason.None and accepted new work, silently
defeating PausePersistence = AcrossReactivations.
Removed [SingleNodeTask] (and the corresponding using) so the task runs on
every node. Expanded the doc <remarks> to call out the per-node requirement
and point at the shell-aware counterpart (InitializePauseStateShellInitializer)
which is correctly per-node by virtue of being an IShellInitializer.
17/17 graceful-shutdown integration tests still pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Potential fix for pull request finding 'CodeQL / Code injection'
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
* chore(deps): bump CShells 0.0.15 → 0.0.17
0.0.17 ships our blueprint-aware-routing PR (valence-works/cshells#93) plus
four follow-up fixes the maintainer added on top:
- e56ebd8 — PreWarmShells removed entirely; ShellMiddleware now does
cold-start endpoint matching (re-runs endpoint resolution after lazy
activation so the very first request to a cold shell hits its endpoint).
- cbe5ee2 — GetCandidateSnapshot returns a bounded ShellRouteCandidateSnapshot
with accurate total counts; sensitive-data redaction in routing logs;
DefaultShellRouteIndex implements IDisposable.
- cd7d4f5 — Last-good snapshot served on rebuild failure (the deferred
Copilot review concern); root-path fallback when path-by-name misses.
- c3679d9 — Cold-start endpoint matching respects inline route constraints;
path-name convention tightening; dead duplicate-detection cleanup.
Net effect for elsa-core:
- Cold blueprints serve their first request via lazy activation, with
endpoints correctly resolved post-activation.
- Reloaded shells re-activate and serve on the next matched request.
- Non-name-mode routing keeps serving the previous snapshot during a
transient blueprint-provider outage.
- No need to call PreWarmShells from Elsa.ModularServer.Web — removed.
The only API removal that touches elsa-core is PreWarmShells. No code
references IShellRouteIndex / ShellRouteCriteria / GetCandidateSnapshot
directly, so the API-shape changes in cbe5ee2 don't ripple here.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(graceful-shutdown): persist Interrupted under non-drain bounded token
Greptile P? finding on the prior force-cancel two-phase fix: Phase B's
inner catch on OperationCanceledException ("drain CT fired — proceed to
persist anyway") was a lie in practice. Phase C immediately passed the
same already-cancelled drain token into PersistInterruptedAsync; the
first DB call (instanceStore.FindAsync) observed the cancellation and
threw OperationCanceledException; the outer non-fatal Exception filter
swallowed it and only logged an error. Net effect: on host shutdown
deadline breach, every burst after cancellation could fail to be
persisted as Interrupted, leaving instances in an unrecovered executing
state.
Phase C now creates a per-handle CancellationTokenSource bounded to a
new PersistInterruptedTimeout (5 s) that is NOT linked to the drain CT.
Each persist gets up to 5 s to land the row update + forensic log entry
even after the drain CT has fired. The bound prevents a stuck DB from
hanging shutdown indefinitely (per-handle worst case is small; total
Phase C upper bound is N × 5 s, but typical persists are millisecond
scale).
Comment expanded to call out why the persist token is independent of
the drain token, so the rationale doesn't drift again.
17/17 graceful-shutdown integration tests still pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(graceful-shutdown): extract PassiveIngressSource base class
The three IIngressSource implementations that ship with this PR
(InternalBookmarkQueueIngressSource, HttpTriggerIngressSource,
ScheduledTriggerIngressSource) were ~25 lines each and ~22 lines of
those were verbatim copies of each other:
- ctor signature `(IQuiescenceSignal signal)`
- `PauseTimeout => TimeSpan.FromMilliseconds(50)`
- `CurrentState => signal.IsAcceptingNewWork ? Running : Paused`
- `PauseAsync` / `ResumeAsync` returning `ValueTask.CompletedTask`
The shared trait is that none of them does any work at pause time —
the actual pause enforcement lives in another layer
(`HttpWorkflowsMiddleware` short-circuits to 503,
`BookmarkQueueProcessor` consults the signal at the top of each
invocation, scheduled triggers dispatch through the bookmark queue and
inherit that behaviour transitively). The IIngressSource adapter is
purely diagnostic: it makes the source visible in
`DrainOutcome.Sources` and the admin status endpoint.
Extracted that pattern into `PassiveIngressSource` (abstract base in
`Elsa.Workflows.Runtime.IngressSources`). Subclasses now provide only
`Name`; `PauseTimeout` is `virtual` with a 50 ms default; everything
else is fixed by the base. The three concretes drop from ~25 lines to
~12 lines each.
The base's XML `<remarks>` calls out when to use it ("your component
already cooperates with IQuiescenceSignal at its hot path") and when
to implement IIngressSource directly ("the source owns concrete
pause/resume behaviour — e.g. a message-queue consumer that calls
Pause() on its underlying client"), so future contributors don't
mis-extend the base for sources that need real work at pause time.
No behavioural change. 18 graceful-shutdown integration + 39 runtime
unit tests still pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(graceful-shutdown): align IIngressSource name to singular
The three IIngressSource names were inconsistent:
http.trigger (singular)
internal.bookmark-queue-worker (singular)
scheduling.triggers (PLURAL — outlier)
The plural slipped in when addressing Greptile's earlier comment to
avoid `scheduling.cron` (which would imply Cron-only coverage). The
right move was to pick a generic word and stay singular like the rest
of the suite — the suite's mental model is "the X source", one
instance per registry slot, regardless of how many triggers or items
it dispatches internally.
Renamed to `scheduling.trigger`. The `<remarks>` block keeps the
"covers Cron, Timer, StartAt, Delay" explanation and now also
explicitly notes the singular convention so future contributors don't
re-pluralize.
Zero test fallout — the literal "scheduling.triggers" only appeared in
the source file itself. Tests of the other two sources all use
singular forms.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(runtime-admin): rename Force endpoint to ForceDrain
"Force" alone is meaningless out of context — force what? — and it
sits oddly next to the verb-named siblings Pause / Resume / Status.
The matching admin service method is already IWorkflowRuntimeAdminService.
ForceDrainAsync, so ForceDrain is the natural pair.
Renamed:
- src/modules/Elsa.Workflows.Api/Endpoints/RuntimeAdmin/Force/ → ForceDrain/
- namespace ...Endpoints.RuntimeAdmin.Force → ...ForceDrain
- class ForceEndpoint → ForceDrainEndpoint
- class ForceRequest → ForceDrainRequest
- class ForceResponse → ForceDrainResponse
- route POST /admin/workflow-runtime/force → /force-drain
Zero external references — no tests, docs, or OpenAPI clients used the
old symbols or the old route literal, so this is a contained pre-ship
rename. Directory move went through `git mv` so commit history follows
the file.
17/17 graceful-shutdown integration tests still pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(ci): correct formatting in checkout step within greploop.yml
Adjusted indentation of environment variables in the checkout step for improved consistency and readability.
* fix(ci): repair malformed Checkout repository step in greploop.yml
The step accumulated stray env keys, an extra `uses:`, and bash commands
that didn't belong inside it (line 65 onward), causing a YAML parse
error on push. The valid structure has two distinct checkout steps:
- Checkout repository : actions/checkout@v4 with fetch-depth: 0
- Checkout PR branch : env: + run: with SHA validation + git fetch
+ git checkout --detach
The PR-branch step (line 78+) was already correct and unchanged. This
fix restores the first step to its intended single-purpose shape (just
checks out the workflow file's commit so the greploop skill is on disk
before the run-greploop step uses it).
No functional change to runtime behaviour or to the security posture
established in the prior hardening commit (0db4ca23e). The PR-branch
checkout still validates HEAD_SHA / HEAD_REPO_OWNER / HEAD_REPO_NAME
shape before they reach a URL or git command.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(ci): remove literal `${{ }}` from greploop.yml comment
GitHub Actions parses `${{ ... }}` workflow expressions across the entire
YAML file, including inside `run:` script comments. The comment that
explained the env-var hardening pattern contained the literal sequence
`${{ }}` (with a space, intended as an English-language description),
which the expression parser rejected as "An expression was expected"
(line 81 col 14).
Reworded the comment to describe the substitution form in prose without
the literal token sequence. Functional behaviour unchanged.
* refactor(graceful-shutdown): rename burst → execution cycle
The graceful-shutdown work introduced "burst of execution" as a
first-class domain concept. The term arrived without rationale and
isn't standard in the workflow-engine domain. Renamed to
"execution cycle" — reads more naturally as the loop-with-commit unit,
is more idiomatic in workflow vocabulary, pairs cleanly with the
existing WorkflowExecutionContext, and avoids collisions with Elsa's
existing terms (Run, Execution, Dispatch, Invocation, Stimulus, Step).
Renamed types
- IBurstRegistry → IExecutionCycleRegistry
- BurstRegistry → ExecutionCycleRegistry
- BurstHandle → ExecutionCycleHandle
- BurstTrackingMiddleware → ExecutionCycleTrackingMiddleware
- BurstAwareCommitStateHandler → ExecutionCycleAwareCommitStateHandler
Renamed members
- BeginBurst → BeginCycle
- ListActiveBursts → ListActiveCycles
- BurstHandleKey constant + value → ExecutionCycleHandleKey
- ActiveBurstCount (IQuiescenceSignal,
RuntimeAdminStatus, StatusResponse) → ActiveExecutionCycleCount
- WaitForBurstsAsync (private) → WaitForCyclesAsync
- ForceCancelActiveBurstsAsync (priv) → ForceCancelActiveCyclesAsync
- UseBurstTracking → UseExecutionCycleTracking
- DrainOutcomeDto.BurstsForceCancelledCount → ExecutionCyclesForceCancelledCount
- _burstRegistry / burstRegistry → _cycleRegistry / cycleRegistry
Backwards-compatibility preservation (the only persisted JSON key)
- WorkflowInterruptedPayload.BurstDuration property → ExecutionCycleDuration
with [JsonPropertyName("BurstDuration")] so the persisted JSON wire
key stays "BurstDuration" forever. Pre-merge testers' log records
still deserialise correctly. The contract test on
WorkflowInterruptedPayloadContractTests still asserts the wire key
"BurstDuration" appears in the serialised JSON — confirms the
guarantee is enforced.
Other unstructured surfaces
- WorkflowExecutionLogRecord.Message text "Workflow burst was force-
cancelled..." now says "Workflow execution cycle was force-cancelled
..." for new records. Old rows keep their old text — purely cosmetic
free-text field.
- Structured log placeholder {BurstId} in DrainOrchestrator log lines
→ {ExecutionCycleId}.
- Lowercase prose / XML doc comments updated throughout.
Test renames
- BurstRegistryTests → ExecutionCycleRegistryTests
- BurstTrackingMiddlewareTests → ExecutionCycleTrackingMiddlewareTests
- Test method names + DisplayName strings updated.
Spec docs (specs/002-graceful-shutdown/) updated to match the new
vocabulary; the historical task records in tasks.md keep the old names
as-is to preserve the audit trail of what was originally built.
Verification
- dotnet build: clean across net8.0 / net9.0 / net10.0.
- 103/103 Elsa.Workflows.Runtime.UnitTests pass.
- 17/17 GracefulShutdown integration tests pass.
- 4/4 WorkflowInterruptedPayloadContractTests pass — confirms the
"BurstDuration" JSON wire-key preservation is intact.
No changes to migrations or DB column names — confirmed via grep.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(ci): give greploop.yml gh-cli a repo context before checkout
The "Resolve PR head metadata" step runs `gh pr view` before
actions/checkout, so there is no `.git` directory and gh's "current
repo" detection fails with `fatal: not a git repository`. The
prior commits to this file masked the runtime failure because the
workflow itself was YAML-invalid — once it became valid, the
workflow_dispatch trigger surfaced this real-world execution bug.
Set GH_REPO=${{ github.repository }} on both `gh pr view` steps. The
gh CLI honours GH_REPO as an explicit repo override, so it no longer
needs git context. Same fix on the "Checkout PR branch" validation
call which also uses gh pr view before the manual fetch.
The error reported as "Invalid workflow file: ... (Line 81 Col 14)"
on PR #7424 was stale from commit ed580682a (which had the bad
comment with literal `${{ }}`); commit 3a8ad0d08 fixed the YAML, but
because greploop's `if:` condition only matches workflow_dispatch /
issue_comment events, push events on later commits were skipped
without re-running validation, so the GitHub UI kept showing the
old error. A workflow_dispatch run on the current SHA now passes
validation and reaches "Resolve PR head metadata", which is what
this commit fixes.
* fix(graceful-shutdown): wire IngressPauseTimeout option to drain orchestrator
GracefulShutdownOptions.IngressPauseTimeout was documented as "Default
per-ingress-source pause timeout" but DrainOrchestrator.PauseOneSourceAsync
read source.PauseTimeout directly and never consulted the option. The
configured value was silently ignored — operators who set
GracefulShutdownOptions:IngressPauseTimeout = 10s were getting whatever
each source's hardcoded value was (50 ms for the three PassiveIngressSource
subclasses we ship), with no way to tune it globally.
Precedence (per the spec's intent of "overridable at registration and by
configuration"):
1. Per-source positive value wins (source.PauseTimeout > Zero).
2. Otherwise fall back to the configured GracefulShutdownOptions.
IngressPauseTimeout default.
3. Resolved value is capped at the overall drain deadline so a single
misbehaving source cannot exceed the host's shutdown budget.
4. 1 ms safety floor remains so a misconfigured zero default still
produces a non-zero CancelAfter.
Changes:
- DrainOrchestrator.PauseOneSourceAsync — adds the precedence above with
a comment block explaining each step.
- IIngressSource.PauseTimeout — XML doc clarifies the Zero-defers-to-
config semantics.
- GracefulShutdownOptions.IngressPauseTimeout — XML doc says it's the
fallback when the source returns Zero; <remarks> spells out the
precedence and the overall-deadline cap.
- PassiveIngressSource.PauseTimeout — virtual property now returns Zero
(was 50 ms). The three shipped subclasses (HttpTriggerIngressSource,
ScheduledTriggerIngressSource, InternalBookmarkQueueIngressSource)
consequently defer to the configured default — flipping the wire-up
bug from "configured value silently ignored" to "configured value
honoured by default for passive sources". Passive subclasses that
want a specific value can still override.
103/103 runtime unit + 17/17 graceful-shutdown integration tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(graceful-shutdown): rename IInterruptedRecoveryScan → IInterruptedRecoveryScanner
The interface had a single verb-method (`ScanAndRequeueAsync`) and its
XML doc described what it *does* ("Scans the workflow instance store
for instances..."). That's an agent role — a scanner that performs a
scan — but the noun-shaped name `IInterruptedRecoveryScan` read as
"the scan itself", which is misleading because the scan results /
event are not first-class types in the codebase.
Renamed to `IInterruptedRecoveryScanner` / `InterruptedRecoveryScanner`
to match the existing `-er` convention in this codebase (Restarter,
Generator, Resolver, etc.). The method stays `ScanAndRequeueAsync` —
the scanner *performs* a scan-and-requeue.
Also renamed the constructor parameter `scan` → `scanner` in
RecoverInterruptedWorkflowsStartupTask, and the local variable `scan`
→ `scanner` in InterruptedRecoveryIntegrationTests, so the "scanner
does the scan" mental model is consistent throughout.
Surface impact (all internal — no API or persistence touch points):
- 2 source files renamed via git mv (interface + implementation)
- 1 test file renamed (InterruptedRecoveryScanTests → ScannerTests)
- DI registrations in both Features/ and ShellFeatures/ WorkflowRuntimeFeature
- 1 startup-task constructor parameter
- Spec doc references under specs/002-graceful-shutdown/
103/103 runtime unit + 17/17 graceful-shutdown integration tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(graceful-shutdown): extract DrainTriggerExecutor
ElsaShellDrainHandler (CShells IDrainHandler) and
DrainOrchestratorHostedService (.NET IHostedService.StopAsync) inlined
near-identical try/catch/log shapes around IDrainOrchestrator.DrainAsync:
- call DrainAsync(<trigger>, ct)
- branch on outcome: DeadlineExceeded/AbortedByUnhandledException →
Warning, otherwise Information
- catch InvalidOperationException (parallel-drain rejected by the
orchestrator) → log Information and swallow
The two had already drifted: host-stop's success log omitted the
paused/waited durations the shell-handler version included, and the
"skipped" message disagreed on the trigger label ("Host-stop drain
skipped" vs "Shell drain skipped"). Centralised the shape in a small
internal static helper so the two — and any future trigger source —
stay uniform.
Both call sites collapse to a single line. Net diff drops 22 lines from
the two consumers and adds a 25-line helper that they both delegate to.
The unified log copy now consistently includes paused/waited durations
on the success path and uses the caller-supplied contextLabel
("Shell drain", "Graceful drain") in all three messages so operators
can attribute log entries by trigger source.
Files:
- src/modules/Elsa.Workflows.Runtime/Services/DrainTriggerExecutor.cs (new)
- src/modules/Elsa.Workflows.Runtime/Lifecycle/ElsaShellDrainHandler.cs
- src/modules/Elsa.Workflows.Runtime/HostedServices/DrainOrchestratorHostedService.cs
103/103 runtime unit + 17/17 graceful-shutdown integration tests pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(workflows-runtime): drop redundant DrainOrchestratorHostedService from CShells path
In CShells deployments, host stop already drives drain via CShellsStartupHostedService → IDrainHandler →
ElsaShellDrainHandler, scoped per shell (FR-027). The additional .AddHostedService<DrainOrchestratorHostedService>()
in ShellFeatures/WorkflowRuntimeFeature was firing a second non-force DrainAsync that the orchestrator rejected
with InvalidOperationException — silently swallowed by DrainTriggerExecutor, but logged on every host stop and
semantically muddled (IHostedService is host-level, not per-shell).
Keep the registration on the IModule path (Features/WorkflowRuntimeFeature) where there is no shell platform
and host-stop is the only available drain trigger. Update ElsaShellDrainHandler XML docs to reflect the now-clean
single-trigger model in CShells.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(workflows-runtime): always dispose ExecutionCycleHandle in tracking middleware
Previously the success path relied on ExecutionCycleAwareCommitStateHandler to dispose the handle after the
runner's commit completed. If a custom dispatcher or test double exited the pipeline without invoking commit
(by design or by accident), the handle stayed registered, IExecutionCycleRegistry.ActiveCount never reached
zero, and drain spun in WaitForExecutionCyclesAsync until the deadline fired — incorrectly force-cancelling
instances that had already finished cleanly.
Collapse the existing try/catch(rethrow) into try/finally so the middleware itself disposes the handle for
both exception and commit-elided paths. Disposal remains idempotent via the ExecutionCycleHandle._disposed
Interlocked guard, so the normal-path dispose by ExecutionCycleAwareCommitStateHandler is a harmless no-op.
Adds an integration regression test that drives the middleware with a stub Next that returns without invoking
commit and asserts ActiveCount returns to zero.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(workflows-runtime): serialize QuiescenceSignal pause-state persistence
Both PauseAsync and ResumeAsync used to release the inner lock before issuing the persistence I/O. A rapid
Pause → Resume sequence could leave the persisted state inconsistent: PauseAsync's slow SaveAsync could land
AFTER ResumeAsync's DeleteAsync, leaving the key present in the store while in-memory state was None. On host
restart, InitializePersistedStateAsync would find the stale key and start the runtime in the paused state
the operator had already cancelled.
Introduce a dedicated SemaphoreSlim that serializes persistence I/O, with each I/O re-reading the live
in-memory state inside the semaphore. N racing Pause/Resume calls now produce N serialized writes, each
reflecting the most recent in-memory transition — so the final persisted state always matches final
in-memory state.
Adds a regression test that gates SaveAsync, races a Resume behind it, and asserts the store is empty after
both complete.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(workflows-runtime): trim verbose comment in ExecutionCycleTrackingMiddleware finally block
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(workflows-runtime): use 'using var' for ExecutionCycleHandle in tracking middleware
Replace the explicit try/finally that only existed to call handle.Dispose() with a `using var` declaration —
identical semantics (compiler-emitted finally with idempotent dispose), more idiomatic. The regression test
HandleReleasedWhenCommitIsElided continues to validate the leak-free property.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(workflows-runtime,api): proper 409 conflict shape + shell-scoped pause-persistence key
ForceDrain endpoint: the 409 path returned `new ForceDrainResponse()` whose `Outcome` was null at runtime
despite the `= null!` annotation, so any strongly-typed client deserializing the conflict body and reading
`Outcome.OverallResult` got an NRE. Switch to the existing `ConflictResponse` shape with
`Code = "DrainInProgress"` and the current runtime status. Routed via HttpContext.Response.WriteAsJsonAsync
because Send.ResponseAsync is constrained to the endpoint's TResponse and cannot send a sibling DTO.
QuiescenceSignal persistence key: the DI-registered `IQuiescenceSignal` was constructed with
`shellName = null` (DI doesn't inject `string?` defaults), so every shell shared the key
`elsa.quiescence.pause.default`. In a CShells multi-shell deployment under
PausePersistencePolicy.AcrossReactivations this caused cross-shell contamination — pausing shell A would
re-pause shell B on its next activation. Replace the simple AddSingleton<IQuiescenceSignal,...> registration
in ShellFeatures/WorkflowRuntimeFeature with a factory that injects `CShells.ShellSettings` and forwards
`Settings.Id` as the shell name. The IModule registration is unchanged (no shell platform; null shellName
remains correct there).
Adds a unit regression test that two QuiescenceSignal instances with different shellNames write to disjoint
persistence keys and never to the legacy "default" key.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(quiescence): use TryGetValue for ContainsKey+indexer assertions
Combines existence check and value retrieval into a single dictionary lookup, addressing the code-quality
bot's repeated suggestion. No behavior change — both PauseWritesKey and PersistenceKeyIncludesShellName
still assert the same keys exist with the same content.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Update specs/002-graceful-shutdown/contracts/admin-endpoints.md
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* fix(workflows-runtime): decouple QuiescenceSignal persistence from caller cancellation
PersistAsync used to forward the caller's CancellationToken to both _persistenceMutex.WaitAsync and the
store I/O. If an HTTP request was cancelled between the in-memory transition (already committed under
_sync) and the persistence call, the I/O was silently skipped — leaving AdministrativePause set in memory
with no persisted record. The idempotent fast-path on subsequent PauseAsync calls (transitioned == false)
meant no retry would happen, so on host restart InitializePersistedStateAsync would find no key and the
runtime would come back unpaused, defeating PausePersistencePolicy.AcrossReactivations.
Drop the parameter from PersistAsync entirely; use CancellationToken.None for both the semaphore wait and
the store I/O. The in-memory transition is already committed by the time PersistAsync runs, so persistence
must complete to keep the store consistent with memory. The public PauseAsync/ResumeAsync methods still
accept a CancellationToken (interface contract) — it just no longer reaches the persistence layer.
Adds a regression test that calls PauseAsync with a pre-cancelled token and asserts both in-memory pause
and the persisted key land correctly.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(workflows-runtime,api,docs): apply Copilot review feedback batch
Code:
- DrainOrchestrator.TryForceStopAsync now bounds force-stop with the *remaining* drain budget
(deadlineAt - now), not the full overall TimeSpan. A per-source pause that already burned the
shutdown window can no longer get another full deadline's worth of force-stop runway.
- DrainOrchestrator catch filter narrowed: drop `ex is not InvalidOperationException` exclusion.
The "drain already in progress / completed" IOEs are thrown outside the protocol's try block,
so they bubble out without entering this handler. Any IOE that lands here is incidental
(e.g., from a store inside the drain) and should now be captured into the outcome rather
than escaping the whole drain.
- ResumeEndpoint 409 path now returns the discriminated ConflictResponse shape (matching
ForceDrain) instead of a plain StatusResponse. Routed via HttpContext.Response.WriteAsJsonAsync
because Send.ResponseAsync is constrained to TResponse.
- Conflict codes aligned to kebab-case across both endpoints to match the contract spec
(`runtime-draining` and `drain-in-progress`).
Spelling sweep — American English per constitution v1.0.1 III:
- DrainOrchestrator.cs: "serialised" → "serialized"
- WorkflowInterruptedPayload.cs: "serialised" / "deserialise" → "serialized" / "deserialize"
- PassiveIngressSource.cs: "behaviour" → "behavior"
- DeadlineBreachEndToEndTests.cs: "serialisable" → "serializable"
- specs/002-graceful-shutdown/quickstart.md: "behaviour" → "behavior"
- specs/002-graceful-shutdown/checklists/requirements.md: "behaviour" → "behavior"
Doc/contract alignment:
- quickstart.md: force route corrected from /force to /force-drain.
- quiescence-signal.md: removed StateChanged event from contract (interface doesn't define it);
corrected persistence section to describe InitializePersistedStateAsync via shell initializer
/ startup task rather than constructor read; added the per-shell key discriminator.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(workflows-runtime): correct DI lifetimes for ExecutionCycleTrackingMiddleware and WorkflowRuntimeAdminService
Two strict-DI-validation failures surfaced in tests using BuildServiceProvider with validate-on-build:
1. ExecutionCycleTrackingMiddleware was registered as AddSingleton<>, but its constructor takes
WorkflowMiddlewareDelegate next — supplied by the workflow execution pipeline builder via
UseMiddleware<>(), not from DI. The registration was both unused (no consumer resolves it through
the container) and broken (DI fails to construct it because next is unregistered). Removing both
registrations.
2. IWorkflowRuntimeAdminService was registered as AddSingleton<> but depends on the scoped
INotificationSender (mediator) — captive-dependency violation. All consumers (Pause/Resume/
Status/ForceDrain endpoints) are FastEndpoints, which are scoped per request, so AddScoped is
the correct alignment. The other deps (IQuiescenceSignal / IIngressSourceRegistry /
IDrainOrchestrator / ISystemClock) are singletons and resolve fine from a scoped consumer.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(workflows-runtime): restore commit-handler-only success disposal + true Cancel idempotency
Two issues raised by Copilot's latest review on commit 32a9c0519:
1. ExecutionCycleTrackingMiddleware was disposing the handle at the end of InvokeAsync (via `using var`),
but WorkflowRunner runs commit AFTER the pipeline returns (WorkflowRunner.cs:235). That meant the handle
was disposed BEFORE the runner's terminal commit, and the drain orchestrator's
`await handle.Disposed` would unblock too early — reintroducing the runner-clobber race the original
design protected against (see ExecutionCycleAwareCommitStateHandler XML doc).
Revert to the original shape: only dispose on exception path. ExecutionCycleAwareCommitStateHandler
remains the SOLE success-path disposer, running in its finally block AFTER the inner commit lands.
The earlier "leak when commit is elided" concern was a non-issue in production (the standard runner
always commits); the buggy `HandleReleasedWhenCommitIsElided` test that specified the wrong contract
is removed. The existing `ActiveCountReturnsToZero` test (which uses the real runner end-to-end)
already verifies success-path disposal.
2. ExecutionCycleHandle.Cancel() was documented as idempotent but only short-circuited via the
_disposed flag. Repeated Cancel() calls before Dispose could trigger the cancel callback multiple
times — easy to accidentally fire non-idempotent cancellation side effects more than once. Add an
Interlocked _cancelled guard so callback + CTS cancellation run at most once. Existing test that
documented the leaky behavior is updated to assert the now-truly-idempotent contract.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Update logging levels and remove unused features in appsettings files
* refactor(workflows-runtime): improve graceful shutdown options handling and cleanup solution
Refactor the handling of `GracefulShutdownOptions` to ensure options are applied correctly without directly invoking the delegate. Update DI registrations to use appropriate lifetimes and remove redundant wrapper services. Additionally, clean up the solution by removing unused projects and documentation folders.
* feat(identity, workflows-runtime): add validation for identity and graceful shutdown options
Introduce validation capabilities for `IdentityTokenOptions` and `GracefulShutdownOptions`. Implement extension methods for option validation, enhance service registration, and add unit tests to ensure configurations are validated at startup. Update solution to include new unit test projects.
* update(docs): clarify shutdown log message expectations and levels in quickstart.md
Optimize explanation of expected log message sequence during graceful shutdown and specify logging levels.
* docs: amend constitution to v1.1.0 (SRP, DRY, KISS, conciseness under Principle VII)
* refactor(multitenancy): rename and restructure TenantTaskManager to TenantTaskLifecycleCoordinator
Rename `TenantTaskManager` to `TenantTaskLifecycleCoordinator` and relocate to a new directory structure, enhancing code organization and test consistency. Retain functional behaviors with no logic alterations. Update unit tests to reflect the naming changes, ensuring consistency with the refactored code structure.
* Update logging levels and dependencies
- Set default logging level to Debug in appsettings.Development.json
- Add missing using directives for Elsa workflows management and runtime features
- Update CShells package versions to 0.0.18-preview.104
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
When using BulkDispatchWorkflows with the Input property set, all dispatched child workflows received the same dictionary reference. This caused the input to be mutated across iterations, resulting in all child workflows seeing the last item's value instead of their own distinct values.
The fix creates a copy of the base input dictionary for each dispatch iteration, ensuring each child workflow receives its own isolated input.
* Add integration tests for enumerables with projection to arrays
Introduce integration tests to verify the conversion of a Select-projected IEnumerable to an array using JavaScript execution. This includes creating new test cases, activities, and workflows to ensure the end-to-end process behaves as expected. Also refined type inference in `ExpressionExecutionContextExtensions` for projection operators.
* Remove unused newline in EnumerableProjectionTests
* Enable multitenancy support and normalize tenant ID handling.
- Activate multitenancy in `Program.cs`.
- Introduce `NormalizeTenantId` method for consistent tenant ID usage.
- Update tenant-related classes and features to support normalization logic.
* Add ADR for adopting empty string as the default tenant ID
- Standardized the tenant ID for the default tenant to use an empty string (`""`) instead of `null`.
- Documented the rationale and migration considerations in ADR 0007.
- Updated ADR table of contents and graph for new entry.
* Apply suggestion from @sfmskywalker
* Update doc/adr/graph.dot
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Normalize spacing and improve readability in `Program.cs`. Fix multitenancy condition formatting.
* Fix ADR numbering and update TOC
* Add ADRs for flowchart execution model, tenant deletion event, merge modes, and default tenant ID
- Introduced ADR 0005: Token-centric flowchart execution model for improved loop and join handling.
- Added ADR 0006: Tenant Deleted event for distinct handling of tenant removal.
- Documented ADR 0007: Explicit merge modes for flowchart joins, improving reliability and configurability.
- Included ADR 0008: Standardization of empty string as the default tenant ID for consistency and clarity.
* Add unit tests for tenant ID normalization and multitenancy pipeline invoker
- Added comprehensive unit tests for tenant ID normalization to ensure consistent handling of null, empty, and valid IDs.
- Introduced tests for the multitenancy pipeline invoker covering various tenant resolution scenarios.
- Updated solution to include new unit testing projects for `Elsa.Tenants` and `Elsa.Common`.
* Update unit tests for `ActivityConstructionResult`
- Refactor test parameterization to verify `HasExceptions` property more explicitly.
- Simplify exception creation logic in helper methods.
- Improve test assertions by combining act and assert phases where applicable.
* Enable configuration-based multitenancy with tenant-specific settings
- Introduced a configuration-based tenant provider to streamline tenant initialization and customization.
- Added tenant ID handling filters to ensure tenant ID is applied and filtered automatically.
- Deprecated the `CommonPersistenceFeature` in favor of modular persistence feature extension.
* Update database indexes to include `TenantId` for multitenancy support
- Added `TenantId` to unique constraints on `Triggers` table across all EFCore providers.
- Adjusted index names to reflect the updated constraints.
- Updated trigger configuration to ensure uniqueness includes `TenantId`.
* Add tenant filtering to `DefaultWorkflowDefinitionStorePopulator`
- Introduced `ITenantAccessor` to support tenant-specific filtering of workflow definitions.
- Updated logic to skip workflows not matching the current tenant.
* Update doc/adr/toc.md
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Remove `CommonPersistenceFeature` as it has been deprecated
* Add tenant-specific filtering to workflow import logic in `DefaultWorkflowDefinitionStorePopulator`
* Replace hardcoded tenant ID with `Tenant.DefaultTenantId` in integration tests
* Update database indexes and migration logic to support `TenantId` for multitenancy
- Added `TenantId` to unique constraints on the `Triggers` table and updated index names.
- Included logic to drop outdated indexes without `TenantId` during migration.
- Adjusted tests to account for `TenantId` in workflow identity and indexing scenarios.
* Remove `TenantId` from workflow identity construction in concurrent trigger indexing tests
* Introduce `SelectiveMockLockProvider` for precise lock mocking in tests
- Added `SelectiveMockLockProvider` to allow targeted lock mocking without affecting unrelated background operations.
- Updated test services to use `SelectiveMockLockProvider` in place of `TestDistributedLockProvider`.
- Refactored `DistributedLockResilienceTests` to support selective mocking for deterministic and reliable assertions.
* Update Elsa.sln
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Normalize tenant ID handling in `DefaultWorkflowDefinitionStorePopulator` for consistent filtering
* Refactor `TenantResolverResult` to support explicit resolved/unresolved state handling
- Updated `TenantResolverResult` to include an explicit `_isResolved` property.
- Adjusted `ResolveTenantId()` and `IsResolved` logic for improved clarity and robustness.
- Simplified tenant resolution invocation in `TenantResolverBase`.
- Removed redundant normalization in `DefaultTenantResolverPipelineInvoker`.
* Normalize tenant ID handling in `DefaultWorkflowDefinitionStorePopulator` and `ClrWorkflowsProvider`.
* Refactor `DefaultWorkflowDefinitionStorePopulatorTests`: streamline object initializations and add tenant-specific test coverage for `PopulateStoreAsync`.
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Add retry mechanism for distributed locks with transient error handling and logging
- Introduced Polly-based retry pipeline for distributed lock acquisition in `DistributedWorkflowClient` to handle transient errors such as network issues or database connection failures.
- Added detailed logging for retry attempts and lock release errors.
- Updated project dependencies to include Polly.
* Refactor transient exception handling to shared resilience module.
Migrated transient exception detection logic from scheduling module to a new shared resilience module. Updated services, jobs, and features to utilize the centralized `ITransientExceptionDetectionService`. This change improves maintainability and promotes reusability across modules.
* Add unit tests for transient exception detection and resilience strategy evaluation.
- Introduced comprehensive unit tests for `DefaultTransientExceptionDetector`, `ResilienceStrategyCatalog`, `ResilienceStrategyConfigEvaluator`, and `TransientExceptionDetectionService`.
- Added helper classes and test data factories to facilitate reusable test patterns for resilience modules.
- Updated solution to include `Elsa.Resilience.Core.UnitTests` project.
* Add component tests for distributed lock resilience
- Introduced new tests to verify retry behavior during transient lock acquisition and release failures.
- Added `TestDistributedLockProvider` and related mocks for simulating transient failures.
- Updated `WorkflowServer` test services to support the new distributed lock test scenarios.
* Refactor distributed lock resilience tests
- Consolidated test logic: streamlined test providers, injected services, and reusable test patterns.
- Simplified `TestDistributedLockProvider` implementation with enhanced initialization and failure simulation.
- Reorganized tests for transient acquisition/release failures to use parameterized `Theory` for improved maintainability.
* Refactor transient exception handling: rename interfaces and classes for consistency, update references across codebase, and improve code readability.
* Apply suggestion from @Copilot
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Simplify WorkflowServer setup and DistributedLockResilienceTests by replacing IDistributedLockProvider with TestDistributedLockProvider.
* Remove unused `using` directives in unit tests to improve code cleanliness.
* Remove `TransientExceptionTypes` helper and inline its usage in tests for improved maintainability.
* Add descriptive `DisplayName` attributes to unit tests for improved test clarity.
* Fix redundant exception checking in TransientExceptionDetector (#7162)
* Initial plan
* Fix redundant exception checking in TransientExceptionDetector
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Extract MaxRetryAttempts constant in DistributedLockResilienceTests (#7164)
* Initial plan
* Extract MaxRetryAttempts constant to eliminate hardcoded magic number
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Make TestDistributedLockProvider thread-safe with Interlocked operations (#7163)
* Initial plan
* Make TestDistributedLockProvider thread-safe using Interlocked operations
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Update src/modules/Elsa.Workflows.Runtime.Distributed/Services/DistributedWorkflowClient.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update test/component/Elsa.Workflows.ComponentTests/Scenarios/DistributedLockResilience/DistributedLockResilienceTests.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update src/modules/Elsa.Workflows.Runtime.Distributed/Services/DistributedWorkflowClient.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update src/modules/Elsa.Resilience.Core/Services/DefaultTransientExceptionStrategy.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Include `CancellationToken` in distributed lock handling methods for improved cancellation support.
* Initial plan (#7168)
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
* Cache detector list in TransientExceptionDetector to avoid repeated allocations (#7167)
* Initial plan
* Cache detector list in field to avoid repeated allocations
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Use IReadOnlyList instead of List for better intent expression
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Fix TestDistributedLockProvider registration to properly decorate IDistributedLockProvider (#7166)
* Initial plan
* Fix TestDistributedLockProvider registration to use Decorate pattern and fix variable reference bug
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Add runtime check for TestDistributedLockProvider registration
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Refactor `DistributedWorkflowClient` to simplify `Lazy<ResiliencePipeline>` initialization.
* Add integration tests for DistributedWorkflowClient lock resilience (#7165)
* Initial plan
* Fix compilation error: use correct parameter name transientExceptionDetector
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Add integration tests for DistributedWorkflowClient lock resilience
- Add SimpleWorkflow for testing distributed lock scenarios
- Add tests exercising RunInstanceAsync with transient lock failures
- Verify retry logic works correctly with actual workflow execution
- Test both acquisition and release failure scenarios
- Decorate IDistributedLockProvider to use TestDistributedLockProvider
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Address code review feedback
- Add explanatory comment for TestDistributedLockProvider cast
- Remove unnecessary blank line for consistent formatting
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
Co-authored-by: Sipke Schoorstra <sipkeschoorstra@outlook.com>
* Simplify transient exception strategy by refactoring message pattern matching logic.
* Enhance distributed lock mock to support per-lock failure configuration and improve resilience tests.
* Refactor `TestDistributedLockProvider` to streamline failure handling logic and improve code clarity.
* Remove unused methods and redundant test case from `DistributedLockResilienceTests`.
* Refactor `DistributedLockResilienceTests` to simplify workflow client creation, consolidate assertion logic, and remove redundant test cases.
* Format `ResilienceStrategyCatalogTests` by removing redundant line breaks in test setup.
* Refactor `TransientExceptionDetectorTests` to simplify test setup, consolidate test cases, and remove redundant logic.
* Handle `InvalidOperationException` in `XunitLogger` to suppress logging errors during inactive tests.
* Update workflows to use .NET 10 and adjust resilience tests project configuration.
* Refactor distributed runtime feature and integrations to improve resilience handling, configure services fluently, and add cancellation safeguards in background services.
* Update `base_version` to `3.7.0` in GitHub workflow configuration.
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Initial plan
* Add Literal handling back to ActivityExecutionContext.TryGet
This restores support for dynamic Literal inputs that was removed in version 3.5.2.
When an Input is created with a Literal, the Literal becomes the MemoryBlockReference.
Since Literals hold values directly rather than in the memory register, they need
special handling in TryGet to return their value.
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Improve test coverage for Literal handling in ActivityExecutionContext
- Removed ineffective unit tests that only checked type relationships
- Added explicit integration test for TryGet with Literal references
- Added integration test for Get with Input<T> containing Literal
- All tests now directly verify the fixed TryGet behavior
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
- Added TaskCompletionSource fields to Spy class for each notification type
- Test handlers now signal completion through TaskCompletionSource
- Tests await TaskCompletionSource instead of arbitrary delays
- Eliminates timing-dependent flakiness in CI systems
Co-authored-by: KnibbsyMan <23156317+KnibbsyMan@users.noreply.github.com>
- Created WorkflowDefinitionDispatching and WorkflowDefinitionDispatched notifications
- Created WorkflowInstanceDispatching and WorkflowInstanceDispatched notifications
- Updated BackgroundWorkflowDispatcher to emit notifications before and after dispatch
- Added integration tests to verify notifications are emitted correctly
Co-authored-by: KnibbsyMan <23156317+KnibbsyMan@users.noreply.github.com>
- Introduced methods to set default workflow and activity commit strategies in CommitStrategiesFeature.
- Updated CommitStateOptions to include properties for default strategies.
- Added extension methods for configuring default strategies in WorkflowsFeature.
- Created usage examples demonstrating how to set and utilize default commit strategies.
- Implemented tests to verify default strategy behavior in various scenarios.
* Introduce `FlowchartExecutionMode` to streamline flowchart execution logic.
- Added `FlowchartExecutionMode` enum to represent execution modes (Default, TokenBased, CounterBased).
- Updated flowchart-related integration and unit tests to use the new execution mode.
- Removed the global `UseTokenFlow` flag in favor of execution-specific configuration via `RunWorkflowOptions`.
- Refactored flowchart-related APIs and test helpers for improved flexibility and modularity.
* Add support for configuring flowchart execution behavior via `FlowchartOptions` and DI.
- Introduced extensions for `FlowchartFeature` to simplify configuration.
- Added DI support for setting default execution modes.
- Refactored flowchart execution logic to prioritize configuration.
* Update default Flowchart execution settings to align with version 3.5.2 behavior
- Changed `DefaultExecutionMode` to `CounterBased`.
- Updated `UseTokenFlow` default to `false`.
* Update src/modules/Elsa.Workflows.Core/Activities/Flowchart/Models/FlowchartExecutionMode.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update src/modules/Elsa.Workflows.Core/Activities/Flowchart/Extensions/FlowchartFeatureExtensions.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update test/integration/Elsa.Workflows.IntegrationTests/Scenarios/FlowchartNextActivity/Tests.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update src/modules/Elsa.Workflows.Core/Activities/Flowchart/Extensions/RunWorkflowOptionsExtensions.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Refactor flowchart integration tests for improved formatting and consistency
* Refactor workflow tests and related services to improve reusability and align with updated Flowchart execution behavior
* Update src/modules/Elsa.Workflows.Core/Activities/Flowchart/Options/FlowchartOptions.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Refactor flowchart execution logic to use `FlowchartExecutionMode` enum, replacing boolean checks for improved clarity and extensibility.
* Refactor tests and workflow logic to replace boolean `useTokenFlow` with `FlowchartExecutionMode` enum for clarity and consistency.
* Refactor flowchart execution logic to centralize mode-based behavior handling and simplify implementation.
* Remove unnecessary whitespace in Flowchart.cs to improve code formatting
* Remove unnecessary whitespace in FlowJoinTests.cs to improve code formatting
* Update src/common/Elsa.Testing.Shared.Integration/WorkflowTestFixture.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update src/modules/Elsa.Workflows.Core/Activities/Flowchart/Activities/Flowchart.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Refactors and clarifies Flowchart merge modes
Improves the clarity and functionality of Flowchart merge modes by:
- Renaming `None` to `Stream` for opportunistic execution.
- Introducing `Merge` for waiting on activated branches only.
- Enhancing `Converge` to be the strictest mode, requiring all inbound connections.
- Providing more detailed descriptions for each mode, emphasizing their behavior and use cases in flow-based terminology.
- Updates default merge mode to Stream
This provides better control over synchronization and execution behavior in workflows.
* Refactor `ActivityExtensions` to improve formatting, fix indentation, and align comments for improved readability and consistency
* Update flowchart tests: replace `SetMergeMode(MergeMode.None)` with `SetMergeMode(null)` and remove unused `Elsa.Workflows` imports.
* Restore None value for Flowchart MergeMode enum (#7107)
* Initial plan
* Add None value to MergeMode enum for backward compatibility
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Fix API client GetMergeMode to maintain non-nullable return type
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* [WIP] Address feedback on flowchart merge modes refactor (#7106)
* Initial plan
* Fix misleading documentation for Merge mode to match actual implementation
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Update test/integration/Elsa.Workflows.IntegrationTests/Scenarios/JoinBehaviors/ForkDecisionJoinTests.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update test/integration/Elsa.Workflows.IntegrationTests/Scenarios/JoinBehaviors/ForkDecisionJoinTests.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update test scenarios for implicit join behavior in `ForkDecisionJoinTests`. Updated file references for merge and stream join modes.
---------
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update packages.yml
* Update elsa-server-and-studio.yml
* Update elsa-server.yml
* Update elsa-studio.yml (#6715)
* Update ListWorkflowDefinitionsRequest.cs (#6761)
Remove unnecessary line breaks
* Correct namespace and import for `ConfigureEngineWithVariableTypes`.
* Resolves build issues, update package versions and restructure project references
- Updated multiple package versions in `Directory.Packages.props` for better dependency management, including `BenchmarkDotNet`, `FastEndpoints`, and `Microsoft.Extensions.Http.Resilience`.
- Minor version upgrade for `System.Formats.Asn1` in `_build.csproj`.
- Replaced project reference to `Elsa.csproj` with `Elsa.IO.Http.csproj` in `Elsa.ServerAndStudio.Web.csproj`, enhancing modularity.
- Added new using directive for `Elsa.IO.Http.Features` in `Program.cs` to support new HTTP functionalities.
* Remove unused project references from Elsa.sln
These changes indicate that the associated projects or dependencies are no longer needed or have been replaced by other components in the solution.
* Rename copilot-setup-steps.yml.yml to copilot-setup-steps.yml
* Update RawStringContent encoding in JsonContentFactory (#6786)
* Update RawStringContent encoding in JsonContentFactory
Modified the instantiation of `RawStringContent` to use a
new `UTF8Encoding` instance with `encoderShouldEmitUTF8Identifier`
set to `false`, affecting the handling of the UTF-8 byte order
mark (BOM) in serialized JSON content. Fixes a bug with content length being different than expected.
* Refactor JsonContentFactory to reuse UTF8Encoding
Introduced a private static readonly field `_utf8Encoding` in the `JsonContentFactory` class to improve code readability and performance. This change replaces the instantiation of `UTF8Encoding` in the `CreateHttpContent` method, allowing for the reuse of the same encoding instance.
---------
Co-authored-by: Max Brooks <Max@compyl.com>
* Enhance thread safety with ConcurrentDictionary usage (#6760)
* Enhance thread safety with ConcurrentDictionary usage
Replaced `IDictionary` with `ConcurrentDictionary` for
both `_scheduledTasks` and `_scheduledTaskKeys` to
improve thread safety in a multi-threaded environment.
Updated methods `RegisterScheduledTask`,
`RemoveScheduledTask`, and `RemoveScheduledTasks` to
utilize the `Remove` method of `ConcurrentDictionary`,
ensuring safe and efficient removal of scheduled tasks.
* Refactor task registration and removal logic
Updated `RegisterScheduledTask` to use `AddOrUpdate` for streamlined task management. This change simplifies the addition and updating of scheduled tasks by consolidating logic into a single operation. Introduced `RemoveScheduledTask` method to handle task removal by name, improving code organization and clarity.
* Improve task removal handling in LocalScheduler
Modified the `LocalScheduler` class to enhance the removal process of scheduled tasks from the `_scheduledTaskKeys` collection. The removal operation now captures the result in a variable and includes a conditional check to log a warning if the task was not found, improving error handling and debugging capabilities.
* Refactor task removal in LocalScheduler
Updated the removal process for scheduled tasks in `_scheduledTasks`.
The new implementation collects all corresponding keys and attempts to remove them individually, logging warnings for any failures. This enhances error handling and provides better debugging information.
---------
Co-authored-by: Max Brooks <Max@compyl.com>
* Add IAsyncEnumerable check to ItemSourceActivityExecutionContextExtensions.GetItemSource (#6897)
* Use FullName in WorkflowDictionary (#6923)
* Fixed ParentWorkflowInstanceId not being set (#7029)
Co-authored-by: Peter Klooster <peter.klooster@autotaalglas.nl>
* Remove unused solution projects and update package references
- Deleted several project references from `Elsa.sln` to clean up the solution.
- Updated `Directory.Packages.props` for consistency and alignment with the latest package versions.
* Simplify CI pipeline by removing `Test` step from `Compile+Test+Pack` process.
* Initial plan
* Add ElsaScript DSL module with parser and compiler
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Add integration tests for ElsaScript DSL
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Add comprehensive documentation for ElsaScript DSL
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Refactor workflow activity instantiation logic
- Removed `ActivityFactory` and its related interfaces and extensions.
- Introduced `ActivityActivator` for handling activity creation.
- Extended AST with support for comprehensive workflow structures:
- Added nodes for flowcharts, if/else, loops, and variable declarations.
- Updated `IElsaScriptCompiler` to use asynchronous methods.
- Expanded `ElsaScriptParser` to simplify syntax for `UseNode` and argument parsing.
- Adjusted compiler and parser for compatibility with new workflow AST model.
* Refactor test method names for clarity and add new compiler and parser tests
- Updated method names in `CompilerTests` and `ParserTests` for better readability and description of test intent.
- Added tests for compiler and parser:
- Support for workflows without the `workflow` keyword.
* Refactor `ElsaScriptParser` to improve statement parsing and introduce a tokenizer
- Added `TokenizeStatements` method to split source into statements for enhanced parsing accuracy.
- Updated logic to process statements instead of raw lines, reducing parsing complexity and improving reliability.
- Improved handling of workflow and statement parsing, including edge cases with braces, parentheses, and string literals.
* Introduce ElsaScript support for BlobStorage workflow provider
- Added the `Elsa.WorkflowProviders.BlobStorage.ElsaScript` module to enable ElsaScript-based workflow definitions for BlobStorage.
- Implemented `ElsaScriptBlobWorkflowFormatHandler` for parsing ElsaScript workflows stored in BlobStorage.
- Extended `ElsaScriptParser` to leverage Parlot for improved DSL parsing.
- Introduced `IBlobWorkflowFormatHandler` to centralize workflow format handling and parsing.
- Updated `Elsa.Server.Web` to reference the new module and include an ElsaScript "Hello World" example workflow.
* Refactor ElsaScript services, update logging, and improve workflow handling
- Changed `ElsaScriptCompiler` service registration from `Singleton` to `Scoped` for better dependency management.
- Enhanced the "Hello World" example workflow and added `CopyToOutputDirectory` configuration.
- Removed unused namespaces and adjusted references in multiple projects to improve maintainability.
- Updated logging levels in `appsettings.json` to reduce unnecessary debug output.
- Improved `PolymorphicObjectConverter` by removing redundant dependencies.
- Added missing references to enhance feature support and ensure compatibility.
* Refactor activity instantiation and improve argument handling in `ElsaScriptCompiler`
- Added support for positional arguments with constructor matching logic.
- Refactored `InstantiateActivityUsingConstructor` to enhance activity creation.
- Updated `ActivityDescriptor` and related types to include `ClrType` for streamlined activity resolution.
- Simplified `TypedActivityProvider` by annotating it with `[UsedImplicitly]`.
- Adjusted `ElsaScriptParser` to remove unnecessary options from string literal definitions.
* Add HTTP-enabled "Hello World" workflow and support for additional HTTP activity constructors
- Introduced a new ElsaScript example workflow `hello-world-http.elsa` with an HTTP endpoint and response.
- Enhanced `HttpEndpoint` and `WriteHttpResponse` activities with additional constructors for improved flexibility.
- Updated project to include the new workflow in the output directory.
* Enhance `ElsaScriptParser` with a custom parser to handle nested raw expressions for ElsaScript workflows
- Introduced `RawExpressionParser` to parse raw text after `=>` up to a matching closing parenthesis.
- Updated `elsaExpressionWithLang` and `elsaExpressionWithoutLang` to use `RawExpressionParser`.
- Trimmed whitespace in parsed expressions.
- Added integration and parser tests for complex workflows with variables and expressions.
- Updated example workflow `hello-world-http.elsa` to demonstrate expression usage.
- Added `Elsa.Http` module reference to enable HTTP-based activities.
* Update "Hello World" workflow to simplify naming and enhance response logic
- Renamed workflow from `HelloWorldHttpDsl2` to `HelloWorldHttpDsl`.
- Updated HTTP endpoint path to `/hello-world-dsl` for consistency.
- Improved response logic by utilizing `getMessage()` JavaScript function.
* Add support for `OriginalSource` in workflow materialization and enhance ElsaScript materializer
- Introduced `OriginalSource` property in `WorkflowDefinition` and `MaterializedWorkflow` for preserving original source representation (e.g., ElsaScript, JSON, YAML).
- Added `ElsaScriptWorkflowMaterializer` implementation to materialize workflows directly from ElsaScript source.
- Updated `DefaultWorkflowDefinitionStorePopulator` to determine `StringData` or `OriginalSource` based on materialized workflow format.
- Enhanced `WorkflowDefinitionMapper` to support symmetric round-tripping with `OriginalSource`.
- Registered `ElsaScriptWorkflowMaterializer` in `ElsaScriptFeature` for dependency injection.
- Updated `JsonBlobWorkflowFormatHandler` and added `OriginalSource` support for round-trip preservation.
- Simplified `ElsaScriptParser` by aligning variable and parser naming.
* Update V3_6 migrations for PostgreSQL, MySQL, and Oracle databases and associated designer files.
* Handle disposal and race conditions in `ScheduledCronTask`
- Added `_disposed` flag to prevent accessing disposed resources.
- Updated `_executionSemaphore` and `_scopeFactory` logic to safely handle `ObjectDisposedException`.
- Enhanced task scheduling and timer disposal with additional safeguards against race conditions.
- Modified tests to ensure proper disposal and logging behavior when handling edge cases.
* Add support for metadata in ElsaScript workflows and enhance parser and compiler functionality
- Introduced metadata syntax in ElsaScript workflows (e.g., `DisplayName`, `Description`, `Version`) to enable metadata-driven behavior.
- Enhanced `ElsaScriptCompiler` to process metadata and properly integrate it into `Workflow` objects.
- Updated `ElsaScriptParser` to parse program-level AST with support for multiple workflows and global use statements.
- Refactored tests to validate metadata parsing and ensure backward compatibility with existing workflows.
- Added new test cases to cover scenarios like metadata parsing, compilation, and multi-workflow programs.
* Add support for `foreach` loops in ElsaScript and remove `let` keyword
- Introduced `foreach` loop syntax in `ElsaScriptParser` and `ElsaScriptCompiler`, enabling iteration over collections with optional variable declaration.
- Updated `ForNode` and `ForEachNode` to include a `DeclaresVariable` flag for improved variable handling.
- Removed support for the `let` keyword in variable declarations, streamlining syntax to use `var` and `const` only.
- Enhanced `for` loop syntax to support optional `var` declaration and block or single-statement bodies.
- Refactored test cases to validate `foreach` and `for` loop enhancements and ensure backward compatibility.
* Simplify ElsaScript workflow syntax by removing redundant quotes in workflow identifiers and updating `for` loop syntax for clarity and consistency.
* Remove redundant quotes from workflow identifiers in integration tests
* Simplify Elsa scripts and improve error handling
- Removed redundant braces in workflow declarations for streamlined syntax.
- Enhanced logging in `JsonBlobWorkflowFormatHandler` and `ElsaScriptBlobWorkflowFormatHandler` to warn on parsing errors and provide context.
- Updated configuration to log errors for `Elsa.Workflows.ActivityRegistry`.
- Refined "Hello World" and "For Loop" workflows for clarity and added improved loop handling.
* Refine Elsa workflows and update compiler logic
- Simplified "Hello World" workflow by adding braces and improving consistency.
- Adjusted "For Loop" workflow to rename and clarify logic, including expression updates and variable handling.
- Fixed compiler mapping of `"cs"` to `"CSharp"` for better clarity.
- Enhanced "Hello World HTTP" workflow to correctly reference `variables.message` in expressions.
* Add flowchart support in ElsaScript parser, compiler, and integration tests
- Introduced `flowchart` syntax in `ElsaScriptParser` to support flowchart-based workflows.
- Updated `ElsaScriptCompiler` to compile `flowchart` nodes with labeled activities, connections, entry points, and variables.
- Added integration tests for parsing and compiling empty and simple flowcharts.
- Enhanced `FlowchartNode` and `LabeledActivityNode` for better representation of flowchart structures.
- Improved error handling and logging for invalid flowchart configurations.
* Add tests for compiling and parsing flowcharts with nodes, connections, and block nodes in ElsaScript
- Added integration tests for compiling and validating flowchart structures, including activities, connections, and entry points.
- Implemented parser tests for parsing flowcharts with node connections and block nodes.
- Updated project files to include new workflow examples for testing.
* Add Parlot package and update project file in integration tests
- Added `Parlot` package version `0.0.27` to `Directory.Packages.props`.
- Updated integration test project file to include a new `Include` directive for better targeting.
* Update Parlot package to version 1.5.2 in Directory.Packages.props
* Remove `elsa-server-and-studio.yml` workflow and update solution file
- Deleted `elsa-server-and-studio.yml` workflow as it's no longer needed.
- Updated `Elsa.sln` to remove reference to the deleted workflow.
* Remove `elsa-studio.yml` workflow and update solution and packages
- Deleted `elsa-studio.yml` workflow as it's no longer used.
- Updated `Elsa.sln` to remove reference to the deleted workflow.
- Changed `base_version` in `packages.yml` from `3.7.0` to `3.6.0`.
* Downgrade Docker image in `elsa-server.yml` workflow from `v3.7.0-preview` to `v3.6.0-preview`
* Update Docker image tag in `elsa-server.yml` workflow from `v3.6.0-preview` to `v3.6-preview`
* Add logging support to `LocalScheduler` and replace `Debug.WriteLine` with `ILogger`
* Remove unused `System.Collections.Generic` and `Elsa.Extensions` imports in `LocalScheduler`
- Cleaned up unnecessary using directives to improve code readability and maintainability.
- Minor whitespace adjustment for consistent formatting.
* Remove unnecessary whitespace in `LocalScheduler` for consistent formatting
* Improve exception handling in blob workflow format handlers
- Updated exception handling in `ElsaScriptBlobWorkflowFormatHandler` and `JsonBlobWorkflowFormatHandler` to gracefully catch and log all exceptions during workflow parsing.
- Adjusted comments to clarify behavior for invalid user-provided files, ensuring the workflow loading process is not disrupted.
* Refactor blob workflow format handlers to use `SupportedExtensions` for improved file filtering
- Added `SupportedExtensions` property to all blob format handlers to optimize blob storage browsing.
- Simplified `CanHandle` logic by removing extension checks, leveraging `SupportedExtensions` for initial filtering.
- Updated comments for clarity and consistency across handlers.
* Refactor `DefaultWorkflowDefinitionStorePopulator` to simplify `stringData` assignment logic and improve readability
* Remove outdated comment in `CompilerTests` about skipped tests
* Apply suggestion from @Copilot
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Refactor `ElsaScriptCompiler` to streamline type conversion logic, improve language mapping, and enhance asynchronous flowchart compilation
* [WIP] Update ParseError printing based on feedback (#7082)
* Initial plan
* Fix ParseError formatting to use Message and Position properties
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Replace `as` casts with direct casts in ParserTests for null safety (#7083)
* Initial plan
* Replace 'as' casts with direct casts in ParserTests for better null safety
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Fix Oracle column types for OriginalSource and other large text fields (#7079)
* Initial plan
* Fix Oracle OriginalSource and StringData column types to handle large data
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Refactor tests to replace type checks with `Assert.IsType` for improved clarity and type safety
* Initial plan (#7080)
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
* Add `Parlot` package reference and update solution structure by removing and reorganizing projects and workflows.
* Set default expression language to "JavaScript" in `ElsaScriptCompiler`.
* Add integration test to verify default expression language resets between ElsaScript compilations
* Simplify UTF-8 encoding in JsonContentFactory (#7081)
* Initial plan
* Remove explicit UTF8Encoding in JsonContentFactory and use Encoding.UTF8
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Fix test to use Encoding.UTF8.GetByteCount for multi-byte character support
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: Ender <37611092+zengande@users.noreply.github.com>
Co-authored-by: Matt <knibbsy10@live.com>
Co-authored-by: Max Brooks <45081361+MaxBrooks114@users.noreply.github.com>
Co-authored-by: Max Brooks <Max@compyl.com>
Co-authored-by: FuJa0815 <30809803+FuJa0815@users.noreply.github.com>
Co-authored-by: Peter Klooster <crashkonijn@gmail.com>
Co-authored-by: Peter Klooster <peter.klooster@autotaalglas.nl>
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Add unit and integration tests for `Container` activity, covering behavior such as variable scoping, child activity execution, and mixed variable types.
* Refactor `RunWorkflowAndCaptureOutput` method in `ContainerTests` for better code organization
* Move `Sequence` activity tests to a dedicated namespace and add new unit and integration tests for enhanced coverage.
- Deleted outdated `SequenceTests` and related workflows.
- Introduced `SequenceActivity` namespace with improved organization.
- Added comprehensive unit and integration test coverage for sequential execution, nested sequences, conditional breaking, variables, and dynamic activities.
* Relocate `SequenceActivity` tests to `Activities` namespace to improve organization and update references in related test classes.
* Refactor `SequenceTests` to move `DynamicSequenceWorkflow` to its own file for better test organization.
* Extract `TestContainer` to `Elsa.Testing.Shared.Activities` for reuse across test projects.
* Add unit and integration tests for `Break` activity; refactor workflow tests for improved organization
- Introduced `BreakInForkWorkflow` and deprecated `BreakWhileForkWorkflow`.
- Added `BreakTests` unit tests to validate behavior of the `Break` activity, including terminal node implementation and execution completion.
- Enhanced integration tests for `Break` activity, covering multiple looping constructs (`ForEach`, `For`, `While`, `Fork`) and nested workflows.
- Updated `Fork` activity to handle the `BreakSignal` asynchronously.
- Simplified workflow definitions by removing redundant constructors and using concise variable initialization syntax.
- Improved test clarity with better organization, comments, and consistent naming conventions.
* Refactor `Fork` activity tests and workflows for improved organization and coverage
- Relocated `BasicForkWorkflow` and `JoinAnyForkWorkflow` to `Fork/Workflows` namespace.
- Introduced `EmptyForkWorkflow` to test Fork behavior with no branches.
- Enhanced `ForkTests` with scenarios for `Fork` execution with different join modes and branch configurations.
- Refactored `Fork` activity to handle empty branches and simplified `BreakSignal` handling.
- Improved consistency and clarity of test cases, including better assertions and comments.
* Update test/integration/Elsa.Workflows.IntegrationTests/Activities/Break/BreakTests.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update test/integration/Elsa.Workflows.IntegrationTests/Activities/Break/BreakTests.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update test/integration/Elsa.Workflows.IntegrationTests/Activities/Break/BreakTests.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update test/integration/Elsa.Workflows.IntegrationTests/Activities/Fork/ForkTests.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update test/integration/Elsa.Workflows.IntegrationTests/Activities/Fork/ForkTests.cs
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* Update GitHub Actions workflow to use .NET 10.x
* Remove outdated GitHub workflows and update configurations to .NET 10.x
* Remove unused result variable assignments in integration tests (#7105)
* Initial plan
* Remove unused result variable assignments in BreakTests and ForkTests
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
---------
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
* Add unit and integration tests for `Container` activity, covering behavior such as variable scoping, child activity execution, and mixed variable types.
* Refactor `RunWorkflowAndCaptureOutput` method in `ContainerTests` for better code organization
* Move `Sequence` activity tests to a dedicated namespace and add new unit and integration tests for enhanced coverage.
- Deleted outdated `SequenceTests` and related workflows.
- Introduced `SequenceActivity` namespace with improved organization.
- Added comprehensive unit and integration test coverage for sequential execution, nested sequences, conditional breaking, variables, and dynamic activities.
* Relocate `SequenceActivity` tests to `Activities` namespace to improve organization and update references in related test classes.
* Refactor `SequenceTests` to move `DynamicSequenceWorkflow` to its own file for better test organization.
* Extract `TestContainer` to `Elsa.Testing.Shared.Activities` for reuse across test projects.
* Add unit and integration tests for `Container` activity, covering behavior such as variable scoping, child activity execution, and mixed variable types.
* Refactor `RunWorkflowAndCaptureOutput` method in `ContainerTests` for better code organization
* Extract `TestContainer` to `Elsa.Testing.Shared.Activities` for reuse across test projects.