Commit graph

7323 commits

Author SHA1 Message Date
Sipke Schoorstra 1c0fb909f9
chore: anchor the Release/ ignore pattern to the repository root
`[Rr]elease*/` is unanchored, so it matched any directory whose name starts with
"release" at any depth -- including documentation and tooling paths that are not
build output. Two directories were silently excluded as a result:
.claude/skills/release-notes/ and .claude/skills/release-announcement/. `git add`
reported nothing and the commit simply omitted them.

bin/ and obj/ are already ignored above, so this pattern only needs to catch a
stray Release/ directory at the repository root. Anchoring it keeps that while
leaving docs/ and tooling directories tracked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-22 23:06:17 +02:00
Sipke Schoorstra dff7d9f987
Harden output converter contracts
(cherry picked from commit 0737def83d)
2026-08-21 22:22:47 +02:00
Sipke Schoorstra 61fc376dac
Add output converter support at binding boundaries
Backport to release/3.8.0 so Elsa.Api.Client 3.8.0-rc2 exposes the
Resources/OutputConverters surface that Elsa Studio's release/3.8.0 branch
already consumes. Without it, Studio cannot build against a released client:
it was green against 3.8.0-preview.5397 (built from main) and broke when its
pin moved to 3.8.0-rc1 (built from this branch).

(cherry picked from commit d698e6b005)

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 22:22:46 +02:00
Sipke Schoorstra cf23279bf1
test(resilience): cover Elsa.Resilience.Core and lift its coverage gate off the Debug/Release seam (#7971)
`dotnet test test/unit/Elsa.Resilience.Core.UnitTests` exited 1 on a clean
checkout with all 56 tests passing. The failure was the coverlet gate, not a
test: the project pinned `<Threshold>49</Threshold>` against 48.17% measured
line coverage in Debug. Release measures slightly differently and cleared it,
so CI (which builds `--configuration Release`) stayed green while every local
run — Debug is the default — went red. A red exit for a suite that passes
trains people to ignore exit codes.

Rather than move the goalposts, cover the code. The gap was concentrated in
`ResilientActivityInvoker`, which had no tests at all, plus the serializer,
the activity-execution extensions and the retry telemetry listener.
`Elsa.Testing.Shared`'s `ActivityTestFixture` was already referenced here and
builds a real `ActivityExecutionContext`, which is what all of them needed.

Adds 40 tests. The invoker ones drive a real zero-delay Polly retry pipeline,
so the telemetry listener is exercised through the actual Polly path rather
than being called directly: pass-through when no strategy is configured, the
applied strategy recorded on the context, retry-then-succeed, one record per
retry carrying identifiers and details, null details dropped, the retries flag
and attempt count, exhausted retries rethrowing, and an unhandled exception
type not being retried. The extensions tests build a three-level context chain
to pin down that the retries flag propagates up the ancestor chain and not
down.

Line coverage goes 48.17% -> 98.17% in Debug and 97.8% in Release; the five
lines still uncovered are defensive early-returns. The threshold moves to 90,
below the lower of the two configurations with enough headroom that the
Debug/Release delta cannot straddle it again. Verified by deleting the invoker
tests once: coverage falls to 68.97% and the gate fails as it should.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 03:01:51 +02:00
Sipke Schoorstra b24beadca9
feat(bpmn): adopt Bpmn.* 0.2.0, and start shipping the two BPMN modules (#7970)
Bpmn.Interchange, Bpmn.Model and Bpmn.Semantics move from 0.1.1-preview.19 to
0.2.0. The three changes below are one unit: the bump is what makes the other
two true.

Retire the private feed. All three packages are on nuget.org at 0.2.0 --
including Bpmn.Semantics, which had no stable release under 0.1.x. The
bpmn-feedz source and its Bpmn.* packageSourceMapping entry both go, which
removes a setup step for every consumer. This is not merely cleanup: the feedz
feed does not carry 0.2.0 stable, so with the mapping left in place the bump
would not restore at all. Restore now resolves Bpmn.* from nuget.org via the
existing `*` mapping.

Lift IsPackable=false. The comment on the flag named its own removal condition
-- the Bpmn.* packages reaching nuget.org -- and that condition is now met, so
Elsa.Bpmn and Elsa.Bpmn.Interchange begin shipping. Both pack with every
dependency publicly restorable, and both emit a package manifest carrying
runtimeKinds ["elsa.server"], so neither is silently excluded from the catalog.
They were the last two IsPackable=false projects under src/.

Flip the compensation pin to the fixed behaviour. 0.2.0 contains the fix for
valence-works/bpmn#13 (filed from here as #7959), so
CompensationRunCancelledMidReplay went red on the bump exactly as it was built
to. Upstream took the wide fix: every token a cancelled transaction abandons now
gets a real teardown. The head handler still starts twice -- that is the release
being real -- but the first run is now torn down, so the scope is left holding
one live record for the slot instead of two. The applier is deliberately
unpatched; the assertions moved to describe the fix, not to accommodate it.

Note on persisted state: 0.2.0 freezes the payload format at 1.0.0 and state
persisted by 0.1.x no longer deserializes. Elsa persists the library's
BpmnExecutionState into workflow state, so an in-flight BPMN instance does not
survive this bump. Neither module has ever been published, so no released
consumer can be holding such state -- which is why this is the moment to take
the break.

BpmnRuntimeCapabilitiesTests stays green: 0.2.0 defines no capability flag Elsa
does not already declare. The new cancel-end-event requirement is
SubtreeCancellation, which BpmnRuntimeCapabilities.Declared already carries.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 01:44:07 +02:00
Sipke Schoorstra 1b38c3511d
fix: stop two silent serialization and test-isolation traps (#7969)
* fix: stop two silent serialization and test-isolation traps

Two follow-ups from #7957.

ExternalAuthentication tests: the same process-global
EndpointSecurityOptions.SecurityIsEnabled race the shells API tests had,
across the six classes in that assembly that build an endpoint host —
five setting it to false and IdentityLinkAuthorizationTests to true.
Unlike the shells case these all call UseAuthorization(), so it does not
surface as a missing-middleware error: anonymous endpoints answer
401/403, and the authorization test's endpoints come back AllowAnonymous
and stop enforcing what it asserts. A module initializer cannot fix it
since the assembly genuinely needs both values, so the six now share one
collection with DisableParallelization. They are also the only six that
build a host, so nothing else can observe a leaked value.

Unaliased payloads: a payload whose type has no registered serialization
alias is written without a _type discriminator and read back as an
ExpandoObject whose keys carry the state serializer's camel-case naming
policy, so a consumer that published Status finds status. The
degradation is deliberate — the alias registry is an allow-list that
keeps arbitrary CLR type names out of deserialization — but it was
silent. It is now reported once per type, naming the type and both
lossless alternatives, and PublishEvent.Payload documents them. Measured
across the integration suite, only genuine user payload types reach this
path, so the warning does not fire for Elsa's own types.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: check the log level before claiming the once-per-type warning slot

WarnAboutUnaliasedType claimed a type's single report via TryAdd before
LogWarning applied its level filter, so a type first serialized while
Warning was disabled spent its slot on a call that logged nothing and
then stayed silent forever, including after the level was raised at
runtime. Check IsEnabled first, so the slot is only consumed by a report
that is actually emitted.

The regression test needs the capture to be the only logging provider:
IsEnabled on the composite logger is an OR across providers, so the test
builder's own xunit provider would otherwise keep Warning enabled
regardless of what the test asked for.

Reported by Greptile on #7969.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-21 00:36:11 +02:00
Sipke Schoorstra d78e1e1277
docs(bpmn): repoint the cancelled-transaction pin at the upstream issue (#7968)
The two-record assertion in CompensationRunCancelledMidReplay pins a bug whose
root cause is in Bpmn.Semantics, not in this host, so the tracker that will
actually move is valence-works/bpmn#13. #7959 stays named as provenance -- it
carries the decompiled evidence and the routing rationale -- but is closed as
routed upstream, so it is no longer the thing to watch.

Also records what the current comment did not: both counts become 1 when the fix
lands, and the flip may not be mechanical. Upstream is choosing between tearing
down only the compensation handler it re-starts and tearing down every abandoned
token, and the second also cancels transaction branches still in flight, which
can move other scenarios in this suite.

Comment-only. The assertions and the applier are deliberately unchanged.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 23:57:01 +02:00
Sipke Schoorstra a818b5110e
fix(features): support features introduced during Module.Apply() (#7966)
Module.Apply() enumerated _features.Values directly while calling
feature.Apply(). A feature whose Apply() introduces another feature —
Module.Configure<T>() directly, or via a helper such as AddActivity<T>()
which configures WorkflowManagementFeature — mutated that collection
mid-enumeration and threw "Collection was modified; enumeration
operation may not execute", naming nothing about features. Whether it
fired depended on whether the other feature happened to be installed
already, so a module built or did not based on unrelated host config.

The module already treats introduction-during-apply as supported: the
ConfigureFeature loop iterates a snapshot for exactly this reason, and
Configure<T>() has an _isApplying branch that creates, resolves and
configures a feature introduced mid-Apply. Only the final apply loop
missed the same treatment, so make it tolerant rather than diagnose a
constraint the code does not hold.

The apply loop now runs in rounds until no new features appear, each
round topologically sorted so a late feature's dependencies apply before
it. Hosted services are registered in a single pass after that loop,
then moved back to the index the block previously occupied: registering
late is needed so features contributed during Apply() are included and
ordered by priority, while keeping the position matters because features
register hosted services directly from Apply() — WorkflowRuntimeFeature
adds DrainOrchestratorHostedService that way — and module-managed
services must keep starting first, or a priority such as ActivateTenants
at -1 would silently start ordering after them.

Adds Elsa.Features.UnitTests, covering the introduced feature applying,
a three-deep introduction chain, dependency ordering, hosted service
registration and priority ordering for late arrivals, the installed-
feature registry, and no double-apply, plus guards for pre-existing
ordering behaviour.

Closes #7944

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 23:53:07 +02:00
Sipke Schoorstra a02ebff129
test: fix two intermittent test failures (#7957) (#7965)
Both tests read a value that is usually one thing and occasionally
another, with a race deciding which.

ReloadTests: EndpointSecurityOptions.SecurityIsEnabled is a process-
global static, and ShellsApiTestBase saved/set/restored it per test
method. ReloadTests and ReloadAllTests carry no [Collection], so xUnit
runs them in parallel. FastEndpoints reads that global once per host
while UseFastEndpoints() configures the endpoints, so when one class's
DisposeAsync restores true inside another class's set-false ->
UseFastEndpoints() window, that host's endpoints get authorization
metadata in a pipeline with no UseAuthorization, and every request to
them throws. Every test in the assembly wants security off, so set it
once in a module initializer and stop mutating it per test.

PublishEvent_WithPayload_TransmitsPayloadToConsumer: the payload's
representation is not stable. While it is still the original CLR object
its properties are PascalCase; once it has been through
JsonWorkflowStateSerializer it is an ExpandoObject whose keys were
camelCased by that serializer's naming policy. Which one the test sees
depends on whether GetSingleWorkflowInstanceAsync returned the live
in-memory instance or one read back from the store, and TryGetProperty
is case-sensitive. Assert the payload's content through a DTO with
PropertyNameCaseInsensitive instead of one of the two representations.

Also require a terminal instance at both exits of
GetSingleWorkflowInstanceAsync: it accepted any save, and an instance is
saved several times over its lifetime, so it could hand a caller that
asserts Finished an instance that is still running.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 23:30:38 +02:00
Sipke Schoorstra 74fc891350
feat(core): let a container withdraw work it scheduled but must not run (#7967)
A container that schedules a child and then decides the child must not run
had no way to withdraw it. `IActivityScheduler` exposed no removal operation,
and `CancelActivityAsync` no-opped on a context whose status was `Pending`,
so a container could tear a branch down and still have an activity from that
branch execute afterwards, side effects and all. Fixes #7943.

- `IActivityScheduler.RemoveWhere` removes work items and keeps the order the
  survivors would have been taken in; implemented in both the FIFO and LIFO
  schedulers.
- `CancelActivityAsync` (both the public extension and the internal one used
  when a container completes) cancels `Pending` contexts as well as running
  ones, and withdraws the work item that would have started the cancelled
  activity plus the items it had scheduled for children with no context yet.

Withdrawal is a real removal rather than a terminal status honoured at dequeue
time, because the scheduler is also read: `Flowchart.HasPendingWork` inspects
it to decide whether it may complete, and the work item list is extracted into
the persisted workflow state — a withdrawn-but-queued item would be persisted
and rehydrated with a fresh context after a suspend/resume.

`StateMachine` had hand-rolled the same operation to drop competing triggers by
clearing the scheduler and re-scheduling everything else; it now calls
`RemoveWhere`. `Elsa.Bpmn` no longer needs to refuse a teardown whose subtree
still has queued work, so `BpmnWorkTeardown` drops the `NotSupportedException`
and records the teardown reason on the torn-down activity's journal instead.

BREAKING: `IActivityScheduler` gains a member; external implementations must
add `RemoveWhere`.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 23:29:46 +02:00
Sipke Schoorstra 0a86a4803e
refactor(bpmn): split BpmnTestProcesses by construct family (#7964)
BpmnTestProcesses had grown to 905 lines as one flat static class holding the
fixtures for every BPMN construct family the runtime slice covers, and two
standards reviews flagged it as Divergent Change while judging the split out of
scope for the issue in hand.

It becomes a partial class across six sibling files, one per family -- boundary
events, compensation and transactions, event subprocesses, flow and gateways,
multi-instance -- with the shared element factories (Timer, Cancel, Compensation,
CompensationBoundary, Escalation, Error, Message, EventSubprocess,
EventSubprocessStart) and the Scope/Immediate/Blocking/Faulting builders left in
one place, so no new file duplicates them. Partial rather than separate types
because every call site says BpmnTestProcesses.X and none of them change.

Pure move: all 47 members were carved out programmatically and diffed back
against HEAD, each present exactly once and byte-identical. In particular
EscalationOutOfSubprocess keeps its leading subFirst work item and the comment
explaining why the nested scope's handle counter must run ahead of its parent's.

The identical Compensation/CompensationBoundary helpers in
Elsa.Bpmn.Interchange.IntegrationTests are deliberately left duplicated: the only
assembly both test projects can see is Elsa.Testing.Shared.Integration, which
ships as a NuGet package, so sharing ten lines of test helper would mean adding
an Elsa.Bpmn reference to a published package's dependency graph.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 22:30:17 +02:00
Sipke Schoorstra fe83d2b385
test(bpmn): event subprocesses, dormant and listener-backed (#7963)
* test(bpmn): event subprocesses, dormant and listener-backed (#7933)

Both flavours, and the three responsibilities the host has for them. The
production diff is empty: the applier and binder need no event-subprocess
code path, which is what #7909 measured and what this pins.

A dormant catcher -- error or escalation -- rides the FaultSignal seam and the
escalation signal path, and reaches the host as an ordinary StartWork for its
body. A listener-backed one gets a second StartWork for its listenerBindingRef
at scope start and a CancelWorkSubtree for it when the scope completes. Both
already apply like any other command.

The host responsibilities, each with the failure it would otherwise hide:

- The listener is armed at scope start and observable in the scope's own ledger
  before anything fires, not inferred from a fire that worked.
- A completing scope retires a still-armed listener. Pinned twice: at the root,
  where Elsa's own container completion would cancel the child regardless and
  only the scope's ledger tells the two apart, and inside a subprocess the
  workflow outlives, where a listener left behind is one something could still
  resume into.
- The start-element hint reaches the body and nothing else inherits it. The
  body's only start event is event-defined, so a body that never received the
  hint faults bpmn.start.none-available rather than starting somewhere
  plausible; the ordinary subprocess inside it faults bpmn.start.unresolved-hint
  if the hint travels where it must not. The scope's invocation correlation is
  read back after the body has run work of its own, because the dictionary is
  fixed for the scope's lifetime and the hint is read from it.
- A non-interrupting listener fired twice re-arms onto the slot the first fire
  vacated, holding one live record and one bookmark at a time. This is the case
  the completed-work-removed-before-the-interpreter-is-asked ordering exists
  for, and it is now observable; BpmnHostInvariantTests points at it.

The library's declaration rules are pinned as refusals rather than gaps: a body
with more than one start event, a second error-triggered event subprocess in a
scope, and a non-interrupting error event subprocess are each refused when the
scope builds its graph, before any work starts. The last is additionally
dropped at import, with the rest of the document reading as written -- the
dropped body carries an undeclared serviceTask, so an import that still
succeeds is what proves the drop took its bindings with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): pin the catch-all escalation event subprocess refusal

Two error-triggered event subprocesses per scope was pinned but its sibling
rule -- at most one code-less catch-all escalation event subprocess -- had
no test. Add TwoCatchAllEscalationEventSubprocesses_AreRefused, mirroring
the error refusal test, and let Escalation() build a code-less definition.

Also record in BpmnCommandApplier why CancelSubtreeAsync's explicit
subtree cancellation is redundant on the scope-completion path (Elsa's own
container-completion behaviour already covers it) while the ledger removal
above it is not, so a future reader does not "simplify" the ledger removal
away.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-20 01:02:51 +02:00
Sipke Schoorstra 37c98b2a78
test(bpmn): compensation, targeted replay and transaction cancellation (#7960)
* test(bpmn): cover compensation, targeted replay and transaction cancellation

W15 turns on compensation boundary events, reverse-order replay, targeted
compensation, transaction subprocesses and cancel end/boundary events. The
measured claim in the design holds: the command applier needs no changes, and
none were made. A compensation handler arrives as an ordinary StartWork
carrying cause=compensation, and the binder already binds it because the reader
emits an ordinary Primary binding for an isForCompensation element.

What is new is the test mass that says so, and each case pins a failure that
otherwise looks like success:

- three registrations replayed in reverse, asserted as an ordered log rather
  than as "all three ran"
- a compensate throw naming an activityRef, where the two unselected handlers
  are bound work that must stay unrun -- which is also where "a handler is
  never scheduled from flow" becomes observable
- a compensation run torn down mid-replay by a cancel end event, so its claimed
  but unrun entry is released back to registered and the cancellation's own
  replay reaches it; leaving it claimed would cancel with nothing to compensate
  and finish looking healthy
- a transaction completing Cancelled with no cancel boundary to route it, which
  must fault rather than take the ordinary sequence flow
- two compensation logs, one per scope, in a subprocess and around it

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): pin the duplicate live-work record a cancelled transaction leaves

CompensationRunCancelledMidReplay asserted only log counts and Finished, so
the two live ledger records/bookmarks it produces for the releaseSeat slot
went unasserted. Add an explicit assertion on the scope's ledger, and correct
the comment that framed the second handler start as evidence only of the
release working, when it is also the symptom of the interpreter defect
tracked in #7959.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 23:25:36 +02:00
Sipke Schoorstra 457e2a6185
feat(bpmn): publish-time validation gate for unbound tasks (#7958)
* feat(bpmn): refuse to publish a definition with an unbound BPMN task

BpmnWorkBinder already refuses an UnboundTask with no elsa:activityBinding
at import - that is the first net. A definition can be edited after
import, through Elsa's own designer rather than the BPMN document, and
that edit can remove the activity a task was bound to without touching
the document snapshot the scope still carries. ValidateBpmnProcessBindings
is the second net: a WorkflowDefinitionValidating handler that walks the
materialized workflow graph (not the inert BpmnProcessDefinition snapshot
or the stored BPMN source, both of which a graph-only edit leaves
untouched) and fails publication for any task-family element whose
binding no longer resolves to an activity in the graph, naming the
offending element id.

Import-time Dropped/Degraded findings from BpmnImportAnalysis are not
persisted anywhere a publish-time handler can reach, and BpmnImportIssue
carries no field distinguishing a Dropped finding that changes executable
meaning from one that does not; WorkflowValidationError has no severity
concept either. Extending this gate to those findings would mean guessing
at a classification the library does not expose, so it is left alone -
see the delivery notes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): refuse to publish a BPMN task bound to a missing activity type

The publish gate only checked that a binding pointed at some activity id;
it did not check the activity actually resolved, so a binding left
pointing at an uninstalled type passed the gate and failed at run time
instead. Report that case distinctly from "not bound at all", name the
containing BpmnProcess node (not the BPMN element id) as the error's
ActivityId to match how the rest of the codebase reports it, and prove
the BpmnProcess-inside-Flowchart graph-walk with a dedicated test. Also
de-duplicate the publish-gate test fixture's binding helpers by deriving
from BpmnBindingTestBase instead of re-declaring them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(bpmn): filter the publish gate's loops explicitly

Replace the implicit-filter foreach/continue pattern in
ValidateBpmnProcessBindings with .OfType/.Where so each loop only
iterates the elements it acts on, without changing behaviour, error
messages, or error ordering.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-19 23:25:31 +02:00
Sipke Schoorstra 063ca814b3
docs: refresh roadmap 2026-08-19 00:06:14 +02:00
Sipke Schoorstra 8955f9ad34
feat(bpmn): interchange endpoints (analyze, import, export) (#7954)
* feat(bpmn): add Analyze/Import/Export endpoints to Elsa.Bpmn.Interchange

Thin FastEndpoints wrappers over Bpmn.Interchange, sharing one
BpmnInterchangeDocumentService so Analyze and Import can never disagree
about what a document costs. Import surfaces capability refusal
(BpmnCapabilityRequirements.Analyze, walked into nested processes) with
the missing capability and offending element ids, and reuses
BpmnWorkBinder to bind the root BpmnProcess scope. Export re-reads the
original XML persisted alongside the workflow definition and re-runs it
through BpmnXmlWriter, so retained extension elements, foreign
attributes and BPMN DI layout survive the round trip without being
reconstructed from the reduced Elsa activity graph.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): make a stale BPMN export refuse instead of mislead

Export now refuses with 422 (naming the reason) when a workflow definition's
BPMN source is missing or no longer matches the definition's version, rather
than exporting stale or absent content while reporting success. Import
records the definition's version alongside the source XML so Export can
detect drift caused by a later save replacing custom properties wholesale.

Also: the interchange package now consumes the runtime host's declared
capability set from a new public Elsa.Bpmn.Hosting.BpmnRuntimeCapabilities
instead of restating it (one value, one home); the Import endpoint's
capability-refusal message no longer misattributes driving elements across
capabilities; the three BPMN REST endpoints get HTTP-level test coverage
(multipart validation, exception-to-status-code mapping, permission gating);
and the wiki documents the endpoints and Export's known limitation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): alert when the library defines a capability Elsa does not declare

Restores the deleted comparison between BpmnRuntimeCapabilities.Declared (ours)
and BpmnHostCapabilities.Full (the library's) — these are two different
constants, not the tautology the earlier deletion assumed. The pinned
Bpmn.Semantics 0.1.1-preview.19 currently defines exactly the four flags Elsa
declares, so capability refusal at import/build is wired but unreachable; this
test is what will say the moment a library bump changes that, and its failure
message names the decision (implement and declare, or leave undeclared on
purpose) rather than just failing silently.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): return 400 for a malformed export version and clarify a partial-import refusal

- Export/Endpoint.cs: a non-numeric or out-of-range VersionOptions query value now returns a
  400 naming the offending value instead of throwing through FromString and bubbling into a 500.
- BpmnInterchangeDocumentService: the message shown when a definition carries BPMN source but not
  its version marker (a second save that never completed after ImportAsync's first) now says so
  explicitly, distinct from "never imported" and "stale".
- BpmnInterchangeDocumentService: replace the implicit filter in EnsureCapabilitiesSatisfied's
  foreach with an explicit .Where(...), same behaviour.
- Test projects: extract the duplicated ReadAsset/Path.Combine helper in
  BpmnInterchangeTestBase and BpmnInterchangeEndpointTests into a single BpmnAssetReader, guarded
  against a rooted or nested file name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): write both BPMN import markers in one save

Move the BPMN source XML off the pre-import model and onto the same
explicit save that already records the definition's version, so a
failed or cancelled post-import save leaves neither custom property
behind instead of a partial, undiagnosable state. Update
BpmnAssetReader to use Path.Join instead of Path.Combine so its
rooted/nested-name guard is defence-in-depth rather than the only
thing standing between the code and a wrong path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:20:42 +02:00
Sipke Schoorstra e447eb2826
feat(bpmn): triggers and correlation for message, signal and timer starts (#7956)
* feat(bpmn): let a BPMN process start from outside via message, signal, and recurring timer starts

BpmnProcess now implements ITrigger: it walks its own event-defined start events
and emits an EventStimulus per resolved message/signal name (matching what
Event/PublishEvent already key bookmarks on) and a TimerTriggerPayload/CronTriggerPayload
per recurring timer start (matching Elsa.Scheduling's own Timer/Cron path), so
correlation is on name for message/signal and reuses the existing scheduling
mechanism for timers, unmodified.

BpmnWorkBinder.Bind now marks the one scope it returns directly as the workflow's
root scope; every nested scope it produces stays off. Because that flag alone
cannot see composition that happens after it is set (e.g. nesting through an
intermediate Flowchart, the gap left open by #7926's applier-level refusal),
BpmnProcess re-derives entry-point status from the whole workflow graph at
trigger-indexing time and refuses to register regardless of what the flag says.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): give each BPMN start trigger its own stimulus name, dedup, and refuse per malformed timer

BpmnProcess.GetTriggerPayloadsAsync now wraps every start-event payload in a
NamedTriggerPayload (the per-payload stimulus naming TriggerIndexer gained in #7950)
instead of the shared TriggerIndexingContext.TriggerName, so a process with both a
message/signal start and a recurring timer start no longer has the last kind processed
claim the name -- and therefore the hash -- for every row.

Two start events (or two event definitions) that resolve to the same stimulus name and
value now collapse to one payload, so StimulusSender no longer starts the workflow twice
for one inbound stimulus. A malformed <timeCycle> interval is refused for its own start
event only, named in the warning like the work binder's own malformed-duration message;
every other valid start event on the process still registers, since letting the exception
propagate would just be swallowed whole by TriggerIndexer's catch-all around
GetTriggerPayloadsAsync, discarding every other start again.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): refuse a non-positive BPMN timer interval

XmlConvert.ToTimeSpan accepts PT0S and negative durations; either
registered as a recurring timer trigger turns into a hot loop, since
the scheduler substitutes a ~1ms delay whenever the next execution
time is non-positive. Refuse it through the same per-start-event
refusal path already used for a malformed interval, so the process's
other start events still register.

Also address two small static-analysis findings in the same method:
filter the start-event loop explicitly with .Where(...) instead of an
implicit continue, and combine two genuinely-simple nested if pairs
(message/signal name resolution, and the cron branch) with &&. The
timer interval's own nested if/try-catch is left alone: combining it
would only get harder to read once the non-positive check joins it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): refuse a BPMN timer interval below the scheduler's resolution

A positive but sub-resolution interval (e.g. PT0.0000001S) passed the
existing non-positive refusal unchanged and rearmed in the same hot loop
that guard was meant to close, one step down: Elsa.Scheduling's
ScheduledRecurringTask.SetupTimer substitutes a 1ms delay for any
non-positive delay it computes, and SchedulingOptions.MinimumPastDueScheduleDelay
defaults to that same 1ms, so 1ms is the scheduler's own resolution floor,
not a guessed constant. Refuse an interval below it through the same
per-start-event path the malformed and non-positive cases already use, so
the offending start event is skipped and the process's other start events
still register.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 05:15:35 +02:00
Sipke Schoorstra f6c35cf1eb
feat: Implement OIDC Trusted Publishing for NuGet
This enhances security by configuring the NuGet publishing workflow to use GitHub OIDC Trusted Publishing. This mechanism exchanges a GitHub OIDC token for a temporary nuget.org API key, eliminating the need for a long-lived API key secret.
2026-08-18 00:10:22 +02:00
Sipke Schoorstra b0ab630a34
fix: quote API key secrets in dotnet nuget push to prevent argument parsing failure (#7951) 2026-08-17 23:53:19 +02:00
Sipke Schoorstra bde8b2ccf2
Add project reference to HTTP Webhooks module 2026-08-17 23:51:52 +02:00
Copilot 0b049fa269
fix: quote API key secrets in dotnet nuget push to prevent argument parsing failure (#7951)
* Initial plan

* Fix: quote API key secrets in dotnet nuget push commands to prevent argument parsing errors when secret is empty

Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
2026-08-17 23:46:25 +02:00
Sipke Schoorstra f9580e7472
fix(runtime): let a trigger index payloads under per-payload stimulus names (#7950)
* fix(runtime): let a trigger index payloads under per-payload stimulus names

TriggerIndexingContext.TriggerName is a single field read once after all
payloads have been collected, so one ITrigger could only ever register its
payloads under one stimulus name. An implementation that assigned the name
more than once - which the stimulus extension methods do as a side effect -
had the last write applied to every row, and since Hash derives from the same
name, the earlier payloads were stored under a hash no publisher computes.

Adds an additive, opt-in path: a payload returned from GetTriggerPayloadsAsync
may be wrapped in NamedTriggerPayload, which carries the stimulus name for that
payload alone. The indexer takes name and payload from the same source, so
Hash always matches the Name stored beside it, and the wrapper is unwrapped
before storage so payload consumers (validators, the trigger diff comparer,
the scheduler) see the payload the trigger produced.

TriggerName keeps its existing meaning as the default for payloads that do not
carry their own, so every existing ITrigger indexes identically: same Name,
same Hash, same Payload, same row count. The empty-payload placeholder row is
left alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(runtime): refuse a nested trigger payload wrapper

NamedTriggerPayload documented that its Payload is never itself a
wrapper, but nothing enforced it. Reject a NamedTriggerPayload whose
payload is another NamedTriggerPayload at construction time, matching
the existing guard against a blank name.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-17 23:43:59 +02:00
Marko Lahma 7328a0f14d
Adopt the Jint 4.15 host-integration surface: lazy type globals, enum names, and register-what-is-referenced (#7895)
* chore(javascript): update Jint to 4.15.3 and stop blocking on promises

`Engine.Evaluate(...).UnwrapIfPromise()` blocks the calling thread while the
engine's event loop drains, which is exactly the wrong thing to do inside an
`async` method — an expression that awaits a .NET `Task`, such as one calling
`getSecret()`, held a thread pool thread for the duration of the I/O.
`Engine.EvaluateAsync` awaits the returned promise instead, and takes the
cancellation token while it is at it.

The Jint version is moved from 4.4.2 to 4.15.3. `EvaluateAsync` arrived in
4.14.0, but the pin lands past 4.15.2 deliberately: once expressions genuinely
suspend and resume instead of draining the event loop on the calling thread,
they exercise the async suspension machinery 4.15.2 corrected — an `await` on a
right-hand side no longer stores the suspension sentinel, async generators and
`for await...of` preserve loop iteration state across a suspension, and a
suspension node is unwrapped correctly. Shipping the non-blocking change on an
earlier 4.14/4.15 would enable exactly the code paths those releases fixed.

One default changed along the way: since 4.14 `Interop.ArrayConversion` defaults
to `LiveView`, so a CLR array reaches script as a live view over the original
array rather than as a copy. That is observable — a script that sorts an array
would now reorder the workflow's own array, and the value round-trips back as
its original element type rather than as `object[]`. The evaluator therefore
pins the previous `Copy` behaviour so the upgrade is not a behavioural change;
hosts that prefer the live view can opt in through
`JintOptions.ConfigureEngineOptions`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

* test(javascript): pin the array copy lane and adopt JsString.Create

The array audit. The parent commit fixes `ArrayConversion` to `Copy`, because
4.14 changed the default to `LiveView` and the two differ in behaviour a script
can observe. The existing tests assert what a copy *produces*, which a live view
also satisfies for a value nothing mutates; asking the engine how many
conversions of each kind it performed (4.15.1's interop conversion counters)
pins the lane itself. The second assertion is the more interesting one: an
ordinary evaluation converts no CLR array at all, because Elsa converts
collection-valued variables itself in `ObjectConverterHelper` long before Jint's
array lane could see them. That makes the `ArrayConversion` setting a narrow
compatibility pin rather than something every evaluation depends on.

`JsString.Create` (public since Jint 4.15.3) is adopted in
`JsonElementConverter`, where the string case was the only one still routed
through `JsValue.FromObject` — re-entering the whole conversion pipeline, the
registered object converters and this one included, to arrive at the same call
the number and boolean cases beside it already make directly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f

* perf(javascript): register .NET type globals lazily

A fresh Jint engine is built for every expression evaluation, and roughly twenty
.NET types are registered on it before the expression runs: the common types
(`DateTime`, `Guid`, `TimeSpan`, …) plus every non-primitive workflow variable
descriptor type. Each registration builds a `TypeReference`, which describes the
type through reflection. The overwhelming majority of expressions reference none
of them.

`Options.AddLazyGlobal` installs a global whose value is produced by a factory
the first time a script reads the name. Both type registration handlers now run
on `CreatingJavaScriptEngine` — which already carries the `Jint.Options` being
built — and register through it, so a type is only described if an expression
actually mentions it. A script that uses `Guid.NewGuid()` still sees `Guid`; a
script that uses none of them pays for none of them.

Types whose name cannot be written as a JavaScript identifier are skipped while
we are here, since no script can reach them. That covers constructed generic
types (``IDictionary`2``, which two different variable descriptors both claimed)
and array types (`Byte[]`).

Moving the registrations to engine construction has a consequence worth pinning
beyond the ordering: a global installed by the host through `configureEngine` is
no longer overwritten by the built-in registration of the same name. The lazy
global for `Guid` is already installed when that callback runs, and
`Engine.SetValue` goes through `[[Set]]` on the global object, which reads the
current value — running the factory once and discarding the result — before
replacing the descriptor. The end state is the host value, but only the test
says so.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f

* perf(javascript): register the common functions lazily

The same argument as the type globals, applied to the other bulk registration a
fresh engine pays for. `ConfigureEngineWithCommonFunctions` installs
twenty-seven functions on every engine — `getVariable`, `toJson`, `newGuid`, the
base64 helpers, the deprecated GUID pair — and each `engine.SetValue(name,
Delegate)` builds an interop function wrapper around the delegate: a JavaScript
function object, plus the signature metadata and invoker lookups Jint resolves
per delegate type. An expression such as `variables.Foo` or a comparison calls
none of them.

This was recorded in the PR as deferred, because a lazy version had to keep the
`NonEnumerable` flag `SetValue(string, Delegate)` applies — otherwise the
functions would start appearing in `Object.keys(globalThis)`, which is
observable. Jint 4.15.3's `Engine.Advanced.AddLazyGlobal` takes a `PropertyFlag`
and is the post-construction counterpart of the options-time API used for the
type globals, so the flag is passed explicitly and the port is a line per
function. The CLR delegate is now created inside the factory as well, so a
function nothing reads costs one closure rather than a closure, a delegate and a
wrapper.

The laziness is invisible, and the tests say so rather than leaving it implied.
`AddLazyGlobal` installs the property itself eagerly and defers only its value,
so `in`, `hasOwnProperty` and `Object.getOwnPropertyNames` answer immediately
without materialising anything; `Object.keys(globalThis)` still omits them; the
descriptor still reports writable, non-enumerable and configurable; two reads of
a function are the same value, which is what says the factory ran once and the
result was stored rather than recomputed per read; and a script can still
overwrite one.

Not separately measured. The measured table in the PR predates this commit and
was not re-run for it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f

* refactor(javascript): let Jint expose enums as names and type its converters

`EnumToStringConverter` turned every CLR enum crossing into JavaScript into
`Enum.ToString()`. Jint 4.15 does that with `Interop.EnumConversion`, so the
converter goes away.

It is not only fewer moving parts. The converter could only see values that
crossed the interop boundary, and a constant read off a registered enum type
does not: `LogPersistenceMode.Include` came back as the underlying number while
the same value held in a workflow variable came back as `"Include"`, so
`mode === LogPersistenceMode.Include` was always false. The built-in switch
covers both directions and they now agree. Values written back to the CLR keep
accepting the member name and the number, as they did before.

That fix changes what an expression using a constant *numerically* produces, and
the hazard worth calling out is persisted state: a workflow variable holding a
number written from a constant before the upgrade no longer compares equal to
that constant after it, which reaches in-flight and resumed instances rather
than only new ones. Typed conversions hold in both directions, so activity
inputs and typed variable reads are unaffected. Recorded in the 3.8.0 changelog.

The two remaining converters register through the overload that declares the CLR
types they handle. A converter that does not declare them has to be offered
every value crossing the boundary, which costs the engine its compiled
member-read and method-invoker lanes for every wrapped .NET object; declaring
`byte[]` and `JsonElement` keeps those lanes for everything that cannot produce
one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f

* perf(javascript): register only the accessors an expression names

Every evaluation registers a getter and a setter for each variable in scope, a
getter for each workflow input, and a getter for each (activity, output) pair in
the enclosing container. The last of those is the expensive one: naming those
accessors walks every node of the workflow and resolves each against the
activity registry, so the cost of setting up an engine grows with the size of
the workflow rather than the size of the expression. Almost no expression uses
any of them.

Jint reports the free identifiers of a prepared program through
`Prepared<T>.ReferencedGlobals`, which is exactly the question being asked here:
an identifier the expression never mentions cannot be read from it. The
evaluator now prepares the script before configuring the engine — the parse was
already cached, so this only reorders work — and passes the set on
`EvaluatingJavaScript`. The accessor handler registers a name only if it is in
the set, and skips the workflow walk entirely when no identifier of the
`get{Output}From{Activity}` shape appears.

The filter is sound for an expression that names an accessor the way one is
meant to be named, and not for one that builds or reaches a name at run time.
Four such forms are detectable and each turns the filter off. A direct `eval`
call is reported as `HasDirectEvalCall`. An *indirect* `eval` call and the
`Function` constructor are deliberately not flagged as direct calls — Jint
reports the identifiers `eval` and `Function` in the set instead, and says so,
because that is the signal a host is meant to act on. Missing that signal is not
theoretical: `new Function('return getMyVariable()')()` would regress from
working to a `ReferenceError`, since Function-constructed code resolves only
against the global scope, and `var e = eval; e('getMyVariable()')` would fail the
same way. The fourth is a reference to `globalThis`, which reaches a global
without naming it. All four are pinned.

What stays undetectable is reaching the global object without naming it at all —
a sloppy-mode top-level `this`, or `[].constructor.constructor(…)`. An
expression written that way loses the generated accessor but not the data:
`getVariable(name)`, `getInput(name)` and `getOutputFrom(activityId, outputName)`
are always registered and reach the same values.

Note this does not replace the regular expression in
`ConfigureEngineWithVariables`, which extracts the member names in
`variables.Foo`. Those are not free identifiers and Jint deliberately does not
report them; the set only says whether `variables` itself is referenced.

The filtering is invisible from inside a script, which also makes it unprovable
from there, so two of the tests hold on to the engine and assert on the globals
directly.

While here, the cancellation token is registered as an engine constraint.
Passing it to `EvaluateAsync` only covers the awaiting part; Jint's own remarks
say the parameter cannot preempt the synchronous evaluation loop and point at
this constraint, so the token read as more coverage than it delivered. A
cancellation constraint is amortizable, so the interpreter keeps its tight-loop
fast path, and a default token registers nothing. #7891 supersedes this with the
full constraint set.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f

* perf(javascript): build marshalled variables in Jint's shaped representation

`ConvertToJsObject` built the JavaScript object a workflow variable is copied
into one `DefineOwnProperty` call at a time, with an explicit descriptor. That
lands the object in Jint's per-object property dictionary, so every marshalled
variable carries its own descriptors even though sibling variables and the
repeated nested payloads of one document present the same key set.

`JsObject.CreateFromEntries` defines the same writable, enumerable and
configurable properties but builds through the shared-layout path. Both wins
land inside a single evaluation — half the descriptor allocations, since the
explicit descriptor caused a second one inside `ValidateAndApplyPropertyDescriptor`,
and one shared layout across objects presenting the same keys. Nothing carries
across evaluations: layouts are interned per engine and per-node inline caches
live in per-engine handler trees that only engage on a second evaluation on the
same engine, which a fresh-engine-per-evaluation host never reaches.

Reaching the shared layout is silent: `CreateFromEntries` falls back to the
ordinary property dictionary whenever a key or a growth guard says the layout
cannot continue, and the object behaves identically either way. A test now asks
the engine whether the object actually got one. `Engine.Advanced.HasSharedShape`,
added in 4.15.3, is the part of that answer Jint documents as a contract — the
finer-grained `GetObjectRepresentation` names an internal representation that may
be renamed or subdivided in any release — and `CreateFromEntries` is one of its
three documented success cases. The assertion is not vacuous: against the
previous property-by-property build it is false, including for the
`CreateDataProperty` variant in #7892, because that object is created through
`Intrinsics.Object.Construct` rather than built as entries.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f

* test(javascript): run the scripting suites with Jint's host-contract verifiers on

Jint has a set of verifiers for the extension points it *trusts* — the ones it
cannot afford to re-check on a hot path, so a violation is otherwise silent.
They used to be compiled out of Release, which made "run your suite against a
Debug Jint" the only way to reach them, and Jint ships Release-only. 4.15.3
moves them behind an AppContext switch read once at type initialization, so a
host that never sets it pays nothing (the JIT folds the guards away) and a test
host can turn them on against the shipped package.

The one Elsa is subject to today is the object-converter type declaration this
PR introduced. Registering a converter with `AddObjectConverter(converter,
handledTypes)` promises the engine the converter produces values only for those
types, and in exchange the compiled interop lanes are kept for every member that
cannot produce one. Nothing links that promise to the converter's own
`TryConvert` switch: a case added there and not added to the registration is
silently skipped on exactly the members the declaration excluded, and honoured
everywhere else. `ByteArrayConverter` and `JsonElementConverter` are consistent
today; this is what would report it if they drifted.

Elsa defines no `ObjectInstance` subclass, so the rest of the verifiers have
nothing to check here yet. They are a standing guard for the day a host handler
or a satellite module adds one.

Wired as a module initializer, because the switch has to be set before the first
use of any Jint type. Duplicated across the three suites that reference
`Elsa.Expressions.JavaScript` rather than shared: `Elsa.Testing.Shared.Integration`
would be the obvious home, but it is a published package, and a module
initializer there would flip a process-wide switch for every external consumer
of it as well. A one-line test pins that the initializer ran, since Elsa
satisfies the contracts it is subject to and nothing else would notice the
checks going away.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Sipke Schoorstra <sipkeschoorstra@outlook.com>
2026-08-17 01:44:03 +02:00
Marko Lahma 58c3c799f8
Stop registering colliding and unreachable type globals in JavaScript expressions (#7893)
* fix(javascript): stop registering colliding and unreachable type globals

`Engine.RegisterType` exposes a .NET type under `Type.Name`. That name is not
always usable, and the type registrations are contributed by several
independent handlers whose sets overlap.

* `IDictionary<string, string>` and `IDictionary<string, object>` are both named
  ``IDictionary`2``, so the two registrations claimed the same global and the
  later one silently won. Neither is reachable from a script: a backtick cannot
  appear in an identifier.
* `byte[]` is named `Byte[]`, which is likewise unreachable.
* `DateTime`, `DateTimeOffset`, `TimeSpan`, `Guid` and `LogPersistenceMode` are
  part of both the common type set and the default workflow variable descriptor
  set, so each was constructed and assigned twice for every expression
  evaluation.

`RegisterType` now skips types whose name is not usable as a JavaScript
identifier, and skips a type that is already registered under that name. Type
aliases used by the TypeScript definition endpoint are unaffected — they are
maintained by `ITypeAliasRegistry` and are independent of this registration.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

* fix(javascript): leave an already-occupied global name alone

`RegisterType` skipped a name only when it already held a `TypeReference` for the
same type, so anything else under that name was replaced. That includes a global
the host installed through the per-evaluation `configureEngine` callback,
`JintOptions.ConfigureEngine` or `JintOptions.RegisterType` — all of which run
before the built-in registrations, since those are contributed by handlers of
`EvaluatingJavaScript`. Silently overwriting a host global is surprising and the
host has no way to win.

`RegisterType` now leaves any occupied name alone. That keeps the duplicate
suppression the check was written for — registering the same type twice is still
a no-op, so the overlapping handlers stop describing the same types through
reflection on every evaluation — and additionally makes the host global win. It
also agrees with #7895, where the registrations move to engine construction and
every host extension point runs after them.

Two tests pin the behaviour: a host value set under a built-in type's name
survives the built-in registrations, and `RegisterType` installs a
`TypeReference` that a second registration leaves untouched.

The remark about unusable type names is tightened while here: ``IDictionary`2``
and `Byte[]` can be reached through bracket notation if they are registered, so
the reason to skip them is that they cannot be written as identifiers, and that
every constructed generic type of the same arity claims the same global.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:23:31 +02:00
Marko Lahma 66b9079d3e
Bound JavaScript expression execution (#7891)
* feat(javascript): bound JavaScript expression execution

JavaScript expressions were evaluated with no execution constraints at all: no
timeout, no statement limit, no memory limit, no recursion limit, and the
ambient `CancellationToken` — already available at the call site and already
passed into `IJavaScriptEvaluator.EvaluateAsync` — was never handed to Jint.
An expression as simple as `while (true) {}` therefore occupied the calling
thread for the lifetime of the process, and cancelling the workflow did not
stop it.

This adds:

* `JintOptions.ExecutionTimeout` — wall-clock limit for a single expression,
  defaulting to 30 seconds. Deliberately generous so that existing expressions
  are unaffected; set to `null` to remove the limit.
* `JintOptions.MaxStatements`, `JintOptions.MemoryLimit` and
  `JintOptions.MaxRecursionDepth` — opt-in resource limits, off by default.
* The cancellation token is now passed to Jint, so cancelling a workflow aborts
  a script that is still running.

The security assessment documents already described a JavaScript execution
timeout as present; they now describe what is actually configurable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

* test(javascript): bound the constraint tests and cover the memory limit

Three of the execution-constraint tests removed the execution timeout entirely
and then ran `while (true) {}`, relying solely on the constraint under test to
stop them. If that constraint regressed, the test did not fail — it ran until
the CI job was killed, taking the rest of the suite with it.

Each such test now registers a generous 30 second failsafe timeout instead of
disabling the timeout. That is two orders of magnitude more than any of these
constraints needs (the slowest aborts in ~270 ms), so it cannot become a flaky
failure on a loaded machine, and `AssertAbortedByAsync` reports a failsafe trip
as exactly that rather than as an unexplained exception type mismatch.

The cancellation test also no longer races a wall-clock timer against engine
construction: the script signals the token itself through a host function, so
cancellation is guaranteed to land while the expression is running. The test
went from a 250 ms wall-clock wait to 5 ms and has no timing dependency left.

Adds the missing `MemoryLimit` test — the one configurable limit the suite did
not exercise. Doubling a string crosses the limit within a couple of dozen
statements, so it asserts `MemoryLimitExceededException` in ~40 ms and bounds
how far past the limit the process can get before the check fires.

Finally, the `ExpressionExecutionContext` is now built on the test host's
`IServiceProvider` rather than a throwaway empty one, matching every other test
in this project. An empty provider does not reflect real evaluation and can hide
failures in notification handlers that resolve services.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:21:19 +02:00
Marko Lahma 2e9a3ad04f
Trim per-evaluation work in the JavaScript evaluator (#7892)
* perf(javascript): trim per-evaluation work in the JavaScript evaluator

A Jint engine is built for every expression evaluation, so anything done during
setup is paid for on every evaluation. Four pieces of that work are avoidable:

* The three `IObjectConverter` implementations are stateless but were allocated
  fresh for every engine. They are now shared static instances.

* Every prepared-script cache lookup — including hits — computed a SHA-256 hash
  of the expression text, base64-encoded it and concatenated a prefix, purely to
  build the cache key. Using a dedicated key type instead keeps the entries
  distinct from other users of the shared cache while letting the expression
  itself be the key, so a hit is a dictionary lookup. Looking the entry up
  directly rather than through `GetOrCreate` also keeps the factory closure off
  the hit path.

* `ObjectConverterHelper.ConvertToJsObject` built an explicit `PropertyDescriptor`
  per property and called `DefineOwnProperty`. `CreateDataProperty` is public,
  produces exactly the same writable/enumerable/configurable descriptor, and is
  the engine's fast path for it.

* The variable write-back resolved the workflow input names — walking the whole
  activity execution context ancestor chain — before checking whether there was
  anything to write back. Only variables the expression actually referenced are
  copied into the engine, so for the common case of an expression that never
  mentions `variables.` the container is empty and all of that work is wasted.
  The input names are also now looked up through a set rather than a list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

* test(javascript): pin the parse-failure test to the same exception every time

Calling `ThrowsAnyAsync<Exception>` twice only proved that both evaluations threw
something, which is exactly the assertion a poisoned cache would still satisfy:
had the failed preparation left a null or half-built entry behind, the second
evaluation would have failed too, just with a different exception. The test now
captures both exceptions and asserts they are the same type with the same
message, so "keeps reporting the same parse failure" is what is actually checked.

The message is stable to compare: it is `Could not prepare script: Unexpected end
of input (1:9)`, and since both evaluations run the identical script literal the
position is identical as well. No file or path detail is involved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:16:38 +02:00
Marko Lahma 14173a19eb
Update Jint to 4.15.3 and stop blocking the calling thread on promises (#7894)
* chore(javascript): update Jint to 4.15.3 and stop blocking on promises

`Engine.Evaluate(...).UnwrapIfPromise()` blocks the calling thread while the
engine's event loop drains, which is exactly the wrong thing to do inside an
`async` method — an expression that awaits a .NET `Task`, such as one calling
`getSecret()`, held a thread pool thread for the duration of the I/O.
`Engine.EvaluateAsync` awaits the returned promise instead, and takes the
cancellation token while it is at it.

The Jint version is moved from 4.4.2 to 4.15.3. `EvaluateAsync` arrived in
4.14.0, but the pin lands past 4.15.2 deliberately: once expressions genuinely
suspend and resume instead of draining the event loop on the calling thread,
they exercise the async suspension machinery 4.15.2 corrected — an `await` on a
right-hand side no longer stores the suspension sentinel, async generators and
`for await...of` preserve loop iteration state across a suspension, and a
suspension node is unwrapped correctly. Shipping the non-blocking change on an
earlier 4.14/4.15 would enable exactly the code paths those releases fixed.

One default changed along the way: since 4.14 `Interop.ArrayConversion` defaults
to `LiveView`, so a CLR array reaches script as a live view over the original
array rather than as a copy. That is observable — a script that sorts an array
would now reorder the workflow's own array, and the value round-trips back as
its original element type rather than as `object[]`. The evaluator therefore
pins the previous `Copy` behaviour so the upgrade is not a behavioural change;
hosts that prefer the live view can opt in through
`JintOptions.ConfigureEngineOptions`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

* test(javascript): pin the array copy lane and adopt JsString.Create

The array audit. The parent commit fixes `ArrayConversion` to `Copy`, because
4.14 changed the default to `LiveView` and the two differ in behaviour a script
can observe. The existing tests assert what a copy *produces*, which a live view
also satisfies for a value nothing mutates; asking the engine how many
conversions of each kind it performed (4.15.1's interop conversion counters)
pins the lane itself. The second assertion is the more interesting one: an
ordinary evaluation converts no CLR array at all, because Elsa converts
collection-valued variables itself in `ObjectConverterHelper` long before Jint's
array lane could see them. That makes the `ArrayConversion` setting a narrow
compatibility pin rather than something every evaluation depends on.

`JsString.Create` (public since Jint 4.15.3) is adopted in
`JsonElementConverter`, where the string case was the only one still routed
through `JsValue.FromObject` — re-entering the whole conversion pipeline, the
registered object converters and this one included, to arrive at the same call
the number and boolean cases beside it already make directly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:02:04 +02:00
github-actions[bot] 30ab056745
Refresh codebase wiki (#7922)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-17 00:58:59 +02:00
Sipke Schoorstra 5fd3a074c9
Resolve NU1903 warning
warning NU1903: Package 'SSH.NET' 2025.1.0 has a known high severity vulnerability, https://github.com/advisories/GHSA-q939-rpr3-3284
2026-08-15 23:52:27 +02:00
Sipke Schoorstra 7eaf056d65
test(bpmn): prove execution state and the work ledger survive real persistence (#7947)
* test(bpmn): prove BpmnExecutionState pruning and ledger rehydration survive a real suspend/resume

Prune() was already being called before every persisted write in BpmnScopeHost, but nothing proved
it, and no BPMN test had ever exercised a genuine rehydration: every existing scenario used the
in-memory WorkflowState object straight from the previous run. Add tests that round-trip a
suspended scope's state through Elsa's own IWorkflowStateSerializer -- the boundary that mangled
values before -- and assert the persisted BpmnExecutionState stays bounded across many evaluations
and that a resumed scope with two live units of work matches each completion back to its binding
through the rehydrated BpmnWorkLedger. Both tests were confirmed red by mutation-testing away
Prune() and by returning an empty ledger from BpmnScopeMemory.Load.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): prove a nested scope's ledger survives real persistence

Both new tests in the prior commit suspended a root scope with live work
outstanding, but neither crossed a nested scope's own ledger through Elsa's
real IWorkflowStateSerializer -- the intersection of the two things that
have actually broken here: the handle-to-context map, and the serializer
boundary. Add a nested parallel split/join, blocking on both branches inside
an embedded subprocess, round-tripped through the serializer between each
branch's completion, and confirmed red by returning an empty ledger from
BpmnScopeMemory.Load and green with it restored.

Also extract the start/split/left/right/join/after/end topology shared
verbatim by ParallelSplitAndJoin and ParallelSplitAndJoinBlocking into one
private builder parameterised by the branches' work, keeping both public
factories unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 20:45:07 +02:00
Sipke Schoorstra 6a3d65ff7e
feat(bpmn): the work binder and the elsa: binding format (#7946)
* feat(bpmn): bind BPMN work declarations to Elsa activities

Turns the reader's BpmnWorkBinding declarations into the activity nodes a
BpmnProcess scope runs. Six of the seven kinds bind automatically: TimerWait to
Delay, MessageWait/SignalWait to Event, MessagePublish to PublishEvent,
CallProcess to DispatchWorkflow, NestedProcess to a nested BpmnProcess. The
seventh, UnboundTask, is an authoring decision and is read from a new elsa:
vendor extension inside the document, so an exported .bpmn is self-contained.

Every binding for a scope is bound whatever its slot, so a ScopeListener needs
no special case. Each binding gets its own freshly built activity with a
scope-qualified id: ActivityVisitor skips an activity it has already collected,
so one instance shared between two scopes would leave the second scope with no
child in Elsa's identity graph.

The binder lives in Elsa.Bpmn.Interchange because BpmnWorkBinding is a
Bpmn.Interchange type; binding it in Elsa.Bpmn would pull the interchange
library into the execution module's closure, which is the split D12 draws.

Every ambiguity resolves loudly: an unbound task, a dead binding declaration, a
malformed ISO-8601 duration, a call activity with nothing to call, and an
activity type nothing registered all refuse at bind time rather than producing a
process that runs to completion doing none of what the document says.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): declare document variables on the bound scope, refuse duplicate input names

BpmnWorkBinder.BindScope never copied BpmnProcessDefinition.Variables onto the
produced BpmnProcess's Elsa Variables, so a document-declared collection variable
was Absent to IBpmnVariableReader and a collection-mode multi-instance over it
faulted the element instead of running once per item. BindScope now declares an
Elsa Variable for each document variable, seeding the declared default as the
JsonElement it already is.

BpmnActivityBindingFormat.Read silently let a second <elsa:input name="..."> with
a duplicate name overwrite the first rather than refusing it, unlike every other
malformed-document case this binder already refuses. It now throws
BpmnBindingException naming the binding and the duplicated input, and the XML doc
now states that rule plus the (verified) XML text-node escaping that already
applies to input JSON.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): carry every declared activity input through the binding format

BpmnActivityBindingFormat.Write only found properties whose CLR type derives
from Input, silently dropping attribute-declared inputs like Switch.Cases from
an export. Read accepted any <elsa:input name="..."> without checking the
activity declares it, so a mistyped or stale name imported silently with the
configuration missing since Elsa's deserializer ignores unknown members. Both
now go through IActivityDescriber.GetInputProperties, the same enumeration
ActivityDescriptor.Inputs is built from, so Write and Read agree on what an
activity's inputs are and Read refuses a name that enumeration does not
report.

Also makes BpmnWorkBinder.RefuseUnusedDeclarations filter its loop explicitly
with .Where(...) instead of an implicit if, per static analysis.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(bpmn): filter undeclared input names explicitly

Express the undeclared-input-name check as an explicit Where filter
instead of an implicit filter inside the loop body, and report every
undeclared name at once rather than only the first. Also fix the
refusal message, which previously named the activity type twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(bpmn): describe the input payload shape accurately

The XML doc on BpmnActivityBindingFormat claimed every <elsa:input> is the
{"typeName":...,"expression":...} wrapper a stored workflow definition uses.
That only holds for Input<T>-typed properties: an [Input]-attributed
plain-typed property such as Switch.Cases is serialized as its own JSON
shape (an array), not the wrapper, which Write already does correctly and
the round-trip test already covers. Correct the doc to describe the payload
as the configured activity serializer's output for that input, dependent on
how the activity declares it, and add a second short example showing the
attribute-declared shape.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 17:38:03 +02:00
Sipke Schoorstra fa989c29d0
test(component): correct coverage gate to 23
The gate was lowered to 24 against a measured 24.74% on the release/3.8.0
merge. Integrating origin/main (#7945, the BpmnProcess container activity)
then added further uncovered production code, taking the measured total to
23.98% and putting it back under the gate.

Set the gate to 23 so it again sits just below actual coverage. 23.98% was
measured with PublishEvent_WithPayload_TransmitsPayloadToConsumer excluded
(see #7749), so it is a lower bound: including that test can only raise the
figure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 16:23:23 +02:00
Sipke Schoorstra 2ed957203a
Merge origin/main into main 2026-08-14 15:38:06 +02:00
Sipke Schoorstra 1ec6f0b1b4
test(component): lower coverage gate to 24 after the 3.8.0 merge
Merging release/3.8.0 pulled ExternalAuthentication, AI and Secrets into this
project's coverage denominator through Elsa.Server.Web, without matching
component-test coverage. Measured total line coverage went from 25.98% on
f99f37407 to 24.74% on the merge, crossing below the 25 gate and failing CI in
both pr.yml (./build.cmd Compile Test, NUKE EnableCollectCoverage) and
packages.yml (/p:CollectCoverage=true).

Lower the gate to 24 so it sits just under actual coverage rather than
disabling it. Raise it back towards 25 as component coverage for the merged-in
modules lands.

Note: Elsa.Resilience.Core.UnitTests is also below its threshold (48.17% vs
49), but that is pre-existing and unrelated -- it measures identically on
f99f37407 and on the merge. Left untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 15:13:21 +02:00
Sipke Schoorstra 65fe688350
feat(bpmn): the BpmnProcess container activity (#7945)
* feat(bpmn): scope variables, trigger opt-out and composability for BpmnProcess

Completes the container W2 left minimal, with the four things it deferred.

Scope variables. BpmnScopeVariables implements IBpmnVariableReader over the
scope's memory register, walking outward so an inner scope sees the enclosing
one's data, and BpmnScopeHost now declares ScopeVariables and hands the reader
to every snapshot. The read is three-valued: false for a name nothing in scope
declares, Null for a declared variable holding nothing, and StoredExternally
for a value JSON cannot carry.

That last case deviates from the issue, deliberately. The issue names the
unmaterialized-driver case, which is not detectable from the container's side:
PersistentVariablesMiddleware loads with no excludeTags, and
VariablePersistenceManager marks a block IsInitialized before testing the
exclusion, so a variable whose driver was never read is indistinguishable from
one whose driver returned null. Closing that needs a change to
Elsa.Workflows.Core, which is out of bounds here, so the reader answers only
what the block actually says and the XML doc records why. The route it does
have is real and in the same spirit: a value the host holds and cannot put on
the wire faults loudly rather than reading as an empty collection.

Trigger opt-out. BpmnProcess.IsRootScope names the BPMN meaning of Elsa's
CanStartWorkflow rather than adding a second flag that could disagree with the
gate TriggerIndexer actually reads. It is off unless something says otherwise,
and the applier refuses to start a BpmnProcess that claims root position as
another scope's work: the damage a mis-flagged subprocess does happens at
publish time, so repairing the object graph at runtime would leave the trigger
registered while every test went green. ITrigger itself remains #7929.

Composability and outcomes. A BpmnProcess in a Flowchart runs and the flowchart
carries on (D11), and a nested transaction completing Cancelled reaches its
parent's completion callback with that outcome intact, which is the only reason
the parent routes the cancel boundary rather than the ordinary sequence flow.

Every guard was mutation-tested red before green: both non-Present answers of
the reader, the reader left unwired, the opt-out's default flipped (7 tests red,
including the pre-existing nested-scope ones), the refusal removed, and the
outcome dropped at each end of the trip to the parent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): apply review findings on scope variables, command batching, and outcome doc

Read a scope variable through Elsa's configured serializer (via IPayloadSerializer,
serialized against the value's own runtime type so a polymorphic value is not wrapped
in Elsa's type-tagged envelope) instead of bare JsonSerializerDefaults, so a value only
Elsa's converters can carry no longer collapses to StoredExternally. Refuse a root-scope
StartWork before any command in the batch is applied, not mid-list, so a refusal cannot
leave scope memory partially mutated under ContinueWithIncidentsStrategy. Document that
BpmnProcess completes with only its interpreter outcome, so a default/null-port
Flowchart connection never fires from it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(bpmn): filter the pre-scan explicitly

Use commands.OfType<BpmnHostCommand.StartWork>() in ApplyAsync's
root-scope pre-scan instead of a foreach + type-check, matching the
static analysis suggestion. The apply loop below is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 12:37:35 +02:00
Sipke Schoorstra 52a7061f89
Merge release/3.8.0 into main
Brings the 3.8.0 release line into main, including the package-manifest
runtime-kind mechanism (src/PackageManifest.props + src/PackageManifestHints.cs)
that main did not have. All 72 manifest-producing packages now declare
compatibility.runtimeKinds = ["elsa.server"]; the two Bpmn modules added on
main pick this up automatically via their ShellFeatures directory.

Conflict resolutions:
- .specify/feature.json, CONTEXT.md, ROADMAP.md, build/_build.csproj: took
  main's, which is newer in every case. Verified byte-identical to main
  afterwards, so nothing from the release branch was dropped.
- NuGet.Config: union of package sources, minus valence-consolelogstream-feedz.
  main removed that feed deliberately in 7389e0a67 and now consumes
  ConsoleLogStreaming 1.1.0 from nuget.org.
- Elsa.sln: union of main's Bpmn projects and the release branch's
  ExternalAuthentication projects; the two sets are disjoint.

Reverted an unintended revert:

release/3.8.0 had lost commit 33181b2c9 ("test: cover Oracle bulk upsert SQL
generation") through an evil merge in c557c455a. That commit is present at the
merge base, so git resolved the release branch's older content as an
intentional change and would have silently undone it on main. It is three
coupled pieces:

  - test/unit/Elsa.Persistence.EFCore.UnitTests (deleted, plus its Elsa.sln
    project declaration and NestedProjects entry)
  - InternalsVisibleTo("Elsa.Persistence.EFCore.UnitTests")
  - the fix itself in BulkUpsertExtensions.GenerateOracleUpsert: internal
    visibility, ISqlGenerationHelper.DelimitIdentifier quoting, and explicit
    CAST(... AS NVARCHAR2(...)) on string columns

Dropping the third would have been an Oracle runtime regression: unquoted
identifiers lose case, and ODP.NET binds .NET strings as VARCHAR2 while Elsa's
Oracle migrations declare NVARCHAR2, causing a datatype mismatch. Merge base
and main are identical for that file and every hunk on the release side is a
revert plus cosmetics, so main's version was kept in full.

Accepted deliberate release-branch changes, verified as real refactors rather
than losses: AI EF Core migrations moved into the provider projects
(5c0d8b0f4), and AIPersistenceFeature.cs renamed to
EFCoreAIPersistenceShellFeatureBase.cs (ShellFeatures/ still present, so
manifest generation is unaffected).

Verified: dotnet build Elsa.sln succeeds with 0 errors and 2 pre-existing
NU1903 warnings; Elsa.Persistence.EFCore.UnitTests passes 1/1; all 72 emitted
manifests declare elsa.server and Elsa.Api.Common emits none.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 11:59:35 +02:00
Sipke Schoorstra 33ad4af835
fix(manifest): stop Elsa.Api.Common publishing a runtime-kind-less manifest
Elsa.Api.Common has a ShellFeatures directory, so it picked up the manifest
generator and emitted an elsa-package.json. It was excluded from the shared
hints file, though, so that manifest carried no compatibility.runtimeKinds.
Runtime kind matching is asymmetric: when an image declares runtime kinds, a
feature whose effective list is empty can never intersect and is excluded.
An empty manifest is therefore worse than no manifest at all.

Elsa.Api.Common is shared API plumbing rather than a selectable catalog
module, so opt it out of manifest generation entirely via
GenerateElsaPackageManifest. That also collapses the two overlapping
ItemGroup conditions into one and removes the hardcoded project name from
them, leaving GenerateElsaPackageManifest as the single opt-out knob.

The remaining 70 ShellFeatures projects are unaffected and continue to
declare elsa.server.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 11:11:58 +02:00
Sipke Schoorstra f99f37407c
feat(bpmn): host port and command applier (#7942)
* feat(bpmn): host-side applier for the Bpmn.Semantics port

Translates the interpreter's three host commands onto ActivityExecutionContext
and feeds its four entry points, plus the minimum BpmnProcess container needed
to exercise them end to end through IWorkflowRunner.

StartWork schedules the bound activity, CancelWorkSubtree calls the public
CancelActivityAsync extension (already recursive), and SignalEnclosingScope
sends a BpmnScopeSignal up the ancestor chain. OnWorkFaulted rides the
FaultSignal seam: it asks the interpreter what BPMN made of the fault and calls
StopPropagation only on a Caught disposition, leaving a Propagated one strictly
alone so an enclosing scope or the incident strategy takes it.

A unit of work is keyed on the child ActivityExecutionContext.Id, recorded in
the scope's own persisted ledger, never on Tag: the completion-callback dispatch
rewrites the receiving context's Tag, so a nested scope wears a different tag
than its parent remembers it by. Interpreter correlation travels on the child's
context rather than on the shared activity instance.

Evaluations go through one queue per workflow instance, so a scope signalled
mid-apply is drained after the command list rather than re-entering the
interpreter. Commands are applied in the order returned.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): cover the teardown refusal path

Adds a focused unit test that drives BpmnWorkTeardown.CancelSubtreeAsync
into the NotSupportedException branch by constructing a real context tree
with a scheduled-but-not-invoked descendant, so a regression that silently
drops the detection is caught. Also records why BpmnWorkLedger's
append-only, handle-keyed Records list cannot strand a context on a
duplicate StartWork for a live (BindingRef, IterationId) slot, a case the
port's own guarantee makes unreachable from this applier.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): keep a refused teardown from stranding ledger state

A subtree cancellation refused with NotSupportedException is absorbed
into an incident under ContinueWithIncidentsStrategy rather than
crashing, so the end-of-command ledger save was being skipped and the
persisted ledger kept claiming work BPMN had just torn down. Save the
ledger removal before the possible throw instead of after, so a later
completion callback for the stranded activity finds no live record and
is discarded instead of being fed to the interpreter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 01:42:31 +02:00
Sipke Schoorstra 7389e0a674
fix(deps): consume ConsoleLogStreaming 1.1.0 from nuget.org (#7941)
Elsa.Diagnostics.ConsoleLogs and Elsa.Dashboard.Api are packable and are
published to the Elsa preview feed on every push to main, and to nuget.org on
release. Both pinned ConsoleLogStreaming.* at 1.0.0-preview.13, a version
published only to the private valence-works feed, so that pin was baked into
the shipped nuspecs.

Restore did not hard-fail for consumers: NuGet floated the >= constraint up to
the nuget.org 1.0.0 and emitted NU1603 for each package, an error under
TreatWarningsAsErrors. The quieter problem was that consumers silently ran a
different build of the library than CI compiled against.

ConsoleLogStreaming 1.1.0 has now been released publicly, carrying the work
that had only ever reached the private feed as previews: the new
ConsoleLogOptions.StreamReleaseInterval option, stream-release gate flood
throttling, and dead subscriber drop accounting. Pinning to it makes the
shipped dependency both resolvable and current, so the private feed and its
package source mapping are no longer needed.

Verified: Elsa.sln restores clean; Elsa.Dashboard.Api builds with 0 warnings;
64 tests pass across the ConsoleLogs and Dashboard.Api unit and integration
projects; the packed nuspec declares 1.1.0 on all three target frameworks; and
a clean consumer project restoring the packed output against nuget.org alone,
with TreatWarningsAsErrors, resolves 1.1.0 with no warnings.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:44:13 +02:00
Sipke Schoorstra 22886809e0
test(bpmn): guard that Elsa.Bpmn never duplicates the library (#7939)
* test(bpmn): guard against Elsa.Bpmn* reimplementing Bpmn.* library types

Adds a reflection-based architecture test asserting no type under Elsa.Bpmn
or Elsa.Bpmn.Interchange shares a type name with Bpmn.Model/Bpmn.Semantics
(and Bpmn.Interchange for the interchange assembly), plus a positive check
that both assemblies still depend on their respective library packages -
closing the gap that let elsa-foundation grow a parallel BpmnElement/BpmnGraph
semantics core undetected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): move the library-duplication guard to the interchange test project

Elsa.Bpmn.UnitTests referencing Elsa.Bpmn.Interchange (plus three
redundant PackageReferences already flowing transitively) inverted the
layering the guard exists to protect. Elsa.Bpmn.Interchange.UnitTests
already sees both assemblies transitively with no new references, so
the guard moves there unchanged apart from namespace and doc. Also
drops the silent null-forgive on Assembly.GetName().Name in favor of
an explained fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): drop the dependency-direction assertions the compiler already enforces

Mutation testing showed AssertDependsOnPackage and the two positive facts using
it can never go red: any state that would trip them fails the test project's
build first (CS0234), because the guard's own typeof bindings already require
the Bpmn.* packages. Delete the decorative facts and the now-unused helper,
and document that the typeof bindings are load-bearing on purpose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): prove the duplication detector actually detects

Split AssertNoTypeNameCollisions into a thin assertion wrapper around a
new FindTypeNameCollisions helper, and add a fact that points the
detector at the test assembly (which carries a deliberately colliding
BpmnGraph fixture) so CI sees the guard fail as well as pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 04:09:53 +02:00
Sipke Schoorstra 735e953dc5
feat(bpmn): scaffold Elsa.Bpmn and Elsa.Bpmn.Interchange modules (#7938)
Adds the valence-works/bpmn package feed and pins Bpmn.Model,
Bpmn.Semantics and Bpmn.Interchange at 0.1.1-preview.19, then wires
two new module projects consuming those libraries without
reimplementing anything they provide (D12): Elsa.Bpmn for BPMN
execution and Elsa.Bpmn.Interchange for XML import/export, kept as a
separate package so hosts that only execute BPMN don't take the XML
reader. Both are marked IsPackable=false until the Bpmn.* packages
are published to nuget.org. Test projects are added for both, unit
and integration, and all four are wired into Elsa.sln so PR CI
discovers and runs them. This unblocks #7925 and the rest of the
BPMN runtime work in #7909.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 01:17:46 +02:00
Sipke Schoorstra 749b6491fc
fix(core): a throwing FaultSignal handler must not escape the middleware (#7924)
* fix(core): a throwing FaultSignal handler must not escape the middleware (#7911)

The signal is sent from inside the catch whose whole job is to stop exceptions
escaping the activity pipeline. A handler that threw went straight through it:
no incident, no strategy, and the original fault lost along with it.

The send is now guarded. A handler that throws is treated as not having handled
the fault, so the incident strategy runs exactly as it would with no handler
present. That is the conservative direction: a handler that failed part way
through may have left the faulted activity in any state, and an incident is a
better answer than silence. Its exception is logged at error level, because a
broken fault handler is a defect in its own right rather than a workflow
outcome.

Covered both ways, since a handler that already claimed the fault before
throwing is the case that could plausibly have been mistaken for success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(core): let cancellation from a fault handler propagate

The handler guard caught OperationCanceledException along with everything else,
so a cancellation raised while an ancestor was being offered a fault was logged
as a broken handler and handed to the incident strategy. A deliberately
cancelled run reported itself faulted.

Cancellation is excluded now, matching how this repository already keeps the two
apart: the workflow-level exception middleware cancels and rethrows before its
general catch, and WorkflowRunner declines to record cancellation as the
workflow's exception.

Caught by review on #7924.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(core): pin that a handled fault stays in the execution log

The justification for dropping the incident is that the journal keeps the
evidence. That was asserted in several places and guarded nowhere.

It holds because ExecutionLogMiddleware writes the Faulted entry from a catch
that rethrows, and it is registered inside ExceptionHandlingMiddleware, so the
entry lands before the fault is ever offered to an ancestor. Swapping those two
registrations would make a handled failure disappear from the record with
nothing failing, which is what this test now prevents.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 05:34:19 +02:00
Sipke Schoorstra 5faba75906
fix(core): a fault a container claimed is not an incident (#7923)
* fix(core): a fault a container claimed is not an incident (#7911)

RecoverFromFault reset the counts and the status but left behind the two other
things Fault recorded: the ActivityIncident and the exception. So a container
that successfully handled a child's fault still left the workflow carrying an
incident.

That is not cosmetic. Code reads a non-empty WorkflowExecutionContext.Incidents
as "this workflow failed" without looking further; HttpWorkflowsMiddleware is
one, and it hands the caller a fault response. A workflow whose container caught
the error and finished normally was reported to its caller as failed.

RecoverFromFault is now the inverse of Fault: it removes the incident Fault
appended, matched on this activity's node id and most recent first so an
activity that faults, recovers and faults again keeps the incident that was
never recovered, and it clears the recorded exception so the activity does not
sit in Running carrying one.

The execution log still records the failure, so nothing is hidden from anyone
reading the journal. Two integration assertions that encoded the old behaviour
are updated; they were written from the reasoning this change corrects.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(core): tie an incident to the execution that raised it, not its node

Recovery matched the incident to remove on ActivityNodeId, which identifies the
static workflow node rather than an execution of it. A node inside a loop,
retried, or run concurrently raises one incident per execution, all under the
same node id, so recovering one execution could remove another's incident and
leave its own behind.

ActivityIncident now carries the ActivityInstanceId of the execution that raised
it, and recovery matches on that. Within a single execution the most recent is
still taken, so fault, recover, fault again keeps the incident that was never
recovered. The property is optional: an incident recorded against the workflow
itself has no execution, and so do incidents persisted before this existed.

Caught by review on #7923.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-client): mirror ActivityInstanceId on the client incident model

The server model gained the property in the previous commit and the API client
carries a hand-maintained copy of it. Left alone, a client deserializing an
incident would silently drop the only field that says which execution raised it.

Also records two consequences of recovery that were implicit: it relies on the
incident collection preserving insertion order to pick an execution's newest
incident, which holds only because the collection is list-backed; and clearing
the exception also clears it from the activity's execution record, which is
intended for the same reason the incident goes, with the journal keeping the
evidence either way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 05:20:22 +02:00
Sipke Schoorstra c58fe8770f
Merge pull request #7917 from Shivamkmr8/improve-revert-version-allocation
fix: use last version for revert version allocation
2026-08-12 01:34:32 +02:00
Sipke Schoorstra 61f9d2c39d
Merge pull request #7921 from elsa-workflows/codex/update-wiki
[codex] Refresh codebase wiki
2026-08-12 00:53:55 +02:00
github-actions[bot] 1754253fb9 Refresh codebase wiki 2026-08-11 22:25:06 +00:00
Sipke Schoorstra 5068e0b631
Merge pull request #7920 from elsa-workflows/claude/fix-wiki-link-check
docs(wiki): point the scheduling link at StartupTasks
2026-08-12 00:13:39 +02:00
Sipke Schoorstra 70690ad3d3
docs: refresh roadmap 2026-08-12 00:09:13 +02:00
Sipke Schoorstra ff72352d8c
docs(wiki): point the scheduling link at StartupTasks
The Update Wiki workflow has failed on every run since 2026-07-28, on one
broken link:

  doc/wiki/http-scheduling-resilience.md:
    ../../src/modules/Elsa.Scheduling/HostedServices -> missing

Elsa.Scheduling/HostedServices was removed in 8e301d4e1, where
HostedServices/CreateSchedulesBackgroundTask.cs became
StartupTasks/CreateSchedulesStartupTask.cs. The wiki page kept pointing at the
old directory.

Point it at StartupTasks, which is where that work lives now.

The workflow's link check only validates wiki files that changed in the push,
so its silence is not evidence the rest is sound. Ran the same validator over
all 22 files in doc/wiki: this was the only break, and everything now resolves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 00:02:13 +02:00
Sipke Schoorstra 0412555b6e
Merge pull request #7913 from elsa-workflows/claude/cranky-boyd-269f7e
feat(core): let a container activity handle a child's fault via FaultSignal
2026-08-11 23:43:26 +02:00