Commit graph

7245 commits

Author SHA1 Message Date
Marko Lahma 2e9a3ad04f
Trim per-evaluation work in the JavaScript evaluator (#7892)
* perf(javascript): trim per-evaluation work in the JavaScript evaluator

A Jint engine is built for every expression evaluation, so anything done during
setup is paid for on every evaluation. Four pieces of that work are avoidable:

* The three `IObjectConverter` implementations are stateless but were allocated
  fresh for every engine. They are now shared static instances.

* Every prepared-script cache lookup — including hits — computed a SHA-256 hash
  of the expression text, base64-encoded it and concatenated a prefix, purely to
  build the cache key. Using a dedicated key type instead keeps the entries
  distinct from other users of the shared cache while letting the expression
  itself be the key, so a hit is a dictionary lookup. Looking the entry up
  directly rather than through `GetOrCreate` also keeps the factory closure off
  the hit path.

* `ObjectConverterHelper.ConvertToJsObject` built an explicit `PropertyDescriptor`
  per property and called `DefineOwnProperty`. `CreateDataProperty` is public,
  produces exactly the same writable/enumerable/configurable descriptor, and is
  the engine's fast path for it.

* The variable write-back resolved the workflow input names — walking the whole
  activity execution context ancestor chain — before checking whether there was
  anything to write back. Only variables the expression actually referenced are
  copied into the engine, so for the common case of an expression that never
  mentions `variables.` the container is empty and all of that work is wasted.
  The input names are also now looked up through a set rather than a list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

* test(javascript): pin the parse-failure test to the same exception every time

Calling `ThrowsAnyAsync<Exception>` twice only proved that both evaluations threw
something, which is exactly the assertion a poisoned cache would still satisfy:
had the failed preparation left a null or half-built entry behind, the second
evaluation would have failed too, just with a different exception. The test now
captures both exceptions and asserts they are the same type with the same
message, so "keeps reporting the same parse failure" is what is actually checked.

The message is stable to compare: it is `Could not prepare script: Unexpected end
of input (1:9)`, and since both evaluations run the identical script literal the
position is identical as well. No file or path detail is involved.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:16:38 +02:00
Marko Lahma 14173a19eb
Update Jint to 4.15.3 and stop blocking the calling thread on promises (#7894)
* chore(javascript): update Jint to 4.15.3 and stop blocking on promises

`Engine.Evaluate(...).UnwrapIfPromise()` blocks the calling thread while the
engine's event loop drains, which is exactly the wrong thing to do inside an
`async` method — an expression that awaits a .NET `Task`, such as one calling
`getSecret()`, held a thread pool thread for the duration of the I/O.
`Engine.EvaluateAsync` awaits the returned promise instead, and takes the
cancellation token while it is at it.

The Jint version is moved from 4.4.2 to 4.15.3. `EvaluateAsync` arrived in
4.14.0, but the pin lands past 4.15.2 deliberately: once expressions genuinely
suspend and resume instead of draining the event loop on the calling thread,
they exercise the async suspension machinery 4.15.2 corrected — an `await` on a
right-hand side no longer stores the suspension sentinel, async generators and
`for await...of` preserve loop iteration state across a suspension, and a
suspension node is unwrapped correctly. Shipping the non-blocking change on an
earlier 4.14/4.15 would enable exactly the code paths those releases fixed.

One default changed along the way: since 4.14 `Interop.ArrayConversion` defaults
to `LiveView`, so a CLR array reaches script as a live view over the original
array rather than as a copy. That is observable — a script that sorts an array
would now reorder the workflow's own array, and the value round-trips back as
its original element type rather than as `object[]`. The evaluator therefore
pins the previous `Copy` behaviour so the upgrade is not a behavioural change;
hosts that prefer the live view can opt in through
`JintOptions.ConfigureEngineOptions`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0179sA2T7HuRfRfSc2JirFik

* test(javascript): pin the array copy lane and adopt JsString.Create

The array audit. The parent commit fixes `ArrayConversion` to `Copy`, because
4.14 changed the default to `LiveView` and the two differ in behaviour a script
can observe. The existing tests assert what a copy *produces*, which a live view
also satisfies for a value nothing mutates; asking the engine how many
conversions of each kind it performed (4.15.1's interop conversion counters)
pins the lane itself. The second assertion is the more interesting one: an
ordinary evaluation converts no CLR array at all, because Elsa converts
collection-valued variables itself in `ObjectConverterHelper` long before Jint's
array lane could see them. That makes the `ArrayConversion` setting a narrow
compatibility pin rather than something every evaluation depends on.

`JsString.Create` (public since Jint 4.15.3) is adopted in
`JsonElementConverter`, where the string case was the only one still routed
through `JsValue.FromObject` — re-entering the whole conversion pipeline, the
registered object converters and this one included, to arrive at the same call
the number and boolean cases beside it already make directly.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016uV6H9cTntzsoKiaJRBn4f

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 01:02:04 +02:00
github-actions[bot] 30ab056745
Refresh codebase wiki (#7922)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-08-17 00:58:59 +02:00
Sipke Schoorstra 7eaf056d65
test(bpmn): prove execution state and the work ledger survive real persistence (#7947)
* test(bpmn): prove BpmnExecutionState pruning and ledger rehydration survive a real suspend/resume

Prune() was already being called before every persisted write in BpmnScopeHost, but nothing proved
it, and no BPMN test had ever exercised a genuine rehydration: every existing scenario used the
in-memory WorkflowState object straight from the previous run. Add tests that round-trip a
suspended scope's state through Elsa's own IWorkflowStateSerializer -- the boundary that mangled
values before -- and assert the persisted BpmnExecutionState stays bounded across many evaluations
and that a resumed scope with two live units of work matches each completion back to its binding
through the rehydrated BpmnWorkLedger. Both tests were confirmed red by mutation-testing away
Prune() and by returning an empty ledger from BpmnScopeMemory.Load.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): prove a nested scope's ledger survives real persistence

Both new tests in the prior commit suspended a root scope with live work
outstanding, but neither crossed a nested scope's own ledger through Elsa's
real IWorkflowStateSerializer -- the intersection of the two things that
have actually broken here: the handle-to-context map, and the serializer
boundary. Add a nested parallel split/join, blocking on both branches inside
an embedded subprocess, round-tripped through the serializer between each
branch's completion, and confirmed red by returning an empty ledger from
BpmnScopeMemory.Load and green with it restored.

Also extract the start/split/left/right/join/after/end topology shared
verbatim by ParallelSplitAndJoin and ParallelSplitAndJoinBlocking into one
private builder parameterised by the branches' work, keeping both public
factories unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 20:45:07 +02:00
Sipke Schoorstra 6a3d65ff7e
feat(bpmn): the work binder and the elsa: binding format (#7946)
* feat(bpmn): bind BPMN work declarations to Elsa activities

Turns the reader's BpmnWorkBinding declarations into the activity nodes a
BpmnProcess scope runs. Six of the seven kinds bind automatically: TimerWait to
Delay, MessageWait/SignalWait to Event, MessagePublish to PublishEvent,
CallProcess to DispatchWorkflow, NestedProcess to a nested BpmnProcess. The
seventh, UnboundTask, is an authoring decision and is read from a new elsa:
vendor extension inside the document, so an exported .bpmn is self-contained.

Every binding for a scope is bound whatever its slot, so a ScopeListener needs
no special case. Each binding gets its own freshly built activity with a
scope-qualified id: ActivityVisitor skips an activity it has already collected,
so one instance shared between two scopes would leave the second scope with no
child in Elsa's identity graph.

The binder lives in Elsa.Bpmn.Interchange because BpmnWorkBinding is a
Bpmn.Interchange type; binding it in Elsa.Bpmn would pull the interchange
library into the execution module's closure, which is the split D12 draws.

Every ambiguity resolves loudly: an unbound task, a dead binding declaration, a
malformed ISO-8601 duration, a call activity with nothing to call, and an
activity type nothing registered all refuse at bind time rather than producing a
process that runs to completion doing none of what the document says.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): declare document variables on the bound scope, refuse duplicate input names

BpmnWorkBinder.BindScope never copied BpmnProcessDefinition.Variables onto the
produced BpmnProcess's Elsa Variables, so a document-declared collection variable
was Absent to IBpmnVariableReader and a collection-mode multi-instance over it
faulted the element instead of running once per item. BindScope now declares an
Elsa Variable for each document variable, seeding the declared default as the
JsonElement it already is.

BpmnActivityBindingFormat.Read silently let a second <elsa:input name="..."> with
a duplicate name overwrite the first rather than refusing it, unlike every other
malformed-document case this binder already refuses. It now throws
BpmnBindingException naming the binding and the duplicated input, and the XML doc
now states that rule plus the (verified) XML text-node escaping that already
applies to input JSON.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): carry every declared activity input through the binding format

BpmnActivityBindingFormat.Write only found properties whose CLR type derives
from Input, silently dropping attribute-declared inputs like Switch.Cases from
an export. Read accepted any <elsa:input name="..."> without checking the
activity declares it, so a mistyped or stale name imported silently with the
configuration missing since Elsa's deserializer ignores unknown members. Both
now go through IActivityDescriber.GetInputProperties, the same enumeration
ActivityDescriptor.Inputs is built from, so Write and Read agree on what an
activity's inputs are and Read refuses a name that enumeration does not
report.

Also makes BpmnWorkBinder.RefuseUnusedDeclarations filter its loop explicitly
with .Where(...) instead of an implicit if, per static analysis.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(bpmn): filter undeclared input names explicitly

Express the undeclared-input-name check as an explicit Where filter
instead of an implicit filter inside the loop body, and report every
undeclared name at once rather than only the first. Also fix the
refusal message, which previously named the activity type twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(bpmn): describe the input payload shape accurately

The XML doc on BpmnActivityBindingFormat claimed every <elsa:input> is the
{"typeName":...,"expression":...} wrapper a stored workflow definition uses.
That only holds for Input<T>-typed properties: an [Input]-attributed
plain-typed property such as Switch.Cases is serialized as its own JSON
shape (an array), not the wrapper, which Write already does correctly and
the round-trip test already covers. Correct the doc to describe the payload
as the configured activity serializer's output for that input, dependent on
how the activity declares it, and add a second short example showing the
attribute-declared shape.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 17:38:03 +02:00
Sipke Schoorstra fa989c29d0
test(component): correct coverage gate to 23
The gate was lowered to 24 against a measured 24.74% on the release/3.8.0
merge. Integrating origin/main (#7945, the BpmnProcess container activity)
then added further uncovered production code, taking the measured total to
23.98% and putting it back under the gate.

Set the gate to 23 so it again sits just below actual coverage. 23.98% was
measured with PublishEvent_WithPayload_TransmitsPayloadToConsumer excluded
(see #7749), so it is a lower bound: including that test can only raise the
figure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 16:23:23 +02:00
Sipke Schoorstra 2ed957203a
Merge origin/main into main 2026-08-14 15:38:06 +02:00
Sipke Schoorstra 1ec6f0b1b4
test(component): lower coverage gate to 24 after the 3.8.0 merge
Merging release/3.8.0 pulled ExternalAuthentication, AI and Secrets into this
project's coverage denominator through Elsa.Server.Web, without matching
component-test coverage. Measured total line coverage went from 25.98% on
f99f37407 to 24.74% on the merge, crossing below the 25 gate and failing CI in
both pr.yml (./build.cmd Compile Test, NUKE EnableCollectCoverage) and
packages.yml (/p:CollectCoverage=true).

Lower the gate to 24 so it sits just under actual coverage rather than
disabling it. Raise it back towards 25 as component coverage for the merged-in
modules lands.

Note: Elsa.Resilience.Core.UnitTests is also below its threshold (48.17% vs
49), but that is pre-existing and unrelated -- it measures identically on
f99f37407 and on the merge. Left untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 15:13:21 +02:00
Sipke Schoorstra 65fe688350
feat(bpmn): the BpmnProcess container activity (#7945)
* feat(bpmn): scope variables, trigger opt-out and composability for BpmnProcess

Completes the container W2 left minimal, with the four things it deferred.

Scope variables. BpmnScopeVariables implements IBpmnVariableReader over the
scope's memory register, walking outward so an inner scope sees the enclosing
one's data, and BpmnScopeHost now declares ScopeVariables and hands the reader
to every snapshot. The read is three-valued: false for a name nothing in scope
declares, Null for a declared variable holding nothing, and StoredExternally
for a value JSON cannot carry.

That last case deviates from the issue, deliberately. The issue names the
unmaterialized-driver case, which is not detectable from the container's side:
PersistentVariablesMiddleware loads with no excludeTags, and
VariablePersistenceManager marks a block IsInitialized before testing the
exclusion, so a variable whose driver was never read is indistinguishable from
one whose driver returned null. Closing that needs a change to
Elsa.Workflows.Core, which is out of bounds here, so the reader answers only
what the block actually says and the XML doc records why. The route it does
have is real and in the same spirit: a value the host holds and cannot put on
the wire faults loudly rather than reading as an empty collection.

Trigger opt-out. BpmnProcess.IsRootScope names the BPMN meaning of Elsa's
CanStartWorkflow rather than adding a second flag that could disagree with the
gate TriggerIndexer actually reads. It is off unless something says otherwise,
and the applier refuses to start a BpmnProcess that claims root position as
another scope's work: the damage a mis-flagged subprocess does happens at
publish time, so repairing the object graph at runtime would leave the trigger
registered while every test went green. ITrigger itself remains #7929.

Composability and outcomes. A BpmnProcess in a Flowchart runs and the flowchart
carries on (D11), and a nested transaction completing Cancelled reaches its
parent's completion callback with that outcome intact, which is the only reason
the parent routes the cancel boundary rather than the ordinary sequence flow.

Every guard was mutation-tested red before green: both non-Present answers of
the reader, the reader left unwired, the opt-out's default flipped (7 tests red,
including the pre-existing nested-scope ones), the refusal removed, and the
outcome dropped at each end of the trip to the parent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): apply review findings on scope variables, command batching, and outcome doc

Read a scope variable through Elsa's configured serializer (via IPayloadSerializer,
serialized against the value's own runtime type so a polymorphic value is not wrapped
in Elsa's type-tagged envelope) instead of bare JsonSerializerDefaults, so a value only
Elsa's converters can carry no longer collapses to StoredExternally. Refuse a root-scope
StartWork before any command in the batch is applied, not mid-list, so a refusal cannot
leave scope memory partially mutated under ContinueWithIncidentsStrategy. Document that
BpmnProcess completes with only its interpreter outcome, so a default/null-port
Flowchart connection never fires from it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* refactor(bpmn): filter the pre-scan explicitly

Use commands.OfType<BpmnHostCommand.StartWork>() in ApplyAsync's
root-scope pre-scan instead of a foreach + type-check, matching the
static analysis suggestion. The apply loop below is unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 12:37:35 +02:00
Sipke Schoorstra 52a7061f89
Merge release/3.8.0 into main
Brings the 3.8.0 release line into main, including the package-manifest
runtime-kind mechanism (src/PackageManifest.props + src/PackageManifestHints.cs)
that main did not have. All 72 manifest-producing packages now declare
compatibility.runtimeKinds = ["elsa.server"]; the two Bpmn modules added on
main pick this up automatically via their ShellFeatures directory.

Conflict resolutions:
- .specify/feature.json, CONTEXT.md, ROADMAP.md, build/_build.csproj: took
  main's, which is newer in every case. Verified byte-identical to main
  afterwards, so nothing from the release branch was dropped.
- NuGet.Config: union of package sources, minus valence-consolelogstream-feedz.
  main removed that feed deliberately in 7389e0a67 and now consumes
  ConsoleLogStreaming 1.1.0 from nuget.org.
- Elsa.sln: union of main's Bpmn projects and the release branch's
  ExternalAuthentication projects; the two sets are disjoint.

Reverted an unintended revert:

release/3.8.0 had lost commit 33181b2c9 ("test: cover Oracle bulk upsert SQL
generation") through an evil merge in c557c455a. That commit is present at the
merge base, so git resolved the release branch's older content as an
intentional change and would have silently undone it on main. It is three
coupled pieces:

  - test/unit/Elsa.Persistence.EFCore.UnitTests (deleted, plus its Elsa.sln
    project declaration and NestedProjects entry)
  - InternalsVisibleTo("Elsa.Persistence.EFCore.UnitTests")
  - the fix itself in BulkUpsertExtensions.GenerateOracleUpsert: internal
    visibility, ISqlGenerationHelper.DelimitIdentifier quoting, and explicit
    CAST(... AS NVARCHAR2(...)) on string columns

Dropping the third would have been an Oracle runtime regression: unquoted
identifiers lose case, and ODP.NET binds .NET strings as VARCHAR2 while Elsa's
Oracle migrations declare NVARCHAR2, causing a datatype mismatch. Merge base
and main are identical for that file and every hunk on the release side is a
revert plus cosmetics, so main's version was kept in full.

Accepted deliberate release-branch changes, verified as real refactors rather
than losses: AI EF Core migrations moved into the provider projects
(5c0d8b0f4), and AIPersistenceFeature.cs renamed to
EFCoreAIPersistenceShellFeatureBase.cs (ShellFeatures/ still present, so
manifest generation is unaffected).

Verified: dotnet build Elsa.sln succeeds with 0 errors and 2 pre-existing
NU1903 warnings; Elsa.Persistence.EFCore.UnitTests passes 1/1; all 72 emitted
manifests declare elsa.server and Elsa.Api.Common emits none.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 11:59:35 +02:00
Sipke Schoorstra 33ad4af835
fix(manifest): stop Elsa.Api.Common publishing a runtime-kind-less manifest
Elsa.Api.Common has a ShellFeatures directory, so it picked up the manifest
generator and emitted an elsa-package.json. It was excluded from the shared
hints file, though, so that manifest carried no compatibility.runtimeKinds.
Runtime kind matching is asymmetric: when an image declares runtime kinds, a
feature whose effective list is empty can never intersect and is excluded.
An empty manifest is therefore worse than no manifest at all.

Elsa.Api.Common is shared API plumbing rather than a selectable catalog
module, so opt it out of manifest generation entirely via
GenerateElsaPackageManifest. That also collapses the two overlapping
ItemGroup conditions into one and removes the hardcoded project name from
them, leaving GenerateElsaPackageManifest as the single opt-out knob.

The remaining 70 ShellFeatures projects are unaffected and continue to
declare elsa.server.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 11:11:58 +02:00
Sipke Schoorstra f99f37407c
feat(bpmn): host port and command applier (#7942)
* feat(bpmn): host-side applier for the Bpmn.Semantics port

Translates the interpreter's three host commands onto ActivityExecutionContext
and feeds its four entry points, plus the minimum BpmnProcess container needed
to exercise them end to end through IWorkflowRunner.

StartWork schedules the bound activity, CancelWorkSubtree calls the public
CancelActivityAsync extension (already recursive), and SignalEnclosingScope
sends a BpmnScopeSignal up the ancestor chain. OnWorkFaulted rides the
FaultSignal seam: it asks the interpreter what BPMN made of the fault and calls
StopPropagation only on a Caught disposition, leaving a Propagated one strictly
alone so an enclosing scope or the incident strategy takes it.

A unit of work is keyed on the child ActivityExecutionContext.Id, recorded in
the scope's own persisted ledger, never on Tag: the completion-callback dispatch
rewrites the receiving context's Tag, so a nested scope wears a different tag
than its parent remembers it by. Interpreter correlation travels on the child's
context rather than on the shared activity instance.

Evaluations go through one queue per workflow instance, so a scope signalled
mid-apply is drained after the command list rather than re-entering the
interpreter. Commands are applied in the order returned.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): cover the teardown refusal path

Adds a focused unit test that drives BpmnWorkTeardown.CancelSubtreeAsync
into the NotSupportedException branch by constructing a real context tree
with a scheduled-but-not-invoked descendant, so a regression that silently
drops the detection is caught. Also records why BpmnWorkLedger's
append-only, handle-keyed Records list cannot strand a context on a
duplicate StartWork for a live (BindingRef, IterationId) slot, a case the
port's own guarantee makes unreachable from this applier.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(bpmn): keep a refused teardown from stranding ledger state

A subtree cancellation refused with NotSupportedException is absorbed
into an incident under ContinueWithIncidentsStrategy rather than
crashing, so the end-of-command ledger save was being skipped and the
persisted ledger kept claiming work BPMN had just torn down. Save the
ledger removal before the possible throw instead of after, so a later
completion callback for the stranded activity finds no live record and
is discarded instead of being fed to the interpreter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 01:42:31 +02:00
Sipke Schoorstra 7389e0a674
fix(deps): consume ConsoleLogStreaming 1.1.0 from nuget.org (#7941)
Elsa.Diagnostics.ConsoleLogs and Elsa.Dashboard.Api are packable and are
published to the Elsa preview feed on every push to main, and to nuget.org on
release. Both pinned ConsoleLogStreaming.* at 1.0.0-preview.13, a version
published only to the private valence-works feed, so that pin was baked into
the shipped nuspecs.

Restore did not hard-fail for consumers: NuGet floated the >= constraint up to
the nuget.org 1.0.0 and emitted NU1603 for each package, an error under
TreatWarningsAsErrors. The quieter problem was that consumers silently ran a
different build of the library than CI compiled against.

ConsoleLogStreaming 1.1.0 has now been released publicly, carrying the work
that had only ever reached the private feed as previews: the new
ConsoleLogOptions.StreamReleaseInterval option, stream-release gate flood
throttling, and dead subscriber drop accounting. Pinning to it makes the
shipped dependency both resolvable and current, so the private feed and its
package source mapping are no longer needed.

Verified: Elsa.sln restores clean; Elsa.Dashboard.Api builds with 0 warnings;
64 tests pass across the ConsoleLogs and Dashboard.Api unit and integration
projects; the packed nuspec declares 1.1.0 on all three target frameworks; and
a clean consumer project restoring the packed output against nuget.org alone,
with TreatWarningsAsErrors, resolves 1.1.0 with no warnings.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 22:44:13 +02:00
Sipke Schoorstra 22886809e0
test(bpmn): guard that Elsa.Bpmn never duplicates the library (#7939)
* test(bpmn): guard against Elsa.Bpmn* reimplementing Bpmn.* library types

Adds a reflection-based architecture test asserting no type under Elsa.Bpmn
or Elsa.Bpmn.Interchange shares a type name with Bpmn.Model/Bpmn.Semantics
(and Bpmn.Interchange for the interchange assembly), plus a positive check
that both assemblies still depend on their respective library packages -
closing the gap that let elsa-foundation grow a parallel BpmnElement/BpmnGraph
semantics core undetected.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): move the library-duplication guard to the interchange test project

Elsa.Bpmn.UnitTests referencing Elsa.Bpmn.Interchange (plus three
redundant PackageReferences already flowing transitively) inverted the
layering the guard exists to protect. Elsa.Bpmn.Interchange.UnitTests
already sees both assemblies transitively with no new references, so
the guard moves there unchanged apart from namespace and doc. Also
drops the silent null-forgive on Assembly.GetName().Name in favor of
an explained fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): drop the dependency-direction assertions the compiler already enforces

Mutation testing showed AssertDependsOnPackage and the two positive facts using
it can never go red: any state that would trip them fails the test project's
build first (CS0234), because the guard's own typeof bindings already require
the Bpmn.* packages. Delete the decorative facts and the now-unused helper,
and document that the typeof bindings are load-bearing on purpose.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(bpmn): prove the duplication detector actually detects

Split AssertNoTypeNameCollisions into a thin assertion wrapper around a
new FindTypeNameCollisions helper, and add a fact that points the
detector at the test assembly (which carries a deliberately colliding
BpmnGraph fixture) so CI sees the guard fail as well as pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 04:09:53 +02:00
Sipke Schoorstra 735e953dc5
feat(bpmn): scaffold Elsa.Bpmn and Elsa.Bpmn.Interchange modules (#7938)
Adds the valence-works/bpmn package feed and pins Bpmn.Model,
Bpmn.Semantics and Bpmn.Interchange at 0.1.1-preview.19, then wires
two new module projects consuming those libraries without
reimplementing anything they provide (D12): Elsa.Bpmn for BPMN
execution and Elsa.Bpmn.Interchange for XML import/export, kept as a
separate package so hosts that only execute BPMN don't take the XML
reader. Both are marked IsPackable=false until the Bpmn.* packages
are published to nuget.org. Test projects are added for both, unit
and integration, and all four are wired into Elsa.sln so PR CI
discovers and runs them. This unblocks #7925 and the rest of the
BPMN runtime work in #7909.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-13 01:17:46 +02:00
Sipke Schoorstra 749b6491fc
fix(core): a throwing FaultSignal handler must not escape the middleware (#7924)
* fix(core): a throwing FaultSignal handler must not escape the middleware (#7911)

The signal is sent from inside the catch whose whole job is to stop exceptions
escaping the activity pipeline. A handler that threw went straight through it:
no incident, no strategy, and the original fault lost along with it.

The send is now guarded. A handler that throws is treated as not having handled
the fault, so the incident strategy runs exactly as it would with no handler
present. That is the conservative direction: a handler that failed part way
through may have left the faulted activity in any state, and an incident is a
better answer than silence. Its exception is logged at error level, because a
broken fault handler is a defect in its own right rather than a workflow
outcome.

Covered both ways, since a handler that already claimed the fault before
throwing is the case that could plausibly have been mistaken for success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(core): let cancellation from a fault handler propagate

The handler guard caught OperationCanceledException along with everything else,
so a cancellation raised while an ancestor was being offered a fault was logged
as a broken handler and handed to the incident strategy. A deliberately
cancelled run reported itself faulted.

Cancellation is excluded now, matching how this repository already keeps the two
apart: the workflow-level exception middleware cancels and rethrows before its
general catch, and WorkflowRunner declines to record cancellation as the
workflow's exception.

Caught by review on #7924.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(core): pin that a handled fault stays in the execution log

The justification for dropping the incident is that the journal keeps the
evidence. That was asserted in several places and guarded nowhere.

It holds because ExecutionLogMiddleware writes the Faulted entry from a catch
that rethrows, and it is registered inside ExceptionHandlingMiddleware, so the
entry lands before the fault is ever offered to an ancestor. Swapping those two
registrations would make a handled failure disappear from the record with
nothing failing, which is what this test now prevents.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 05:34:19 +02:00
Sipke Schoorstra 5faba75906
fix(core): a fault a container claimed is not an incident (#7923)
* fix(core): a fault a container claimed is not an incident (#7911)

RecoverFromFault reset the counts and the status but left behind the two other
things Fault recorded: the ActivityIncident and the exception. So a container
that successfully handled a child's fault still left the workflow carrying an
incident.

That is not cosmetic. Code reads a non-empty WorkflowExecutionContext.Incidents
as "this workflow failed" without looking further; HttpWorkflowsMiddleware is
one, and it hands the caller a fault response. A workflow whose container caught
the error and finished normally was reported to its caller as failed.

RecoverFromFault is now the inverse of Fault: it removes the incident Fault
appended, matched on this activity's node id and most recent first so an
activity that faults, recovers and faults again keeps the incident that was
never recovered, and it clears the recorded exception so the activity does not
sit in Running carrying one.

The execution log still records the failure, so nothing is hidden from anyone
reading the journal. Two integration assertions that encoded the old behaviour
are updated; they were written from the reasoning this change corrects.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(core): tie an incident to the execution that raised it, not its node

Recovery matched the incident to remove on ActivityNodeId, which identifies the
static workflow node rather than an execution of it. A node inside a loop,
retried, or run concurrently raises one incident per execution, all under the
same node id, so recovering one execution could remove another's incident and
leave its own behind.

ActivityIncident now carries the ActivityInstanceId of the execution that raised
it, and recovery matches on that. Within a single execution the most recent is
still taken, so fault, recover, fault again keeps the incident that was never
recovered. The property is optional: an incident recorded against the workflow
itself has no execution, and so do incidents persisted before this existed.

Caught by review on #7923.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(api-client): mirror ActivityInstanceId on the client incident model

The server model gained the property in the previous commit and the API client
carries a hand-maintained copy of it. Left alone, a client deserializing an
incident would silently drop the only field that says which execution raised it.

Also records two consequences of recovery that were implicit: it relies on the
incident collection preserving insertion order to pick an execution's newest
incident, which holds only because the collection is list-backed; and clearing
the exception also clears it from the activity's execution record, which is
intended for the same reason the incident goes, with the journal keeping the
evidence either way.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 05:20:22 +02:00
Sipke Schoorstra 61f9d2c39d
Merge pull request #7921 from elsa-workflows/codex/update-wiki
[codex] Refresh codebase wiki
2026-08-12 00:53:55 +02:00
github-actions[bot] 1754253fb9 Refresh codebase wiki 2026-08-11 22:25:06 +00:00
Sipke Schoorstra 5068e0b631
Merge pull request #7920 from elsa-workflows/claude/fix-wiki-link-check
docs(wiki): point the scheduling link at StartupTasks
2026-08-12 00:13:39 +02:00
Sipke Schoorstra 70690ad3d3
docs: refresh roadmap 2026-08-12 00:09:13 +02:00
Sipke Schoorstra ff72352d8c
docs(wiki): point the scheduling link at StartupTasks
The Update Wiki workflow has failed on every run since 2026-07-28, on one
broken link:

  doc/wiki/http-scheduling-resilience.md:
    ../../src/modules/Elsa.Scheduling/HostedServices -> missing

Elsa.Scheduling/HostedServices was removed in 8e301d4e1, where
HostedServices/CreateSchedulesBackgroundTask.cs became
StartupTasks/CreateSchedulesStartupTask.cs. The wiki page kept pointing at the
old directory.

Point it at StartupTasks, which is where that work lives now.

The workflow's link check only validates wiki files that changed in the push,
so its silence is not evidence the rest is sound. Ran the same validator over
all 22 files in doc/wiki: this was the only break, and everything now resolves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-12 00:02:13 +02:00
Sipke Schoorstra 0412555b6e
Merge pull request #7913 from elsa-workflows/claude/cranky-boyd-269f7e
feat(core): let a container activity handle a child's fault via FaultSignal
2026-08-11 23:43:26 +02:00
Sipke Schoorstra 5e56161031
Merge pull request #7919 from elsa-workflows/claude/fix-nuke-nuget-frameworks
build: bump NuGet.Packaging to 7.9.0 to unbreak the NUKE build on SDK 10.0.400
2026-08-11 23:36:35 +02:00
Sipke Schoorstra fe6601ab5b
build: bump NuGet.Packaging to 7.9.0 to unbreak the NUKE build on SDK 10.0.400
CI started failing on every branch with an assembly load error out of NUKE's
project parsing, before any test ran:

  InvalidProjectFileException: The expression
  "[MSBuild]::GetTargetFrameworkIdentifier(net10.0)" cannot be evaluated.
  Could not load file or assembly 'NuGet.Frameworks, Version=7.9.0.0'.
  The located assembly's manifest definition does not match the assembly reference.
     at Nuke.Common.ProjectModel.ProjectModelTasks.ParseProject

Nothing in the repo changed to cause it. pr.yml requests dotnet-version 10.x,
and the hosted runner moved from SDK 10.0.302 to 10.0.400. Measured, the two
SDKs ship different NuGet.Frameworks:

  SDK 10.0.300 / 10.0.302 -> NuGet.Frameworks 7.6.0
  SDK 10.0.400            -> NuGet.Frameworks 7.9.0

_build.csproj pinned NuGet.Packaging 7.6.0, which puts NuGet.Frameworks 7.6.0
in the NUKE output directory, where it shadows the SDK's own copy. The loader
accepts an assembly newer than the reference but not older, so once MSBuild
started asking for 7.9.0.0 the app-local 7.6.0 no longer satisfied it. The same
commits passed 19 hours earlier on 10.0.302.

The reference is not used by the build itself - there are no NuGet.* usages
anywhere in build/*.cs. It exists only as a transitive vulnerability override,
added in f5dc29cdc, so bumping it preserves the original intent while matching
what the current SDK ships. Staying loadable on 10.0.3xx follows from the same
newer-than-reference rule that broke the old pin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 23:25:33 +02:00
Sipke Schoorstra a618337923
docs(core): give FaultSignal handlers a working completion path
Review caught that the contract advertised something that cannot work. It
offered handlers three ways to terminalize the faulted activity - cancel,
complete, or reschedule - but CompleteActivityAsync returns immediately unless
the activity is Running, and throughout the handler it is still Faulted, since
recovery runs only after the handler returns. Completing inline did nothing at
all, silently, leaving the child Running.

Measured, same container, handler completing the faulted child:

  inline complete                    -> child Running,   no output, Running/Suspended
  TransitionTo(Running), complete    -> child Completed, "after",   Finished/Finished

So a supported path exists; it just needed writing down. Document it on
FaultSignal, note that it is not licence to call RecoverFromFault (which also
rewrites the fault counts), and note that cancelling and rescheduling need no
equivalent step. Cover it with an integration test asserting that completing the
child with a substitute result fires the container's completion callback and
resumes its sequencing.

Refs #7911

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 22:08:44 +02:00
Sipke Schoorstra d085b536f0
docs(core): document FaultSignal self-receipt and pin it with a test
Review raised that TrySendSignalAsync delivers to the faulting activity before
walking ancestors, so an activity that throws and also handles FaultSignal can
claim its own fault and suppress the incident strategy.

That is real, but it is the channel's existing dispatch, which #7911 chose
deliberately over a variant of it, and SignalContext.IsSelf exists so handlers
can discriminate. It also grants no capability: an activity that catches its own
exception never faults at all, ending Finished/Finished with zero incidents,
which is a cleaner suppression than self-handling (incident still recorded,
activity left Running, workflow suspended).

So dispatch is unchanged. What was missing is that none of this was written
down: the contract describes the handler as an enclosing container and never
mentioned self-receipt. Document it on FaultSignal, including how a handler
that wants ancestors-only semantics opts out, and add a test so the behavior is
pinned rather than incidental.

Refs #7911

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 03:04:49 +02:00
Sipke Schoorstra 13eb48e002
feat(core): let a container handle a child's fault via FaultSignal
A container activity had no way to learn that one of its children faulted.
ExceptionHandlingMiddleware caught the exception, called context.Fault(e) and
handed off to the workflow-global IIncidentStrategy; the container's completion
callback never fired, because the child never completed.

Add a seam on the ancestor-bubbling signal channel that already exists:

- FaultSignal(Exception, ActivityExecutionContext), beside CancelSignal. Its XML
  doc carries the contract, including why a handler must not call
  RecoverFromFault and why the CompleteActivityAsync sweep is a backstop rather
  than the mechanism.
- An internal bool-returning TrySendSignalAsync, since SignalContext
  .StopPropagationRequested is internal and SendSignalAsync reported nothing.
  SendSignalAsync keeps its public signature and delegates to it.
- ExceptionHandlingMiddleware sends the signal after faulting and, when an
  ancestor stops propagation, calls RecoverFromFault once and returns instead of
  raising an incident.

RecoverFromFault now transitions to Running only when the activity is still
Faulted. It is called after the handler runs, so the unconditional transition
would otherwise undo a handler that cancelled or completed the faulted child.
The counts are still reset unconditionally, and the one pre-existing caller is
unaffected.

Behavior is unchanged when nobody handles the signal: verified by running the
new unhandled-fault theory against the pre-change middleware, and by
IncidentStrategyTests and Primitives/FaultTests passing unmodified.

Refs #7911

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-11 01:11:25 +02:00
Sipke Schoorstra bd903f63dc
docs: refresh roadmap 2026-08-05 00:05:27 +02:00
Sipke Schoorstra 5429008d98
Add architecture and practices documentation
Introduce comprehensive documentation covering architecture (ARCHITECTURE.md), concerns (CONCERNS.md), coding conventions (CONVENTIONS.md), integrations (INTEGRATIONS.md), technology stack (STACK.md), codebase structure (STRUCTURE.md), and testing patterns (TESTING.md). Enhance testing with additional test cases for external sign-in flows, ensuring accurate timestamp recording for identity links.
2026-08-03 23:46:38 +02:00
Sipke Schoorstra ff10a68100
refactor: update package reference configuration
- Updated `CShells.Abstractions` package reference to use the default version in the unit test project.
2026-08-03 13:15:18 +02:00
Sipke Schoorstra e59a0f1721
Merge remote-tracking branch 'origin/release/3.8.0' into release/3.8.0 2026-08-03 02:08:23 +02:00
Sipke Schoorstra 103028452f
feat: introduce HTTP webhooks module
Adds a new `Elsa.Http.Webhooks` module, enabling workflows to receive incoming webhook events and dispatch outgoing webhooks. This integrates the WebhooksCore library.

Further improvements include:
- Enhanced validation for configured application instance names, providing clearer feedback, especially regarding Azure Service Bus entity name limits.
- Improved API error reporting for shell reload operations, distinguishing between blueprint not found (404) and other failures (503).
- Updated release announcement rendering to dynamically reference the correct major.minor release line for feedback messages.
2026-08-03 02:08:15 +02:00
Sipke Schoorstra 46b29d6efd
Merge pull request #7901 from DenDeline/bugfix/7900-tenant-request-services
fix: restore request services after tenant middleware exceptions
2026-08-03 01:06:56 +02:00
Sipke Schoorstra c557c455ad
Merge release/3.8.0 into bugfix/7900-tenant-request-services 2026-08-03 00:48:56 +02:00
Sipke Schoorstra dacad14643
Merge pull request #7905 from elsa-workflows/codex/pr-7903-review-fixes
Harden external identity race compensation
2026-08-02 23:44:31 +02:00
Sipke Schoorstra 286a0d83d1
Harden external identity race compensation 2026-08-02 23:19:49 +02:00
Sipke Schoorstra d10dbdd482
Merge branch 'codex/pr-7903-review-fixes' into release/3.8.0 2026-08-02 22:40:23 +02:00
Sipke Schoorstra f4d749d3f1
Merge branch 'codex/auth-refinements' into release/3.8.0 2026-08-02 22:39:54 +02:00
Sipke Schoorstra fe125ac336
Fix external authentication review findings 2026-08-02 03:35:39 +02:00
Sipke Schoorstra ed50e1cc96
Merge pull request #7904 from elsa-workflows/codex/fix-7306-tenant-agnostic-registry-population
Avoid repeated tenant-agnostic registry population
2026-08-02 03:25:43 +02:00
Sipke Schoorstra 1b74bb94c0
fix: validate external login methods before discovery
Use the same structural and secret-binding assessment for management, discovery, and initiation so incomplete overrides are never advertised as available sign-in methods.
2026-08-02 02:47:28 +02:00
Sipke Schoorstra 542058971b
Avoid repeated tenant-agnostic registry population 2026-08-02 01:59:45 +02:00
Sipke Schoorstra a24a6fe267
Add test to validate complete configuration requirements for connections
Introduce a new test `ValidateRequiresCompleteConfigurationAndReturnsMissingSecretDetails` to verify that a connection requires a complete configuration, including handling missing secret details. Adjust configuration to enforce `RequiresClientSecret`.
2026-08-02 01:43:30 +02:00
Sipke Schoorstra a60b5a36b2
Refine test for archived connections in shadow relationships; introduce IsAvailableForAuthentication helper for improved connection filtering logic. 2026-08-02 00:49:26 +02:00
Sipke Schoorstra 572cd325b5
Merge compact external authentication secret IDs 2026-08-02 00:22:29 +02:00
Sipke Schoorstra 119f49f7a2
fix: simplify external authentication secret IDs
Use a compact namespaced GUID for newly managed secrets while preserving unique staging and opaque existing references.
2026-08-02 00:22:16 +02:00
Sipke Schoorstra 9f09aca3f6
Refine test to ensure archived connections do not participate as active shadows; update shadow relationship management to exclude archived entries. 2026-07-31 23:07:08 +02:00
Sipke Schoorstra 14f373528d
Add test for handling archived database overrides in connection shadows 2026-07-31 21:34:35 +02:00
Sipke Schoorstra 6e3ed5c4e0
Remove migration for external authentication in EFCore.Sqlite module 2026-07-31 14:31:18 +02:00