Commit graph

27 commits

Author SHA1 Message Date
Sipke Schoorstra ec9acd4f3f
Fix tenant service mutation race (#7898)
Serialize tenant lifecycle mutations and keep synchronization available through shutdown.

Closes #7771.
2026-07-30 02:47:44 +02:00
Sipke Schoorstra 8e301d4e1e
[codex] Scope console logs to workflow instances (#7535)
* Avoid null endpoint DTO metadata in tests

* Enforce console logs hub read permission

* Remove unused console logs hub import

* Support mapped endpoint metadata in auth tests

* Reduce console log capture throughput impact

* Address Copilot console logs review

* Refactor task scheduling to support tenant-level background work and enhance logging functionality.

* Introduce ConsoleStreamHook for stdout/stderr tee and enhance logging validation. Adjust test cases and startup warnings for distributed lock provider usage.

* Refactor console logging pipeline with capture optimization and new ConsoleLogsHost; update tests accordingly.

* Add Ansi SGR parser for console logs and associated unit tests

* Remove ANSI color renderings and parsers; integrate ConsoleLogScopeAccessor for improved logging context with workflow instance ID support.

* Address console logs code quality feedback

* Address PR review feedback

* Preserve console logs extension points

* Stabilize console logs host lifecycle

* Address final automated review comments

* Tighten console log capture shutdown

* Address console log review feedback

* Address follow-up review feedback

* Cover final review feedback

* Avoid recursive console provider initialization

* Guard console host lease shutdown

* Preserve console log scope and provider lifetime

* Correlate console log scope fallback

* Tighten console scope correlation

* Expose host services during provider construction

* Redact ANSI-normalized console lines
2026-05-25 11:49:51 +02:00
Sipke Schoorstra 26b17e35e2
Refactor QuiescenceSignal to inject IServiceScopeFactory, enhance tenant ID handling, and expand unit tests with DI capabilities. 2026-05-03 17:56:37 +02:00
Sipke Schoorstra d7bdbfb26d
Graceful shutdown for the workflow runtime (drain, pause, recover) (#7424)
* feat(workflows-runtime): add quiescence machinery foundation for graceful shutdown

Introduces the container-scoped quiescence signal, ingress-source contract,
burst registry, and the Interrupted workflow sub-status — the foundational
primitives the drain orchestrator and admin endpoints will build on. No
behaviour change yet: workflows continue to run and shut down exactly as
before. The new types are registered but no host-stop or pause path drives
them.

Highlights:
* IQuiescenceSignal — composable Drain + AdministrativePause flags;
  forward-only drain, reversible pause, idempotent transitions, optional
  persistence via IKeyValueStore.
* IIngressSource + IForceStoppable — uniform contract for components that
  inject external events (HTTP, schedulers, message consumers, internal
  workers, third-party modules); IIngressSourceRegistry collects and
  surfaces their states.
* IBurstRegistry — atomic counter for in-flight workflow execution
  bursts, with per-burst ingress attribution and FR-018 inconsistency
  detection (a source claiming Paused but starting bursts is flipped to
  PauseFailed).
* WorkflowSubStatus.Interrupted — new value distinct from Suspended,
  Cancelled, Faulted; semantics: "last burst force-cancelled by graceful
  drain; resumable on next runtime generation". Mirrored on the API client
  enum.
* GracefulShutdownOptions — drain deadline, per-source pause timeout,
  stimulus-queue back-pressure policy, pause-persistence policy.
  Configurable via UseWorkflowRuntime(...).ConfigureGracefulShutdown(...).
* PermissionNames.ManageWorkflowRuntime — single permission for the
  forthcoming admin pause/resume/status/force endpoints.

Implements 31 of 77 tasks for the graceful-shutdown feature
(specs/002-graceful-shutdown). Subsequent commits add the drain
orchestrator (US1 / MVP), Interrupted recovery scan (US3), admin
endpoints (US2), and first-party ingress adapters.

Tests: 25 new xUnit unit tests; 100/100 runtime unit tests pass; all
existing tests continue to pass on net8.0/net9.0/net10.0.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(workflows-runtime): add drain orchestrator + host-stop integration (US1, MVP)

When the host receives a stop signal (SIGTERM, Ctrl+C, orchestrator
rollout), the runtime now drains gracefully: ingress sources are paused
in parallel, in-flight workflow bursts run to their next natural
persistence boundary within a configurable deadline, and any burst that
breaches the deadline is force-cancelled and persisted with the
Interrupted sub-status plus a forensic WorkflowInterrupted log entry.

This is the MVP — without the activation-time recovery scan (PR 3) the
existing timeout-based RestartInterruptedWorkflowsTask still picks up
Interrupted instances, just on its periodic cadence. No regression in
that recovery path (SC-008).

Highlights:
* IDrainOrchestrator + DrainOrchestrator — protocol per the contract:
  BeginDrainAsync → parallel ingress pause with per-source timeouts +
  IForceStoppable escalation → poll BurstRegistry.ActiveCount until zero
  or deadline → on breach iterate live handles, cancel, persist
  Interrupted, write log entry. All exceptions are captured into the
  returned DrainOutcome; only second-invocation throws.
* Deadline clamping: effective deadline is min(GracefulShutdownOptions.
  DrainDeadline, HostOptions.ShutdownTimeout - 500ms safety epsilon),
  so the runtime never outlives its host process.
* DrainOrchestratorHostedService — IHostedService.StopAsync wakes the
  orchestrator on host stop. Registered AFTER the heartbeat
  (Elsa.Hosting.Management) so reverse-order shutdown keeps the
  heartbeat alive throughout drain. Prevents sibling-node crash recovery
  from false-positive-recovering instances we are gracefully handling
  here (FR-029).
* BurstTrackingMiddleware — workflow-execution-pipeline middleware that
  registers a BurstHandle for the lifetime of every burst. All nine
  IWorkflowRunner.RunAsync overloads ultimately funnel into
  pipeline.ExecuteAsync(context), so this single middleware covers the
  three "burst choke points" the spec references without nine separate
  decorators. Added to UseDefaultPipeline().
* Ingress attribution: optional IngressSourceName property on
  DispatchStimulusRequest, DispatchWorkflowDefinitionRequest, and
  DispatchWorkflowInstanceRequest. Adapters set it; the middleware reads
  it via WorkflowExecutionContext.TransientProperties (helpers in
  IngressAttributionExtensions). The BurstRegistry uses the name to
  detect the FR-018 invariant violation — a source that reports Paused
  but starts a burst is flipped to PauseFailed.
* InterruptedLogExtensions — the LogWorkflowInterruptedAsync helper
  that the orchestrator calls when persisting the forensic record.

Tests: 11 new unit tests (DrainOrchestrator parallel-pause +
wait-for-bursts + idempotency + persistence-failure path); 5 new
integration tests (full DI graph resolves, burst-tracking middleware
registers handles end-to-end, no-op drain returns
CompletedWithinDeadline). 100/100 runtime unit tests pass; all
existing tests continue to pass on net8.0/net9.0/net10.0.

Implements 14 of 77 tasks (T032–T045). Subsequent commits add
Interrupted recovery scan (US3), admin endpoints (US2), and
first-party ingress adapters.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(admin-endpoints): add admin endpoints for workflow runtime control

Introduced admin endpoints to manage workflow runtime: `/pause`, `/resume`, `/status`, and `/force` with full authentication and audit logging. Integrated idempotency checks and error handling to ensure reliable runtime control. Added corresponding integration tests for verification.

* fix(workflows-runtime): propagate drain cancellation into running workflows + initialize persisted pause + fix log-comment

Addresses three findings from the PR #7424 code review:

1. **HIGH — Cancellation now propagates into the running workflow.**
   `BurstHandle.Cancel()` previously cancelled only its own linked CTS,
   which the workflow runner never observes (the runner reads from
   `WorkflowExecutionContext.CancellationToken`, captured at context
   construction and not part of the linked chain). On deadline breach
   the orchestrator would persist `Interrupted`, but the workflow
   continued executing and could overwrite the sub-status with whatever
   terminal state it eventually reached.

   Fix: `BurstHandle` accepts an optional cancel callback at construction.
   `BurstTrackingMiddleware` wires it to `context.Cancel()` so the burst's
   cancellation triggers the workflow's own cancellation chain — the
   workflow transitions to `Cancelled` and stops scheduling new
   activities. The orchestrator then awaits `BurstHandle.Disposed` (with
   a 2 s settle timeout) before persisting `Interrupted`, ensuring the
   runner's terminal commit completes BEFORE the orchestrator overwrites
   the sub-status. Race resolved.

   The settle timeout is bounded so a non-cancellable activity (genuinely
   pathological case) does not block drain — on timeout the orchestrator
   logs and proceeds, accepting the runner-clobber for that one
   instance, which the existing timeout-based RestartInterruptedWorkflows
   recovery picks up afterwards.

2. **MEDIUM — Pause persistence is now actually wired.**
   `QuiescenceSignal.InitializePersistedStateAsync` was implemented but
   nothing called it on host startup. A host configured with
   `PausePersistence = AcrossReactivations` would write the persisted
   key on pause, but on subsequent activation the new
   `QuiescenceSignal` instance would never read it, so the runtime would
   resume dispatching despite the operator having paused.

   Fix: `InitializePauseStateStartupTask : IStartupTask` reads the policy
   and calls `InitializePersistedStateAsync` once per activation when the
   policy demands it. Registered in both `WorkflowRuntimeFeature`
   flavours alongside the other graceful-shutdown services.

3. **MEDIUM — Comment in `DrainOrchestrator.PersistInterruptedAsync` no
   longer lies.** The previous comment promised a "synthetic log entry"
   that the next statement (`return`) prevented from being written. The
   comment is now honest about what actually happens: when no instance
   row exists, no log entry is emitted, but the burst metadata is still
   captured in the drain outcome's logged warning so operators have a
   forensic trail.

Tests:
* New unit tests on `BurstRegistry` (now 9, was 6): cancel-callback is
  invoked, callback exceptions are swallowed (drain remains best-effort),
  `BurstHandle.Disposed` completes on dispose.
* Full suites continue to pass: 103/103 runtime unit tests; 247/247
  workflow integration tests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(workflows-runtime): address PR review feedback + close runner-clobber race + add e2e drain test

Six issues raised in PR #7424 review (one e2e gap, five comments inline):

1. **Runner-clobber race closed via ICommitStateHandler decorator.**
   The previous fix wired `BurstHandle.Cancel()` to `WorkflowExecutionContext.Cancel()`
   so the workflow's cancellation chain fires on deadline breach, but the e2e
   test exposed that `BurstHandle` disposed at the END of the pipeline middleware
   (i.e., BEFORE `WorkflowRunner` calls `commitStateHandler.CommitAsync`). The
   orchestrator's `await handle.Disposed` therefore returned too early, the
   instance row didn't yet exist, and the orchestrator's Interrupted write was
   either a no-op (no row) or got clobbered by the runner's subsequent Cancelled
   commit.

   Fix: `BurstAwareCommitStateHandler` decorates `ICommitStateHandler`. The
   middleware no longer disposes the handle in the success path — it stores
   the handle in `WorkflowExecutionContext.TransientProperties`, and the
   decorator disposes it AFTER `inner.CommitAsync` completes. The exception
   path in the middleware still disposes for safety. Result: the orchestrator's
   await-disposed sequencing now correctly lands the Interrupted write last.

2. **C1: Null-instance log entry.** `DrainOrchestrator.PersistInterruptedAsync`
   now writes a synthetic `WorkflowInterrupted` log entry directly when no
   instance row exists, populating only the fields it knows. Previously the
   forensic trail was lost.

3. **C2: Force endpoint cached-outcome audit.** Added `WasCached` flag to
   `DrainOutcome` (default false). The orchestrator sets it on the cached
   return path (`_previousOutcome with { WasCached = true }`). The force
   endpoint now skips the audit notification when the flag is true, so
   repeated force calls no longer emit spurious `RuntimeForceRequested` events
   (SC-007 idempotency restored).

4. **C3: `StateChanged` raised under lock — deadlock risk closed.**
   `QuiescenceSignal.BeginDrainAsync`/`PauseAsync`/`ResumeAsync` now do their
   transitions under the lock, capture whether a transition occurred, release
   the lock, and only then invoke `RaiseStateChanged`. Subscribers that
   synchronously call back into the signal can no longer deadlock.

5. **C4: Scheduling source name.** Renamed `scheduling.cron` → `scheduling.triggers`
   to honestly reflect the four trigger types the adapter covers (Cron, Timer,
   StartAt, Delay). The name is surfaced verbatim in admin status responses.

6. **C5: Hardcoded Retry-After.** `HttpWorkflowsMiddleware`'s 503 response now
   sets a reason-aware `Retry-After`: 5 s during drain (host is exiting and
   will be replaced shortly), 60 s during administrative pause (indefinite,
   so a longer back-off avoids tight retry loops).

Tests:
* New e2e `DeadlineBreachEndToEndTests` (2 tests): verifies that drain
  against a real running workflow detects the in-flight burst, force-cancels
  it, persists the instance as `Interrupted`, and writes a `WorkflowInterrupted`
  log entry — closing the test gap that hid the cancellation-propagation
  issue identified in the previous review pass.
* Updated `OperatorForceAfterPreviousReturnsCachedOutcome` to assert
  value-equality + the `WasCached` flag instead of reference-equality
  (records use `with` for the cached return path).

Full suites pass: 103/103 runtime unit, 249/249 workflow integration.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(workflows-api): move runtime admin endpoints into Elsa.Workflows.Api

Per PR review feedback: rather than introducing a new sub-module
(Elsa.Workflows.Runtime.Admin) for the four pause/resume/status/force
endpoints, fold them into the existing Elsa.Workflows.Api project. That
project already references both Elsa.Workflows.Runtime and
Elsa.Api.Common (FastEndpoints) and is the established home for
client-facing workflow APIs — so the admin endpoints belong there.

Changes:
* New folder src/modules/Elsa.Workflows.Api/Endpoints/RuntimeAdmin/ with
  Models.cs and Pause/Resume/Status/Force/Endpoint.cs. Namespaces moved
  from `Elsa.Workflows.Runtime.Admin` → `Elsa.Workflows.Api.Endpoints.RuntimeAdmin`.
* Deleted src/modules/Elsa.Workflows.Runtime.Admin/ entirely and removed
  it from Elsa.sln. The ShellFeature marker class
  (WorkflowRuntimeAdminFeature) is no longer needed — the existing
  WorkflowsApiFeature already discovers FastEndpoints in the Workflows.Api
  assembly.
* No consumer changes: the endpoints sit in the same routes
  (/admin/workflow-runtime/*) and behave identically.

Note on the second architectural point ("update IShellFeature if cleaner"):
the IShellFeature contract is defined in the external CShells NuGet
package, not in this repo, so we cannot add a DeactivateAsync hook
without an upstream CShells change. The current IHostedService.StopAsync
hook continues to work correctly for the host-stop path; per-shell
deactivation would require either a CShells upstream addition or a
separate Elsa-owned shell-feature variant — neither lighter than what we
have today.

Tests: 103/103 runtime unit + 249/249 workflow integration pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(common): introduce IsFatal exception extension + apply to drain best-effort catches

Per PR review feedback on the static-analyzer "Generic catch clause"
comments: rather than catching ALL exceptions in best-effort drain code
paths, narrow the swallow to non-fatal exceptions. Process-fatal
conditions (StackOverflowException, AccessViolationException,
SEHException, ThreadAbortException, OutOfMemoryException) propagate so
the host's failure-fast policy can act on them, while normal failures
(InvalidOperationException, IOException, etc.) continue to be logged
and allowed through so a single misbehaving ingress source / activity
cannot abort the overall drain.

Highlights:
* New `Elsa.Common.Extensions.ExceptionExtensions.IsFatal` utility:
  classifies fatal conditions, unwraps reflection-style wrappers
  (TypeInitializationException, TargetInvocationException) before
  classification, and treats InsufficientMemoryException (the
  recoverable OOM subclass) as non-fatal.
* Applied as a `when (!ex.IsFatal())` filter to:
    - BurstHandle.Cancel (cancel callback try/catch)
    - DrainOrchestrator.PauseOneSourceAsync (per-source exception path)
    - DrainOrchestrator.TryForceStopAsync
    - DrainOrchestrator.ForceCancelActiveBurstsAsync (per-burst loop)
    - DrainOrchestrator.PersistInterruptedAsync (orphan log write,
      instance save, log write)
    - DrainOrchestrator.DrainAsync outer catch (existing
      `not InvalidOperationException` filter extended)
    - InterruptedRecoveryScan (per-instance restart loop)

Tests: 7 new unit tests for IsFatal classification (fatal types,
recoverable types, wrapped causes, null tolerance). Full suites:
103/103 runtime unit (incl. 14/14 in Common.UnitTests including new
tests) + 249/249 workflow integration pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(workflows-runtime): integrate CShells 0.0.15 lifecycle hooks (IDrainHandler + IShellInitializer)

CShells 0.0.15 ships the lifecycle framework needed for first-class per-shell
graceful shutdown — IDrainHandler / IShellInitializer / IShellLifecycleSubscriber.
This commit bumps the package, migrates Elsa's existing usage of the removed
0.0.14 API, and registers the runtime drain orchestrator + pause-state initializer
through the new primitives.

Highlights:
* `ElsaShellDrainHandler : IDrainHandler` — bridges per-shell drain into
  `IDrainOrchestrator.DrainAsync(DrainTrigger.ShellDeactivation, ct)`. Invoked
  by CShells when a shell enters `ShellLifecycleState.Draining`; the drain
  handler's CancellationToken is signalled when the per-shell deadline elapses,
  so the orchestrator's own deadline-bounded protocol nests cleanly.
  Coexists with the host-stop `DrainOrchestratorHostedService`; the
  orchestrator's `DrainAsync` is idempotent — second invocations log and skip.
* `InitializePauseStateShellInitializer : IShellInitializer` — replaces the
  IStartupTask variant in shell-aware deployments. IShellInitializer fires on
  EVERY shell (re)activation, including reactivations after a reload — exactly
  what FR-028 requires. The IStartupTask remains for IModule consumers where
  there is no shell platform.

Migrations (CShells 0.0.14 → 0.0.15 breaking changes):
* `ActivateShellTenants`: was `IShellActivatedHandler` + `IShellDeactivatingHandler`,
  now `IShellInitializer` + `IDrainHandler`.
* `MultitenancyFeature`: registrations updated to the new transient interface,
  `using CShells.Hosting` → `using CShells.Lifecycle`.
* `Reload/Endpoint`, `ReloadAll/Endpoint`: `IShellManager` → `IShellRegistry`,
  `ReloadShellAsync` → `ReloadAsync` (returns `ReloadResult` with `Error`),
  `ReloadAllShellsAsync` → `ReloadActiveAsync` (returns
  `IReadOnlyList<ReloadResult>` with per-shell errors aggregated into 503).

Build + restore:
* `Directory.Packages.props`: all CShells.* packages bumped to 0.0.15.
* `NuGet.Config`: added `cshells-feedz` source
  (https://f.feedz.io/sfmskywalker/cshells/nuget/index.json) and split the
  package-source-mapping pattern into `CShells` (exact) + `CShells.*`
  (prefix). Single-pattern `CShells*` does NOT match correctly under
  PackageSourceMapping.

Tests: 103/103 runtime unit + 249/249 workflow integration pass on the new
package version.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(workflows-runtime): apply PR #7424 review feedback

Consolidates the architectural fixes asked for during /review:

- Extract IWorkflowRuntimeAdminService to back the four /admin/workflow-runtime endpoints with a single domain service; thin Pause/Resume/Status/Force endpoints to delegating shells.
- Remove StateChanged C# event from IQuiescenceSignal (Constitution VII: no external subscribers existed; mediator was suggested as the alternative if/when it's needed).
- Promote InitializePersistedStateAsync to IQuiescenceSignal, dropping the concrete-cast in both InitializePauseStateStartupTask and InitializePauseStateShellInitializer.
- Invert ingress-source DI to Lazy<IEnumerable<IIngressSource>> to break the cycle through IQuiescenceSignal; ingress adapters take the signal directly via primary constructor.
- Replace Guid.NewGuid().ToString("N") with IIdentityGenerator in InterruptedLogExtensions and DrainOrchestrator.
- Switch admin-audit timestamps to ISystemClock in WorkflowRuntimeAdminService.
- Make GracefulShutdownOptions.StimulusQueueMaxDepthWhilePaused nullable (null = unlimited).
- Rename RuntimeForceRequested → RuntimeForceDrainRequested.
- Apply IsFatal exception filter to drain best-effort catches.
- Rename IBurstRegistry.EnumerateActive → ListActiveBursts.
- Refresh "Phase X" comments to user-story (USx) references.
- Delete unused IngressAttributionExtensions, IngressSourceServiceCollectionExtensions, IngressSourceRegistrationOptions.
- Migrate Elsa.Shells.Api.Tests to CShells 0.0.15 IShellRegistry / ReloadResult surface.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(graceful-shutdown): apply IsFatal filter to deadline-breach test catch

Aligns the test scaffolding's swallow-everything catch with the project standard introduced in c00eee80c so the analyzer no longer flags the bare `catch` clause. The semantics are unchanged — non-fatal exceptions (OCE, TimeoutException, workflow exceptions) are still acceptable test outcomes.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(workflows-runtime): close IngressSourceRegistry first-access race + spelling

Replaces the non-atomic `_entries.Count > 0` early-return guard in
EnsureMaterialized with a double-checked lock against a volatile
`_materialized` flag, so concurrent first callers can no longer both
iterate the source factory and crash one of them with a "Duplicate ingress
source registration" InvalidOperationException. Adds a regression test that
launches 16 readers behind a TaskCompletionSource gate and asserts every
reader observes the full source set without throwing.

Also flips the British spellings introduced in this PR's scope to American
English (materialize/behavior) — project convention going forward.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(constitution): require American English for new code (v1.0.1)

Adds a "Spelling & language" bullet under principle III (Convention-Driven Design): every newly-introduced symbol, comment, identifier, error message, XML doc, commit message, and Speckit artifact uses American English. Established public API symbols (e.g. WorkflowSubStatus.Cancelled) are not renamed retroactively. PATCH bump because this is a clarification of an existing principle, not a new principle.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Add Greploop skill and workflow for GitLab, GitHub, and Perforce integration

- Introduced Greploop, an iterative optimization and review workflow for GitLab MRs, GitHub PRs, and Perforce changelists.
- Added API and GraphQL references for fetching and resolving review skill.

* Remove GenerateWorkflowVariableAccessorsTests; redundant ExpandoObject type check in handlers

* Potential fix for pull request finding 'CodeQL / Untrusted Checkout TOCTOU'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* fix(graceful-shutdown): apply PR #7424 review feedback round 2

Two P1 findings from Greptile:

1. DrainOrchestrator.ForceCancelActiveBurstsAsync was sequential — each
   burst was cancelled, awaited up to ForceCancelSettleTimeout (2 s), and
   persisted before the next burst's Cancel() ran. Total wall time was
   O(N × 2 s) and bursts 2..N kept executing at full speed during prior
   bursts' settle waits, defeating the intent of force-cancel under
   concurrency.

   Refactored to three phases:
   - Phase A — cancel every handle synchronously (cheap CTS.Cancel calls)
     so all runners observe cancellation simultaneously.
   - Phase B — await every Disposed signal in parallel under a single
     shared ForceCancelSettleTimeout. Total wall time bounded regardless
     of N.
   - Phase C — persist Interrupted for each handle sequentially (keeps
     DbContext usage single-threaded; per-handle work is small).

   Per-phase failures are caught with !ex.IsFatal() and logged so a single
   misbehaving handle doesn't abort the rest of the batch.

2. ShellFeatures/WorkflowRuntimeFeature.ConfigureServices was missing the
   IWorkflowRuntimeAdminService registration that Features/WorkflowRuntimeFeature
   already had. Any CShells deployment that includes the Pause / Resume /
   Status / Force admin endpoints (in Elsa.Workflows.Api) would throw
   InvalidOperationException at endpoint construction. Added the singleton
   alongside the other graceful-shutdown registrations with a comment
   pointing out the symmetry with the IModule path.

17/17 graceful-shutdown integration tests pass; 103/103 runtime unit tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Potential fix for pull request finding 'CodeQL / Untrusted Checkout TOCTOU'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* fix(ci): harden greploop.yml against CodeQL Actions findings

CodeQL flagged 12 findings on .github/workflows/greploop.yml after the
prior commit (c08183a3c) addressed an earlier round. Two distinct issue
classes remain:

1. Code injection (× ~10): step-output values
   (steps.pr_head.outputs.head_sha / head_repo_owner / head_repo_name /
   head_ref) and inputs.pr_number were interpolated directly into shell
   `run:` blocks via `${{ ... }}`. Because PR author controls the branch
   name and the manual-dispatch input, those values can carry shell
   metacharacters. Standard fix: route every such interpolation through
   an `env:` block on the step, then reference $VAR inside the script.
   Applied to the Resolve, Resolve PR head metadata, and Checkout PR
   branch steps.

2. Untrusted Checkout TOCTOU + Checkout of untrusted code in trusted
   context: the workflow runs on `issue_comment` (a privileged trigger)
   and checks out PR-author code. Mitigations stacked here:
   - Author-association gate already restricts the trigger to OWNER /
     MEMBER / COLLABORATOR (existing).
   - Step-output values now travel via env vars (above).
   - Resolve step rejects pr_number that isn't ^[0-9]{1,10}$ — so
     downstream `gh pr view` and the prompt argument can't be hijacked.
   - Checkout step now validates HEAD_SHA matches ^[0-9a-f]{40}$ and the
     repo owner/name match ^[A-Za-z0-9_.-]+$ before either reaches a
     URL or a git command.
   - Existing TOCTOU guard preserved: re-fetch head SHA at checkout
     time and abort if it changed since the initial resolve.

These match the canonical "Securing your GitHub Actions workflows"
patterns recommended by CodeQL.

No functional change to greploop's runtime behaviour.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(graceful-shutdown): drop [SingleNodeTask] from InitializePauseStateStartupTask

Greptile P1: [SingleNodeTask] gates the task to a single cluster winner via
distributed lock, but IQuiescenceSignal is a singleton scoped to each node's
DI container — each node holds its own in-memory QuiescenceState. With the
attribute, only the winning node restored the persisted pause; every other
node started with QuiescenceReason.None and accepted new work, silently
defeating PausePersistence = AcrossReactivations.

Removed [SingleNodeTask] (and the corresponding using) so the task runs on
every node. Expanded the doc <remarks> to call out the per-node requirement
and point at the shell-aware counterpart (InitializePauseStateShellInitializer)
which is correctly per-node by virtue of being an IShellInitializer.

17/17 graceful-shutdown integration tests still pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Potential fix for pull request finding 'CodeQL / Code injection'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* chore(deps): bump CShells 0.0.15 → 0.0.17

0.0.17 ships our blueprint-aware-routing PR (valence-works/cshells#93) plus
four follow-up fixes the maintainer added on top:

- e56ebd8 — PreWarmShells removed entirely; ShellMiddleware now does
  cold-start endpoint matching (re-runs endpoint resolution after lazy
  activation so the very first request to a cold shell hits its endpoint).
- cbe5ee2 — GetCandidateSnapshot returns a bounded ShellRouteCandidateSnapshot
  with accurate total counts; sensitive-data redaction in routing logs;
  DefaultShellRouteIndex implements IDisposable.
- cd7d4f5 — Last-good snapshot served on rebuild failure (the deferred
  Copilot review concern); root-path fallback when path-by-name misses.
- c3679d9 — Cold-start endpoint matching respects inline route constraints;
  path-name convention tightening; dead duplicate-detection cleanup.

Net effect for elsa-core:
- Cold blueprints serve their first request via lazy activation, with
  endpoints correctly resolved post-activation.
- Reloaded shells re-activate and serve on the next matched request.
- Non-name-mode routing keeps serving the previous snapshot during a
  transient blueprint-provider outage.
- No need to call PreWarmShells from Elsa.ModularServer.Web — removed.

The only API removal that touches elsa-core is PreWarmShells. No code
references IShellRouteIndex / ShellRouteCriteria / GetCandidateSnapshot
directly, so the API-shape changes in cbe5ee2 don't ripple here.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(graceful-shutdown): persist Interrupted under non-drain bounded token

Greptile P? finding on the prior force-cancel two-phase fix: Phase B's
inner catch on OperationCanceledException ("drain CT fired — proceed to
persist anyway") was a lie in practice. Phase C immediately passed the
same already-cancelled drain token into PersistInterruptedAsync; the
first DB call (instanceStore.FindAsync) observed the cancellation and
threw OperationCanceledException; the outer non-fatal Exception filter
swallowed it and only logged an error. Net effect: on host shutdown
deadline breach, every burst after cancellation could fail to be
persisted as Interrupted, leaving instances in an unrecovered executing
state.

Phase C now creates a per-handle CancellationTokenSource bounded to a
new PersistInterruptedTimeout (5 s) that is NOT linked to the drain CT.
Each persist gets up to 5 s to land the row update + forensic log entry
even after the drain CT has fired. The bound prevents a stuck DB from
hanging shutdown indefinitely (per-handle worst case is small; total
Phase C upper bound is N × 5 s, but typical persists are millisecond
scale).

Comment expanded to call out why the persist token is independent of
the drain token, so the rationale doesn't drift again.

17/17 graceful-shutdown integration tests still pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(graceful-shutdown): extract PassiveIngressSource base class

The three IIngressSource implementations that ship with this PR
(InternalBookmarkQueueIngressSource, HttpTriggerIngressSource,
ScheduledTriggerIngressSource) were ~25 lines each and ~22 lines of
those were verbatim copies of each other:

- ctor signature `(IQuiescenceSignal signal)`
- `PauseTimeout => TimeSpan.FromMilliseconds(50)`
- `CurrentState => signal.IsAcceptingNewWork ? Running : Paused`
- `PauseAsync` / `ResumeAsync` returning `ValueTask.CompletedTask`

The shared trait is that none of them does any work at pause time —
the actual pause enforcement lives in another layer
(`HttpWorkflowsMiddleware` short-circuits to 503,
`BookmarkQueueProcessor` consults the signal at the top of each
invocation, scheduled triggers dispatch through the bookmark queue and
inherit that behaviour transitively). The IIngressSource adapter is
purely diagnostic: it makes the source visible in
`DrainOutcome.Sources` and the admin status endpoint.

Extracted that pattern into `PassiveIngressSource` (abstract base in
`Elsa.Workflows.Runtime.IngressSources`). Subclasses now provide only
`Name`; `PauseTimeout` is `virtual` with a 50 ms default; everything
else is fixed by the base. The three concretes drop from ~25 lines to
~12 lines each.

The base's XML `<remarks>` calls out when to use it ("your component
already cooperates with IQuiescenceSignal at its hot path") and when
to implement IIngressSource directly ("the source owns concrete
pause/resume behaviour — e.g. a message-queue consumer that calls
Pause() on its underlying client"), so future contributors don't
mis-extend the base for sources that need real work at pause time.

No behavioural change. 18 graceful-shutdown integration + 39 runtime
unit tests still pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(graceful-shutdown): align IIngressSource name to singular

The three IIngressSource names were inconsistent:

  http.trigger                       (singular)
  internal.bookmark-queue-worker     (singular)
  scheduling.triggers                (PLURAL — outlier)

The plural slipped in when addressing Greptile's earlier comment to
avoid `scheduling.cron` (which would imply Cron-only coverage). The
right move was to pick a generic word and stay singular like the rest
of the suite — the suite's mental model is "the X source", one
instance per registry slot, regardless of how many triggers or items
it dispatches internally.

Renamed to `scheduling.trigger`. The `<remarks>` block keeps the
"covers Cron, Timer, StartAt, Delay" explanation and now also
explicitly notes the singular convention so future contributors don't
re-pluralize.

Zero test fallout — the literal "scheduling.triggers" only appeared in
the source file itself. Tests of the other two sources all use
singular forms.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(runtime-admin): rename Force endpoint to ForceDrain

"Force" alone is meaningless out of context — force what? — and it
sits oddly next to the verb-named siblings Pause / Resume / Status.
The matching admin service method is already IWorkflowRuntimeAdminService.
ForceDrainAsync, so ForceDrain is the natural pair.

Renamed:
- src/modules/Elsa.Workflows.Api/Endpoints/RuntimeAdmin/Force/        → ForceDrain/
- namespace ...Endpoints.RuntimeAdmin.Force                            → ...ForceDrain
- class ForceEndpoint                                                  → ForceDrainEndpoint
- class ForceRequest                                                   → ForceDrainRequest
- class ForceResponse                                                  → ForceDrainResponse
- route  POST /admin/workflow-runtime/force                            → /force-drain

Zero external references — no tests, docs, or OpenAPI clients used the
old symbols or the old route literal, so this is a contained pre-ship
rename. Directory move went through `git mv` so commit history follows
the file.

17/17 graceful-shutdown integration tests still pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(ci): correct formatting in checkout step within greploop.yml

Adjusted indentation of environment variables in the checkout step for improved consistency and readability.

* fix(ci): repair malformed Checkout repository step in greploop.yml

The step accumulated stray env keys, an extra `uses:`, and bash commands
that didn't belong inside it (line 65 onward), causing a YAML parse
error on push. The valid structure has two distinct checkout steps:

  - Checkout repository  : actions/checkout@v4 with fetch-depth: 0
  - Checkout PR branch   : env: + run: with SHA validation + git fetch
                            + git checkout --detach

The PR-branch step (line 78+) was already correct and unchanged. This
fix restores the first step to its intended single-purpose shape (just
checks out the workflow file's commit so the greploop skill is on disk
before the run-greploop step uses it).

No functional change to runtime behaviour or to the security posture
established in the prior hardening commit (0db4ca23e). The PR-branch
checkout still validates HEAD_SHA / HEAD_REPO_OWNER / HEAD_REPO_NAME
shape before they reach a URL or git command.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(ci): remove literal `${{ }}` from greploop.yml comment

GitHub Actions parses `${{ ... }}` workflow expressions across the entire
YAML file, including inside `run:` script comments. The comment that
explained the env-var hardening pattern contained the literal sequence
`${{ }}` (with a space, intended as an English-language description),
which the expression parser rejected as "An expression was expected"
(line 81 col 14).

Reworded the comment to describe the substitution form in prose without
the literal token sequence. Functional behaviour unchanged.

* refactor(graceful-shutdown): rename burst → execution cycle

The graceful-shutdown work introduced "burst of execution" as a
first-class domain concept. The term arrived without rationale and
isn't standard in the workflow-engine domain. Renamed to
"execution cycle" — reads more naturally as the loop-with-commit unit,
is more idiomatic in workflow vocabulary, pairs cleanly with the
existing WorkflowExecutionContext, and avoids collisions with Elsa's
existing terms (Run, Execution, Dispatch, Invocation, Stimulus, Step).

Renamed types
- IBurstRegistry              → IExecutionCycleRegistry
- BurstRegistry               → ExecutionCycleRegistry
- BurstHandle                 → ExecutionCycleHandle
- BurstTrackingMiddleware     → ExecutionCycleTrackingMiddleware
- BurstAwareCommitStateHandler → ExecutionCycleAwareCommitStateHandler

Renamed members
- BeginBurst                          → BeginCycle
- ListActiveBursts                    → ListActiveCycles
- BurstHandleKey constant + value     → ExecutionCycleHandleKey
- ActiveBurstCount (IQuiescenceSignal,
  RuntimeAdminStatus, StatusResponse) → ActiveExecutionCycleCount
- WaitForBurstsAsync (private)        → WaitForCyclesAsync
- ForceCancelActiveBurstsAsync (priv) → ForceCancelActiveCyclesAsync
- UseBurstTracking                    → UseExecutionCycleTracking
- DrainOutcomeDto.BurstsForceCancelledCount → ExecutionCyclesForceCancelledCount
- _burstRegistry / burstRegistry      → _cycleRegistry / cycleRegistry

Backwards-compatibility preservation (the only persisted JSON key)
- WorkflowInterruptedPayload.BurstDuration property → ExecutionCycleDuration
  with [JsonPropertyName("BurstDuration")] so the persisted JSON wire
  key stays "BurstDuration" forever. Pre-merge testers' log records
  still deserialise correctly. The contract test on
  WorkflowInterruptedPayloadContractTests still asserts the wire key
  "BurstDuration" appears in the serialised JSON — confirms the
  guarantee is enforced.

Other unstructured surfaces
- WorkflowExecutionLogRecord.Message text "Workflow burst was force-
  cancelled..." now says "Workflow execution cycle was force-cancelled
  ..." for new records. Old rows keep their old text — purely cosmetic
  free-text field.
- Structured log placeholder {BurstId} in DrainOrchestrator log lines
  → {ExecutionCycleId}.
- Lowercase prose / XML doc comments updated throughout.

Test renames
- BurstRegistryTests          → ExecutionCycleRegistryTests
- BurstTrackingMiddlewareTests → ExecutionCycleTrackingMiddlewareTests
- Test method names + DisplayName strings updated.

Spec docs (specs/002-graceful-shutdown/) updated to match the new
vocabulary; the historical task records in tasks.md keep the old names
as-is to preserve the audit trail of what was originally built.

Verification
- dotnet build: clean across net8.0 / net9.0 / net10.0.
- 103/103 Elsa.Workflows.Runtime.UnitTests pass.
- 17/17 GracefulShutdown integration tests pass.
- 4/4 WorkflowInterruptedPayloadContractTests pass — confirms the
  "BurstDuration" JSON wire-key preservation is intact.

No changes to migrations or DB column names — confirmed via grep.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(ci): give greploop.yml gh-cli a repo context before checkout

The "Resolve PR head metadata" step runs `gh pr view` before
actions/checkout, so there is no `.git` directory and gh's "current
repo" detection fails with `fatal: not a git repository`. The
prior commits to this file masked the runtime failure because the
workflow itself was YAML-invalid — once it became valid, the
workflow_dispatch trigger surfaced this real-world execution bug.

Set GH_REPO=${{ github.repository }} on both `gh pr view` steps. The
gh CLI honours GH_REPO as an explicit repo override, so it no longer
needs git context. Same fix on the "Checkout PR branch" validation
call which also uses gh pr view before the manual fetch.

The error reported as "Invalid workflow file: ... (Line 81 Col 14)"
on PR #7424 was stale from commit ed580682a (which had the bad
comment with literal `${{ }}`); commit 3a8ad0d08 fixed the YAML, but
because greploop's `if:` condition only matches workflow_dispatch /
issue_comment events, push events on later commits were skipped
without re-running validation, so the GitHub UI kept showing the
old error. A workflow_dispatch run on the current SHA now passes
validation and reaches "Resolve PR head metadata", which is what
this commit fixes.

* fix(graceful-shutdown): wire IngressPauseTimeout option to drain orchestrator

GracefulShutdownOptions.IngressPauseTimeout was documented as "Default
per-ingress-source pause timeout" but DrainOrchestrator.PauseOneSourceAsync
read source.PauseTimeout directly and never consulted the option. The
configured value was silently ignored — operators who set
GracefulShutdownOptions:IngressPauseTimeout = 10s were getting whatever
each source's hardcoded value was (50 ms for the three PassiveIngressSource
subclasses we ship), with no way to tune it globally.

Precedence (per the spec's intent of "overridable at registration and by
configuration"):

  1. Per-source positive value wins (source.PauseTimeout > Zero).
  2. Otherwise fall back to the configured GracefulShutdownOptions.
     IngressPauseTimeout default.
  3. Resolved value is capped at the overall drain deadline so a single
     misbehaving source cannot exceed the host's shutdown budget.
  4. 1 ms safety floor remains so a misconfigured zero default still
     produces a non-zero CancelAfter.

Changes:

- DrainOrchestrator.PauseOneSourceAsync — adds the precedence above with
  a comment block explaining each step.
- IIngressSource.PauseTimeout — XML doc clarifies the Zero-defers-to-
  config semantics.
- GracefulShutdownOptions.IngressPauseTimeout — XML doc says it's the
  fallback when the source returns Zero; <remarks> spells out the
  precedence and the overall-deadline cap.
- PassiveIngressSource.PauseTimeout — virtual property now returns Zero
  (was 50 ms). The three shipped subclasses (HttpTriggerIngressSource,
  ScheduledTriggerIngressSource, InternalBookmarkQueueIngressSource)
  consequently defer to the configured default — flipping the wire-up
  bug from "configured value silently ignored" to "configured value
  honoured by default for passive sources". Passive subclasses that
  want a specific value can still override.

103/103 runtime unit + 17/17 graceful-shutdown integration tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(graceful-shutdown): rename IInterruptedRecoveryScan → IInterruptedRecoveryScanner

The interface had a single verb-method (`ScanAndRequeueAsync`) and its
XML doc described what it *does* ("Scans the workflow instance store
for instances..."). That's an agent role — a scanner that performs a
scan — but the noun-shaped name `IInterruptedRecoveryScan` read as
"the scan itself", which is misleading because the scan results /
event are not first-class types in the codebase.

Renamed to `IInterruptedRecoveryScanner` / `InterruptedRecoveryScanner`
to match the existing `-er` convention in this codebase (Restarter,
Generator, Resolver, etc.). The method stays `ScanAndRequeueAsync` —
the scanner *performs* a scan-and-requeue.

Also renamed the constructor parameter `scan` → `scanner` in
RecoverInterruptedWorkflowsStartupTask, and the local variable `scan`
→ `scanner` in InterruptedRecoveryIntegrationTests, so the "scanner
does the scan" mental model is consistent throughout.

Surface impact (all internal — no API or persistence touch points):
- 2 source files renamed via git mv (interface + implementation)
- 1 test file renamed (InterruptedRecoveryScanTests → ScannerTests)
- DI registrations in both Features/ and ShellFeatures/ WorkflowRuntimeFeature
- 1 startup-task constructor parameter
- Spec doc references under specs/002-graceful-shutdown/

103/103 runtime unit + 17/17 graceful-shutdown integration tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(graceful-shutdown): extract DrainTriggerExecutor

ElsaShellDrainHandler (CShells IDrainHandler) and
DrainOrchestratorHostedService (.NET IHostedService.StopAsync) inlined
near-identical try/catch/log shapes around IDrainOrchestrator.DrainAsync:

  - call DrainAsync(<trigger>, ct)
  - branch on outcome: DeadlineExceeded/AbortedByUnhandledException →
    Warning, otherwise Information
  - catch InvalidOperationException (parallel-drain rejected by the
    orchestrator) → log Information and swallow

The two had already drifted: host-stop's success log omitted the
paused/waited durations the shell-handler version included, and the
"skipped" message disagreed on the trigger label ("Host-stop drain
skipped" vs "Shell drain skipped"). Centralised the shape in a small
internal static helper so the two — and any future trigger source —
stay uniform.

Both call sites collapse to a single line. Net diff drops 22 lines from
the two consumers and adds a 25-line helper that they both delegate to.
The unified log copy now consistently includes paused/waited durations
on the success path and uses the caller-supplied contextLabel
("Shell drain", "Graceful drain") in all three messages so operators
can attribute log entries by trigger source.

Files:
- src/modules/Elsa.Workflows.Runtime/Services/DrainTriggerExecutor.cs (new)
- src/modules/Elsa.Workflows.Runtime/Lifecycle/ElsaShellDrainHandler.cs
- src/modules/Elsa.Workflows.Runtime/HostedServices/DrainOrchestratorHostedService.cs

103/103 runtime unit + 17/17 graceful-shutdown integration tests pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(workflows-runtime): drop redundant DrainOrchestratorHostedService from CShells path

In CShells deployments, host stop already drives drain via CShellsStartupHostedService → IDrainHandler →
ElsaShellDrainHandler, scoped per shell (FR-027). The additional .AddHostedService<DrainOrchestratorHostedService>()
in ShellFeatures/WorkflowRuntimeFeature was firing a second non-force DrainAsync that the orchestrator rejected
with InvalidOperationException — silently swallowed by DrainTriggerExecutor, but logged on every host stop and
semantically muddled (IHostedService is host-level, not per-shell).

Keep the registration on the IModule path (Features/WorkflowRuntimeFeature) where there is no shell platform
and host-stop is the only available drain trigger. Update ElsaShellDrainHandler XML docs to reflect the now-clean
single-trigger model in CShells.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(workflows-runtime): always dispose ExecutionCycleHandle in tracking middleware

Previously the success path relied on ExecutionCycleAwareCommitStateHandler to dispose the handle after the
runner's commit completed. If a custom dispatcher or test double exited the pipeline without invoking commit
(by design or by accident), the handle stayed registered, IExecutionCycleRegistry.ActiveCount never reached
zero, and drain spun in WaitForExecutionCyclesAsync until the deadline fired — incorrectly force-cancelling
instances that had already finished cleanly.

Collapse the existing try/catch(rethrow) into try/finally so the middleware itself disposes the handle for
both exception and commit-elided paths. Disposal remains idempotent via the ExecutionCycleHandle._disposed
Interlocked guard, so the normal-path dispose by ExecutionCycleAwareCommitStateHandler is a harmless no-op.

Adds an integration regression test that drives the middleware with a stub Next that returns without invoking
commit and asserts ActiveCount returns to zero.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(workflows-runtime): serialize QuiescenceSignal pause-state persistence

Both PauseAsync and ResumeAsync used to release the inner lock before issuing the persistence I/O. A rapid
Pause → Resume sequence could leave the persisted state inconsistent: PauseAsync's slow SaveAsync could land
AFTER ResumeAsync's DeleteAsync, leaving the key present in the store while in-memory state was None. On host
restart, InitializePersistedStateAsync would find the stale key and start the runtime in the paused state
the operator had already cancelled.

Introduce a dedicated SemaphoreSlim that serializes persistence I/O, with each I/O re-reading the live
in-memory state inside the semaphore. N racing Pause/Resume calls now produce N serialized writes, each
reflecting the most recent in-memory transition — so the final persisted state always matches final
in-memory state.

Adds a regression test that gates SaveAsync, races a Resume behind it, and asserts the store is empty after
both complete.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(workflows-runtime): trim verbose comment in ExecutionCycleTrackingMiddleware finally block

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(workflows-runtime): use 'using var' for ExecutionCycleHandle in tracking middleware

Replace the explicit try/finally that only existed to call handle.Dispose() with a `using var` declaration —
identical semantics (compiler-emitted finally with idempotent dispose), more idiomatic. The regression test
HandleReleasedWhenCommitIsElided continues to validate the leak-free property.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(workflows-runtime,api): proper 409 conflict shape + shell-scoped pause-persistence key

ForceDrain endpoint: the 409 path returned `new ForceDrainResponse()` whose `Outcome` was null at runtime
despite the `= null!` annotation, so any strongly-typed client deserializing the conflict body and reading
`Outcome.OverallResult` got an NRE. Switch to the existing `ConflictResponse` shape with
`Code = "DrainInProgress"` and the current runtime status. Routed via HttpContext.Response.WriteAsJsonAsync
because Send.ResponseAsync is constrained to the endpoint's TResponse and cannot send a sibling DTO.

QuiescenceSignal persistence key: the DI-registered `IQuiescenceSignal` was constructed with
`shellName = null` (DI doesn't inject `string?` defaults), so every shell shared the key
`elsa.quiescence.pause.default`. In a CShells multi-shell deployment under
PausePersistencePolicy.AcrossReactivations this caused cross-shell contamination — pausing shell A would
re-pause shell B on its next activation. Replace the simple AddSingleton<IQuiescenceSignal,...> registration
in ShellFeatures/WorkflowRuntimeFeature with a factory that injects `CShells.ShellSettings` and forwards
`Settings.Id` as the shell name. The IModule registration is unchanged (no shell platform; null shellName
remains correct there).

Adds a unit regression test that two QuiescenceSignal instances with different shellNames write to disjoint
persistence keys and never to the legacy "default" key.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(quiescence): use TryGetValue for ContainsKey+indexer assertions

Combines existence check and value retrieval into a single dictionary lookup, addressing the code-quality
bot's repeated suggestion. No behavior change — both PauseWritesKey and PersistenceKeyIncludesShellName
still assert the same keys exist with the same content.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Update specs/002-graceful-shutdown/contracts/admin-endpoints.md

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* fix(workflows-runtime): decouple QuiescenceSignal persistence from caller cancellation

PersistAsync used to forward the caller's CancellationToken to both _persistenceMutex.WaitAsync and the
store I/O. If an HTTP request was cancelled between the in-memory transition (already committed under
_sync) and the persistence call, the I/O was silently skipped — leaving AdministrativePause set in memory
with no persisted record. The idempotent fast-path on subsequent PauseAsync calls (transitioned == false)
meant no retry would happen, so on host restart InitializePersistedStateAsync would find no key and the
runtime would come back unpaused, defeating PausePersistencePolicy.AcrossReactivations.

Drop the parameter from PersistAsync entirely; use CancellationToken.None for both the semaphore wait and
the store I/O. The in-memory transition is already committed by the time PersistAsync runs, so persistence
must complete to keep the store consistent with memory. The public PauseAsync/ResumeAsync methods still
accept a CancellationToken (interface contract) — it just no longer reaches the persistence layer.

Adds a regression test that calls PauseAsync with a pre-cancelled token and asserts both in-memory pause
and the persisted key land correctly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(workflows-runtime,api,docs): apply Copilot review feedback batch

Code:
- DrainOrchestrator.TryForceStopAsync now bounds force-stop with the *remaining* drain budget
  (deadlineAt - now), not the full overall TimeSpan. A per-source pause that already burned the
  shutdown window can no longer get another full deadline's worth of force-stop runway.
- DrainOrchestrator catch filter narrowed: drop `ex is not InvalidOperationException` exclusion.
  The "drain already in progress / completed" IOEs are thrown outside the protocol's try block,
  so they bubble out without entering this handler. Any IOE that lands here is incidental
  (e.g., from a store inside the drain) and should now be captured into the outcome rather
  than escaping the whole drain.
- ResumeEndpoint 409 path now returns the discriminated ConflictResponse shape (matching
  ForceDrain) instead of a plain StatusResponse. Routed via HttpContext.Response.WriteAsJsonAsync
  because Send.ResponseAsync is constrained to TResponse.
- Conflict codes aligned to kebab-case across both endpoints to match the contract spec
  (`runtime-draining` and `drain-in-progress`).

Spelling sweep — American English per constitution v1.0.1 III:
- DrainOrchestrator.cs: "serialised" → "serialized"
- WorkflowInterruptedPayload.cs: "serialised" / "deserialise" → "serialized" / "deserialize"
- PassiveIngressSource.cs: "behaviour" → "behavior"
- DeadlineBreachEndToEndTests.cs: "serialisable" → "serializable"
- specs/002-graceful-shutdown/quickstart.md: "behaviour" → "behavior"
- specs/002-graceful-shutdown/checklists/requirements.md: "behaviour" → "behavior"

Doc/contract alignment:
- quickstart.md: force route corrected from /force to /force-drain.
- quiescence-signal.md: removed StateChanged event from contract (interface doesn't define it);
  corrected persistence section to describe InitializePersistedStateAsync via shell initializer
  / startup task rather than constructor read; added the per-shell key discriminator.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(workflows-runtime): correct DI lifetimes for ExecutionCycleTrackingMiddleware and WorkflowRuntimeAdminService

Two strict-DI-validation failures surfaced in tests using BuildServiceProvider with validate-on-build:

1. ExecutionCycleTrackingMiddleware was registered as AddSingleton<>, but its constructor takes
   WorkflowMiddlewareDelegate next — supplied by the workflow execution pipeline builder via
   UseMiddleware<>(), not from DI. The registration was both unused (no consumer resolves it through
   the container) and broken (DI fails to construct it because next is unregistered). Removing both
   registrations.

2. IWorkflowRuntimeAdminService was registered as AddSingleton<> but depends on the scoped
   INotificationSender (mediator) — captive-dependency violation. All consumers (Pause/Resume/
   Status/ForceDrain endpoints) are FastEndpoints, which are scoped per request, so AddScoped is
   the correct alignment. The other deps (IQuiescenceSignal / IIngressSourceRegistry /
   IDrainOrchestrator / ISystemClock) are singletons and resolve fine from a scoped consumer.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(workflows-runtime): restore commit-handler-only success disposal + true Cancel idempotency

Two issues raised by Copilot's latest review on commit 32a9c0519:

1. ExecutionCycleTrackingMiddleware was disposing the handle at the end of InvokeAsync (via `using var`),
   but WorkflowRunner runs commit AFTER the pipeline returns (WorkflowRunner.cs:235). That meant the handle
   was disposed BEFORE the runner's terminal commit, and the drain orchestrator's
   `await handle.Disposed` would unblock too early — reintroducing the runner-clobber race the original
   design protected against (see ExecutionCycleAwareCommitStateHandler XML doc).

   Revert to the original shape: only dispose on exception path. ExecutionCycleAwareCommitStateHandler
   remains the SOLE success-path disposer, running in its finally block AFTER the inner commit lands.
   The earlier "leak when commit is elided" concern was a non-issue in production (the standard runner
   always commits); the buggy `HandleReleasedWhenCommitIsElided` test that specified the wrong contract
   is removed. The existing `ActiveCountReturnsToZero` test (which uses the real runner end-to-end)
   already verifies success-path disposal.

2. ExecutionCycleHandle.Cancel() was documented as idempotent but only short-circuited via the
   _disposed flag. Repeated Cancel() calls before Dispose could trigger the cancel callback multiple
   times — easy to accidentally fire non-idempotent cancellation side effects more than once. Add an
   Interlocked _cancelled guard so callback + CTS cancellation run at most once. Existing test that
   documented the leaky behavior is updated to assert the now-truly-idempotent contract.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Update logging levels and remove unused features in appsettings files

* refactor(workflows-runtime): improve graceful shutdown options handling and cleanup solution

Refactor the handling of `GracefulShutdownOptions` to ensure options are applied correctly without directly invoking the delegate. Update DI registrations to use appropriate lifetimes and remove redundant wrapper services. Additionally, clean up the solution by removing unused projects and documentation folders.

* feat(identity, workflows-runtime): add validation for identity and graceful shutdown options

Introduce validation capabilities for `IdentityTokenOptions` and `GracefulShutdownOptions`. Implement extension methods for option validation, enhance service registration, and add unit tests to ensure configurations are validated at startup. Update solution to include new unit test projects.

* update(docs): clarify shutdown log message expectations and levels in quickstart.md

Optimize explanation of expected log message sequence during graceful shutdown and specify logging levels.

* docs: amend constitution to v1.1.0 (SRP, DRY, KISS, conciseness under Principle VII)

* refactor(multitenancy): rename and restructure TenantTaskManager to TenantTaskLifecycleCoordinator

Rename `TenantTaskManager` to `TenantTaskLifecycleCoordinator` and relocate to a new directory structure, enhancing code organization and test consistency. Retain functional behaviors with no logic alterations. Update unit tests to reflect the naming changes, ensuring consistency with the refactored code structure.

* Update logging levels and dependencies

- Set default logging level to Debug in appsettings.Development.json
- Add missing using directives for Elsa workflows management and runtime features
- Update CShells package versions to 0.0.18-preview.104

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-05-02 19:27:08 +02:00
Sipke Schoorstra 3d8d3b7de2
Merge remote-tracking branch 'origin/release/3.6.1' 2026-04-20 15:08:10 +02:00
Sipke Schoorstra 3decb12680
feat: extend shells integration and modular server support (#7399)
* refactor(deps): use local CShells project refs

Replace CShells NuGet package references with direct project references to the local CShells source to enable developing and testing against local changes and simplify build integration across modules.

* Handle assembly load errors in feature discovery

Added error handling for assembly load failures in feature discovery to improve resilience. Also updated configuration for identity token options and removed unused service bus consumer dependencies. Simplified project structure by moving and cleaning up `Directory.Build.targets` files.

* Refactor configuration and service extension methods.

Moved `ShellSettingsExtensions` and `ShellConfiguration` to `CShells.Abstractions` for better modularity. Added new `ServiceCollectionFeatureExtensions` to improve options registration. Updated appsettings and references to support these changes.

* Introduce ManagementServiceCollectionExtensions to streamline activity and variable registration

Added `ManagementServiceCollectionExtensions` for registering Elsa activity types and variable descriptors, providing a modular and shell-feature-compatible approach to configuration. Updated relevant features to utilize these new extension methods, enhancing code modularity and reducing redundancy.

* Add resilience strategy registration to HTTP feature

Introduced `ResilienceServiceCollectionExtensions` to register resilience strategies within the `Elsa.Resilience.Core` module. Updated `HttpFeature` to incorporate resilience strategies, enhancing HTTP-related resilience configuration leveraging the new extension methods.

* Add new configuration options to JavaScriptFeature

Implemented multiple properties in `JavaScriptFeature` to enhance JavaScript execution: `AllowClrAccess`, `AllowConfigurationAccess`, `ScriptCacheTimeout`, `DisableWrappers`, and `DisableVariableCopying`. These additions enable more flexible and secure configuration of the Jint JavaScript engine.

* refactor(workflows): unify graph caching

Resolve workflow definitions first and store graphs under stable per-version-ID cache keys so different lookup paths share entries.
Centralize cache creation and change-token registration to remove duplicated caching logic.
Skip materializer-unavailable definitions to avoid caching null graphs and simplify flow.

* refactor(tests): centralize default IDs and materializer setup

Introduce constants for default definition and version IDs, and materializer name. Refactor tests to use these constants, streamline graph and definition resolution, and improve cache key creation by sharing logic across tests. Extend tests to check scenarios with unavailable materializers, ensuring caching only occurs for valid cases.

* extend(tests): enhance cache key verification in AutoUpdateTests

Added checks for both workflow definition and version cache keys in AutoUpdateTests to ensure comprehensive cache validation, improving test reliability and coverage.

* refactor(projects): update CShells project paths and solution configuration

Revised project reference paths in `Elsa.ModularServer.Web.csproj` for CShells projects and updated `Elsa.sln` to include new CShells projects, streamlining project organization and build configuration.

* Add `IWorkflowReferenceGraphBuilder` to `WorkflowManagementFeature`; rename `ResilienceShellFeature` to `ResilienceFeature`.

* Refactor `HttpFeature` to use `IMiddlewareShellFeature`, include `HttpWorkflowsMiddleware`, and update `HttpActivityOptions` defaults.

* Add `AddTypeAlias` and `AddVariableTypeAndAlias` extension methods to service collections

- Introduced `AddTypeAlias<T>` method in `ServiceCollectionExtensions.cs` for adding type aliases.
- Added `AddVariableTypeAndAlias<T>` method in `ManagementServiceCollectionExtensions.cs` to add variable types with aliases.

* Remove shell reload API endpoints, orchestrator, and associated tests from the codebase.

* Introduce `DefaultAdminUser` options and refactor `AdminUserInitializer` to use `IOptions`.

* Add user management endpoints: Delete, List, Update with enhanced user store functionality

* Implement `DefaultAdminUser` feature for initial admin bootstrap, decouple `SecurityRoot` from user management endpoints, update related documentation and permissions.

* Add role management endpoints: Delete, List, and Update, including role data models and handle obsolete SecurityRoot policy.

* Update CShells package references to version 0.0.12-preview.66 and refactor `TenantTaskManager` for improved task lifecycle management.

* Replace project references with package references in csproj files and remove unused folders.

* Integrate Nuplane features, add sample packages, and update dependency handling within ModularServer Web.

* Improve `CShells` startup endpoint registration and resolver handling

- Address duplicate endpoint registration by adding state-aware tracking and deduplication
- Resolve `WebRoutingShellResolver` constructor ambiguity by switching to factory-based registration
- Implement a startup-specific filter to prevent redundant endpoint remapping during `ShellsReloaded`
- Update project to use project references for `CShells` and `Nuplane` components in csproj files.

* Update logging configuration in appsettings for Development and Production

- Change default log level to 'Warning' in Development settings
- Adjust Microsoft.Hosting and Elsa.SamplePackage log levels to 'Information'
- Remove Microsoft.EntityFrameworkCore log level entry from Production settings

* Refactor assembly retrieval methods and update endpoint calls for consistency.

* Add `SampleEndpointFeature` and enhance logging and service integration

- Implement `SampleEndpointFeature` with a new endpoint for handling requests.
- Log endpoint access and integrate `ISampleService` with method `DoSomething`.
- Update logging configuration to include `CShells` and `Nuplane` log levels in Development settings.
- Update `Elsa.SamplePackage` to version 1.0.1 and manage dependencies with project and assembly references.
- Modify JSON configuration for `SampleEndpoint`.

* Update package versions for `CShells` to 0.0.13 and `Nuplane` to 0.0.1-preview.15 in props file.

* Refactor `DefaultAdminUserFeature` by renaming `ConfigureServices` to `Apply` and adjusting service registration method.

* Replace project references with package references across multiple projects and remove obsolete cshells-related solution entries.

* Remove `SampleCatalogEndpointExtensions.cs` and related endpoint mappings.

* Improve `TenantTaskManager` by using `TryRemove` for state clean-up and clarify `SemaphoreSlim` disposal behavior.

* Remove hardcoded default admin credentials and add warning for unconfigured AdminRoleName in admin user setup.

* Address unresolved review comments: fix doc comments, security defaults, compilation issue, and restore reload response contracts

Agent-Logs-Url: https://github.com/elsa-workflows/elsa-core/sessions/34eb1e13-833f-4b3c-9db6-2e9221d221b9

Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

* Refine reload endpoints: use specific exceptions, add error messages, rename ReloadedAt to Timestamp, remove unused model

Agent-Logs-Url: https://github.com/elsa-workflows/elsa-core/sessions/34eb1e13-833f-4b3c-9db6-2e9221d221b9

Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

* Potential fix for pull request finding 'Generic catch clause'

Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com>

* Potential fix for pull request finding 'Generic catch clause'

Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com>

* Potential fix for pull request finding 'Generic catch clause'

Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com>

* Add `ExceptionExtensions` with `IsFatal` method and simplify exception handling in `TenantTaskManager`. Remove unused properties from `Directory.Build.props`.

* Add unit tests for `TenantTaskManager` and fix potential state orphaning issue.

* Potential fix for pull request finding 'Generic catch clause'

Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com>

* Fix logger dependency in `SampleEndpointFeature` constructor to use correct type.

* Add unit tests for Elsa Shells API endpoints and update solution configuration.

* Refactor ShellReload models: remove ShellReloadItemResult, update ShellReloadResponse properties.

* Potential fix for pull request finding 'Generic catch clause'

Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com>
2026-04-18 14:33:34 +02:00
RalfvandenBurg c0e37e38d0
Fix/startuptask activate default tenant when empty (#7305)
* fix: activate Tenant.Default when no tenants configured (Option C)

Ensures IStartupTask implementations (e.g., PopulateRegistriesStartupTask,
RunMigrationsStartupTask) run when multitenancy is enabled but the tenant
provider returns an empty list.

- In DefaultTenantService: treat empty provider response as [Tenant.Default]
  in GetTenantsDictionaryAsync (initial load) and RefreshAsync
- Logic is internal to tenant service; no explicit call required

Co-authored-by: Cursor <cursoragent@cursor.com>

* test: add DefaultTenantService tests for empty-provider fallback

- ActivateTenantsAsync_WhenProviderReturnsEmpty_ActivatesDefaultTenant
- ListAsync_WhenProviderReturnsEmpty_ReturnsDefaultTenant
- ActivateTenantsAsync_WhenProviderReturnsTenants_ReturnsThoseTenants
- RefreshAsync_WhenProviderChangesFromTenantsToEmpty_KeepsDefaultTenant

Co-authored-by: Cursor <cursoragent@cursor.com>

* PR feedback disposes serviceprovider also

---------

Co-authored-by: Ralf <Ralf@Careconnections.nl>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-04-11 11:00:12 +02:00
Sipke Schoorstra 7bc9035f5e
Implement null TenantId for tenant-agnostic entities (#7226)
* Add ADR for Null Tenant ID, implement tenant-agnostic logic

Introduce ADR-0009 to document the use of `null` for tenant-agnostic entities, enhancing multitenancy handling. Update multitenancy features across the codebase, including EF Core query filters and ActivityRegistry, to handle null as a tenant ID, ensuring tenant-agnostic entities are accessible across all tenants.

* Add multitenancy support in `ActivityTestFixture` by registering `ITenantAccessor`.

* Update src/modules/Elsa.Common/Multitenancy/Implementations/DefaultTenantAccessor.cs

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Remove unused `using Elsa.Common.Multitenancy;` from WorkflowDefinitionActivityDescriptorFactory (#7230)

* Initial plan

* Remove unused using Elsa.Common.Multitenancy statement

Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

* Optimize ActivityRegistry.Find to prefer tenant-specific descriptors without performance regression (#7227)

* Initial plan

* Optimize Find(string type) to prefer tenant-specific descriptors with single-pass iteration

Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

* Apply review feedback: combine if statements and add comprehensive unit tests

Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

* Refactor tests for DRYness using theories and helper methods

Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

* Clean up extra whitespace in test file

Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

* Fix default tenant data visibility leak by removing NullIfEmpty conversion (#7229)

* Initial plan

* Remove NullIfEmpty conversion to align with ADR-0008 and ADR-0009

- Updated ElsaDbContextBase to use empty string for default tenant
- Updated ApplyTenantId to stop converting empty string to null
- Updated TenantAwareDbContextFactory to preserve empty string for default tenant
- Updated Store.cs to preserve empty string for default tenant
- This ensures: null = tenant-agnostic (visible to all), "" = default tenant

Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

* Add database migration to convert null TenantId to empty string for SqlServer

- Added Management migration to convert null to "" for WorkflowDefinitions and WorkflowInstances
- Added Runtime migration to convert null to "" for all runtime entities
- This ensures existing default tenant data is properly migrated per ADR-0008
- Note: Similar migrations needed for PostgreSql, MySql, Sqlite, and Oracle providers

Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>

* Clarify tenant handling logic in `ElsaDbContextBase` with new ADR references.

* Add tenant ID convention analysis documents and quick reference guide

* Implement tenant-agnostic functionality across modules

- Introduce `AgnosticTenantId` constant to manage tenant-agnostic entities.
- Modify entity handling logic to respect tenant-agnostic designations.
- Adjust workflow processing to include tenant-agnostic workflows.
- Update caching and activity descriptor logic to accommodate the `AgnosticTenantId`.

* Refactor tenant management and registry logic in `ActivityRegistry` for improved clarity and separation of tenant-specific and tenant-agnostic activity descriptors. Remove `TestTenantResolver` and update workflow definition handling for tenant support.

* Refactor `ActivityRegistry`: prioritize tenant-specific descriptors over tenant-agnostic and simplify descriptor retrieval logic.

* Improve async handling in `CommandHandlerInvokerMiddleware` to await tasks without blocking

* Update ADR to use asterisk as sentinel value for tenant-agnostic entities

Replace the previous convention of using `null` for tenant-agnostic entities with an asterisk (`"*"`) for improved clarity and system architecture. Updated ADR documentation, TOC, and dependency graph accordingly.

* Remove migration `ConvertNullTenantIdToEmptyString` and its associated designer file to clean up the codebase.

* Refactor `ActivityRegistry`: streamline activity descriptor removal logic and simplify tenant ID checks.

* Update Elsa.sln

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Simplify `RefreshDescriptorsAsync` by removing unnecessary local variable `currentTenantId`.

* Remove unused `currentTenantId` variable from `ActivityRegistry`.

* Add detailed semantic flow and key points to ADR 0009

Document the tenant ID flow from entity creation to query, emphasizing normalization and tenant-agnostic workflows. Update semantic flow diagrams and provide testing considerations for preserving `"*"` values in multi-tenant scenarios.

* Remove outdated Tenant ID Analysis and associated documents

* Add security-by-default design for tenant-agnostic entities in ADR

Enhance Architecture Decision Record to detail explicit requirements for tenant-agnostic database entities, highlighting differences between in-memory activity descriptors and persistent entities. Emphasize importance of setting `TenantId = "*"` to prevent accidental data leakage.

* Normalize tenant ID grouping in `ActivityRegistry` to unify null and agnostic IDs, reducing redundant processing.

* Refactor `SignalManager`: improve timeout handling and streamline signal task cancellation.

* Update src/modules/Elsa.Workflows.Core/Models/TenantRegistryData.cs

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Update src/modules/Elsa.Workflows.Core/Services/ActivityRegistry.cs

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Refactor tests to use `Tenant.AgnosticTenantId` instead of `null` for tenant-agnostic descriptors.

---------

Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
Co-authored-by: Sipke Schoorstra <sipkeschoorstra@outlook.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Enhance logging in recurring tasks: add error handling and logger support to prevent crashes in scheduled timers.

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <198982749+Copilot@users.noreply.github.com>
Co-authored-by: sfmskywalker <938393+sfmskywalker@users.noreply.github.com>
2026-02-02 10:59:01 +01:00
Sipke Schoorstra b09a564812
Fix Multitenancy Support and Normalize Tenant ID Handling (#7217)
* Enable multitenancy support and normalize tenant ID handling.

- Activate multitenancy in `Program.cs`.
- Introduce `NormalizeTenantId` method for consistent tenant ID usage.
- Update tenant-related classes and features to support normalization logic.

* Add ADR for adopting empty string as the default tenant ID

- Standardized the tenant ID for the default tenant to use an empty string (`""`) instead of `null`.
- Documented the rationale and migration considerations in ADR 0007.
- Updated ADR table of contents and graph for new entry.

* Apply suggestion from @sfmskywalker

* Update doc/adr/graph.dot

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Normalize spacing and improve readability in `Program.cs`. Fix multitenancy condition formatting.

* Fix ADR numbering and update TOC

* Add ADRs for flowchart execution model, tenant deletion event, merge modes, and default tenant ID

- Introduced ADR 0005: Token-centric flowchart execution model for improved loop and join handling.
- Added ADR 0006: Tenant Deleted event for distinct handling of tenant removal.
- Documented ADR 0007: Explicit merge modes for flowchart joins, improving reliability and configurability.
- Included ADR 0008: Standardization of empty string as the default tenant ID for consistency and clarity.

* Add unit tests for tenant ID normalization and multitenancy pipeline invoker

- Added comprehensive unit tests for tenant ID normalization to ensure consistent handling of null, empty, and valid IDs.
- Introduced tests for the multitenancy pipeline invoker covering various tenant resolution scenarios.
- Updated solution to include new unit testing projects for `Elsa.Tenants` and `Elsa.Common`.

* Update unit tests for `ActivityConstructionResult`

- Refactor test parameterization to verify `HasExceptions` property more explicitly.
- Simplify exception creation logic in helper methods.
- Improve test assertions by combining act and assert phases where applicable.

* Enable configuration-based multitenancy with tenant-specific settings

- Introduced a configuration-based tenant provider to streamline tenant initialization and customization.
- Added tenant ID handling filters to ensure tenant ID is applied and filtered automatically.
- Deprecated the `CommonPersistenceFeature` in favor of modular persistence feature extension.

* Update database indexes to include `TenantId` for multitenancy support

- Added `TenantId` to unique constraints on `Triggers` table across all EFCore providers.
- Adjusted index names to reflect the updated constraints.
- Updated trigger configuration to ensure uniqueness includes `TenantId`.

* Add tenant filtering to `DefaultWorkflowDefinitionStorePopulator`

- Introduced `ITenantAccessor` to support tenant-specific filtering of workflow definitions.
- Updated logic to skip workflows not matching the current tenant.

* Update doc/adr/toc.md

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Remove `CommonPersistenceFeature` as it has been deprecated

* Add tenant-specific filtering to workflow import logic in `DefaultWorkflowDefinitionStorePopulator`

* Replace hardcoded tenant ID with `Tenant.DefaultTenantId` in integration tests

* Update database indexes and migration logic to support `TenantId` for multitenancy

- Added `TenantId` to unique constraints on the `Triggers` table and updated index names.
- Included logic to drop outdated indexes without `TenantId` during migration.
- Adjusted tests to account for `TenantId` in workflow identity and indexing scenarios.

* Remove `TenantId` from workflow identity construction in concurrent trigger indexing tests

* Introduce `SelectiveMockLockProvider` for precise lock mocking in tests

- Added `SelectiveMockLockProvider` to allow targeted lock mocking without affecting unrelated background operations.
- Updated test services to use `SelectiveMockLockProvider` in place of `TestDistributedLockProvider`.
- Refactored `DistributedLockResilienceTests` to support selective mocking for deterministic and reliable assertions.

* Update Elsa.sln

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Normalize tenant ID handling in `DefaultWorkflowDefinitionStorePopulator` for consistent filtering

* Refactor `TenantResolverResult` to support explicit resolved/unresolved state handling

- Updated `TenantResolverResult` to include an explicit `_isResolved` property.
- Adjusted `ResolveTenantId()` and `IsResolved` logic for improved clarity and robustness.
- Simplified tenant resolution invocation in `TenantResolverBase`.
- Removed redundant normalization in `DefaultTenantResolverPipelineInvoker`.

* Normalize tenant ID handling in `DefaultWorkflowDefinitionStorePopulator` and `ClrWorkflowsProvider`.

* Refactor `DefaultWorkflowDefinitionStorePopulatorTests`: streamline object initializations and add tenant-specific test coverage for `PopulateStoreAsync`.

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-01-30 19:54:13 +01:00
Sipke Schoorstra b577279321
Improves tenant task management with dependencies (#7174)
* Consolidate tenant task lifecycle logic into `TenantTaskManager` and remove obsolete task event handlers (`RunStartupTasks`, `RunBackgroundTasks`, `StartRecurringTasks`). Introduce `TopologicalTaskSorter` for dependency-based task execution.

* Handles multiple tasks of the same type

Updates the topological task sorter to handle multiple tasks of the same type.

Previously, the sorter assumed a one-to-one mapping between task types and task instances, which caused issues when multiple tasks of the same type were present.
Now, it groups tasks by type and adds them to the result in the correct order.

* Add unit tests for `TopologicalTaskSorter`

Introduce `Elsa.Common.UnitTests` project with comprehensive test coverage for `TopologicalTaskSorter`, including dependency resolution, circular dependency handling, and task ordering scenarios. Update `Elsa.sln` to include the new test project.

* Update src/modules/Elsa.Common/Multitenancy/EventHandlers/TenantTaskManager.cs

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Unwrap background task continuation in `TenantTaskManager` to ensure proper task execution tracking.

* Update src/modules/Elsa.Common/Helpers/TopologicalTaskSorter.cs

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Refactor `TryGet` method to prioritize memory register lookup and update `Output` constructor to use `MemoryBlockReference`.

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-12-29 20:49:24 +01:00
Sipke Schoorstra 133e8dfc63
Introduce TenantDeleted event to handle tenant cleanup (#6843)
* Introduce `TenantDeleted` event to handle tenant cleanup

Added a `TenantDeleted` event to differentiate between tenant deactivation and deletion. Updated event handlers and services to support unregistering resources only during tenant deletion, ensuring clearer separation of responsibilities. Included ADR documentation for the new event.

* Remove unnecessary blank line in tenant deactivation logic
2025-08-06 08:43:39 +02:00
Sipke Schoorstra 2eec542f50
Merge remote-tracking branch 'origin/rc/3.4.0' 2025-04-16 08:41:04 +02:00
Sipke Schoorstra e2288d0b34
Fix infinitely waiting Alterations Workflow (#6561)
* Refactor TenantId string handling and nullability checks.

Moved `StringExtensions` to a common module for reuse. Updated tenant-related logic to utilize null-safe string extensions, enhancing consistency and simplifying nullability handling across the codebase.

* Set StrictMode to false by default in ObjectConverter

Modified the default value of StrictMode to `false` to enable the original flexible behavior. Developers can opt into strict mode by explicitly setting it to `true`. This change aims to enhance backward compatibility and minimize unexpected strict conversions.

* Add tenant ID retrieval to ElsaDbContextBase constructor

Retrieve the current tenant ID if available using ITenantAccessor and assign it to the TenantId property. This ensures proper handling of multi-tenancy scenarios in the database context initialization.

* Disable DbContext pooling, manual OTEL instrumentation, and strict mode.

DbContext pooling is turned off to prevent potential issues with shared context instances. Manual OpenTelemetry instrumentation is disabled to rely on automatic instrumentation instead. Strict mode is also disabled to allow more flexibility in object conversion.

* Fix infinitely waiting Alterations Workflow

Replaced workflow dispatch logic with BookmarkQueue and StimulusHasher for triggering workflows. This fixes the issue where the Alterations workflow would signal completion while a later step awaits a completion bookmark. The Bookmark Queue now handles this.
2025-04-05 10:37:05 +02:00
Sipke Schoorstra 51dff1a061
Support for Writing Custom Trigger Activities Using Existing Trigger Infrastructure (#6527)
* Refactor bookmark naming to use "Name" instead of "ActivityTypeName".

Replaces usages of "ActivityTypeName" with "Name" across relevant classes, filters, and database mappings for clarity and consistency. Maintains backward compatibility where necessary and updates corresponding indices, filters, and methods for proper functionality.

* Add migration for V3_5 with schema updates for EF Core

This migration modifies the `Triggers` and `Bookmarks` tables by adding the `Name` column, updating its nullable state, and creating corresponding indexes. Changes apply to both SQLite and MySQL contexts, ensuring compatibility across databases.

* Refactor bookmark filtering and tenant events handling.

Unified the bookmark filtering logic with overloads that accept multiple names, improving flexibility and reusability. Simplified object initializations in tenant events to enhance code readability and reduce verbosity. Added task continuation logic to background task execution for better task sequencing and error handling.

* Refactor bookmark creation to use target-typed `new` expressions.

Replaces explicit `CreateBookmarkArgs` instantiations with concise target-typed `new` expressions for improved readability and reduced redundancy. This does not alter functionality but simplifies the code structure.

* Make callback parameters optional and adjust OTEL settings

Updated methods to allow optional callbacks for improved flexibility. Refactored Delay activity to reuse helper methods. Adjusted OTEL instrumentation settings to enable console exporter and disable manual instrumentation.

* Refactor stimulus handling and streamline event workflows.

Introduces `WaitForEvent` and `GetEventInput` extensions to improve ActivityExecutionContext usability. Replaces generic filter methods with targeted single-name filtering, simplifying scheduling logic. Updates stimulus names for consistency and deprecates obsolete methods to enhance clarity and maintainability.

* Refactor `Event` activity handling and event stimulus logic.

Replaces inline event stimulus creation with a new `GetEventStimulus` helper method for cleaner code. Simplifies event execution handling by removing redundant logic in `ExecuteAsync`. Extends `WaitForEvent` to handle workflow triggers more efficiently.

* Refactor and enhance Timer and Delay execution logic

Introduced `TimerBase` for shared timer functionality and refactored `Timer` to extend it. Improved method names for clarity, replacing `ResumeIn`/`ResumeAt` with `DelayFor`/`DelayUntil`. Enhanced flexibility in bookmark handling and activity execution context extensions.

* Add custom activities and refactor HTTP stimulus handling

Introduce new custom activities (CustomDelay, CustomEvent, CustomHttpEndpoint, CustomTimer) to enhance workflow functionality. Refactor HTTP stimulus handling by replacing activity type names with a new centralized HttpStimulusNames constant, improving consistency and maintainability. Additionally, streamline HTTP endpoint logic with new helper extensions and simplify related services to reduce redundancy.

* Add HttpEndpointBase abstraction to simplify HTTP endpoints

Introduce a new `HttpEndpointBase` class to centralize common logic for HTTP endpoint activities. Refactored `CustomHttpEndpoint` to inherit from this new base class, reducing redundancy and improving maintainability.

* Refactor events framework with base class for event activities

Introduce `EventBase` to streamline implementations of event-driven activities. Updated `CustomEvent` to inherit from `EventBase`, reducing duplicate logic and improving maintainability. Removed unnecessary dependencies in `CustomTimer`.

* Refactor bookmark creation to use target-typed `new` expressions.

Replaces explicit `CreateBookmarkArgs` instantiations with concise target-typed `new` expressions for improved readability and reduced redundancy. This does not alter functionality but simplifies the code structure.

Make callback parameters optional and adjust OTEL settings

Updated methods to allow optional callbacks for improved flexibility. Refactored Delay activity to reuse helper methods. Adjusted OTEL instrumentation settings to enable console exporter and disable manual instrumentation.

Refactor stimulus handling and streamline event workflows.

Introduces `WaitForEvent` and `GetEventInput` extensions to improve ActivityExecutionContext usability. Replaces generic filter methods with targeted single-name filtering, simplifying scheduling logic. Updates stimulus names for consistency and deprecates obsolete methods to enhance clarity and maintainability.

Refactor `Event` activity handling and event stimulus logic.

Replaces inline event stimulus creation with a new `GetEventStimulus` helper method for cleaner code. Simplifies event execution handling by removing redundant logic in `ExecuteAsync`. Extends `WaitForEvent` to handle workflow triggers more efficiently.

Refactor and enhance Timer and Delay execution logic

Introduced `TimerBase` for shared timer functionality and refactored `Timer` to extend it. Improved method names for clarity, replacing `ResumeIn`/`ResumeAt` with `DelayFor`/`DelayUntil`. Enhanced flexibility in bookmark handling and activity execution context extensions.

Add custom activities and refactor HTTP stimulus handling

Introduce new custom activities (CustomDelay, CustomEvent, CustomHttpEndpoint, CustomTimer) to enhance workflow functionality. Refactor HTTP stimulus handling by replacing activity type names with a new centralized HttpStimulusNames constant, improving consistency and maintainability. Additionally, streamline HTTP endpoint logic with new helper extensions and simplify related services to reduce redundancy.

Add HttpEndpointBase abstraction to simplify HTTP endpoints

Introduce a new `HttpEndpointBase` class to centralize common logic for HTTP endpoint activities. Refactored `CustomHttpEndpoint` to inherit from this new base class, reducing redundancy and improving maintainability.

Refactor events framework with base class for event activities

Introduce `EventBase` to streamline implementations of event-driven activities. Updated `CustomEvent` to inherit from `EventBase`, reducing duplicate logic and improving maintainability. Removed unnecessary dependencies in `CustomTimer`.

* Move HttpEndpointOptions model to its own file

The HttpEndpointOptions class was moved from an extension file to its own dedicated file for better organization and modularity. This model defines HTTP endpoint properties such as path, methods, authorization, policies, request timeout, and size limit. The change improves code clarity and structure.

* Fix unnecessary whitespace in Timer.cs

Removed an extra whitespace line in the Timer.cs file to maintain code formatting consistency. No functional changes were made to the code.

* Remove extraneous whitespace in IStimulusSender.cs file

Eliminate unnecessary blank line in the IStimulusSender interface for improved code cleanliness. This change enhances readability and aligns with coding standards.
2025-03-21 23:16:56 +01:00
Sipke Schoorstra 7a758f8850 Fix multitenancy support in Hangfire services and jobs
Refactored Hangfire-related services and jobs to include tenant context management using ITenantAccessor and ITenantFinder. Updated job constructors and method signatures to support tenant-specific execution. Improved exception-throwing syntax for better readability in BackgroundActivityInvoker.
2025-02-06 16:39:38 +01:00
Sipke Schoorstra 6c0f7a48e7 Refactor recurring tasks scheduling logic.
Removed `ConfigureRecurringTasksScheduleStartupTask` and moved its functionality into `RecurringTaskScheduleManager`. Simplified recurring task configuration and streamlined dependencies, improving maintainability. Added retention policies to `Elsa.Server.Web`.
2025-01-02 11:11:42 +01:00
Sipke Schoorstra a9c8ad64f1 Add concurrency locks to DefaultTenantService operations
Introduced SemaphoreSlim for initialization and refresh methods to ensure thread safety in DefaultTenantService. Improved tenant unregistration to handle scope cleanup only when mappings exist. These changes enhance reliability and prevent race conditions during tenant operations.
2024-12-12 11:20:45 +01:00
Sipke Schoorstra 692937d7d7
Enhance Multitenancy with Runtime Tenant Management and Task Handling (#6173)
* Work in progress: Add DefaultTenantService for tenant management

Introduce `DefaultTenantService` and its corresponding interface `ITenantService` to manage tenant operations such as finding, getting, and listing tenants. Update `MultitenantBackgroundService` to utilize `DefaultTenantService` for handling tenant lifecycle events. This enhancement standardizes tenant operations and improves the maintainability of the multitenancy feature.

* WIP

* Add multitenancy event handlers and task interfaces

Implemented new interfaces IBackgroundTaskStarter and ITaskExecutor to manage task lifecycle events efficiently. Introduced new classes such as RunBackgroundTasks, RunStartupTasks, and StartRecurringTasks for handling tenant activation and deactivation events. Modified TaskExecutor to implement these interfaces and adjusted tenant registration logic to invoke these new handlers.

* Refactor multitenancy and task management services.

Remove background and recurring task runners, and integrate tenant activation and deactivation into the multitenancy feature. Enable multitenancy in the server application and create a new service for tenant activation and deactivation. This refactor simplifies the management of tenant-specific tasks and enhances the modularity of the platform.

* Refactor background service to use startup tasks

Replaced hosted service implementation with startup tasks for executing multi-tenant tasks and EF Core migrations. Introduced `PriorityAttribute` to manage task execution order, ensuring migrations run before other services that require database access. This simplifies tenant activation with an ordered task execution and removes redundant classes.

* Refactor MultitenancyFeature service registrations

Reorganized service registrations for better clarity and maintainability. Changed the registration of some services to use factory delegates for retrieving existing services to ensure correct dependencies. This refactor improves the flexibility of the tenant lifecycle event handling.

* Update V3_3 migration files

* Add tenant management endpoints and enhance tenant handling

Implemented tenant management endpoints including Add, Get, List, and Update. Enhanced tenant handling by introducing configuration and store-based providers, and improved error logging for tenant updates. Adjusted various internal functionalities to better support multitenancy features through different persistence providers.

* Implement tenant deletion endpoint and refactor migration setup.

Introduce a new API endpoint to handle tenant deletions while providing appropriate responses based on successful or unsuccessful attempts. Refactor migration handling by replacing startup tasks with hosted services across various modules to streamline the migration execution process.

* Add and integrate ConfigurationJsonConverter

Introduce a `ConfigurationJsonConverter` to handle JSON serialization and deserialization of `IConfiguration` objects. This change centralizes configuration serialization logic, leading to cleaner and more maintainable code. Updated various parts of the codebase to use the new serialization utility, ensuring a consistent approach throughout the application.

* Refactor JSON conversion and update tenant endpoint.

Removed unused workflow references and streamlined JSON handling in `ConfigurationJsonConverter`. Simplified tenant ID handling by removing `IIdentityGenerator` and setting a default value for `UpdatedTenant.Id`.

* Add logging for cancelled recurring tasks

Integrated ILogger to log information when a recurring task is canceled due to an OperationCanceledException. This change enhances troubleshooting by providing clearer insights into task cancellations and their underlying reasons, improving maintainability and observability of the task execution process.

* Disable multitenancy support and adjust default Tenant ID.

Multitenancy is now disabled by setting 'useMultitenancy' to false in the configuration. Additionally, the default Tenant's ID has been changed from null to an empty string to prevent potential null reference issues.

* Remove MultitenantHostedService abstraction file

The MultitenantHostedService.cs file was removed as it is no longer necessary. Its responsibilities have likely been refactored or integrated into another service, indicating a simplification or restructuring of the multitenancy handling in the codebase.

* Rename PriorityAttribute to OrderAttribute for clarity.

This change improves the clarity of the code by renaming PriorityAttribute to OrderAttribute, reflecting its actual purpose. All occurrences of the attribute in the codebase have been updated accordingly to maintain consistency. This makes the intent of the code more understandable for future maintenance and development.

* Fix message key retrieval in ProduceMessage activity

Update the ProduceMessage activity to use GetOrDefault for retrieving the message key. This change ensures that a null key is used if no explicit key is provided or if the key is empty or whitespace, preventing potential errors during message production.

* Refactor multitenancy and scheduling services.

Removed DefaultTenantContextInitializer interface and class, refactored tenant activation/deactivation to use try-catch logging, and updated tenant context handling to use IDisposable for context push. New activities and workflows added in Elsa.Server.Web, and scheduling services enhanced to schedule jobs with explicit job keys and groups. Also, adjusted configurations to enable multitenancy, providing improved maintainability and flexibility.

* Remove Example1 activities and disable multitenancy

Deleted Example1Activity, Example1Workflow, and FirstActivity classes to clean up unused code and simplify the codebase. Disabled multitenancy by setting useMultitenancy to false, likely to streamline configuration and resource utilization.

* Fix and normalize URL path concatenation.

Ensure that the base URLs in both base path providers consistently end with a forward slash. This normalization prevents potential issues with endpoint routing and path concatenation, improving overall URL construction robustness.
2024-12-03 15:07:55 +01:00
Sipke Schoorstra 09a5c79211
Fix Tenant ID propagation for MassTransit (#6144)
* Refactor tenant middleware and consumer implementations

Update `TenantConsumeMiddleware` to utilize `ITenantFinder` and `ITenantScopeFactory` for tenant context management. Refactor `WorkflowMessageConsumer` to use constructor parameters directly for improved dependency injection. Additionally, replace private scope variable in `TenantScope` with a public property to enhance code clarity.

* Refactor to replace IWorkflowDispatcher with IStimulusSender

Updated WorkflowMessageConsumer to utilize IStimulusSender for handling message-triggered activities. Simplified dependencies and method calls to align with the new interface, improving the code's clarity and maintainability.
2024-11-24 23:28:38 +01:00
Sipke Schoorstra 939fb95a97
Add multitenancy support for background tasks (#6059)
* Remove initial migrations

Deleted obsolete initial migration files from multiple databases: MySQL, SQL Server, SQLite, and PostgreSQL. This cleanup helps maintain a streamlined and updated migration history.

* Add Document base class and create tenant-specific indices

Introduced a new abstract `Document` base class to unify common properties. Implemented tenant-specific unique indices across multiple collections by including `TenantId` alongside `Id` to ensure uniqueness within tenant scopes.

* Remove outdated migration files

Deleted various migration files under MySql, PostgreSql, Sqlite, and SqlServer directories. This cleanup removes unnecessary schema definitions and helps to streamline the codebase.

* Refactor workflow identity assignment logic

Streamline workflow identity handling to ensure consistent assignment of Id, DefinitionId, and TenantId values. This change integrates tenant prefix and version suffix cleanly, enhancing clarity and maintainability.

* Enable multitenancy support

Added configuration for a new tenant (tenant-1) in appsettings.json and enabled multitenancy feature in Program.cs. This change allows the application to support multiple tenants, with specific configurations for each.

* Refactor route table update to run as startup task

Replaced `UpdateRouteTableHostedService` with `UpdateRouteTableStartupTask` to ensure route table updates are executed during application startup instead of as a hosted service. Updated configuration in `HttpFeature` and adjusted trigger validation logic in `ValidateWorkflowRequestHandler`.

* Add recurring task scheduling and single-node task support.

Introduce `IntervalExpressionType`, recurring task scheduling classes, and `SingleNodeTaskAttribute`. Update `RecurringTasksRunner` to handle schedules and add single-node task logic to `StartupTasksRunner`. Ensure proper namespace changes and configure sample recurring tasks.

* Refactor recurring tasks scheduling system

Replaced existing scheduling classes with a more modular and granular approach. Introduced new classes and interfaces like `ISchedule`, `CronSchedule`, `IntervalSchedule`, and `RecurringTaskScheduleManager`. Updated related methods and code to comply with the new design.

* Refactor background task management

Removed `ExpiredSecretsHostedService` and refactored it into a recurring task. Introduced `TaskExecutor` for shared task execution logic. Updated and renamed feature classes to better represent their purpose, improving task scheduling and execution management.

* Add BackgroundTask abstract class to Elsa.Common module

This new abstract class implements the IBackgroundTask interface with default methods for executing, starting, and stopping tasks asynchronously. It provides a basic framework for background task management in the Elsa.Common module.

* Switch to CreateAsyncScope in DefaultTenantScopeFactory

Updated the CreateScope method to use CreateAsyncScope instead of CreateScope. This change improves asynchronous handling of service scopes within the DefaultTenantScopeFactory class.

* Add tenant handling and move StartWorkers background task

Introduce ITenantAccessor in Worker class for multitenancy support. Rename and relocate StartWorkers service to BackgroundTask, ensuring smoother workflow initialization. Also, update the configuration to support Azure Service Bus connection string.

* Add tenant support and refactor ProtoActor client

Integrated ITenantAccessor in ProtoActorWorkflowClient class to handle multi-tenancy. Refactored methods in the client to support custom headers and added async disposable pattern in various services for proper resource management. Additionally, enabled Azure Service Bus and updated related documentation.

* Add support for custom headers in ProtoActor grain methods

Introduced a T4 template to generate grain methods with custom headers, enabling the use of tenant ID in requests. Updated `ProtoActorWorkflowClient` to employ these methods, removing redundant code and directly utilizing the client for various workflow operations.

* Add tenant middleware to MassTransit configurations

Introduced multitenancy middleware for MassTransit message handling. Added new message type `OrderReceived` and updated RabbitMQ setup in Elsa Server. Applied middleware to configure tenant data on send, publish, and consume operations.

* Add new product workflow and streamline ID handling

Introduced a new `RequestResponseWorkflow` for handling product requests. Simplified ID handling in `WorkflowBuilder` and `ClrWorkflowsProvider` by defaulting to empty strings and adding a version prefix. Enhanced `HttpWorkflowsMiddleware` to correctly parse full request paths.

* Remove redundant files and update configuration

Deleted unused files `Product.cs` and `RequestResponseWorkflow.cs` to clean up the codebase. Updated `Program.cs` configuration: switched MassTransitBroker to Memory and disabled multitenancy.

* Remove MultitenantRecurringTaskService and update AzureServiceBus

Removed `MultitenantRecurringTaskService` and adjusted related code for Azure Service Bus to work without it. This includes removal of tenant accessor dependency from `Worker` and cleanup of service configuration flags in `Program.cs`.

* Increase signal wait timeout to 10000 milliseconds.

Extended the default timeout for signal awaiting methods from 8000 to 10000 milliseconds. This change ensures more flexible and resilient waiting periods, reducing timeout occurrences in scenarios with longer processing times.

* Refactor scheduling service to be a background task

Renamed `CreateSchedulesHostedService` to `CreateSchedulesBackgroundTask` and refactored it to inherit from `BackgroundTask` instead of `BackgroundService`. Simplified the constructor by injecting the required dependencies directly, eliminating the need for a scoped factory.

* Refactor workflow version suffix formatting

Changed the version suffix format from `:v{version}` to `v{version}` and adjusted the ID concatenation accordingly. This improves consistency and readability of workflow IDs.

* Enable multitenancy support in Quartz scheduler

Added `TenantJobListener` to inject tenant context into jobs. Modified `QuartzWorkflowScheduler` to incorporate tenant IDs into job data maps and adjusted the configuration to acknowledge multitenancy settings.

* Remove ConfigureSchedulerHostedService and TenantJobListener

Consolidated tenant resolution logic into JobExecutionExtensions class. Updated ResumeWorkflowJob and RunWorkflowJob to use the new extension method for tenant retrieval. This simplifies the QuartzSchedulerFeature setup by removing the hosted service configuration.

* Refactor HTTP feature and update route table task

Move 'UpdateRouteTableStartupTask' from 'HostedServices' to 'Tasks' and update dependency injection configurations accordingly. Simplify 'DefaultRouteTableUpdater' by removing unnecessary options and tenant-agnostic settings from filters.

* Disable multitenancy in Program.cs

The useMultitenancy flag has been changed from true to false. This update affects the Elsa.Server.Web application configuration.

* Simplify variable usage in HttpWorkflowsMiddleware

Replaced 'fullPath' variable with 'path' to streamline code. This change enhances readability by reducing redundancy and ensures consistency in variable naming throughout the method.

* Enable multitenancy and refactor tenant handling logic

Enable multitenancy in the application and refactor tenant handling logic to use ITenantFinder and ITenantContextInitializer interfaces. Added header constants, updated middleware to use these interfaces, and moved extension methods to the appropriate namespace.

* Add input validation to user registration form

Implemented checks to ensure all required fields are filled and that input data adheres to format requirements. This change reduces errors and enhances form reliability.

* Remove unused import from TenantPrefixHttpEndpointRoutesProvider

This change cleans up the code by removing an unnecessary import statement. It improves code readability and reduces clutter, making future maintenance easier. The functionality remains unchanged.

* Rename filter scope to "tenantPublish" in Probe method

Updated the Probe method in TenantPublishMiddleware.cs to use "tenantPublish" instead of "tenantSend" for better clarity. Ensures consistency with the method's context and aligns with naming conventions.

* Refactor: Remove extraneous whitespace

Eliminate unnecessary whitespace in ProtoActorWorkflowClient.cs for cleaner code. This change helps maintain consistent formatting and improves readability.

* Refactor DefaultRegistriesPopulator for cleaner initialization

Converted constructor to use read-only fields directly, removing unnecessary instance variables. This change simplifies the code by reducing redundancy and making the constructor cleaner.
2024-10-28 19:38:24 +01:00
Sipke Schoorstra bdaae64e72 Add base path providers and standardize tenant accessor property
Introduce DefaultHttpEndpointBasePathProvider and TenantPrefixHttpEndpointBasePathProvider to manage HTTP endpoint base paths. Update ITenantAccessor property from CurrentTenant to Tenant for consistency across the codebase.
2024-10-17 19:47:55 +02:00
Sipke Schoorstra 99ab5e4ba1 Refactor StopAsync method to return completed task
Modified the StopAsync method to directly return a completed task instead of running through tenants asynchronously. This change simplifies the shutdown process and removes redundant tenant processing during service stop.
2024-10-15 23:47:17 +02:00
Sipke Schoorstra 225ad49ea8
Improved multitenancy support for HTTP workflows with per-tenant DbContext (#6032)
* Refactor: Update namespaces and add TenantExtensions

Updated namespaces throughout the project to improve clarity and consistency by moving from 'Common' to appropriate modules. Added TenantExtensions class to simplify fetching connection strings for tenants.

* Implement multitenant DB connection strings

Redesign tenant-specific classes to support multitenancy more effectively. Introduce `MultitenantBackgroundService` and `MultitenantHostedService` for handling tasks per tenant.

* Refactor constructors and remove redundant code

Simplified the constructor parameters for `MultitenantBackgroundService` and `List` class. Removed the unused parameter in `MultitenantBackgroundService` and redundant folder inclusion in the project file. Updated the method calls to use direct parameters in `List` class.

* Remove unused import in ActivityDescriptors Endpoint

The Elsa.Common.Multitenancy import was removed as it is unused in the List/Endpoint.cs file. Removing unused imports helps to improve code readability and maintainability. This change does not affect functionality.
2024-10-15 23:16:50 +02:00
Sipke Schoorstra 7e7a899bbf
Implement multitenant HTTP routing (#6031)
* Add tenant awareness to bookmark handling and route resolution

Added tenant ID support across various components, including bookmark updates, route resolution, and middleware processing. This ensures that bookmark and route operations can now appropriately handle tenant-specific data, improving the system's multitenancy capabilities.

* Add Multitenant HTTP Routing feature to Tenants module

Introduced a new MultitenantHttpRoutingFeature class to the Elsa.Tenants.AspNetCore module, enhancing the tenant resolution capabilities. Moved RoutePrefixTenantResolver from Elsa.Http to Elsa.Tenants.AspNetCore and updated relevant project references and namespaces accordingly. This refactor improves modularity and separation of concerns between HTTP and tenancy features.

* Refactor route handling and tenant configuration

Removed redundant `RouteTableExtensions` and replaced with new route providers and updaters, enhancing flexibility and modularity. Introduced tenant-specific HTTP endpoint configurations for better customization and configuration management.

* Rename HttpEndpointBookmarkStimulus to HttpEndpointBookmarkPayload

Refactor various classes and methods to reflect the renaming from `HttpEndpointBookmarkStimulus` to `HttpEndpointBookmarkPayload`. Add and configure new extension methods for tenant route handling, update the route provider to support multi-tenancy, and adjust the tenants provider to bind configuration properly.

* Add HeaderTenantResolver and refactor Http namespace.

Introduce HeaderTenantResolver to resolve tenants via HTTP headers. Refactor multiple classes and interfaces to move from the Elsa.Http.Models namespace directly into Elsa.Http for clarity and consistency.

* Add Host-based tenant resolution

Implemented a HostTenantResolver to resolve tenants based on the request's host and updated tenant configurations with host information. Modified the tenant resolver pipeline and added the new host resolver to the service registrations.

* Add tenant-aware caching and accessor support

Enhanced caching by incorporating tenant identifiers into cache keys for more granular cache management. Introduced ITenantAccessor dependencies in various services to retrieve the current tenant information. This ensures that cache entries are correctly isolated per tenant.

* Reorder tenant resolvers for pipeline setup.

Reordered the tenant resolvers in the pipeline to prioritize HostTenantResolver before RoutePrefixTenantResolver. This ensures that tenant resolution is correctly aligned with host-based resolving before checking the route prefix.

* Remove unused imports

This commit eliminates redundant `using` directives across multiple files to streamline the codebase. This cleanup helps improve code readability and maintainability by removing unnecessary dependencies.
2024-10-14 21:27:11 +02:00
Sipke Schoorstra a5cc3fc9e8
Refactor Tenant Resolution to Use Async Local Storage for Operation-wide Access (#6022)
* Remove obsolete tenant-related classes and add ASP.NET Core middleware

Refactored tenant resolution by removing obsolete interfaces and classes, such as `IAmbientTenantAccessor` and `ITenantResolutionStrategy`. Introduced new ASP.NET Core middleware for tenant resolution, encapsulated in the new `Elsa.Tenants.AspNetCore` project. Updated related usage in various parts of the application to align with these changes.

* Remove HttpContextTenantResolver.

Removed HttpContextTenantResolver from the multitenancy pipeline and related service registrations. This simplifies the tenant resolution by relying on remaining resolvers like ClaimsTenantResolver and RoutePrefixTenantResolver.

* Add Elsa solution definition file

This commit introduces the main solution file, Elsa.slnx, defining the folder structure, projects, and configuration for the Elsa repository. This includes folders for Docker, documentation, pipelines, samples, scripts, source code, and tests.

* Refactor DefaultAccessTokenIssuer for clarity and efficiency

Refactored the DefaultAccessTokenIssuer class by simplifying its constructor and utilizing scoped variables for token options. Improved token creation logic by adding a dedicated method to configure token options, enhancing code readability and maintainability.

* Remove Elsa.slnx solution file

No dotnet build support yet.

* Refactor tenant resolver service registrations

Updated the service registrations to use interfaces for DefaultTenantResolver and DefaultTenantResolverPipelineInvoker. This improves the code's flexibility, making it easier to replace or extend these implementations in the future.

* Add multitenancy support and tenant scope management

Introduced ITenantScopeFactory and related implementations for tenant scope management across the application. Enhanced the HTTP workflows middleware to handle tenants and updated relevant configurations and extension methods to support tenant resolution.

* Remove unnecessary folder inclusion

The <Folder> tag for "Modules\Modules\" was redundant and has been removed to clean up the project file. This change will not affect the existing functionality or project structure.

* Rename Create to CreateScope and improve authorization.

Updated the method name from Create to CreateScope for better clarity in the TenantScopeFactory. Fixed a logical error in the authorization process, ensuring proper status code setting for unauthorized requests, and refactored token expiration calculation for clarity.

* Add tenant agnostic filters and remove tenant setup

This commit introduces tenant agnostic filters in AutoUpdateTests to ensure workflows can trigger regardless of tenant. Additionally, it removes tenant configuration from WorkflowServer setup as it is no longer required for the current tests.
2024-10-12 12:08:09 +02:00
Sipke Schoorstra e1d0cc2e27 Add DefaultTenantAccessor and ITenantAccessor interface
Introduce ITenantAccessor interface and its implementation, DefaultTenantAccessor, for tenant management in the multitenancy module. This setup will allow future enhancements for tenant-specific functionality.
2024-10-06 14:49:02 +02:00
Sipke Schoorstra dd812645c7 Refactor imports to reduce use of Elsa.Common.Contracts
Consolidate imports by replacing Elsa.Common.Contracts with Elsa.Common and Elsa.Common.Multitenancy. This update streamlines import statements across various modules, improving code readability and maintainability.
2024-10-05 18:30:45 +02:00