elsa-core/test/unit/Elsa.Common.UnitTests/ExceptionExtensionsTests.cs

60 lines
2.3 KiB
C#
Raw Permalink Normal View History

Graceful shutdown for the workflow runtime (drain, pause, recover) (#7424) * feat(workflows-runtime): add quiescence machinery foundation for graceful shutdown Introduces the container-scoped quiescence signal, ingress-source contract, burst registry, and the Interrupted workflow sub-status — the foundational primitives the drain orchestrator and admin endpoints will build on. No behaviour change yet: workflows continue to run and shut down exactly as before. The new types are registered but no host-stop or pause path drives them. Highlights: * IQuiescenceSignal — composable Drain + AdministrativePause flags; forward-only drain, reversible pause, idempotent transitions, optional persistence via IKeyValueStore. * IIngressSource + IForceStoppable — uniform contract for components that inject external events (HTTP, schedulers, message consumers, internal workers, third-party modules); IIngressSourceRegistry collects and surfaces their states. * IBurstRegistry — atomic counter for in-flight workflow execution bursts, with per-burst ingress attribution and FR-018 inconsistency detection (a source claiming Paused but starting bursts is flipped to PauseFailed). * WorkflowSubStatus.Interrupted — new value distinct from Suspended, Cancelled, Faulted; semantics: "last burst force-cancelled by graceful drain; resumable on next runtime generation". Mirrored on the API client enum. * GracefulShutdownOptions — drain deadline, per-source pause timeout, stimulus-queue back-pressure policy, pause-persistence policy. Configurable via UseWorkflowRuntime(...).ConfigureGracefulShutdown(...). * PermissionNames.ManageWorkflowRuntime — single permission for the forthcoming admin pause/resume/status/force endpoints. Implements 31 of 77 tasks for the graceful-shutdown feature (specs/002-graceful-shutdown). Subsequent commits add the drain orchestrator (US1 / MVP), Interrupted recovery scan (US3), admin endpoints (US2), and first-party ingress adapters. Tests: 25 new xUnit unit tests; 100/100 runtime unit tests pass; all existing tests continue to pass on net8.0/net9.0/net10.0. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(workflows-runtime): add drain orchestrator + host-stop integration (US1, MVP) When the host receives a stop signal (SIGTERM, Ctrl+C, orchestrator rollout), the runtime now drains gracefully: ingress sources are paused in parallel, in-flight workflow bursts run to their next natural persistence boundary within a configurable deadline, and any burst that breaches the deadline is force-cancelled and persisted with the Interrupted sub-status plus a forensic WorkflowInterrupted log entry. This is the MVP — without the activation-time recovery scan (PR 3) the existing timeout-based RestartInterruptedWorkflowsTask still picks up Interrupted instances, just on its periodic cadence. No regression in that recovery path (SC-008). Highlights: * IDrainOrchestrator + DrainOrchestrator — protocol per the contract: BeginDrainAsync → parallel ingress pause with per-source timeouts + IForceStoppable escalation → poll BurstRegistry.ActiveCount until zero or deadline → on breach iterate live handles, cancel, persist Interrupted, write log entry. All exceptions are captured into the returned DrainOutcome; only second-invocation throws. * Deadline clamping: effective deadline is min(GracefulShutdownOptions. DrainDeadline, HostOptions.ShutdownTimeout - 500ms safety epsilon), so the runtime never outlives its host process. * DrainOrchestratorHostedService — IHostedService.StopAsync wakes the orchestrator on host stop. Registered AFTER the heartbeat (Elsa.Hosting.Management) so reverse-order shutdown keeps the heartbeat alive throughout drain. Prevents sibling-node crash recovery from false-positive-recovering instances we are gracefully handling here (FR-029). * BurstTrackingMiddleware — workflow-execution-pipeline middleware that registers a BurstHandle for the lifetime of every burst. All nine IWorkflowRunner.RunAsync overloads ultimately funnel into pipeline.ExecuteAsync(context), so this single middleware covers the three "burst choke points" the spec references without nine separate decorators. Added to UseDefaultPipeline(). * Ingress attribution: optional IngressSourceName property on DispatchStimulusRequest, DispatchWorkflowDefinitionRequest, and DispatchWorkflowInstanceRequest. Adapters set it; the middleware reads it via WorkflowExecutionContext.TransientProperties (helpers in IngressAttributionExtensions). The BurstRegistry uses the name to detect the FR-018 invariant violation — a source that reports Paused but starts a burst is flipped to PauseFailed. * InterruptedLogExtensions — the LogWorkflowInterruptedAsync helper that the orchestrator calls when persisting the forensic record. Tests: 11 new unit tests (DrainOrchestrator parallel-pause + wait-for-bursts + idempotency + persistence-failure path); 5 new integration tests (full DI graph resolves, burst-tracking middleware registers handles end-to-end, no-op drain returns CompletedWithinDeadline). 100/100 runtime unit tests pass; all existing tests continue to pass on net8.0/net9.0/net10.0. Implements 14 of 77 tasks (T032–T045). Subsequent commits add Interrupted recovery scan (US3), admin endpoints (US2), and first-party ingress adapters. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(admin-endpoints): add admin endpoints for workflow runtime control Introduced admin endpoints to manage workflow runtime: `/pause`, `/resume`, `/status`, and `/force` with full authentication and audit logging. Integrated idempotency checks and error handling to ensure reliable runtime control. Added corresponding integration tests for verification. * fix(workflows-runtime): propagate drain cancellation into running workflows + initialize persisted pause + fix log-comment Addresses three findings from the PR #7424 code review: 1. **HIGH — Cancellation now propagates into the running workflow.** `BurstHandle.Cancel()` previously cancelled only its own linked CTS, which the workflow runner never observes (the runner reads from `WorkflowExecutionContext.CancellationToken`, captured at context construction and not part of the linked chain). On deadline breach the orchestrator would persist `Interrupted`, but the workflow continued executing and could overwrite the sub-status with whatever terminal state it eventually reached. Fix: `BurstHandle` accepts an optional cancel callback at construction. `BurstTrackingMiddleware` wires it to `context.Cancel()` so the burst's cancellation triggers the workflow's own cancellation chain — the workflow transitions to `Cancelled` and stops scheduling new activities. The orchestrator then awaits `BurstHandle.Disposed` (with a 2 s settle timeout) before persisting `Interrupted`, ensuring the runner's terminal commit completes BEFORE the orchestrator overwrites the sub-status. Race resolved. The settle timeout is bounded so a non-cancellable activity (genuinely pathological case) does not block drain — on timeout the orchestrator logs and proceeds, accepting the runner-clobber for that one instance, which the existing timeout-based RestartInterruptedWorkflows recovery picks up afterwards. 2. **MEDIUM — Pause persistence is now actually wired.** `QuiescenceSignal.InitializePersistedStateAsync` was implemented but nothing called it on host startup. A host configured with `PausePersistence = AcrossReactivations` would write the persisted key on pause, but on subsequent activation the new `QuiescenceSignal` instance would never read it, so the runtime would resume dispatching despite the operator having paused. Fix: `InitializePauseStateStartupTask : IStartupTask` reads the policy and calls `InitializePersistedStateAsync` once per activation when the policy demands it. Registered in both `WorkflowRuntimeFeature` flavours alongside the other graceful-shutdown services. 3. **MEDIUM — Comment in `DrainOrchestrator.PersistInterruptedAsync` no longer lies.** The previous comment promised a "synthetic log entry" that the next statement (`return`) prevented from being written. The comment is now honest about what actually happens: when no instance row exists, no log entry is emitted, but the burst metadata is still captured in the drain outcome's logged warning so operators have a forensic trail. Tests: * New unit tests on `BurstRegistry` (now 9, was 6): cancel-callback is invoked, callback exceptions are swallowed (drain remains best-effort), `BurstHandle.Disposed` completes on dispose. * Full suites continue to pass: 103/103 runtime unit tests; 247/247 workflow integration tests. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): address PR review feedback + close runner-clobber race + add e2e drain test Six issues raised in PR #7424 review (one e2e gap, five comments inline): 1. **Runner-clobber race closed via ICommitStateHandler decorator.** The previous fix wired `BurstHandle.Cancel()` to `WorkflowExecutionContext.Cancel()` so the workflow's cancellation chain fires on deadline breach, but the e2e test exposed that `BurstHandle` disposed at the END of the pipeline middleware (i.e., BEFORE `WorkflowRunner` calls `commitStateHandler.CommitAsync`). The orchestrator's `await handle.Disposed` therefore returned too early, the instance row didn't yet exist, and the orchestrator's Interrupted write was either a no-op (no row) or got clobbered by the runner's subsequent Cancelled commit. Fix: `BurstAwareCommitStateHandler` decorates `ICommitStateHandler`. The middleware no longer disposes the handle in the success path — it stores the handle in `WorkflowExecutionContext.TransientProperties`, and the decorator disposes it AFTER `inner.CommitAsync` completes. The exception path in the middleware still disposes for safety. Result: the orchestrator's await-disposed sequencing now correctly lands the Interrupted write last. 2. **C1: Null-instance log entry.** `DrainOrchestrator.PersistInterruptedAsync` now writes a synthetic `WorkflowInterrupted` log entry directly when no instance row exists, populating only the fields it knows. Previously the forensic trail was lost. 3. **C2: Force endpoint cached-outcome audit.** Added `WasCached` flag to `DrainOutcome` (default false). The orchestrator sets it on the cached return path (`_previousOutcome with { WasCached = true }`). The force endpoint now skips the audit notification when the flag is true, so repeated force calls no longer emit spurious `RuntimeForceRequested` events (SC-007 idempotency restored). 4. **C3: `StateChanged` raised under lock — deadlock risk closed.** `QuiescenceSignal.BeginDrainAsync`/`PauseAsync`/`ResumeAsync` now do their transitions under the lock, capture whether a transition occurred, release the lock, and only then invoke `RaiseStateChanged`. Subscribers that synchronously call back into the signal can no longer deadlock. 5. **C4: Scheduling source name.** Renamed `scheduling.cron` → `scheduling.triggers` to honestly reflect the four trigger types the adapter covers (Cron, Timer, StartAt, Delay). The name is surfaced verbatim in admin status responses. 6. **C5: Hardcoded Retry-After.** `HttpWorkflowsMiddleware`'s 503 response now sets a reason-aware `Retry-After`: 5 s during drain (host is exiting and will be replaced shortly), 60 s during administrative pause (indefinite, so a longer back-off avoids tight retry loops). Tests: * New e2e `DeadlineBreachEndToEndTests` (2 tests): verifies that drain against a real running workflow detects the in-flight burst, force-cancels it, persists the instance as `Interrupted`, and writes a `WorkflowInterrupted` log entry — closing the test gap that hid the cancellation-propagation issue identified in the previous review pass. * Updated `OperatorForceAfterPreviousReturnsCachedOutcome` to assert value-equality + the `WasCached` flag instead of reference-equality (records use `with` for the cached return path). Full suites pass: 103/103 runtime unit, 249/249 workflow integration. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-api): move runtime admin endpoints into Elsa.Workflows.Api Per PR review feedback: rather than introducing a new sub-module (Elsa.Workflows.Runtime.Admin) for the four pause/resume/status/force endpoints, fold them into the existing Elsa.Workflows.Api project. That project already references both Elsa.Workflows.Runtime and Elsa.Api.Common (FastEndpoints) and is the established home for client-facing workflow APIs — so the admin endpoints belong there. Changes: * New folder src/modules/Elsa.Workflows.Api/Endpoints/RuntimeAdmin/ with Models.cs and Pause/Resume/Status/Force/Endpoint.cs. Namespaces moved from `Elsa.Workflows.Runtime.Admin` → `Elsa.Workflows.Api.Endpoints.RuntimeAdmin`. * Deleted src/modules/Elsa.Workflows.Runtime.Admin/ entirely and removed it from Elsa.sln. The ShellFeature marker class (WorkflowRuntimeAdminFeature) is no longer needed — the existing WorkflowsApiFeature already discovers FastEndpoints in the Workflows.Api assembly. * No consumer changes: the endpoints sit in the same routes (/admin/workflow-runtime/*) and behave identically. Note on the second architectural point ("update IShellFeature if cleaner"): the IShellFeature contract is defined in the external CShells NuGet package, not in this repo, so we cannot add a DeactivateAsync hook without an upstream CShells change. The current IHostedService.StopAsync hook continues to work correctly for the host-stop path; per-shell deactivation would require either a CShells upstream addition or a separate Elsa-owned shell-feature variant — neither lighter than what we have today. Tests: 103/103 runtime unit + 249/249 workflow integration pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(common): introduce IsFatal exception extension + apply to drain best-effort catches Per PR review feedback on the static-analyzer "Generic catch clause" comments: rather than catching ALL exceptions in best-effort drain code paths, narrow the swallow to non-fatal exceptions. Process-fatal conditions (StackOverflowException, AccessViolationException, SEHException, ThreadAbortException, OutOfMemoryException) propagate so the host's failure-fast policy can act on them, while normal failures (InvalidOperationException, IOException, etc.) continue to be logged and allowed through so a single misbehaving ingress source / activity cannot abort the overall drain. Highlights: * New `Elsa.Common.Extensions.ExceptionExtensions.IsFatal` utility: classifies fatal conditions, unwraps reflection-style wrappers (TypeInitializationException, TargetInvocationException) before classification, and treats InsufficientMemoryException (the recoverable OOM subclass) as non-fatal. * Applied as a `when (!ex.IsFatal())` filter to: - BurstHandle.Cancel (cancel callback try/catch) - DrainOrchestrator.PauseOneSourceAsync (per-source exception path) - DrainOrchestrator.TryForceStopAsync - DrainOrchestrator.ForceCancelActiveBurstsAsync (per-burst loop) - DrainOrchestrator.PersistInterruptedAsync (orphan log write, instance save, log write) - DrainOrchestrator.DrainAsync outer catch (existing `not InvalidOperationException` filter extended) - InterruptedRecoveryScan (per-instance restart loop) Tests: 7 new unit tests for IsFatal classification (fatal types, recoverable types, wrapped causes, null tolerance). Full suites: 103/103 runtime unit (incl. 14/14 in Common.UnitTests including new tests) + 249/249 workflow integration pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(workflows-runtime): integrate CShells 0.0.15 lifecycle hooks (IDrainHandler + IShellInitializer) CShells 0.0.15 ships the lifecycle framework needed for first-class per-shell graceful shutdown — IDrainHandler / IShellInitializer / IShellLifecycleSubscriber. This commit bumps the package, migrates Elsa's existing usage of the removed 0.0.14 API, and registers the runtime drain orchestrator + pause-state initializer through the new primitives. Highlights: * `ElsaShellDrainHandler : IDrainHandler` — bridges per-shell drain into `IDrainOrchestrator.DrainAsync(DrainTrigger.ShellDeactivation, ct)`. Invoked by CShells when a shell enters `ShellLifecycleState.Draining`; the drain handler's CancellationToken is signalled when the per-shell deadline elapses, so the orchestrator's own deadline-bounded protocol nests cleanly. Coexists with the host-stop `DrainOrchestratorHostedService`; the orchestrator's `DrainAsync` is idempotent — second invocations log and skip. * `InitializePauseStateShellInitializer : IShellInitializer` — replaces the IStartupTask variant in shell-aware deployments. IShellInitializer fires on EVERY shell (re)activation, including reactivations after a reload — exactly what FR-028 requires. The IStartupTask remains for IModule consumers where there is no shell platform. Migrations (CShells 0.0.14 → 0.0.15 breaking changes): * `ActivateShellTenants`: was `IShellActivatedHandler` + `IShellDeactivatingHandler`, now `IShellInitializer` + `IDrainHandler`. * `MultitenancyFeature`: registrations updated to the new transient interface, `using CShells.Hosting` → `using CShells.Lifecycle`. * `Reload/Endpoint`, `ReloadAll/Endpoint`: `IShellManager` → `IShellRegistry`, `ReloadShellAsync` → `ReloadAsync` (returns `ReloadResult` with `Error`), `ReloadAllShellsAsync` → `ReloadActiveAsync` (returns `IReadOnlyList<ReloadResult>` with per-shell errors aggregated into 503). Build + restore: * `Directory.Packages.props`: all CShells.* packages bumped to 0.0.15. * `NuGet.Config`: added `cshells-feedz` source (https://f.feedz.io/sfmskywalker/cshells/nuget/index.json) and split the package-source-mapping pattern into `CShells` (exact) + `CShells.*` (prefix). Single-pattern `CShells*` does NOT match correctly under PackageSourceMapping. Tests: 103/103 runtime unit + 249/249 workflow integration pass on the new package version. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-runtime): apply PR #7424 review feedback Consolidates the architectural fixes asked for during /review: - Extract IWorkflowRuntimeAdminService to back the four /admin/workflow-runtime endpoints with a single domain service; thin Pause/Resume/Status/Force endpoints to delegating shells. - Remove StateChanged C# event from IQuiescenceSignal (Constitution VII: no external subscribers existed; mediator was suggested as the alternative if/when it's needed). - Promote InitializePersistedStateAsync to IQuiescenceSignal, dropping the concrete-cast in both InitializePauseStateStartupTask and InitializePauseStateShellInitializer. - Invert ingress-source DI to Lazy<IEnumerable<IIngressSource>> to break the cycle through IQuiescenceSignal; ingress adapters take the signal directly via primary constructor. - Replace Guid.NewGuid().ToString("N") with IIdentityGenerator in InterruptedLogExtensions and DrainOrchestrator. - Switch admin-audit timestamps to ISystemClock in WorkflowRuntimeAdminService. - Make GracefulShutdownOptions.StimulusQueueMaxDepthWhilePaused nullable (null = unlimited). - Rename RuntimeForceRequested → RuntimeForceDrainRequested. - Apply IsFatal exception filter to drain best-effort catches. - Rename IBurstRegistry.EnumerateActive → ListActiveBursts. - Refresh "Phase X" comments to user-story (USx) references. - Delete unused IngressAttributionExtensions, IngressSourceServiceCollectionExtensions, IngressSourceRegistrationOptions. - Migrate Elsa.Shells.Api.Tests to CShells 0.0.15 IShellRegistry / ReloadResult surface. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(graceful-shutdown): apply IsFatal filter to deadline-breach test catch Aligns the test scaffolding's swallow-everything catch with the project standard introduced in c00eee80c so the analyzer no longer flags the bare `catch` clause. The semantics are unchanged — non-fatal exceptions (OCE, TimeoutException, workflow exceptions) are still acceptable test outcomes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): close IngressSourceRegistry first-access race + spelling Replaces the non-atomic `_entries.Count > 0` early-return guard in EnsureMaterialized with a double-checked lock against a volatile `_materialized` flag, so concurrent first callers can no longer both iterate the source factory and crash one of them with a "Duplicate ingress source registration" InvalidOperationException. Adds a regression test that launches 16 readers behind a TaskCompletionSource gate and asserts every reader observes the full source set without throwing. Also flips the British spellings introduced in this PR's scope to American English (materialize/behavior) — project convention going forward. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(constitution): require American English for new code (v1.0.1) Adds a "Spelling & language" bullet under principle III (Convention-Driven Design): every newly-introduced symbol, comment, identifier, error message, XML doc, commit message, and Speckit artifact uses American English. Established public API symbols (e.g. WorkflowSubStatus.Cancelled) are not renamed retroactively. PATCH bump because this is a clarification of an existing principle, not a new principle. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Add Greploop skill and workflow for GitLab, GitHub, and Perforce integration - Introduced Greploop, an iterative optimization and review workflow for GitLab MRs, GitHub PRs, and Perforce changelists. - Added API and GraphQL references for fetching and resolving review skill. * Remove GenerateWorkflowVariableAccessorsTests; redundant ExpandoObject type check in handlers * Potential fix for pull request finding 'CodeQL / Untrusted Checkout TOCTOU' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(graceful-shutdown): apply PR #7424 review feedback round 2 Two P1 findings from Greptile: 1. DrainOrchestrator.ForceCancelActiveBurstsAsync was sequential — each burst was cancelled, awaited up to ForceCancelSettleTimeout (2 s), and persisted before the next burst's Cancel() ran. Total wall time was O(N × 2 s) and bursts 2..N kept executing at full speed during prior bursts' settle waits, defeating the intent of force-cancel under concurrency. Refactored to three phases: - Phase A — cancel every handle synchronously (cheap CTS.Cancel calls) so all runners observe cancellation simultaneously. - Phase B — await every Disposed signal in parallel under a single shared ForceCancelSettleTimeout. Total wall time bounded regardless of N. - Phase C — persist Interrupted for each handle sequentially (keeps DbContext usage single-threaded; per-handle work is small). Per-phase failures are caught with !ex.IsFatal() and logged so a single misbehaving handle doesn't abort the rest of the batch. 2. ShellFeatures/WorkflowRuntimeFeature.ConfigureServices was missing the IWorkflowRuntimeAdminService registration that Features/WorkflowRuntimeFeature already had. Any CShells deployment that includes the Pause / Resume / Status / Force admin endpoints (in Elsa.Workflows.Api) would throw InvalidOperationException at endpoint construction. Added the singleton alongside the other graceful-shutdown registrations with a comment pointing out the symmetry with the IModule path. 17/17 graceful-shutdown integration tests pass; 103/103 runtime unit tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Untrusted Checkout TOCTOU' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(ci): harden greploop.yml against CodeQL Actions findings CodeQL flagged 12 findings on .github/workflows/greploop.yml after the prior commit (c08183a3c) addressed an earlier round. Two distinct issue classes remain: 1. Code injection (× ~10): step-output values (steps.pr_head.outputs.head_sha / head_repo_owner / head_repo_name / head_ref) and inputs.pr_number were interpolated directly into shell `run:` blocks via `${{ ... }}`. Because PR author controls the branch name and the manual-dispatch input, those values can carry shell metacharacters. Standard fix: route every such interpolation through an `env:` block on the step, then reference $VAR inside the script. Applied to the Resolve, Resolve PR head metadata, and Checkout PR branch steps. 2. Untrusted Checkout TOCTOU + Checkout of untrusted code in trusted context: the workflow runs on `issue_comment` (a privileged trigger) and checks out PR-author code. Mitigations stacked here: - Author-association gate already restricts the trigger to OWNER / MEMBER / COLLABORATOR (existing). - Step-output values now travel via env vars (above). - Resolve step rejects pr_number that isn't ^[0-9]{1,10}$ — so downstream `gh pr view` and the prompt argument can't be hijacked. - Checkout step now validates HEAD_SHA matches ^[0-9a-f]{40}$ and the repo owner/name match ^[A-Za-z0-9_.-]+$ before either reaches a URL or a git command. - Existing TOCTOU guard preserved: re-fetch head SHA at checkout time and abort if it changed since the initial resolve. These match the canonical "Securing your GitHub Actions workflows" patterns recommended by CodeQL. No functional change to greploop's runtime behaviour. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(graceful-shutdown): drop [SingleNodeTask] from InitializePauseStateStartupTask Greptile P1: [SingleNodeTask] gates the task to a single cluster winner via distributed lock, but IQuiescenceSignal is a singleton scoped to each node's DI container — each node holds its own in-memory QuiescenceState. With the attribute, only the winning node restored the persisted pause; every other node started with QuiescenceReason.None and accepted new work, silently defeating PausePersistence = AcrossReactivations. Removed [SingleNodeTask] (and the corresponding using) so the task runs on every node. Expanded the doc <remarks> to call out the per-node requirement and point at the shell-aware counterpart (InitializePauseStateShellInitializer) which is correctly per-node by virtue of being an IShellInitializer. 17/17 graceful-shutdown integration tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Code injection' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * chore(deps): bump CShells 0.0.15 → 0.0.17 0.0.17 ships our blueprint-aware-routing PR (valence-works/cshells#93) plus four follow-up fixes the maintainer added on top: - e56ebd8 — PreWarmShells removed entirely; ShellMiddleware now does cold-start endpoint matching (re-runs endpoint resolution after lazy activation so the very first request to a cold shell hits its endpoint). - cbe5ee2 — GetCandidateSnapshot returns a bounded ShellRouteCandidateSnapshot with accurate total counts; sensitive-data redaction in routing logs; DefaultShellRouteIndex implements IDisposable. - cd7d4f5 — Last-good snapshot served on rebuild failure (the deferred Copilot review concern); root-path fallback when path-by-name misses. - c3679d9 — Cold-start endpoint matching respects inline route constraints; path-name convention tightening; dead duplicate-detection cleanup. Net effect for elsa-core: - Cold blueprints serve their first request via lazy activation, with endpoints correctly resolved post-activation. - Reloaded shells re-activate and serve on the next matched request. - Non-name-mode routing keeps serving the previous snapshot during a transient blueprint-provider outage. - No need to call PreWarmShells from Elsa.ModularServer.Web — removed. The only API removal that touches elsa-core is PreWarmShells. No code references IShellRouteIndex / ShellRouteCriteria / GetCandidateSnapshot directly, so the API-shape changes in cbe5ee2 don't ripple here. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(graceful-shutdown): persist Interrupted under non-drain bounded token Greptile P? finding on the prior force-cancel two-phase fix: Phase B's inner catch on OperationCanceledException ("drain CT fired — proceed to persist anyway") was a lie in practice. Phase C immediately passed the same already-cancelled drain token into PersistInterruptedAsync; the first DB call (instanceStore.FindAsync) observed the cancellation and threw OperationCanceledException; the outer non-fatal Exception filter swallowed it and only logged an error. Net effect: on host shutdown deadline breach, every burst after cancellation could fail to be persisted as Interrupted, leaving instances in an unrecovered executing state. Phase C now creates a per-handle CancellationTokenSource bounded to a new PersistInterruptedTimeout (5 s) that is NOT linked to the drain CT. Each persist gets up to 5 s to land the row update + forensic log entry even after the drain CT has fired. The bound prevents a stuck DB from hanging shutdown indefinitely (per-handle worst case is small; total Phase C upper bound is N × 5 s, but typical persists are millisecond scale). Comment expanded to call out why the persist token is independent of the drain token, so the rationale doesn't drift again. 17/17 graceful-shutdown integration tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): extract PassiveIngressSource base class The three IIngressSource implementations that ship with this PR (InternalBookmarkQueueIngressSource, HttpTriggerIngressSource, ScheduledTriggerIngressSource) were ~25 lines each and ~22 lines of those were verbatim copies of each other: - ctor signature `(IQuiescenceSignal signal)` - `PauseTimeout => TimeSpan.FromMilliseconds(50)` - `CurrentState => signal.IsAcceptingNewWork ? Running : Paused` - `PauseAsync` / `ResumeAsync` returning `ValueTask.CompletedTask` The shared trait is that none of them does any work at pause time — the actual pause enforcement lives in another layer (`HttpWorkflowsMiddleware` short-circuits to 503, `BookmarkQueueProcessor` consults the signal at the top of each invocation, scheduled triggers dispatch through the bookmark queue and inherit that behaviour transitively). The IIngressSource adapter is purely diagnostic: it makes the source visible in `DrainOutcome.Sources` and the admin status endpoint. Extracted that pattern into `PassiveIngressSource` (abstract base in `Elsa.Workflows.Runtime.IngressSources`). Subclasses now provide only `Name`; `PauseTimeout` is `virtual` with a 50 ms default; everything else is fixed by the base. The three concretes drop from ~25 lines to ~12 lines each. The base's XML `<remarks>` calls out when to use it ("your component already cooperates with IQuiescenceSignal at its hot path") and when to implement IIngressSource directly ("the source owns concrete pause/resume behaviour — e.g. a message-queue consumer that calls Pause() on its underlying client"), so future contributors don't mis-extend the base for sources that need real work at pause time. No behavioural change. 18 graceful-shutdown integration + 39 runtime unit tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): align IIngressSource name to singular The three IIngressSource names were inconsistent: http.trigger (singular) internal.bookmark-queue-worker (singular) scheduling.triggers (PLURAL — outlier) The plural slipped in when addressing Greptile's earlier comment to avoid `scheduling.cron` (which would imply Cron-only coverage). The right move was to pick a generic word and stay singular like the rest of the suite — the suite's mental model is "the X source", one instance per registry slot, regardless of how many triggers or items it dispatches internally. Renamed to `scheduling.trigger`. The `<remarks>` block keeps the "covers Cron, Timer, StartAt, Delay" explanation and now also explicitly notes the singular convention so future contributors don't re-pluralize. Zero test fallout — the literal "scheduling.triggers" only appeared in the source file itself. Tests of the other two sources all use singular forms. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(runtime-admin): rename Force endpoint to ForceDrain "Force" alone is meaningless out of context — force what? — and it sits oddly next to the verb-named siblings Pause / Resume / Status. The matching admin service method is already IWorkflowRuntimeAdminService. ForceDrainAsync, so ForceDrain is the natural pair. Renamed: - src/modules/Elsa.Workflows.Api/Endpoints/RuntimeAdmin/Force/ → ForceDrain/ - namespace ...Endpoints.RuntimeAdmin.Force → ...ForceDrain - class ForceEndpoint → ForceDrainEndpoint - class ForceRequest → ForceDrainRequest - class ForceResponse → ForceDrainResponse - route POST /admin/workflow-runtime/force → /force-drain Zero external references — no tests, docs, or OpenAPI clients used the old symbols or the old route literal, so this is a contained pre-ship rename. Directory move went through `git mv` so commit history follows the file. 17/17 graceful-shutdown integration tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): correct formatting in checkout step within greploop.yml Adjusted indentation of environment variables in the checkout step for improved consistency and readability. * fix(ci): repair malformed Checkout repository step in greploop.yml The step accumulated stray env keys, an extra `uses:`, and bash commands that didn't belong inside it (line 65 onward), causing a YAML parse error on push. The valid structure has two distinct checkout steps: - Checkout repository : actions/checkout@v4 with fetch-depth: 0 - Checkout PR branch : env: + run: with SHA validation + git fetch + git checkout --detach The PR-branch step (line 78+) was already correct and unchanged. This fix restores the first step to its intended single-purpose shape (just checks out the workflow file's commit so the greploop skill is on disk before the run-greploop step uses it). No functional change to runtime behaviour or to the security posture established in the prior hardening commit (0db4ca23e). The PR-branch checkout still validates HEAD_SHA / HEAD_REPO_OWNER / HEAD_REPO_NAME shape before they reach a URL or git command. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): remove literal `${{ }}` from greploop.yml comment GitHub Actions parses `${{ ... }}` workflow expressions across the entire YAML file, including inside `run:` script comments. The comment that explained the env-var hardening pattern contained the literal sequence `${{ }}` (with a space, intended as an English-language description), which the expression parser rejected as "An expression was expected" (line 81 col 14). Reworded the comment to describe the substitution form in prose without the literal token sequence. Functional behaviour unchanged. * refactor(graceful-shutdown): rename burst → execution cycle The graceful-shutdown work introduced "burst of execution" as a first-class domain concept. The term arrived without rationale and isn't standard in the workflow-engine domain. Renamed to "execution cycle" — reads more naturally as the loop-with-commit unit, is more idiomatic in workflow vocabulary, pairs cleanly with the existing WorkflowExecutionContext, and avoids collisions with Elsa's existing terms (Run, Execution, Dispatch, Invocation, Stimulus, Step). Renamed types - IBurstRegistry → IExecutionCycleRegistry - BurstRegistry → ExecutionCycleRegistry - BurstHandle → ExecutionCycleHandle - BurstTrackingMiddleware → ExecutionCycleTrackingMiddleware - BurstAwareCommitStateHandler → ExecutionCycleAwareCommitStateHandler Renamed members - BeginBurst → BeginCycle - ListActiveBursts → ListActiveCycles - BurstHandleKey constant + value → ExecutionCycleHandleKey - ActiveBurstCount (IQuiescenceSignal, RuntimeAdminStatus, StatusResponse) → ActiveExecutionCycleCount - WaitForBurstsAsync (private) → WaitForCyclesAsync - ForceCancelActiveBurstsAsync (priv) → ForceCancelActiveCyclesAsync - UseBurstTracking → UseExecutionCycleTracking - DrainOutcomeDto.BurstsForceCancelledCount → ExecutionCyclesForceCancelledCount - _burstRegistry / burstRegistry → _cycleRegistry / cycleRegistry Backwards-compatibility preservation (the only persisted JSON key) - WorkflowInterruptedPayload.BurstDuration property → ExecutionCycleDuration with [JsonPropertyName("BurstDuration")] so the persisted JSON wire key stays "BurstDuration" forever. Pre-merge testers' log records still deserialise correctly. The contract test on WorkflowInterruptedPayloadContractTests still asserts the wire key "BurstDuration" appears in the serialised JSON — confirms the guarantee is enforced. Other unstructured surfaces - WorkflowExecutionLogRecord.Message text "Workflow burst was force- cancelled..." now says "Workflow execution cycle was force-cancelled ..." for new records. Old rows keep their old text — purely cosmetic free-text field. - Structured log placeholder {BurstId} in DrainOrchestrator log lines → {ExecutionCycleId}. - Lowercase prose / XML doc comments updated throughout. Test renames - BurstRegistryTests → ExecutionCycleRegistryTests - BurstTrackingMiddlewareTests → ExecutionCycleTrackingMiddlewareTests - Test method names + DisplayName strings updated. Spec docs (specs/002-graceful-shutdown/) updated to match the new vocabulary; the historical task records in tasks.md keep the old names as-is to preserve the audit trail of what was originally built. Verification - dotnet build: clean across net8.0 / net9.0 / net10.0. - 103/103 Elsa.Workflows.Runtime.UnitTests pass. - 17/17 GracefulShutdown integration tests pass. - 4/4 WorkflowInterruptedPayloadContractTests pass — confirms the "BurstDuration" JSON wire-key preservation is intact. No changes to migrations or DB column names — confirmed via grep. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): give greploop.yml gh-cli a repo context before checkout The "Resolve PR head metadata" step runs `gh pr view` before actions/checkout, so there is no `.git` directory and gh's "current repo" detection fails with `fatal: not a git repository`. The prior commits to this file masked the runtime failure because the workflow itself was YAML-invalid — once it became valid, the workflow_dispatch trigger surfaced this real-world execution bug. Set GH_REPO=${{ github.repository }} on both `gh pr view` steps. The gh CLI honours GH_REPO as an explicit repo override, so it no longer needs git context. Same fix on the "Checkout PR branch" validation call which also uses gh pr view before the manual fetch. The error reported as "Invalid workflow file: ... (Line 81 Col 14)" on PR #7424 was stale from commit ed580682a (which had the bad comment with literal `${{ }}`); commit 3a8ad0d08 fixed the YAML, but because greploop's `if:` condition only matches workflow_dispatch / issue_comment events, push events on later commits were skipped without re-running validation, so the GitHub UI kept showing the old error. A workflow_dispatch run on the current SHA now passes validation and reaches "Resolve PR head metadata", which is what this commit fixes. * fix(graceful-shutdown): wire IngressPauseTimeout option to drain orchestrator GracefulShutdownOptions.IngressPauseTimeout was documented as "Default per-ingress-source pause timeout" but DrainOrchestrator.PauseOneSourceAsync read source.PauseTimeout directly and never consulted the option. The configured value was silently ignored — operators who set GracefulShutdownOptions:IngressPauseTimeout = 10s were getting whatever each source's hardcoded value was (50 ms for the three PassiveIngressSource subclasses we ship), with no way to tune it globally. Precedence (per the spec's intent of "overridable at registration and by configuration"): 1. Per-source positive value wins (source.PauseTimeout > Zero). 2. Otherwise fall back to the configured GracefulShutdownOptions. IngressPauseTimeout default. 3. Resolved value is capped at the overall drain deadline so a single misbehaving source cannot exceed the host's shutdown budget. 4. 1 ms safety floor remains so a misconfigured zero default still produces a non-zero CancelAfter. Changes: - DrainOrchestrator.PauseOneSourceAsync — adds the precedence above with a comment block explaining each step. - IIngressSource.PauseTimeout — XML doc clarifies the Zero-defers-to- config semantics. - GracefulShutdownOptions.IngressPauseTimeout — XML doc says it's the fallback when the source returns Zero; <remarks> spells out the precedence and the overall-deadline cap. - PassiveIngressSource.PauseTimeout — virtual property now returns Zero (was 50 ms). The three shipped subclasses (HttpTriggerIngressSource, ScheduledTriggerIngressSource, InternalBookmarkQueueIngressSource) consequently defer to the configured default — flipping the wire-up bug from "configured value silently ignored" to "configured value honoured by default for passive sources". Passive subclasses that want a specific value can still override. 103/103 runtime unit + 17/17 graceful-shutdown integration tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): rename IInterruptedRecoveryScan → IInterruptedRecoveryScanner The interface had a single verb-method (`ScanAndRequeueAsync`) and its XML doc described what it *does* ("Scans the workflow instance store for instances..."). That's an agent role — a scanner that performs a scan — but the noun-shaped name `IInterruptedRecoveryScan` read as "the scan itself", which is misleading because the scan results / event are not first-class types in the codebase. Renamed to `IInterruptedRecoveryScanner` / `InterruptedRecoveryScanner` to match the existing `-er` convention in this codebase (Restarter, Generator, Resolver, etc.). The method stays `ScanAndRequeueAsync` — the scanner *performs* a scan-and-requeue. Also renamed the constructor parameter `scan` → `scanner` in RecoverInterruptedWorkflowsStartupTask, and the local variable `scan` → `scanner` in InterruptedRecoveryIntegrationTests, so the "scanner does the scan" mental model is consistent throughout. Surface impact (all internal — no API or persistence touch points): - 2 source files renamed via git mv (interface + implementation) - 1 test file renamed (InterruptedRecoveryScanTests → ScannerTests) - DI registrations in both Features/ and ShellFeatures/ WorkflowRuntimeFeature - 1 startup-task constructor parameter - Spec doc references under specs/002-graceful-shutdown/ 103/103 runtime unit + 17/17 graceful-shutdown integration tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): extract DrainTriggerExecutor ElsaShellDrainHandler (CShells IDrainHandler) and DrainOrchestratorHostedService (.NET IHostedService.StopAsync) inlined near-identical try/catch/log shapes around IDrainOrchestrator.DrainAsync: - call DrainAsync(<trigger>, ct) - branch on outcome: DeadlineExceeded/AbortedByUnhandledException → Warning, otherwise Information - catch InvalidOperationException (parallel-drain rejected by the orchestrator) → log Information and swallow The two had already drifted: host-stop's success log omitted the paused/waited durations the shell-handler version included, and the "skipped" message disagreed on the trigger label ("Host-stop drain skipped" vs "Shell drain skipped"). Centralised the shape in a small internal static helper so the two — and any future trigger source — stay uniform. Both call sites collapse to a single line. Net diff drops 22 lines from the two consumers and adds a 25-line helper that they both delegate to. The unified log copy now consistently includes paused/waited durations on the success path and uses the caller-supplied contextLabel ("Shell drain", "Graceful drain") in all three messages so operators can attribute log entries by trigger source. Files: - src/modules/Elsa.Workflows.Runtime/Services/DrainTriggerExecutor.cs (new) - src/modules/Elsa.Workflows.Runtime/Lifecycle/ElsaShellDrainHandler.cs - src/modules/Elsa.Workflows.Runtime/HostedServices/DrainOrchestratorHostedService.cs 103/103 runtime unit + 17/17 graceful-shutdown integration tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-runtime): drop redundant DrainOrchestratorHostedService from CShells path In CShells deployments, host stop already drives drain via CShellsStartupHostedService → IDrainHandler → ElsaShellDrainHandler, scoped per shell (FR-027). The additional .AddHostedService<DrainOrchestratorHostedService>() in ShellFeatures/WorkflowRuntimeFeature was firing a second non-force DrainAsync that the orchestrator rejected with InvalidOperationException — silently swallowed by DrainTriggerExecutor, but logged on every host stop and semantically muddled (IHostedService is host-level, not per-shell). Keep the registration on the IModule path (Features/WorkflowRuntimeFeature) where there is no shell platform and host-stop is the only available drain trigger. Update ElsaShellDrainHandler XML docs to reflect the now-clean single-trigger model in CShells. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): always dispose ExecutionCycleHandle in tracking middleware Previously the success path relied on ExecutionCycleAwareCommitStateHandler to dispose the handle after the runner's commit completed. If a custom dispatcher or test double exited the pipeline without invoking commit (by design or by accident), the handle stayed registered, IExecutionCycleRegistry.ActiveCount never reached zero, and drain spun in WaitForExecutionCyclesAsync until the deadline fired — incorrectly force-cancelling instances that had already finished cleanly. Collapse the existing try/catch(rethrow) into try/finally so the middleware itself disposes the handle for both exception and commit-elided paths. Disposal remains idempotent via the ExecutionCycleHandle._disposed Interlocked guard, so the normal-path dispose by ExecutionCycleAwareCommitStateHandler is a harmless no-op. Adds an integration regression test that drives the middleware with a stub Next that returns without invoking commit and asserts ActiveCount returns to zero. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): serialize QuiescenceSignal pause-state persistence Both PauseAsync and ResumeAsync used to release the inner lock before issuing the persistence I/O. A rapid Pause → Resume sequence could leave the persisted state inconsistent: PauseAsync's slow SaveAsync could land AFTER ResumeAsync's DeleteAsync, leaving the key present in the store while in-memory state was None. On host restart, InitializePersistedStateAsync would find the stale key and start the runtime in the paused state the operator had already cancelled. Introduce a dedicated SemaphoreSlim that serializes persistence I/O, with each I/O re-reading the live in-memory state inside the semaphore. N racing Pause/Resume calls now produce N serialized writes, each reflecting the most recent in-memory transition — so the final persisted state always matches final in-memory state. Adds a regression test that gates SaveAsync, races a Resume behind it, and asserts the store is empty after both complete. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(workflows-runtime): trim verbose comment in ExecutionCycleTrackingMiddleware finally block Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-runtime): use 'using var' for ExecutionCycleHandle in tracking middleware Replace the explicit try/finally that only existed to call handle.Dispose() with a `using var` declaration — identical semantics (compiler-emitted finally with idempotent dispose), more idiomatic. The regression test HandleReleasedWhenCommitIsElided continues to validate the leak-free property. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime,api): proper 409 conflict shape + shell-scoped pause-persistence key ForceDrain endpoint: the 409 path returned `new ForceDrainResponse()` whose `Outcome` was null at runtime despite the `= null!` annotation, so any strongly-typed client deserializing the conflict body and reading `Outcome.OverallResult` got an NRE. Switch to the existing `ConflictResponse` shape with `Code = "DrainInProgress"` and the current runtime status. Routed via HttpContext.Response.WriteAsJsonAsync because Send.ResponseAsync is constrained to the endpoint's TResponse and cannot send a sibling DTO. QuiescenceSignal persistence key: the DI-registered `IQuiescenceSignal` was constructed with `shellName = null` (DI doesn't inject `string?` defaults), so every shell shared the key `elsa.quiescence.pause.default`. In a CShells multi-shell deployment under PausePersistencePolicy.AcrossReactivations this caused cross-shell contamination — pausing shell A would re-pause shell B on its next activation. Replace the simple AddSingleton<IQuiescenceSignal,...> registration in ShellFeatures/WorkflowRuntimeFeature with a factory that injects `CShells.ShellSettings` and forwards `Settings.Id` as the shell name. The IModule registration is unchanged (no shell platform; null shellName remains correct there). Adds a unit regression test that two QuiescenceSignal instances with different shellNames write to disjoint persistence keys and never to the legacy "default" key. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(quiescence): use TryGetValue for ContainsKey+indexer assertions Combines existence check and value retrieval into a single dictionary lookup, addressing the code-quality bot's repeated suggestion. No behavior change — both PauseWritesKey and PersistenceKeyIncludesShellName still assert the same keys exist with the same content. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Update specs/002-graceful-shutdown/contracts/admin-endpoints.md Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * fix(workflows-runtime): decouple QuiescenceSignal persistence from caller cancellation PersistAsync used to forward the caller's CancellationToken to both _persistenceMutex.WaitAsync and the store I/O. If an HTTP request was cancelled between the in-memory transition (already committed under _sync) and the persistence call, the I/O was silently skipped — leaving AdministrativePause set in memory with no persisted record. The idempotent fast-path on subsequent PauseAsync calls (transitioned == false) meant no retry would happen, so on host restart InitializePersistedStateAsync would find no key and the runtime would come back unpaused, defeating PausePersistencePolicy.AcrossReactivations. Drop the parameter from PersistAsync entirely; use CancellationToken.None for both the semaphore wait and the store I/O. The in-memory transition is already committed by the time PersistAsync runs, so persistence must complete to keep the store consistent with memory. The public PauseAsync/ResumeAsync methods still accept a CancellationToken (interface contract) — it just no longer reaches the persistence layer. Adds a regression test that calls PauseAsync with a pre-cancelled token and asserts both in-memory pause and the persisted key land correctly. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime,api,docs): apply Copilot review feedback batch Code: - DrainOrchestrator.TryForceStopAsync now bounds force-stop with the *remaining* drain budget (deadlineAt - now), not the full overall TimeSpan. A per-source pause that already burned the shutdown window can no longer get another full deadline's worth of force-stop runway. - DrainOrchestrator catch filter narrowed: drop `ex is not InvalidOperationException` exclusion. The "drain already in progress / completed" IOEs are thrown outside the protocol's try block, so they bubble out without entering this handler. Any IOE that lands here is incidental (e.g., from a store inside the drain) and should now be captured into the outcome rather than escaping the whole drain. - ResumeEndpoint 409 path now returns the discriminated ConflictResponse shape (matching ForceDrain) instead of a plain StatusResponse. Routed via HttpContext.Response.WriteAsJsonAsync because Send.ResponseAsync is constrained to TResponse. - Conflict codes aligned to kebab-case across both endpoints to match the contract spec (`runtime-draining` and `drain-in-progress`). Spelling sweep — American English per constitution v1.0.1 III: - DrainOrchestrator.cs: "serialised" → "serialized" - WorkflowInterruptedPayload.cs: "serialised" / "deserialise" → "serialized" / "deserialize" - PassiveIngressSource.cs: "behaviour" → "behavior" - DeadlineBreachEndToEndTests.cs: "serialisable" → "serializable" - specs/002-graceful-shutdown/quickstart.md: "behaviour" → "behavior" - specs/002-graceful-shutdown/checklists/requirements.md: "behaviour" → "behavior" Doc/contract alignment: - quickstart.md: force route corrected from /force to /force-drain. - quiescence-signal.md: removed StateChanged event from contract (interface doesn't define it); corrected persistence section to describe InitializePersistedStateAsync via shell initializer / startup task rather than constructor read; added the per-shell key discriminator. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): correct DI lifetimes for ExecutionCycleTrackingMiddleware and WorkflowRuntimeAdminService Two strict-DI-validation failures surfaced in tests using BuildServiceProvider with validate-on-build: 1. ExecutionCycleTrackingMiddleware was registered as AddSingleton<>, but its constructor takes WorkflowMiddlewareDelegate next — supplied by the workflow execution pipeline builder via UseMiddleware<>(), not from DI. The registration was both unused (no consumer resolves it through the container) and broken (DI fails to construct it because next is unregistered). Removing both registrations. 2. IWorkflowRuntimeAdminService was registered as AddSingleton<> but depends on the scoped INotificationSender (mediator) — captive-dependency violation. All consumers (Pause/Resume/ Status/ForceDrain endpoints) are FastEndpoints, which are scoped per request, so AddScoped is the correct alignment. The other deps (IQuiescenceSignal / IIngressSourceRegistry / IDrainOrchestrator / ISystemClock) are singletons and resolve fine from a scoped consumer. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): restore commit-handler-only success disposal + true Cancel idempotency Two issues raised by Copilot's latest review on commit 32a9c0519: 1. ExecutionCycleTrackingMiddleware was disposing the handle at the end of InvokeAsync (via `using var`), but WorkflowRunner runs commit AFTER the pipeline returns (WorkflowRunner.cs:235). That meant the handle was disposed BEFORE the runner's terminal commit, and the drain orchestrator's `await handle.Disposed` would unblock too early — reintroducing the runner-clobber race the original design protected against (see ExecutionCycleAwareCommitStateHandler XML doc). Revert to the original shape: only dispose on exception path. ExecutionCycleAwareCommitStateHandler remains the SOLE success-path disposer, running in its finally block AFTER the inner commit lands. The earlier "leak when commit is elided" concern was a non-issue in production (the standard runner always commits); the buggy `HandleReleasedWhenCommitIsElided` test that specified the wrong contract is removed. The existing `ActiveCountReturnsToZero` test (which uses the real runner end-to-end) already verifies success-path disposal. 2. ExecutionCycleHandle.Cancel() was documented as idempotent but only short-circuited via the _disposed flag. Repeated Cancel() calls before Dispose could trigger the cancel callback multiple times — easy to accidentally fire non-idempotent cancellation side effects more than once. Add an Interlocked _cancelled guard so callback + CTS cancellation run at most once. Existing test that documented the leaky behavior is updated to assert the now-truly-idempotent contract. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Update logging levels and remove unused features in appsettings files * refactor(workflows-runtime): improve graceful shutdown options handling and cleanup solution Refactor the handling of `GracefulShutdownOptions` to ensure options are applied correctly without directly invoking the delegate. Update DI registrations to use appropriate lifetimes and remove redundant wrapper services. Additionally, clean up the solution by removing unused projects and documentation folders. * feat(identity, workflows-runtime): add validation for identity and graceful shutdown options Introduce validation capabilities for `IdentityTokenOptions` and `GracefulShutdownOptions`. Implement extension methods for option validation, enhance service registration, and add unit tests to ensure configurations are validated at startup. Update solution to include new unit test projects. * update(docs): clarify shutdown log message expectations and levels in quickstart.md Optimize explanation of expected log message sequence during graceful shutdown and specify logging levels. * docs: amend constitution to v1.1.0 (SRP, DRY, KISS, conciseness under Principle VII) * refactor(multitenancy): rename and restructure TenantTaskManager to TenantTaskLifecycleCoordinator Rename `TenantTaskManager` to `TenantTaskLifecycleCoordinator` and relocate to a new directory structure, enhancing code organization and test consistency. Retain functional behaviors with no logic alterations. Update unit tests to reflect the naming changes, ensuring consistency with the refactored code structure. * Update logging levels and dependencies - Set default logging level to Debug in appsettings.Development.json - Add missing using directives for Elsa workflows management and runtime features - Update CShells package versions to 0.0.18-preview.104 --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-05-02 17:27:08 +00:00
using Elsa.Common;
using System.Reflection;
using System.Runtime.InteropServices;
using Xunit;
namespace Elsa.Common.UnitTests;
public class ExceptionExtensionsTests
{
[Fact(DisplayName = "Null is not fatal")]
public void NullNotFatal() => Assert.False(((Exception?)null).IsFatal());
[Theory(DisplayName = "Process-fatal exceptions are classified fatal")]
[InlineData(typeof(StackOverflowException))]
[InlineData(typeof(AccessViolationException))]
[InlineData(typeof(SEHException))]
[InlineData(typeof(ThreadAbortException))]
public void FatalExceptions(Type exceptionType)
{
#pragma warning disable SYSLIB0050 // FormatterServices is needed to instantiate fatal exception types without throwing them.
Graceful shutdown for the workflow runtime (drain, pause, recover) (#7424) * feat(workflows-runtime): add quiescence machinery foundation for graceful shutdown Introduces the container-scoped quiescence signal, ingress-source contract, burst registry, and the Interrupted workflow sub-status — the foundational primitives the drain orchestrator and admin endpoints will build on. No behaviour change yet: workflows continue to run and shut down exactly as before. The new types are registered but no host-stop or pause path drives them. Highlights: * IQuiescenceSignal — composable Drain + AdministrativePause flags; forward-only drain, reversible pause, idempotent transitions, optional persistence via IKeyValueStore. * IIngressSource + IForceStoppable — uniform contract for components that inject external events (HTTP, schedulers, message consumers, internal workers, third-party modules); IIngressSourceRegistry collects and surfaces their states. * IBurstRegistry — atomic counter for in-flight workflow execution bursts, with per-burst ingress attribution and FR-018 inconsistency detection (a source claiming Paused but starting bursts is flipped to PauseFailed). * WorkflowSubStatus.Interrupted — new value distinct from Suspended, Cancelled, Faulted; semantics: "last burst force-cancelled by graceful drain; resumable on next runtime generation". Mirrored on the API client enum. * GracefulShutdownOptions — drain deadline, per-source pause timeout, stimulus-queue back-pressure policy, pause-persistence policy. Configurable via UseWorkflowRuntime(...).ConfigureGracefulShutdown(...). * PermissionNames.ManageWorkflowRuntime — single permission for the forthcoming admin pause/resume/status/force endpoints. Implements 31 of 77 tasks for the graceful-shutdown feature (specs/002-graceful-shutdown). Subsequent commits add the drain orchestrator (US1 / MVP), Interrupted recovery scan (US3), admin endpoints (US2), and first-party ingress adapters. Tests: 25 new xUnit unit tests; 100/100 runtime unit tests pass; all existing tests continue to pass on net8.0/net9.0/net10.0. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(workflows-runtime): add drain orchestrator + host-stop integration (US1, MVP) When the host receives a stop signal (SIGTERM, Ctrl+C, orchestrator rollout), the runtime now drains gracefully: ingress sources are paused in parallel, in-flight workflow bursts run to their next natural persistence boundary within a configurable deadline, and any burst that breaches the deadline is force-cancelled and persisted with the Interrupted sub-status plus a forensic WorkflowInterrupted log entry. This is the MVP — without the activation-time recovery scan (PR 3) the existing timeout-based RestartInterruptedWorkflowsTask still picks up Interrupted instances, just on its periodic cadence. No regression in that recovery path (SC-008). Highlights: * IDrainOrchestrator + DrainOrchestrator — protocol per the contract: BeginDrainAsync → parallel ingress pause with per-source timeouts + IForceStoppable escalation → poll BurstRegistry.ActiveCount until zero or deadline → on breach iterate live handles, cancel, persist Interrupted, write log entry. All exceptions are captured into the returned DrainOutcome; only second-invocation throws. * Deadline clamping: effective deadline is min(GracefulShutdownOptions. DrainDeadline, HostOptions.ShutdownTimeout - 500ms safety epsilon), so the runtime never outlives its host process. * DrainOrchestratorHostedService — IHostedService.StopAsync wakes the orchestrator on host stop. Registered AFTER the heartbeat (Elsa.Hosting.Management) so reverse-order shutdown keeps the heartbeat alive throughout drain. Prevents sibling-node crash recovery from false-positive-recovering instances we are gracefully handling here (FR-029). * BurstTrackingMiddleware — workflow-execution-pipeline middleware that registers a BurstHandle for the lifetime of every burst. All nine IWorkflowRunner.RunAsync overloads ultimately funnel into pipeline.ExecuteAsync(context), so this single middleware covers the three "burst choke points" the spec references without nine separate decorators. Added to UseDefaultPipeline(). * Ingress attribution: optional IngressSourceName property on DispatchStimulusRequest, DispatchWorkflowDefinitionRequest, and DispatchWorkflowInstanceRequest. Adapters set it; the middleware reads it via WorkflowExecutionContext.TransientProperties (helpers in IngressAttributionExtensions). The BurstRegistry uses the name to detect the FR-018 invariant violation — a source that reports Paused but starts a burst is flipped to PauseFailed. * InterruptedLogExtensions — the LogWorkflowInterruptedAsync helper that the orchestrator calls when persisting the forensic record. Tests: 11 new unit tests (DrainOrchestrator parallel-pause + wait-for-bursts + idempotency + persistence-failure path); 5 new integration tests (full DI graph resolves, burst-tracking middleware registers handles end-to-end, no-op drain returns CompletedWithinDeadline). 100/100 runtime unit tests pass; all existing tests continue to pass on net8.0/net9.0/net10.0. Implements 14 of 77 tasks (T032–T045). Subsequent commits add Interrupted recovery scan (US3), admin endpoints (US2), and first-party ingress adapters. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(admin-endpoints): add admin endpoints for workflow runtime control Introduced admin endpoints to manage workflow runtime: `/pause`, `/resume`, `/status`, and `/force` with full authentication and audit logging. Integrated idempotency checks and error handling to ensure reliable runtime control. Added corresponding integration tests for verification. * fix(workflows-runtime): propagate drain cancellation into running workflows + initialize persisted pause + fix log-comment Addresses three findings from the PR #7424 code review: 1. **HIGH — Cancellation now propagates into the running workflow.** `BurstHandle.Cancel()` previously cancelled only its own linked CTS, which the workflow runner never observes (the runner reads from `WorkflowExecutionContext.CancellationToken`, captured at context construction and not part of the linked chain). On deadline breach the orchestrator would persist `Interrupted`, but the workflow continued executing and could overwrite the sub-status with whatever terminal state it eventually reached. Fix: `BurstHandle` accepts an optional cancel callback at construction. `BurstTrackingMiddleware` wires it to `context.Cancel()` so the burst's cancellation triggers the workflow's own cancellation chain — the workflow transitions to `Cancelled` and stops scheduling new activities. The orchestrator then awaits `BurstHandle.Disposed` (with a 2 s settle timeout) before persisting `Interrupted`, ensuring the runner's terminal commit completes BEFORE the orchestrator overwrites the sub-status. Race resolved. The settle timeout is bounded so a non-cancellable activity (genuinely pathological case) does not block drain — on timeout the orchestrator logs and proceeds, accepting the runner-clobber for that one instance, which the existing timeout-based RestartInterruptedWorkflows recovery picks up afterwards. 2. **MEDIUM — Pause persistence is now actually wired.** `QuiescenceSignal.InitializePersistedStateAsync` was implemented but nothing called it on host startup. A host configured with `PausePersistence = AcrossReactivations` would write the persisted key on pause, but on subsequent activation the new `QuiescenceSignal` instance would never read it, so the runtime would resume dispatching despite the operator having paused. Fix: `InitializePauseStateStartupTask : IStartupTask` reads the policy and calls `InitializePersistedStateAsync` once per activation when the policy demands it. Registered in both `WorkflowRuntimeFeature` flavours alongside the other graceful-shutdown services. 3. **MEDIUM — Comment in `DrainOrchestrator.PersistInterruptedAsync` no longer lies.** The previous comment promised a "synthetic log entry" that the next statement (`return`) prevented from being written. The comment is now honest about what actually happens: when no instance row exists, no log entry is emitted, but the burst metadata is still captured in the drain outcome's logged warning so operators have a forensic trail. Tests: * New unit tests on `BurstRegistry` (now 9, was 6): cancel-callback is invoked, callback exceptions are swallowed (drain remains best-effort), `BurstHandle.Disposed` completes on dispose. * Full suites continue to pass: 103/103 runtime unit tests; 247/247 workflow integration tests. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): address PR review feedback + close runner-clobber race + add e2e drain test Six issues raised in PR #7424 review (one e2e gap, five comments inline): 1. **Runner-clobber race closed via ICommitStateHandler decorator.** The previous fix wired `BurstHandle.Cancel()` to `WorkflowExecutionContext.Cancel()` so the workflow's cancellation chain fires on deadline breach, but the e2e test exposed that `BurstHandle` disposed at the END of the pipeline middleware (i.e., BEFORE `WorkflowRunner` calls `commitStateHandler.CommitAsync`). The orchestrator's `await handle.Disposed` therefore returned too early, the instance row didn't yet exist, and the orchestrator's Interrupted write was either a no-op (no row) or got clobbered by the runner's subsequent Cancelled commit. Fix: `BurstAwareCommitStateHandler` decorates `ICommitStateHandler`. The middleware no longer disposes the handle in the success path — it stores the handle in `WorkflowExecutionContext.TransientProperties`, and the decorator disposes it AFTER `inner.CommitAsync` completes. The exception path in the middleware still disposes for safety. Result: the orchestrator's await-disposed sequencing now correctly lands the Interrupted write last. 2. **C1: Null-instance log entry.** `DrainOrchestrator.PersistInterruptedAsync` now writes a synthetic `WorkflowInterrupted` log entry directly when no instance row exists, populating only the fields it knows. Previously the forensic trail was lost. 3. **C2: Force endpoint cached-outcome audit.** Added `WasCached` flag to `DrainOutcome` (default false). The orchestrator sets it on the cached return path (`_previousOutcome with { WasCached = true }`). The force endpoint now skips the audit notification when the flag is true, so repeated force calls no longer emit spurious `RuntimeForceRequested` events (SC-007 idempotency restored). 4. **C3: `StateChanged` raised under lock — deadlock risk closed.** `QuiescenceSignal.BeginDrainAsync`/`PauseAsync`/`ResumeAsync` now do their transitions under the lock, capture whether a transition occurred, release the lock, and only then invoke `RaiseStateChanged`. Subscribers that synchronously call back into the signal can no longer deadlock. 5. **C4: Scheduling source name.** Renamed `scheduling.cron` → `scheduling.triggers` to honestly reflect the four trigger types the adapter covers (Cron, Timer, StartAt, Delay). The name is surfaced verbatim in admin status responses. 6. **C5: Hardcoded Retry-After.** `HttpWorkflowsMiddleware`'s 503 response now sets a reason-aware `Retry-After`: 5 s during drain (host is exiting and will be replaced shortly), 60 s during administrative pause (indefinite, so a longer back-off avoids tight retry loops). Tests: * New e2e `DeadlineBreachEndToEndTests` (2 tests): verifies that drain against a real running workflow detects the in-flight burst, force-cancels it, persists the instance as `Interrupted`, and writes a `WorkflowInterrupted` log entry — closing the test gap that hid the cancellation-propagation issue identified in the previous review pass. * Updated `OperatorForceAfterPreviousReturnsCachedOutcome` to assert value-equality + the `WasCached` flag instead of reference-equality (records use `with` for the cached return path). Full suites pass: 103/103 runtime unit, 249/249 workflow integration. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-api): move runtime admin endpoints into Elsa.Workflows.Api Per PR review feedback: rather than introducing a new sub-module (Elsa.Workflows.Runtime.Admin) for the four pause/resume/status/force endpoints, fold them into the existing Elsa.Workflows.Api project. That project already references both Elsa.Workflows.Runtime and Elsa.Api.Common (FastEndpoints) and is the established home for client-facing workflow APIs — so the admin endpoints belong there. Changes: * New folder src/modules/Elsa.Workflows.Api/Endpoints/RuntimeAdmin/ with Models.cs and Pause/Resume/Status/Force/Endpoint.cs. Namespaces moved from `Elsa.Workflows.Runtime.Admin` → `Elsa.Workflows.Api.Endpoints.RuntimeAdmin`. * Deleted src/modules/Elsa.Workflows.Runtime.Admin/ entirely and removed it from Elsa.sln. The ShellFeature marker class (WorkflowRuntimeAdminFeature) is no longer needed — the existing WorkflowsApiFeature already discovers FastEndpoints in the Workflows.Api assembly. * No consumer changes: the endpoints sit in the same routes (/admin/workflow-runtime/*) and behave identically. Note on the second architectural point ("update IShellFeature if cleaner"): the IShellFeature contract is defined in the external CShells NuGet package, not in this repo, so we cannot add a DeactivateAsync hook without an upstream CShells change. The current IHostedService.StopAsync hook continues to work correctly for the host-stop path; per-shell deactivation would require either a CShells upstream addition or a separate Elsa-owned shell-feature variant — neither lighter than what we have today. Tests: 103/103 runtime unit + 249/249 workflow integration pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(common): introduce IsFatal exception extension + apply to drain best-effort catches Per PR review feedback on the static-analyzer "Generic catch clause" comments: rather than catching ALL exceptions in best-effort drain code paths, narrow the swallow to non-fatal exceptions. Process-fatal conditions (StackOverflowException, AccessViolationException, SEHException, ThreadAbortException, OutOfMemoryException) propagate so the host's failure-fast policy can act on them, while normal failures (InvalidOperationException, IOException, etc.) continue to be logged and allowed through so a single misbehaving ingress source / activity cannot abort the overall drain. Highlights: * New `Elsa.Common.Extensions.ExceptionExtensions.IsFatal` utility: classifies fatal conditions, unwraps reflection-style wrappers (TypeInitializationException, TargetInvocationException) before classification, and treats InsufficientMemoryException (the recoverable OOM subclass) as non-fatal. * Applied as a `when (!ex.IsFatal())` filter to: - BurstHandle.Cancel (cancel callback try/catch) - DrainOrchestrator.PauseOneSourceAsync (per-source exception path) - DrainOrchestrator.TryForceStopAsync - DrainOrchestrator.ForceCancelActiveBurstsAsync (per-burst loop) - DrainOrchestrator.PersistInterruptedAsync (orphan log write, instance save, log write) - DrainOrchestrator.DrainAsync outer catch (existing `not InvalidOperationException` filter extended) - InterruptedRecoveryScan (per-instance restart loop) Tests: 7 new unit tests for IsFatal classification (fatal types, recoverable types, wrapped causes, null tolerance). Full suites: 103/103 runtime unit (incl. 14/14 in Common.UnitTests including new tests) + 249/249 workflow integration pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(workflows-runtime): integrate CShells 0.0.15 lifecycle hooks (IDrainHandler + IShellInitializer) CShells 0.0.15 ships the lifecycle framework needed for first-class per-shell graceful shutdown — IDrainHandler / IShellInitializer / IShellLifecycleSubscriber. This commit bumps the package, migrates Elsa's existing usage of the removed 0.0.14 API, and registers the runtime drain orchestrator + pause-state initializer through the new primitives. Highlights: * `ElsaShellDrainHandler : IDrainHandler` — bridges per-shell drain into `IDrainOrchestrator.DrainAsync(DrainTrigger.ShellDeactivation, ct)`. Invoked by CShells when a shell enters `ShellLifecycleState.Draining`; the drain handler's CancellationToken is signalled when the per-shell deadline elapses, so the orchestrator's own deadline-bounded protocol nests cleanly. Coexists with the host-stop `DrainOrchestratorHostedService`; the orchestrator's `DrainAsync` is idempotent — second invocations log and skip. * `InitializePauseStateShellInitializer : IShellInitializer` — replaces the IStartupTask variant in shell-aware deployments. IShellInitializer fires on EVERY shell (re)activation, including reactivations after a reload — exactly what FR-028 requires. The IStartupTask remains for IModule consumers where there is no shell platform. Migrations (CShells 0.0.14 → 0.0.15 breaking changes): * `ActivateShellTenants`: was `IShellActivatedHandler` + `IShellDeactivatingHandler`, now `IShellInitializer` + `IDrainHandler`. * `MultitenancyFeature`: registrations updated to the new transient interface, `using CShells.Hosting` → `using CShells.Lifecycle`. * `Reload/Endpoint`, `ReloadAll/Endpoint`: `IShellManager` → `IShellRegistry`, `ReloadShellAsync` → `ReloadAsync` (returns `ReloadResult` with `Error`), `ReloadAllShellsAsync` → `ReloadActiveAsync` (returns `IReadOnlyList<ReloadResult>` with per-shell errors aggregated into 503). Build + restore: * `Directory.Packages.props`: all CShells.* packages bumped to 0.0.15. * `NuGet.Config`: added `cshells-feedz` source (https://f.feedz.io/sfmskywalker/cshells/nuget/index.json) and split the package-source-mapping pattern into `CShells` (exact) + `CShells.*` (prefix). Single-pattern `CShells*` does NOT match correctly under PackageSourceMapping. Tests: 103/103 runtime unit + 249/249 workflow integration pass on the new package version. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-runtime): apply PR #7424 review feedback Consolidates the architectural fixes asked for during /review: - Extract IWorkflowRuntimeAdminService to back the four /admin/workflow-runtime endpoints with a single domain service; thin Pause/Resume/Status/Force endpoints to delegating shells. - Remove StateChanged C# event from IQuiescenceSignal (Constitution VII: no external subscribers existed; mediator was suggested as the alternative if/when it's needed). - Promote InitializePersistedStateAsync to IQuiescenceSignal, dropping the concrete-cast in both InitializePauseStateStartupTask and InitializePauseStateShellInitializer. - Invert ingress-source DI to Lazy<IEnumerable<IIngressSource>> to break the cycle through IQuiescenceSignal; ingress adapters take the signal directly via primary constructor. - Replace Guid.NewGuid().ToString("N") with IIdentityGenerator in InterruptedLogExtensions and DrainOrchestrator. - Switch admin-audit timestamps to ISystemClock in WorkflowRuntimeAdminService. - Make GracefulShutdownOptions.StimulusQueueMaxDepthWhilePaused nullable (null = unlimited). - Rename RuntimeForceRequested → RuntimeForceDrainRequested. - Apply IsFatal exception filter to drain best-effort catches. - Rename IBurstRegistry.EnumerateActive → ListActiveBursts. - Refresh "Phase X" comments to user-story (USx) references. - Delete unused IngressAttributionExtensions, IngressSourceServiceCollectionExtensions, IngressSourceRegistrationOptions. - Migrate Elsa.Shells.Api.Tests to CShells 0.0.15 IShellRegistry / ReloadResult surface. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(graceful-shutdown): apply IsFatal filter to deadline-breach test catch Aligns the test scaffolding's swallow-everything catch with the project standard introduced in c00eee80c so the analyzer no longer flags the bare `catch` clause. The semantics are unchanged — non-fatal exceptions (OCE, TimeoutException, workflow exceptions) are still acceptable test outcomes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): close IngressSourceRegistry first-access race + spelling Replaces the non-atomic `_entries.Count > 0` early-return guard in EnsureMaterialized with a double-checked lock against a volatile `_materialized` flag, so concurrent first callers can no longer both iterate the source factory and crash one of them with a "Duplicate ingress source registration" InvalidOperationException. Adds a regression test that launches 16 readers behind a TaskCompletionSource gate and asserts every reader observes the full source set without throwing. Also flips the British spellings introduced in this PR's scope to American English (materialize/behavior) — project convention going forward. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(constitution): require American English for new code (v1.0.1) Adds a "Spelling & language" bullet under principle III (Convention-Driven Design): every newly-introduced symbol, comment, identifier, error message, XML doc, commit message, and Speckit artifact uses American English. Established public API symbols (e.g. WorkflowSubStatus.Cancelled) are not renamed retroactively. PATCH bump because this is a clarification of an existing principle, not a new principle. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Add Greploop skill and workflow for GitLab, GitHub, and Perforce integration - Introduced Greploop, an iterative optimization and review workflow for GitLab MRs, GitHub PRs, and Perforce changelists. - Added API and GraphQL references for fetching and resolving review skill. * Remove GenerateWorkflowVariableAccessorsTests; redundant ExpandoObject type check in handlers * Potential fix for pull request finding 'CodeQL / Untrusted Checkout TOCTOU' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(graceful-shutdown): apply PR #7424 review feedback round 2 Two P1 findings from Greptile: 1. DrainOrchestrator.ForceCancelActiveBurstsAsync was sequential — each burst was cancelled, awaited up to ForceCancelSettleTimeout (2 s), and persisted before the next burst's Cancel() ran. Total wall time was O(N × 2 s) and bursts 2..N kept executing at full speed during prior bursts' settle waits, defeating the intent of force-cancel under concurrency. Refactored to three phases: - Phase A — cancel every handle synchronously (cheap CTS.Cancel calls) so all runners observe cancellation simultaneously. - Phase B — await every Disposed signal in parallel under a single shared ForceCancelSettleTimeout. Total wall time bounded regardless of N. - Phase C — persist Interrupted for each handle sequentially (keeps DbContext usage single-threaded; per-handle work is small). Per-phase failures are caught with !ex.IsFatal() and logged so a single misbehaving handle doesn't abort the rest of the batch. 2. ShellFeatures/WorkflowRuntimeFeature.ConfigureServices was missing the IWorkflowRuntimeAdminService registration that Features/WorkflowRuntimeFeature already had. Any CShells deployment that includes the Pause / Resume / Status / Force admin endpoints (in Elsa.Workflows.Api) would throw InvalidOperationException at endpoint construction. Added the singleton alongside the other graceful-shutdown registrations with a comment pointing out the symmetry with the IModule path. 17/17 graceful-shutdown integration tests pass; 103/103 runtime unit tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Untrusted Checkout TOCTOU' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(ci): harden greploop.yml against CodeQL Actions findings CodeQL flagged 12 findings on .github/workflows/greploop.yml after the prior commit (c08183a3c) addressed an earlier round. Two distinct issue classes remain: 1. Code injection (× ~10): step-output values (steps.pr_head.outputs.head_sha / head_repo_owner / head_repo_name / head_ref) and inputs.pr_number were interpolated directly into shell `run:` blocks via `${{ ... }}`. Because PR author controls the branch name and the manual-dispatch input, those values can carry shell metacharacters. Standard fix: route every such interpolation through an `env:` block on the step, then reference $VAR inside the script. Applied to the Resolve, Resolve PR head metadata, and Checkout PR branch steps. 2. Untrusted Checkout TOCTOU + Checkout of untrusted code in trusted context: the workflow runs on `issue_comment` (a privileged trigger) and checks out PR-author code. Mitigations stacked here: - Author-association gate already restricts the trigger to OWNER / MEMBER / COLLABORATOR (existing). - Step-output values now travel via env vars (above). - Resolve step rejects pr_number that isn't ^[0-9]{1,10}$ — so downstream `gh pr view` and the prompt argument can't be hijacked. - Checkout step now validates HEAD_SHA matches ^[0-9a-f]{40}$ and the repo owner/name match ^[A-Za-z0-9_.-]+$ before either reaches a URL or a git command. - Existing TOCTOU guard preserved: re-fetch head SHA at checkout time and abort if it changed since the initial resolve. These match the canonical "Securing your GitHub Actions workflows" patterns recommended by CodeQL. No functional change to greploop's runtime behaviour. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(graceful-shutdown): drop [SingleNodeTask] from InitializePauseStateStartupTask Greptile P1: [SingleNodeTask] gates the task to a single cluster winner via distributed lock, but IQuiescenceSignal is a singleton scoped to each node's DI container — each node holds its own in-memory QuiescenceState. With the attribute, only the winning node restored the persisted pause; every other node started with QuiescenceReason.None and accepted new work, silently defeating PausePersistence = AcrossReactivations. Removed [SingleNodeTask] (and the corresponding using) so the task runs on every node. Expanded the doc <remarks> to call out the per-node requirement and point at the shell-aware counterpart (InitializePauseStateShellInitializer) which is correctly per-node by virtue of being an IShellInitializer. 17/17 graceful-shutdown integration tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Code injection' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * chore(deps): bump CShells 0.0.15 → 0.0.17 0.0.17 ships our blueprint-aware-routing PR (valence-works/cshells#93) plus four follow-up fixes the maintainer added on top: - e56ebd8 — PreWarmShells removed entirely; ShellMiddleware now does cold-start endpoint matching (re-runs endpoint resolution after lazy activation so the very first request to a cold shell hits its endpoint). - cbe5ee2 — GetCandidateSnapshot returns a bounded ShellRouteCandidateSnapshot with accurate total counts; sensitive-data redaction in routing logs; DefaultShellRouteIndex implements IDisposable. - cd7d4f5 — Last-good snapshot served on rebuild failure (the deferred Copilot review concern); root-path fallback when path-by-name misses. - c3679d9 — Cold-start endpoint matching respects inline route constraints; path-name convention tightening; dead duplicate-detection cleanup. Net effect for elsa-core: - Cold blueprints serve their first request via lazy activation, with endpoints correctly resolved post-activation. - Reloaded shells re-activate and serve on the next matched request. - Non-name-mode routing keeps serving the previous snapshot during a transient blueprint-provider outage. - No need to call PreWarmShells from Elsa.ModularServer.Web — removed. The only API removal that touches elsa-core is PreWarmShells. No code references IShellRouteIndex / ShellRouteCriteria / GetCandidateSnapshot directly, so the API-shape changes in cbe5ee2 don't ripple here. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(graceful-shutdown): persist Interrupted under non-drain bounded token Greptile P? finding on the prior force-cancel two-phase fix: Phase B's inner catch on OperationCanceledException ("drain CT fired — proceed to persist anyway") was a lie in practice. Phase C immediately passed the same already-cancelled drain token into PersistInterruptedAsync; the first DB call (instanceStore.FindAsync) observed the cancellation and threw OperationCanceledException; the outer non-fatal Exception filter swallowed it and only logged an error. Net effect: on host shutdown deadline breach, every burst after cancellation could fail to be persisted as Interrupted, leaving instances in an unrecovered executing state. Phase C now creates a per-handle CancellationTokenSource bounded to a new PersistInterruptedTimeout (5 s) that is NOT linked to the drain CT. Each persist gets up to 5 s to land the row update + forensic log entry even after the drain CT has fired. The bound prevents a stuck DB from hanging shutdown indefinitely (per-handle worst case is small; total Phase C upper bound is N × 5 s, but typical persists are millisecond scale). Comment expanded to call out why the persist token is independent of the drain token, so the rationale doesn't drift again. 17/17 graceful-shutdown integration tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): extract PassiveIngressSource base class The three IIngressSource implementations that ship with this PR (InternalBookmarkQueueIngressSource, HttpTriggerIngressSource, ScheduledTriggerIngressSource) were ~25 lines each and ~22 lines of those were verbatim copies of each other: - ctor signature `(IQuiescenceSignal signal)` - `PauseTimeout => TimeSpan.FromMilliseconds(50)` - `CurrentState => signal.IsAcceptingNewWork ? Running : Paused` - `PauseAsync` / `ResumeAsync` returning `ValueTask.CompletedTask` The shared trait is that none of them does any work at pause time — the actual pause enforcement lives in another layer (`HttpWorkflowsMiddleware` short-circuits to 503, `BookmarkQueueProcessor` consults the signal at the top of each invocation, scheduled triggers dispatch through the bookmark queue and inherit that behaviour transitively). The IIngressSource adapter is purely diagnostic: it makes the source visible in `DrainOutcome.Sources` and the admin status endpoint. Extracted that pattern into `PassiveIngressSource` (abstract base in `Elsa.Workflows.Runtime.IngressSources`). Subclasses now provide only `Name`; `PauseTimeout` is `virtual` with a 50 ms default; everything else is fixed by the base. The three concretes drop from ~25 lines to ~12 lines each. The base's XML `<remarks>` calls out when to use it ("your component already cooperates with IQuiescenceSignal at its hot path") and when to implement IIngressSource directly ("the source owns concrete pause/resume behaviour — e.g. a message-queue consumer that calls Pause() on its underlying client"), so future contributors don't mis-extend the base for sources that need real work at pause time. No behavioural change. 18 graceful-shutdown integration + 39 runtime unit tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): align IIngressSource name to singular The three IIngressSource names were inconsistent: http.trigger (singular) internal.bookmark-queue-worker (singular) scheduling.triggers (PLURAL — outlier) The plural slipped in when addressing Greptile's earlier comment to avoid `scheduling.cron` (which would imply Cron-only coverage). The right move was to pick a generic word and stay singular like the rest of the suite — the suite's mental model is "the X source", one instance per registry slot, regardless of how many triggers or items it dispatches internally. Renamed to `scheduling.trigger`. The `<remarks>` block keeps the "covers Cron, Timer, StartAt, Delay" explanation and now also explicitly notes the singular convention so future contributors don't re-pluralize. Zero test fallout — the literal "scheduling.triggers" only appeared in the source file itself. Tests of the other two sources all use singular forms. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(runtime-admin): rename Force endpoint to ForceDrain "Force" alone is meaningless out of context — force what? — and it sits oddly next to the verb-named siblings Pause / Resume / Status. The matching admin service method is already IWorkflowRuntimeAdminService. ForceDrainAsync, so ForceDrain is the natural pair. Renamed: - src/modules/Elsa.Workflows.Api/Endpoints/RuntimeAdmin/Force/ → ForceDrain/ - namespace ...Endpoints.RuntimeAdmin.Force → ...ForceDrain - class ForceEndpoint → ForceDrainEndpoint - class ForceRequest → ForceDrainRequest - class ForceResponse → ForceDrainResponse - route POST /admin/workflow-runtime/force → /force-drain Zero external references — no tests, docs, or OpenAPI clients used the old symbols or the old route literal, so this is a contained pre-ship rename. Directory move went through `git mv` so commit history follows the file. 17/17 graceful-shutdown integration tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): correct formatting in checkout step within greploop.yml Adjusted indentation of environment variables in the checkout step for improved consistency and readability. * fix(ci): repair malformed Checkout repository step in greploop.yml The step accumulated stray env keys, an extra `uses:`, and bash commands that didn't belong inside it (line 65 onward), causing a YAML parse error on push. The valid structure has two distinct checkout steps: - Checkout repository : actions/checkout@v4 with fetch-depth: 0 - Checkout PR branch : env: + run: with SHA validation + git fetch + git checkout --detach The PR-branch step (line 78+) was already correct and unchanged. This fix restores the first step to its intended single-purpose shape (just checks out the workflow file's commit so the greploop skill is on disk before the run-greploop step uses it). No functional change to runtime behaviour or to the security posture established in the prior hardening commit (0db4ca23e). The PR-branch checkout still validates HEAD_SHA / HEAD_REPO_OWNER / HEAD_REPO_NAME shape before they reach a URL or git command. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): remove literal `${{ }}` from greploop.yml comment GitHub Actions parses `${{ ... }}` workflow expressions across the entire YAML file, including inside `run:` script comments. The comment that explained the env-var hardening pattern contained the literal sequence `${{ }}` (with a space, intended as an English-language description), which the expression parser rejected as "An expression was expected" (line 81 col 14). Reworded the comment to describe the substitution form in prose without the literal token sequence. Functional behaviour unchanged. * refactor(graceful-shutdown): rename burst → execution cycle The graceful-shutdown work introduced "burst of execution" as a first-class domain concept. The term arrived without rationale and isn't standard in the workflow-engine domain. Renamed to "execution cycle" — reads more naturally as the loop-with-commit unit, is more idiomatic in workflow vocabulary, pairs cleanly with the existing WorkflowExecutionContext, and avoids collisions with Elsa's existing terms (Run, Execution, Dispatch, Invocation, Stimulus, Step). Renamed types - IBurstRegistry → IExecutionCycleRegistry - BurstRegistry → ExecutionCycleRegistry - BurstHandle → ExecutionCycleHandle - BurstTrackingMiddleware → ExecutionCycleTrackingMiddleware - BurstAwareCommitStateHandler → ExecutionCycleAwareCommitStateHandler Renamed members - BeginBurst → BeginCycle - ListActiveBursts → ListActiveCycles - BurstHandleKey constant + value → ExecutionCycleHandleKey - ActiveBurstCount (IQuiescenceSignal, RuntimeAdminStatus, StatusResponse) → ActiveExecutionCycleCount - WaitForBurstsAsync (private) → WaitForCyclesAsync - ForceCancelActiveBurstsAsync (priv) → ForceCancelActiveCyclesAsync - UseBurstTracking → UseExecutionCycleTracking - DrainOutcomeDto.BurstsForceCancelledCount → ExecutionCyclesForceCancelledCount - _burstRegistry / burstRegistry → _cycleRegistry / cycleRegistry Backwards-compatibility preservation (the only persisted JSON key) - WorkflowInterruptedPayload.BurstDuration property → ExecutionCycleDuration with [JsonPropertyName("BurstDuration")] so the persisted JSON wire key stays "BurstDuration" forever. Pre-merge testers' log records still deserialise correctly. The contract test on WorkflowInterruptedPayloadContractTests still asserts the wire key "BurstDuration" appears in the serialised JSON — confirms the guarantee is enforced. Other unstructured surfaces - WorkflowExecutionLogRecord.Message text "Workflow burst was force- cancelled..." now says "Workflow execution cycle was force-cancelled ..." for new records. Old rows keep their old text — purely cosmetic free-text field. - Structured log placeholder {BurstId} in DrainOrchestrator log lines → {ExecutionCycleId}. - Lowercase prose / XML doc comments updated throughout. Test renames - BurstRegistryTests → ExecutionCycleRegistryTests - BurstTrackingMiddlewareTests → ExecutionCycleTrackingMiddlewareTests - Test method names + DisplayName strings updated. Spec docs (specs/002-graceful-shutdown/) updated to match the new vocabulary; the historical task records in tasks.md keep the old names as-is to preserve the audit trail of what was originally built. Verification - dotnet build: clean across net8.0 / net9.0 / net10.0. - 103/103 Elsa.Workflows.Runtime.UnitTests pass. - 17/17 GracefulShutdown integration tests pass. - 4/4 WorkflowInterruptedPayloadContractTests pass — confirms the "BurstDuration" JSON wire-key preservation is intact. No changes to migrations or DB column names — confirmed via grep. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): give greploop.yml gh-cli a repo context before checkout The "Resolve PR head metadata" step runs `gh pr view` before actions/checkout, so there is no `.git` directory and gh's "current repo" detection fails with `fatal: not a git repository`. The prior commits to this file masked the runtime failure because the workflow itself was YAML-invalid — once it became valid, the workflow_dispatch trigger surfaced this real-world execution bug. Set GH_REPO=${{ github.repository }} on both `gh pr view` steps. The gh CLI honours GH_REPO as an explicit repo override, so it no longer needs git context. Same fix on the "Checkout PR branch" validation call which also uses gh pr view before the manual fetch. The error reported as "Invalid workflow file: ... (Line 81 Col 14)" on PR #7424 was stale from commit ed580682a (which had the bad comment with literal `${{ }}`); commit 3a8ad0d08 fixed the YAML, but because greploop's `if:` condition only matches workflow_dispatch / issue_comment events, push events on later commits were skipped without re-running validation, so the GitHub UI kept showing the old error. A workflow_dispatch run on the current SHA now passes validation and reaches "Resolve PR head metadata", which is what this commit fixes. * fix(graceful-shutdown): wire IngressPauseTimeout option to drain orchestrator GracefulShutdownOptions.IngressPauseTimeout was documented as "Default per-ingress-source pause timeout" but DrainOrchestrator.PauseOneSourceAsync read source.PauseTimeout directly and never consulted the option. The configured value was silently ignored — operators who set GracefulShutdownOptions:IngressPauseTimeout = 10s were getting whatever each source's hardcoded value was (50 ms for the three PassiveIngressSource subclasses we ship), with no way to tune it globally. Precedence (per the spec's intent of "overridable at registration and by configuration"): 1. Per-source positive value wins (source.PauseTimeout > Zero). 2. Otherwise fall back to the configured GracefulShutdownOptions. IngressPauseTimeout default. 3. Resolved value is capped at the overall drain deadline so a single misbehaving source cannot exceed the host's shutdown budget. 4. 1 ms safety floor remains so a misconfigured zero default still produces a non-zero CancelAfter. Changes: - DrainOrchestrator.PauseOneSourceAsync — adds the precedence above with a comment block explaining each step. - IIngressSource.PauseTimeout — XML doc clarifies the Zero-defers-to- config semantics. - GracefulShutdownOptions.IngressPauseTimeout — XML doc says it's the fallback when the source returns Zero; <remarks> spells out the precedence and the overall-deadline cap. - PassiveIngressSource.PauseTimeout — virtual property now returns Zero (was 50 ms). The three shipped subclasses (HttpTriggerIngressSource, ScheduledTriggerIngressSource, InternalBookmarkQueueIngressSource) consequently defer to the configured default — flipping the wire-up bug from "configured value silently ignored" to "configured value honoured by default for passive sources". Passive subclasses that want a specific value can still override. 103/103 runtime unit + 17/17 graceful-shutdown integration tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): rename IInterruptedRecoveryScan → IInterruptedRecoveryScanner The interface had a single verb-method (`ScanAndRequeueAsync`) and its XML doc described what it *does* ("Scans the workflow instance store for instances..."). That's an agent role — a scanner that performs a scan — but the noun-shaped name `IInterruptedRecoveryScan` read as "the scan itself", which is misleading because the scan results / event are not first-class types in the codebase. Renamed to `IInterruptedRecoveryScanner` / `InterruptedRecoveryScanner` to match the existing `-er` convention in this codebase (Restarter, Generator, Resolver, etc.). The method stays `ScanAndRequeueAsync` — the scanner *performs* a scan-and-requeue. Also renamed the constructor parameter `scan` → `scanner` in RecoverInterruptedWorkflowsStartupTask, and the local variable `scan` → `scanner` in InterruptedRecoveryIntegrationTests, so the "scanner does the scan" mental model is consistent throughout. Surface impact (all internal — no API or persistence touch points): - 2 source files renamed via git mv (interface + implementation) - 1 test file renamed (InterruptedRecoveryScanTests → ScannerTests) - DI registrations in both Features/ and ShellFeatures/ WorkflowRuntimeFeature - 1 startup-task constructor parameter - Spec doc references under specs/002-graceful-shutdown/ 103/103 runtime unit + 17/17 graceful-shutdown integration tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): extract DrainTriggerExecutor ElsaShellDrainHandler (CShells IDrainHandler) and DrainOrchestratorHostedService (.NET IHostedService.StopAsync) inlined near-identical try/catch/log shapes around IDrainOrchestrator.DrainAsync: - call DrainAsync(<trigger>, ct) - branch on outcome: DeadlineExceeded/AbortedByUnhandledException → Warning, otherwise Information - catch InvalidOperationException (parallel-drain rejected by the orchestrator) → log Information and swallow The two had already drifted: host-stop's success log omitted the paused/waited durations the shell-handler version included, and the "skipped" message disagreed on the trigger label ("Host-stop drain skipped" vs "Shell drain skipped"). Centralised the shape in a small internal static helper so the two — and any future trigger source — stay uniform. Both call sites collapse to a single line. Net diff drops 22 lines from the two consumers and adds a 25-line helper that they both delegate to. The unified log copy now consistently includes paused/waited durations on the success path and uses the caller-supplied contextLabel ("Shell drain", "Graceful drain") in all three messages so operators can attribute log entries by trigger source. Files: - src/modules/Elsa.Workflows.Runtime/Services/DrainTriggerExecutor.cs (new) - src/modules/Elsa.Workflows.Runtime/Lifecycle/ElsaShellDrainHandler.cs - src/modules/Elsa.Workflows.Runtime/HostedServices/DrainOrchestratorHostedService.cs 103/103 runtime unit + 17/17 graceful-shutdown integration tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-runtime): drop redundant DrainOrchestratorHostedService from CShells path In CShells deployments, host stop already drives drain via CShellsStartupHostedService → IDrainHandler → ElsaShellDrainHandler, scoped per shell (FR-027). The additional .AddHostedService<DrainOrchestratorHostedService>() in ShellFeatures/WorkflowRuntimeFeature was firing a second non-force DrainAsync that the orchestrator rejected with InvalidOperationException — silently swallowed by DrainTriggerExecutor, but logged on every host stop and semantically muddled (IHostedService is host-level, not per-shell). Keep the registration on the IModule path (Features/WorkflowRuntimeFeature) where there is no shell platform and host-stop is the only available drain trigger. Update ElsaShellDrainHandler XML docs to reflect the now-clean single-trigger model in CShells. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): always dispose ExecutionCycleHandle in tracking middleware Previously the success path relied on ExecutionCycleAwareCommitStateHandler to dispose the handle after the runner's commit completed. If a custom dispatcher or test double exited the pipeline without invoking commit (by design or by accident), the handle stayed registered, IExecutionCycleRegistry.ActiveCount never reached zero, and drain spun in WaitForExecutionCyclesAsync until the deadline fired — incorrectly force-cancelling instances that had already finished cleanly. Collapse the existing try/catch(rethrow) into try/finally so the middleware itself disposes the handle for both exception and commit-elided paths. Disposal remains idempotent via the ExecutionCycleHandle._disposed Interlocked guard, so the normal-path dispose by ExecutionCycleAwareCommitStateHandler is a harmless no-op. Adds an integration regression test that drives the middleware with a stub Next that returns without invoking commit and asserts ActiveCount returns to zero. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): serialize QuiescenceSignal pause-state persistence Both PauseAsync and ResumeAsync used to release the inner lock before issuing the persistence I/O. A rapid Pause → Resume sequence could leave the persisted state inconsistent: PauseAsync's slow SaveAsync could land AFTER ResumeAsync's DeleteAsync, leaving the key present in the store while in-memory state was None. On host restart, InitializePersistedStateAsync would find the stale key and start the runtime in the paused state the operator had already cancelled. Introduce a dedicated SemaphoreSlim that serializes persistence I/O, with each I/O re-reading the live in-memory state inside the semaphore. N racing Pause/Resume calls now produce N serialized writes, each reflecting the most recent in-memory transition — so the final persisted state always matches final in-memory state. Adds a regression test that gates SaveAsync, races a Resume behind it, and asserts the store is empty after both complete. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(workflows-runtime): trim verbose comment in ExecutionCycleTrackingMiddleware finally block Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-runtime): use 'using var' for ExecutionCycleHandle in tracking middleware Replace the explicit try/finally that only existed to call handle.Dispose() with a `using var` declaration — identical semantics (compiler-emitted finally with idempotent dispose), more idiomatic. The regression test HandleReleasedWhenCommitIsElided continues to validate the leak-free property. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime,api): proper 409 conflict shape + shell-scoped pause-persistence key ForceDrain endpoint: the 409 path returned `new ForceDrainResponse()` whose `Outcome` was null at runtime despite the `= null!` annotation, so any strongly-typed client deserializing the conflict body and reading `Outcome.OverallResult` got an NRE. Switch to the existing `ConflictResponse` shape with `Code = "DrainInProgress"` and the current runtime status. Routed via HttpContext.Response.WriteAsJsonAsync because Send.ResponseAsync is constrained to the endpoint's TResponse and cannot send a sibling DTO. QuiescenceSignal persistence key: the DI-registered `IQuiescenceSignal` was constructed with `shellName = null` (DI doesn't inject `string?` defaults), so every shell shared the key `elsa.quiescence.pause.default`. In a CShells multi-shell deployment under PausePersistencePolicy.AcrossReactivations this caused cross-shell contamination — pausing shell A would re-pause shell B on its next activation. Replace the simple AddSingleton<IQuiescenceSignal,...> registration in ShellFeatures/WorkflowRuntimeFeature with a factory that injects `CShells.ShellSettings` and forwards `Settings.Id` as the shell name. The IModule registration is unchanged (no shell platform; null shellName remains correct there). Adds a unit regression test that two QuiescenceSignal instances with different shellNames write to disjoint persistence keys and never to the legacy "default" key. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(quiescence): use TryGetValue for ContainsKey+indexer assertions Combines existence check and value retrieval into a single dictionary lookup, addressing the code-quality bot's repeated suggestion. No behavior change — both PauseWritesKey and PersistenceKeyIncludesShellName still assert the same keys exist with the same content. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Update specs/002-graceful-shutdown/contracts/admin-endpoints.md Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * fix(workflows-runtime): decouple QuiescenceSignal persistence from caller cancellation PersistAsync used to forward the caller's CancellationToken to both _persistenceMutex.WaitAsync and the store I/O. If an HTTP request was cancelled between the in-memory transition (already committed under _sync) and the persistence call, the I/O was silently skipped — leaving AdministrativePause set in memory with no persisted record. The idempotent fast-path on subsequent PauseAsync calls (transitioned == false) meant no retry would happen, so on host restart InitializePersistedStateAsync would find no key and the runtime would come back unpaused, defeating PausePersistencePolicy.AcrossReactivations. Drop the parameter from PersistAsync entirely; use CancellationToken.None for both the semaphore wait and the store I/O. The in-memory transition is already committed by the time PersistAsync runs, so persistence must complete to keep the store consistent with memory. The public PauseAsync/ResumeAsync methods still accept a CancellationToken (interface contract) — it just no longer reaches the persistence layer. Adds a regression test that calls PauseAsync with a pre-cancelled token and asserts both in-memory pause and the persisted key land correctly. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime,api,docs): apply Copilot review feedback batch Code: - DrainOrchestrator.TryForceStopAsync now bounds force-stop with the *remaining* drain budget (deadlineAt - now), not the full overall TimeSpan. A per-source pause that already burned the shutdown window can no longer get another full deadline's worth of force-stop runway. - DrainOrchestrator catch filter narrowed: drop `ex is not InvalidOperationException` exclusion. The "drain already in progress / completed" IOEs are thrown outside the protocol's try block, so they bubble out without entering this handler. Any IOE that lands here is incidental (e.g., from a store inside the drain) and should now be captured into the outcome rather than escaping the whole drain. - ResumeEndpoint 409 path now returns the discriminated ConflictResponse shape (matching ForceDrain) instead of a plain StatusResponse. Routed via HttpContext.Response.WriteAsJsonAsync because Send.ResponseAsync is constrained to TResponse. - Conflict codes aligned to kebab-case across both endpoints to match the contract spec (`runtime-draining` and `drain-in-progress`). Spelling sweep — American English per constitution v1.0.1 III: - DrainOrchestrator.cs: "serialised" → "serialized" - WorkflowInterruptedPayload.cs: "serialised" / "deserialise" → "serialized" / "deserialize" - PassiveIngressSource.cs: "behaviour" → "behavior" - DeadlineBreachEndToEndTests.cs: "serialisable" → "serializable" - specs/002-graceful-shutdown/quickstart.md: "behaviour" → "behavior" - specs/002-graceful-shutdown/checklists/requirements.md: "behaviour" → "behavior" Doc/contract alignment: - quickstart.md: force route corrected from /force to /force-drain. - quiescence-signal.md: removed StateChanged event from contract (interface doesn't define it); corrected persistence section to describe InitializePersistedStateAsync via shell initializer / startup task rather than constructor read; added the per-shell key discriminator. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): correct DI lifetimes for ExecutionCycleTrackingMiddleware and WorkflowRuntimeAdminService Two strict-DI-validation failures surfaced in tests using BuildServiceProvider with validate-on-build: 1. ExecutionCycleTrackingMiddleware was registered as AddSingleton<>, but its constructor takes WorkflowMiddlewareDelegate next — supplied by the workflow execution pipeline builder via UseMiddleware<>(), not from DI. The registration was both unused (no consumer resolves it through the container) and broken (DI fails to construct it because next is unregistered). Removing both registrations. 2. IWorkflowRuntimeAdminService was registered as AddSingleton<> but depends on the scoped INotificationSender (mediator) — captive-dependency violation. All consumers (Pause/Resume/ Status/ForceDrain endpoints) are FastEndpoints, which are scoped per request, so AddScoped is the correct alignment. The other deps (IQuiescenceSignal / IIngressSourceRegistry / IDrainOrchestrator / ISystemClock) are singletons and resolve fine from a scoped consumer. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): restore commit-handler-only success disposal + true Cancel idempotency Two issues raised by Copilot's latest review on commit 32a9c0519: 1. ExecutionCycleTrackingMiddleware was disposing the handle at the end of InvokeAsync (via `using var`), but WorkflowRunner runs commit AFTER the pipeline returns (WorkflowRunner.cs:235). That meant the handle was disposed BEFORE the runner's terminal commit, and the drain orchestrator's `await handle.Disposed` would unblock too early — reintroducing the runner-clobber race the original design protected against (see ExecutionCycleAwareCommitStateHandler XML doc). Revert to the original shape: only dispose on exception path. ExecutionCycleAwareCommitStateHandler remains the SOLE success-path disposer, running in its finally block AFTER the inner commit lands. The earlier "leak when commit is elided" concern was a non-issue in production (the standard runner always commits); the buggy `HandleReleasedWhenCommitIsElided` test that specified the wrong contract is removed. The existing `ActiveCountReturnsToZero` test (which uses the real runner end-to-end) already verifies success-path disposal. 2. ExecutionCycleHandle.Cancel() was documented as idempotent but only short-circuited via the _disposed flag. Repeated Cancel() calls before Dispose could trigger the cancel callback multiple times — easy to accidentally fire non-idempotent cancellation side effects more than once. Add an Interlocked _cancelled guard so callback + CTS cancellation run at most once. Existing test that documented the leaky behavior is updated to assert the now-truly-idempotent contract. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Update logging levels and remove unused features in appsettings files * refactor(workflows-runtime): improve graceful shutdown options handling and cleanup solution Refactor the handling of `GracefulShutdownOptions` to ensure options are applied correctly without directly invoking the delegate. Update DI registrations to use appropriate lifetimes and remove redundant wrapper services. Additionally, clean up the solution by removing unused projects and documentation folders. * feat(identity, workflows-runtime): add validation for identity and graceful shutdown options Introduce validation capabilities for `IdentityTokenOptions` and `GracefulShutdownOptions`. Implement extension methods for option validation, enhance service registration, and add unit tests to ensure configurations are validated at startup. Update solution to include new unit test projects. * update(docs): clarify shutdown log message expectations and levels in quickstart.md Optimize explanation of expected log message sequence during graceful shutdown and specify logging levels. * docs: amend constitution to v1.1.0 (SRP, DRY, KISS, conciseness under Principle VII) * refactor(multitenancy): rename and restructure TenantTaskManager to TenantTaskLifecycleCoordinator Rename `TenantTaskManager` to `TenantTaskLifecycleCoordinator` and relocate to a new directory structure, enhancing code organization and test consistency. Retain functional behaviors with no logic alterations. Update unit tests to reflect the naming changes, ensuring consistency with the refactored code structure. * Update logging levels and dependencies - Set default logging level to Debug in appsettings.Development.json - Add missing using directives for Elsa workflows management and runtime features - Update CShells package versions to 0.0.18-preview.104 --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-05-02 17:27:08 +00:00
var ex = (Exception)System.Runtime.Serialization.FormatterServices.GetUninitializedObject(exceptionType);
#pragma warning restore SYSLIB0050
Graceful shutdown for the workflow runtime (drain, pause, recover) (#7424) * feat(workflows-runtime): add quiescence machinery foundation for graceful shutdown Introduces the container-scoped quiescence signal, ingress-source contract, burst registry, and the Interrupted workflow sub-status — the foundational primitives the drain orchestrator and admin endpoints will build on. No behaviour change yet: workflows continue to run and shut down exactly as before. The new types are registered but no host-stop or pause path drives them. Highlights: * IQuiescenceSignal — composable Drain + AdministrativePause flags; forward-only drain, reversible pause, idempotent transitions, optional persistence via IKeyValueStore. * IIngressSource + IForceStoppable — uniform contract for components that inject external events (HTTP, schedulers, message consumers, internal workers, third-party modules); IIngressSourceRegistry collects and surfaces their states. * IBurstRegistry — atomic counter for in-flight workflow execution bursts, with per-burst ingress attribution and FR-018 inconsistency detection (a source claiming Paused but starting bursts is flipped to PauseFailed). * WorkflowSubStatus.Interrupted — new value distinct from Suspended, Cancelled, Faulted; semantics: "last burst force-cancelled by graceful drain; resumable on next runtime generation". Mirrored on the API client enum. * GracefulShutdownOptions — drain deadline, per-source pause timeout, stimulus-queue back-pressure policy, pause-persistence policy. Configurable via UseWorkflowRuntime(...).ConfigureGracefulShutdown(...). * PermissionNames.ManageWorkflowRuntime — single permission for the forthcoming admin pause/resume/status/force endpoints. Implements 31 of 77 tasks for the graceful-shutdown feature (specs/002-graceful-shutdown). Subsequent commits add the drain orchestrator (US1 / MVP), Interrupted recovery scan (US3), admin endpoints (US2), and first-party ingress adapters. Tests: 25 new xUnit unit tests; 100/100 runtime unit tests pass; all existing tests continue to pass on net8.0/net9.0/net10.0. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(workflows-runtime): add drain orchestrator + host-stop integration (US1, MVP) When the host receives a stop signal (SIGTERM, Ctrl+C, orchestrator rollout), the runtime now drains gracefully: ingress sources are paused in parallel, in-flight workflow bursts run to their next natural persistence boundary within a configurable deadline, and any burst that breaches the deadline is force-cancelled and persisted with the Interrupted sub-status plus a forensic WorkflowInterrupted log entry. This is the MVP — without the activation-time recovery scan (PR 3) the existing timeout-based RestartInterruptedWorkflowsTask still picks up Interrupted instances, just on its periodic cadence. No regression in that recovery path (SC-008). Highlights: * IDrainOrchestrator + DrainOrchestrator — protocol per the contract: BeginDrainAsync → parallel ingress pause with per-source timeouts + IForceStoppable escalation → poll BurstRegistry.ActiveCount until zero or deadline → on breach iterate live handles, cancel, persist Interrupted, write log entry. All exceptions are captured into the returned DrainOutcome; only second-invocation throws. * Deadline clamping: effective deadline is min(GracefulShutdownOptions. DrainDeadline, HostOptions.ShutdownTimeout - 500ms safety epsilon), so the runtime never outlives its host process. * DrainOrchestratorHostedService — IHostedService.StopAsync wakes the orchestrator on host stop. Registered AFTER the heartbeat (Elsa.Hosting.Management) so reverse-order shutdown keeps the heartbeat alive throughout drain. Prevents sibling-node crash recovery from false-positive-recovering instances we are gracefully handling here (FR-029). * BurstTrackingMiddleware — workflow-execution-pipeline middleware that registers a BurstHandle for the lifetime of every burst. All nine IWorkflowRunner.RunAsync overloads ultimately funnel into pipeline.ExecuteAsync(context), so this single middleware covers the three "burst choke points" the spec references without nine separate decorators. Added to UseDefaultPipeline(). * Ingress attribution: optional IngressSourceName property on DispatchStimulusRequest, DispatchWorkflowDefinitionRequest, and DispatchWorkflowInstanceRequest. Adapters set it; the middleware reads it via WorkflowExecutionContext.TransientProperties (helpers in IngressAttributionExtensions). The BurstRegistry uses the name to detect the FR-018 invariant violation — a source that reports Paused but starts a burst is flipped to PauseFailed. * InterruptedLogExtensions — the LogWorkflowInterruptedAsync helper that the orchestrator calls when persisting the forensic record. Tests: 11 new unit tests (DrainOrchestrator parallel-pause + wait-for-bursts + idempotency + persistence-failure path); 5 new integration tests (full DI graph resolves, burst-tracking middleware registers handles end-to-end, no-op drain returns CompletedWithinDeadline). 100/100 runtime unit tests pass; all existing tests continue to pass on net8.0/net9.0/net10.0. Implements 14 of 77 tasks (T032–T045). Subsequent commits add Interrupted recovery scan (US3), admin endpoints (US2), and first-party ingress adapters. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(admin-endpoints): add admin endpoints for workflow runtime control Introduced admin endpoints to manage workflow runtime: `/pause`, `/resume`, `/status`, and `/force` with full authentication and audit logging. Integrated idempotency checks and error handling to ensure reliable runtime control. Added corresponding integration tests for verification. * fix(workflows-runtime): propagate drain cancellation into running workflows + initialize persisted pause + fix log-comment Addresses three findings from the PR #7424 code review: 1. **HIGH — Cancellation now propagates into the running workflow.** `BurstHandle.Cancel()` previously cancelled only its own linked CTS, which the workflow runner never observes (the runner reads from `WorkflowExecutionContext.CancellationToken`, captured at context construction and not part of the linked chain). On deadline breach the orchestrator would persist `Interrupted`, but the workflow continued executing and could overwrite the sub-status with whatever terminal state it eventually reached. Fix: `BurstHandle` accepts an optional cancel callback at construction. `BurstTrackingMiddleware` wires it to `context.Cancel()` so the burst's cancellation triggers the workflow's own cancellation chain — the workflow transitions to `Cancelled` and stops scheduling new activities. The orchestrator then awaits `BurstHandle.Disposed` (with a 2 s settle timeout) before persisting `Interrupted`, ensuring the runner's terminal commit completes BEFORE the orchestrator overwrites the sub-status. Race resolved. The settle timeout is bounded so a non-cancellable activity (genuinely pathological case) does not block drain — on timeout the orchestrator logs and proceeds, accepting the runner-clobber for that one instance, which the existing timeout-based RestartInterruptedWorkflows recovery picks up afterwards. 2. **MEDIUM — Pause persistence is now actually wired.** `QuiescenceSignal.InitializePersistedStateAsync` was implemented but nothing called it on host startup. A host configured with `PausePersistence = AcrossReactivations` would write the persisted key on pause, but on subsequent activation the new `QuiescenceSignal` instance would never read it, so the runtime would resume dispatching despite the operator having paused. Fix: `InitializePauseStateStartupTask : IStartupTask` reads the policy and calls `InitializePersistedStateAsync` once per activation when the policy demands it. Registered in both `WorkflowRuntimeFeature` flavours alongside the other graceful-shutdown services. 3. **MEDIUM — Comment in `DrainOrchestrator.PersistInterruptedAsync` no longer lies.** The previous comment promised a "synthetic log entry" that the next statement (`return`) prevented from being written. The comment is now honest about what actually happens: when no instance row exists, no log entry is emitted, but the burst metadata is still captured in the drain outcome's logged warning so operators have a forensic trail. Tests: * New unit tests on `BurstRegistry` (now 9, was 6): cancel-callback is invoked, callback exceptions are swallowed (drain remains best-effort), `BurstHandle.Disposed` completes on dispose. * Full suites continue to pass: 103/103 runtime unit tests; 247/247 workflow integration tests. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): address PR review feedback + close runner-clobber race + add e2e drain test Six issues raised in PR #7424 review (one e2e gap, five comments inline): 1. **Runner-clobber race closed via ICommitStateHandler decorator.** The previous fix wired `BurstHandle.Cancel()` to `WorkflowExecutionContext.Cancel()` so the workflow's cancellation chain fires on deadline breach, but the e2e test exposed that `BurstHandle` disposed at the END of the pipeline middleware (i.e., BEFORE `WorkflowRunner` calls `commitStateHandler.CommitAsync`). The orchestrator's `await handle.Disposed` therefore returned too early, the instance row didn't yet exist, and the orchestrator's Interrupted write was either a no-op (no row) or got clobbered by the runner's subsequent Cancelled commit. Fix: `BurstAwareCommitStateHandler` decorates `ICommitStateHandler`. The middleware no longer disposes the handle in the success path — it stores the handle in `WorkflowExecutionContext.TransientProperties`, and the decorator disposes it AFTER `inner.CommitAsync` completes. The exception path in the middleware still disposes for safety. Result: the orchestrator's await-disposed sequencing now correctly lands the Interrupted write last. 2. **C1: Null-instance log entry.** `DrainOrchestrator.PersistInterruptedAsync` now writes a synthetic `WorkflowInterrupted` log entry directly when no instance row exists, populating only the fields it knows. Previously the forensic trail was lost. 3. **C2: Force endpoint cached-outcome audit.** Added `WasCached` flag to `DrainOutcome` (default false). The orchestrator sets it on the cached return path (`_previousOutcome with { WasCached = true }`). The force endpoint now skips the audit notification when the flag is true, so repeated force calls no longer emit spurious `RuntimeForceRequested` events (SC-007 idempotency restored). 4. **C3: `StateChanged` raised under lock — deadlock risk closed.** `QuiescenceSignal.BeginDrainAsync`/`PauseAsync`/`ResumeAsync` now do their transitions under the lock, capture whether a transition occurred, release the lock, and only then invoke `RaiseStateChanged`. Subscribers that synchronously call back into the signal can no longer deadlock. 5. **C4: Scheduling source name.** Renamed `scheduling.cron` → `scheduling.triggers` to honestly reflect the four trigger types the adapter covers (Cron, Timer, StartAt, Delay). The name is surfaced verbatim in admin status responses. 6. **C5: Hardcoded Retry-After.** `HttpWorkflowsMiddleware`'s 503 response now sets a reason-aware `Retry-After`: 5 s during drain (host is exiting and will be replaced shortly), 60 s during administrative pause (indefinite, so a longer back-off avoids tight retry loops). Tests: * New e2e `DeadlineBreachEndToEndTests` (2 tests): verifies that drain against a real running workflow detects the in-flight burst, force-cancels it, persists the instance as `Interrupted`, and writes a `WorkflowInterrupted` log entry — closing the test gap that hid the cancellation-propagation issue identified in the previous review pass. * Updated `OperatorForceAfterPreviousReturnsCachedOutcome` to assert value-equality + the `WasCached` flag instead of reference-equality (records use `with` for the cached return path). Full suites pass: 103/103 runtime unit, 249/249 workflow integration. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-api): move runtime admin endpoints into Elsa.Workflows.Api Per PR review feedback: rather than introducing a new sub-module (Elsa.Workflows.Runtime.Admin) for the four pause/resume/status/force endpoints, fold them into the existing Elsa.Workflows.Api project. That project already references both Elsa.Workflows.Runtime and Elsa.Api.Common (FastEndpoints) and is the established home for client-facing workflow APIs — so the admin endpoints belong there. Changes: * New folder src/modules/Elsa.Workflows.Api/Endpoints/RuntimeAdmin/ with Models.cs and Pause/Resume/Status/Force/Endpoint.cs. Namespaces moved from `Elsa.Workflows.Runtime.Admin` → `Elsa.Workflows.Api.Endpoints.RuntimeAdmin`. * Deleted src/modules/Elsa.Workflows.Runtime.Admin/ entirely and removed it from Elsa.sln. The ShellFeature marker class (WorkflowRuntimeAdminFeature) is no longer needed — the existing WorkflowsApiFeature already discovers FastEndpoints in the Workflows.Api assembly. * No consumer changes: the endpoints sit in the same routes (/admin/workflow-runtime/*) and behave identically. Note on the second architectural point ("update IShellFeature if cleaner"): the IShellFeature contract is defined in the external CShells NuGet package, not in this repo, so we cannot add a DeactivateAsync hook without an upstream CShells change. The current IHostedService.StopAsync hook continues to work correctly for the host-stop path; per-shell deactivation would require either a CShells upstream addition or a separate Elsa-owned shell-feature variant — neither lighter than what we have today. Tests: 103/103 runtime unit + 249/249 workflow integration pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(common): introduce IsFatal exception extension + apply to drain best-effort catches Per PR review feedback on the static-analyzer "Generic catch clause" comments: rather than catching ALL exceptions in best-effort drain code paths, narrow the swallow to non-fatal exceptions. Process-fatal conditions (StackOverflowException, AccessViolationException, SEHException, ThreadAbortException, OutOfMemoryException) propagate so the host's failure-fast policy can act on them, while normal failures (InvalidOperationException, IOException, etc.) continue to be logged and allowed through so a single misbehaving ingress source / activity cannot abort the overall drain. Highlights: * New `Elsa.Common.Extensions.ExceptionExtensions.IsFatal` utility: classifies fatal conditions, unwraps reflection-style wrappers (TypeInitializationException, TargetInvocationException) before classification, and treats InsufficientMemoryException (the recoverable OOM subclass) as non-fatal. * Applied as a `when (!ex.IsFatal())` filter to: - BurstHandle.Cancel (cancel callback try/catch) - DrainOrchestrator.PauseOneSourceAsync (per-source exception path) - DrainOrchestrator.TryForceStopAsync - DrainOrchestrator.ForceCancelActiveBurstsAsync (per-burst loop) - DrainOrchestrator.PersistInterruptedAsync (orphan log write, instance save, log write) - DrainOrchestrator.DrainAsync outer catch (existing `not InvalidOperationException` filter extended) - InterruptedRecoveryScan (per-instance restart loop) Tests: 7 new unit tests for IsFatal classification (fatal types, recoverable types, wrapped causes, null tolerance). Full suites: 103/103 runtime unit (incl. 14/14 in Common.UnitTests including new tests) + 249/249 workflow integration pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(workflows-runtime): integrate CShells 0.0.15 lifecycle hooks (IDrainHandler + IShellInitializer) CShells 0.0.15 ships the lifecycle framework needed for first-class per-shell graceful shutdown — IDrainHandler / IShellInitializer / IShellLifecycleSubscriber. This commit bumps the package, migrates Elsa's existing usage of the removed 0.0.14 API, and registers the runtime drain orchestrator + pause-state initializer through the new primitives. Highlights: * `ElsaShellDrainHandler : IDrainHandler` — bridges per-shell drain into `IDrainOrchestrator.DrainAsync(DrainTrigger.ShellDeactivation, ct)`. Invoked by CShells when a shell enters `ShellLifecycleState.Draining`; the drain handler's CancellationToken is signalled when the per-shell deadline elapses, so the orchestrator's own deadline-bounded protocol nests cleanly. Coexists with the host-stop `DrainOrchestratorHostedService`; the orchestrator's `DrainAsync` is idempotent — second invocations log and skip. * `InitializePauseStateShellInitializer : IShellInitializer` — replaces the IStartupTask variant in shell-aware deployments. IShellInitializer fires on EVERY shell (re)activation, including reactivations after a reload — exactly what FR-028 requires. The IStartupTask remains for IModule consumers where there is no shell platform. Migrations (CShells 0.0.14 → 0.0.15 breaking changes): * `ActivateShellTenants`: was `IShellActivatedHandler` + `IShellDeactivatingHandler`, now `IShellInitializer` + `IDrainHandler`. * `MultitenancyFeature`: registrations updated to the new transient interface, `using CShells.Hosting` → `using CShells.Lifecycle`. * `Reload/Endpoint`, `ReloadAll/Endpoint`: `IShellManager` → `IShellRegistry`, `ReloadShellAsync` → `ReloadAsync` (returns `ReloadResult` with `Error`), `ReloadAllShellsAsync` → `ReloadActiveAsync` (returns `IReadOnlyList<ReloadResult>` with per-shell errors aggregated into 503). Build + restore: * `Directory.Packages.props`: all CShells.* packages bumped to 0.0.15. * `NuGet.Config`: added `cshells-feedz` source (https://f.feedz.io/sfmskywalker/cshells/nuget/index.json) and split the package-source-mapping pattern into `CShells` (exact) + `CShells.*` (prefix). Single-pattern `CShells*` does NOT match correctly under PackageSourceMapping. Tests: 103/103 runtime unit + 249/249 workflow integration pass on the new package version. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-runtime): apply PR #7424 review feedback Consolidates the architectural fixes asked for during /review: - Extract IWorkflowRuntimeAdminService to back the four /admin/workflow-runtime endpoints with a single domain service; thin Pause/Resume/Status/Force endpoints to delegating shells. - Remove StateChanged C# event from IQuiescenceSignal (Constitution VII: no external subscribers existed; mediator was suggested as the alternative if/when it's needed). - Promote InitializePersistedStateAsync to IQuiescenceSignal, dropping the concrete-cast in both InitializePauseStateStartupTask and InitializePauseStateShellInitializer. - Invert ingress-source DI to Lazy<IEnumerable<IIngressSource>> to break the cycle through IQuiescenceSignal; ingress adapters take the signal directly via primary constructor. - Replace Guid.NewGuid().ToString("N") with IIdentityGenerator in InterruptedLogExtensions and DrainOrchestrator. - Switch admin-audit timestamps to ISystemClock in WorkflowRuntimeAdminService. - Make GracefulShutdownOptions.StimulusQueueMaxDepthWhilePaused nullable (null = unlimited). - Rename RuntimeForceRequested → RuntimeForceDrainRequested. - Apply IsFatal exception filter to drain best-effort catches. - Rename IBurstRegistry.EnumerateActive → ListActiveBursts. - Refresh "Phase X" comments to user-story (USx) references. - Delete unused IngressAttributionExtensions, IngressSourceServiceCollectionExtensions, IngressSourceRegistrationOptions. - Migrate Elsa.Shells.Api.Tests to CShells 0.0.15 IShellRegistry / ReloadResult surface. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(graceful-shutdown): apply IsFatal filter to deadline-breach test catch Aligns the test scaffolding's swallow-everything catch with the project standard introduced in c00eee80c so the analyzer no longer flags the bare `catch` clause. The semantics are unchanged — non-fatal exceptions (OCE, TimeoutException, workflow exceptions) are still acceptable test outcomes. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): close IngressSourceRegistry first-access race + spelling Replaces the non-atomic `_entries.Count > 0` early-return guard in EnsureMaterialized with a double-checked lock against a volatile `_materialized` flag, so concurrent first callers can no longer both iterate the source factory and crash one of them with a "Duplicate ingress source registration" InvalidOperationException. Adds a regression test that launches 16 readers behind a TaskCompletionSource gate and asserts every reader observes the full source set without throwing. Also flips the British spellings introduced in this PR's scope to American English (materialize/behavior) — project convention going forward. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * docs(constitution): require American English for new code (v1.0.1) Adds a "Spelling & language" bullet under principle III (Convention-Driven Design): every newly-introduced symbol, comment, identifier, error message, XML doc, commit message, and Speckit artifact uses American English. Established public API symbols (e.g. WorkflowSubStatus.Cancelled) are not renamed retroactively. PATCH bump because this is a clarification of an existing principle, not a new principle. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Add Greploop skill and workflow for GitLab, GitHub, and Perforce integration - Introduced Greploop, an iterative optimization and review workflow for GitLab MRs, GitHub PRs, and Perforce changelists. - Added API and GraphQL references for fetching and resolving review skill. * Remove GenerateWorkflowVariableAccessorsTests; redundant ExpandoObject type check in handlers * Potential fix for pull request finding 'CodeQL / Untrusted Checkout TOCTOU' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(graceful-shutdown): apply PR #7424 review feedback round 2 Two P1 findings from Greptile: 1. DrainOrchestrator.ForceCancelActiveBurstsAsync was sequential — each burst was cancelled, awaited up to ForceCancelSettleTimeout (2 s), and persisted before the next burst's Cancel() ran. Total wall time was O(N × 2 s) and bursts 2..N kept executing at full speed during prior bursts' settle waits, defeating the intent of force-cancel under concurrency. Refactored to three phases: - Phase A — cancel every handle synchronously (cheap CTS.Cancel calls) so all runners observe cancellation simultaneously. - Phase B — await every Disposed signal in parallel under a single shared ForceCancelSettleTimeout. Total wall time bounded regardless of N. - Phase C — persist Interrupted for each handle sequentially (keeps DbContext usage single-threaded; per-handle work is small). Per-phase failures are caught with !ex.IsFatal() and logged so a single misbehaving handle doesn't abort the rest of the batch. 2. ShellFeatures/WorkflowRuntimeFeature.ConfigureServices was missing the IWorkflowRuntimeAdminService registration that Features/WorkflowRuntimeFeature already had. Any CShells deployment that includes the Pause / Resume / Status / Force admin endpoints (in Elsa.Workflows.Api) would throw InvalidOperationException at endpoint construction. Added the singleton alongside the other graceful-shutdown registrations with a comment pointing out the symmetry with the IModule path. 17/17 graceful-shutdown integration tests pass; 103/103 runtime unit tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Untrusted Checkout TOCTOU' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * fix(ci): harden greploop.yml against CodeQL Actions findings CodeQL flagged 12 findings on .github/workflows/greploop.yml after the prior commit (c08183a3c) addressed an earlier round. Two distinct issue classes remain: 1. Code injection (× ~10): step-output values (steps.pr_head.outputs.head_sha / head_repo_owner / head_repo_name / head_ref) and inputs.pr_number were interpolated directly into shell `run:` blocks via `${{ ... }}`. Because PR author controls the branch name and the manual-dispatch input, those values can carry shell metacharacters. Standard fix: route every such interpolation through an `env:` block on the step, then reference $VAR inside the script. Applied to the Resolve, Resolve PR head metadata, and Checkout PR branch steps. 2. Untrusted Checkout TOCTOU + Checkout of untrusted code in trusted context: the workflow runs on `issue_comment` (a privileged trigger) and checks out PR-author code. Mitigations stacked here: - Author-association gate already restricts the trigger to OWNER / MEMBER / COLLABORATOR (existing). - Step-output values now travel via env vars (above). - Resolve step rejects pr_number that isn't ^[0-9]{1,10}$ — so downstream `gh pr view` and the prompt argument can't be hijacked. - Checkout step now validates HEAD_SHA matches ^[0-9a-f]{40}$ and the repo owner/name match ^[A-Za-z0-9_.-]+$ before either reaches a URL or a git command. - Existing TOCTOU guard preserved: re-fetch head SHA at checkout time and abort if it changed since the initial resolve. These match the canonical "Securing your GitHub Actions workflows" patterns recommended by CodeQL. No functional change to greploop's runtime behaviour. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(graceful-shutdown): drop [SingleNodeTask] from InitializePauseStateStartupTask Greptile P1: [SingleNodeTask] gates the task to a single cluster winner via distributed lock, but IQuiescenceSignal is a singleton scoped to each node's DI container — each node holds its own in-memory QuiescenceState. With the attribute, only the winning node restored the persisted pause; every other node started with QuiescenceReason.None and accepted new work, silently defeating PausePersistence = AcrossReactivations. Removed [SingleNodeTask] (and the corresponding using) so the task runs on every node. Expanded the doc <remarks> to call out the per-node requirement and point at the shell-aware counterpart (InitializePauseStateShellInitializer) which is correctly per-node by virtue of being an IShellInitializer. 17/17 graceful-shutdown integration tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Potential fix for pull request finding 'CodeQL / Code injection' Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> * chore(deps): bump CShells 0.0.15 → 0.0.17 0.0.17 ships our blueprint-aware-routing PR (valence-works/cshells#93) plus four follow-up fixes the maintainer added on top: - e56ebd8 — PreWarmShells removed entirely; ShellMiddleware now does cold-start endpoint matching (re-runs endpoint resolution after lazy activation so the very first request to a cold shell hits its endpoint). - cbe5ee2 — GetCandidateSnapshot returns a bounded ShellRouteCandidateSnapshot with accurate total counts; sensitive-data redaction in routing logs; DefaultShellRouteIndex implements IDisposable. - cd7d4f5 — Last-good snapshot served on rebuild failure (the deferred Copilot review concern); root-path fallback when path-by-name misses. - c3679d9 — Cold-start endpoint matching respects inline route constraints; path-name convention tightening; dead duplicate-detection cleanup. Net effect for elsa-core: - Cold blueprints serve their first request via lazy activation, with endpoints correctly resolved post-activation. - Reloaded shells re-activate and serve on the next matched request. - Non-name-mode routing keeps serving the previous snapshot during a transient blueprint-provider outage. - No need to call PreWarmShells from Elsa.ModularServer.Web — removed. The only API removal that touches elsa-core is PreWarmShells. No code references IShellRouteIndex / ShellRouteCriteria / GetCandidateSnapshot directly, so the API-shape changes in cbe5ee2 don't ripple here. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(graceful-shutdown): persist Interrupted under non-drain bounded token Greptile P? finding on the prior force-cancel two-phase fix: Phase B's inner catch on OperationCanceledException ("drain CT fired — proceed to persist anyway") was a lie in practice. Phase C immediately passed the same already-cancelled drain token into PersistInterruptedAsync; the first DB call (instanceStore.FindAsync) observed the cancellation and threw OperationCanceledException; the outer non-fatal Exception filter swallowed it and only logged an error. Net effect: on host shutdown deadline breach, every burst after cancellation could fail to be persisted as Interrupted, leaving instances in an unrecovered executing state. Phase C now creates a per-handle CancellationTokenSource bounded to a new PersistInterruptedTimeout (5 s) that is NOT linked to the drain CT. Each persist gets up to 5 s to land the row update + forensic log entry even after the drain CT has fired. The bound prevents a stuck DB from hanging shutdown indefinitely (per-handle worst case is small; total Phase C upper bound is N × 5 s, but typical persists are millisecond scale). Comment expanded to call out why the persist token is independent of the drain token, so the rationale doesn't drift again. 17/17 graceful-shutdown integration tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): extract PassiveIngressSource base class The three IIngressSource implementations that ship with this PR (InternalBookmarkQueueIngressSource, HttpTriggerIngressSource, ScheduledTriggerIngressSource) were ~25 lines each and ~22 lines of those were verbatim copies of each other: - ctor signature `(IQuiescenceSignal signal)` - `PauseTimeout => TimeSpan.FromMilliseconds(50)` - `CurrentState => signal.IsAcceptingNewWork ? Running : Paused` - `PauseAsync` / `ResumeAsync` returning `ValueTask.CompletedTask` The shared trait is that none of them does any work at pause time — the actual pause enforcement lives in another layer (`HttpWorkflowsMiddleware` short-circuits to 503, `BookmarkQueueProcessor` consults the signal at the top of each invocation, scheduled triggers dispatch through the bookmark queue and inherit that behaviour transitively). The IIngressSource adapter is purely diagnostic: it makes the source visible in `DrainOutcome.Sources` and the admin status endpoint. Extracted that pattern into `PassiveIngressSource` (abstract base in `Elsa.Workflows.Runtime.IngressSources`). Subclasses now provide only `Name`; `PauseTimeout` is `virtual` with a 50 ms default; everything else is fixed by the base. The three concretes drop from ~25 lines to ~12 lines each. The base's XML `<remarks>` calls out when to use it ("your component already cooperates with IQuiescenceSignal at its hot path") and when to implement IIngressSource directly ("the source owns concrete pause/resume behaviour — e.g. a message-queue consumer that calls Pause() on its underlying client"), so future contributors don't mis-extend the base for sources that need real work at pause time. No behavioural change. 18 graceful-shutdown integration + 39 runtime unit tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): align IIngressSource name to singular The three IIngressSource names were inconsistent: http.trigger (singular) internal.bookmark-queue-worker (singular) scheduling.triggers (PLURAL — outlier) The plural slipped in when addressing Greptile's earlier comment to avoid `scheduling.cron` (which would imply Cron-only coverage). The right move was to pick a generic word and stay singular like the rest of the suite — the suite's mental model is "the X source", one instance per registry slot, regardless of how many triggers or items it dispatches internally. Renamed to `scheduling.trigger`. The `<remarks>` block keeps the "covers Cron, Timer, StartAt, Delay" explanation and now also explicitly notes the singular convention so future contributors don't re-pluralize. Zero test fallout — the literal "scheduling.triggers" only appeared in the source file itself. Tests of the other two sources all use singular forms. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(runtime-admin): rename Force endpoint to ForceDrain "Force" alone is meaningless out of context — force what? — and it sits oddly next to the verb-named siblings Pause / Resume / Status. The matching admin service method is already IWorkflowRuntimeAdminService. ForceDrainAsync, so ForceDrain is the natural pair. Renamed: - src/modules/Elsa.Workflows.Api/Endpoints/RuntimeAdmin/Force/ → ForceDrain/ - namespace ...Endpoints.RuntimeAdmin.Force → ...ForceDrain - class ForceEndpoint → ForceDrainEndpoint - class ForceRequest → ForceDrainRequest - class ForceResponse → ForceDrainResponse - route POST /admin/workflow-runtime/force → /force-drain Zero external references — no tests, docs, or OpenAPI clients used the old symbols or the old route literal, so this is a contained pre-ship rename. Directory move went through `git mv` so commit history follows the file. 17/17 graceful-shutdown integration tests still pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): correct formatting in checkout step within greploop.yml Adjusted indentation of environment variables in the checkout step for improved consistency and readability. * fix(ci): repair malformed Checkout repository step in greploop.yml The step accumulated stray env keys, an extra `uses:`, and bash commands that didn't belong inside it (line 65 onward), causing a YAML parse error on push. The valid structure has two distinct checkout steps: - Checkout repository : actions/checkout@v4 with fetch-depth: 0 - Checkout PR branch : env: + run: with SHA validation + git fetch + git checkout --detach The PR-branch step (line 78+) was already correct and unchanged. This fix restores the first step to its intended single-purpose shape (just checks out the workflow file's commit so the greploop skill is on disk before the run-greploop step uses it). No functional change to runtime behaviour or to the security posture established in the prior hardening commit (0db4ca23e). The PR-branch checkout still validates HEAD_SHA / HEAD_REPO_OWNER / HEAD_REPO_NAME shape before they reach a URL or git command. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): remove literal `${{ }}` from greploop.yml comment GitHub Actions parses `${{ ... }}` workflow expressions across the entire YAML file, including inside `run:` script comments. The comment that explained the env-var hardening pattern contained the literal sequence `${{ }}` (with a space, intended as an English-language description), which the expression parser rejected as "An expression was expected" (line 81 col 14). Reworded the comment to describe the substitution form in prose without the literal token sequence. Functional behaviour unchanged. * refactor(graceful-shutdown): rename burst → execution cycle The graceful-shutdown work introduced "burst of execution" as a first-class domain concept. The term arrived without rationale and isn't standard in the workflow-engine domain. Renamed to "execution cycle" — reads more naturally as the loop-with-commit unit, is more idiomatic in workflow vocabulary, pairs cleanly with the existing WorkflowExecutionContext, and avoids collisions with Elsa's existing terms (Run, Execution, Dispatch, Invocation, Stimulus, Step). Renamed types - IBurstRegistry → IExecutionCycleRegistry - BurstRegistry → ExecutionCycleRegistry - BurstHandle → ExecutionCycleHandle - BurstTrackingMiddleware → ExecutionCycleTrackingMiddleware - BurstAwareCommitStateHandler → ExecutionCycleAwareCommitStateHandler Renamed members - BeginBurst → BeginCycle - ListActiveBursts → ListActiveCycles - BurstHandleKey constant + value → ExecutionCycleHandleKey - ActiveBurstCount (IQuiescenceSignal, RuntimeAdminStatus, StatusResponse) → ActiveExecutionCycleCount - WaitForBurstsAsync (private) → WaitForCyclesAsync - ForceCancelActiveBurstsAsync (priv) → ForceCancelActiveCyclesAsync - UseBurstTracking → UseExecutionCycleTracking - DrainOutcomeDto.BurstsForceCancelledCount → ExecutionCyclesForceCancelledCount - _burstRegistry / burstRegistry → _cycleRegistry / cycleRegistry Backwards-compatibility preservation (the only persisted JSON key) - WorkflowInterruptedPayload.BurstDuration property → ExecutionCycleDuration with [JsonPropertyName("BurstDuration")] so the persisted JSON wire key stays "BurstDuration" forever. Pre-merge testers' log records still deserialise correctly. The contract test on WorkflowInterruptedPayloadContractTests still asserts the wire key "BurstDuration" appears in the serialised JSON — confirms the guarantee is enforced. Other unstructured surfaces - WorkflowExecutionLogRecord.Message text "Workflow burst was force- cancelled..." now says "Workflow execution cycle was force-cancelled ..." for new records. Old rows keep their old text — purely cosmetic free-text field. - Structured log placeholder {BurstId} in DrainOrchestrator log lines → {ExecutionCycleId}. - Lowercase prose / XML doc comments updated throughout. Test renames - BurstRegistryTests → ExecutionCycleRegistryTests - BurstTrackingMiddlewareTests → ExecutionCycleTrackingMiddlewareTests - Test method names + DisplayName strings updated. Spec docs (specs/002-graceful-shutdown/) updated to match the new vocabulary; the historical task records in tasks.md keep the old names as-is to preserve the audit trail of what was originally built. Verification - dotnet build: clean across net8.0 / net9.0 / net10.0. - 103/103 Elsa.Workflows.Runtime.UnitTests pass. - 17/17 GracefulShutdown integration tests pass. - 4/4 WorkflowInterruptedPayloadContractTests pass — confirms the "BurstDuration" JSON wire-key preservation is intact. No changes to migrations or DB column names — confirmed via grep. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): give greploop.yml gh-cli a repo context before checkout The "Resolve PR head metadata" step runs `gh pr view` before actions/checkout, so there is no `.git` directory and gh's "current repo" detection fails with `fatal: not a git repository`. The prior commits to this file masked the runtime failure because the workflow itself was YAML-invalid — once it became valid, the workflow_dispatch trigger surfaced this real-world execution bug. Set GH_REPO=${{ github.repository }} on both `gh pr view` steps. The gh CLI honours GH_REPO as an explicit repo override, so it no longer needs git context. Same fix on the "Checkout PR branch" validation call which also uses gh pr view before the manual fetch. The error reported as "Invalid workflow file: ... (Line 81 Col 14)" on PR #7424 was stale from commit ed580682a (which had the bad comment with literal `${{ }}`); commit 3a8ad0d08 fixed the YAML, but because greploop's `if:` condition only matches workflow_dispatch / issue_comment events, push events on later commits were skipped without re-running validation, so the GitHub UI kept showing the old error. A workflow_dispatch run on the current SHA now passes validation and reaches "Resolve PR head metadata", which is what this commit fixes. * fix(graceful-shutdown): wire IngressPauseTimeout option to drain orchestrator GracefulShutdownOptions.IngressPauseTimeout was documented as "Default per-ingress-source pause timeout" but DrainOrchestrator.PauseOneSourceAsync read source.PauseTimeout directly and never consulted the option. The configured value was silently ignored — operators who set GracefulShutdownOptions:IngressPauseTimeout = 10s were getting whatever each source's hardcoded value was (50 ms for the three PassiveIngressSource subclasses we ship), with no way to tune it globally. Precedence (per the spec's intent of "overridable at registration and by configuration"): 1. Per-source positive value wins (source.PauseTimeout > Zero). 2. Otherwise fall back to the configured GracefulShutdownOptions. IngressPauseTimeout default. 3. Resolved value is capped at the overall drain deadline so a single misbehaving source cannot exceed the host's shutdown budget. 4. 1 ms safety floor remains so a misconfigured zero default still produces a non-zero CancelAfter. Changes: - DrainOrchestrator.PauseOneSourceAsync — adds the precedence above with a comment block explaining each step. - IIngressSource.PauseTimeout — XML doc clarifies the Zero-defers-to- config semantics. - GracefulShutdownOptions.IngressPauseTimeout — XML doc says it's the fallback when the source returns Zero; <remarks> spells out the precedence and the overall-deadline cap. - PassiveIngressSource.PauseTimeout — virtual property now returns Zero (was 50 ms). The three shipped subclasses (HttpTriggerIngressSource, ScheduledTriggerIngressSource, InternalBookmarkQueueIngressSource) consequently defer to the configured default — flipping the wire-up bug from "configured value silently ignored" to "configured value honoured by default for passive sources". Passive subclasses that want a specific value can still override. 103/103 runtime unit + 17/17 graceful-shutdown integration tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): rename IInterruptedRecoveryScan → IInterruptedRecoveryScanner The interface had a single verb-method (`ScanAndRequeueAsync`) and its XML doc described what it *does* ("Scans the workflow instance store for instances..."). That's an agent role — a scanner that performs a scan — but the noun-shaped name `IInterruptedRecoveryScan` read as "the scan itself", which is misleading because the scan results / event are not first-class types in the codebase. Renamed to `IInterruptedRecoveryScanner` / `InterruptedRecoveryScanner` to match the existing `-er` convention in this codebase (Restarter, Generator, Resolver, etc.). The method stays `ScanAndRequeueAsync` — the scanner *performs* a scan-and-requeue. Also renamed the constructor parameter `scan` → `scanner` in RecoverInterruptedWorkflowsStartupTask, and the local variable `scan` → `scanner` in InterruptedRecoveryIntegrationTests, so the "scanner does the scan" mental model is consistent throughout. Surface impact (all internal — no API or persistence touch points): - 2 source files renamed via git mv (interface + implementation) - 1 test file renamed (InterruptedRecoveryScanTests → ScannerTests) - DI registrations in both Features/ and ShellFeatures/ WorkflowRuntimeFeature - 1 startup-task constructor parameter - Spec doc references under specs/002-graceful-shutdown/ 103/103 runtime unit + 17/17 graceful-shutdown integration tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(graceful-shutdown): extract DrainTriggerExecutor ElsaShellDrainHandler (CShells IDrainHandler) and DrainOrchestratorHostedService (.NET IHostedService.StopAsync) inlined near-identical try/catch/log shapes around IDrainOrchestrator.DrainAsync: - call DrainAsync(<trigger>, ct) - branch on outcome: DeadlineExceeded/AbortedByUnhandledException → Warning, otherwise Information - catch InvalidOperationException (parallel-drain rejected by the orchestrator) → log Information and swallow The two had already drifted: host-stop's success log omitted the paused/waited durations the shell-handler version included, and the "skipped" message disagreed on the trigger label ("Host-stop drain skipped" vs "Shell drain skipped"). Centralised the shape in a small internal static helper so the two — and any future trigger source — stay uniform. Both call sites collapse to a single line. Net diff drops 22 lines from the two consumers and adds a 25-line helper that they both delegate to. The unified log copy now consistently includes paused/waited durations on the success path and uses the caller-supplied contextLabel ("Shell drain", "Graceful drain") in all three messages so operators can attribute log entries by trigger source. Files: - src/modules/Elsa.Workflows.Runtime/Services/DrainTriggerExecutor.cs (new) - src/modules/Elsa.Workflows.Runtime/Lifecycle/ElsaShellDrainHandler.cs - src/modules/Elsa.Workflows.Runtime/HostedServices/DrainOrchestratorHostedService.cs 103/103 runtime unit + 17/17 graceful-shutdown integration tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-runtime): drop redundant DrainOrchestratorHostedService from CShells path In CShells deployments, host stop already drives drain via CShellsStartupHostedService → IDrainHandler → ElsaShellDrainHandler, scoped per shell (FR-027). The additional .AddHostedService<DrainOrchestratorHostedService>() in ShellFeatures/WorkflowRuntimeFeature was firing a second non-force DrainAsync that the orchestrator rejected with InvalidOperationException — silently swallowed by DrainTriggerExecutor, but logged on every host stop and semantically muddled (IHostedService is host-level, not per-shell). Keep the registration on the IModule path (Features/WorkflowRuntimeFeature) where there is no shell platform and host-stop is the only available drain trigger. Update ElsaShellDrainHandler XML docs to reflect the now-clean single-trigger model in CShells. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): always dispose ExecutionCycleHandle in tracking middleware Previously the success path relied on ExecutionCycleAwareCommitStateHandler to dispose the handle after the runner's commit completed. If a custom dispatcher or test double exited the pipeline without invoking commit (by design or by accident), the handle stayed registered, IExecutionCycleRegistry.ActiveCount never reached zero, and drain spun in WaitForExecutionCyclesAsync until the deadline fired — incorrectly force-cancelling instances that had already finished cleanly. Collapse the existing try/catch(rethrow) into try/finally so the middleware itself disposes the handle for both exception and commit-elided paths. Disposal remains idempotent via the ExecutionCycleHandle._disposed Interlocked guard, so the normal-path dispose by ExecutionCycleAwareCommitStateHandler is a harmless no-op. Adds an integration regression test that drives the middleware with a stub Next that returns without invoking commit and asserts ActiveCount returns to zero. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): serialize QuiescenceSignal pause-state persistence Both PauseAsync and ResumeAsync used to release the inner lock before issuing the persistence I/O. A rapid Pause → Resume sequence could leave the persisted state inconsistent: PauseAsync's slow SaveAsync could land AFTER ResumeAsync's DeleteAsync, leaving the key present in the store while in-memory state was None. On host restart, InitializePersistedStateAsync would find the stale key and start the runtime in the paused state the operator had already cancelled. Introduce a dedicated SemaphoreSlim that serializes persistence I/O, with each I/O re-reading the live in-memory state inside the semaphore. N racing Pause/Resume calls now produce N serialized writes, each reflecting the most recent in-memory transition — so the final persisted state always matches final in-memory state. Adds a regression test that gates SaveAsync, races a Resume behind it, and asserts the store is empty after both complete. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(workflows-runtime): trim verbose comment in ExecutionCycleTrackingMiddleware finally block Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(workflows-runtime): use 'using var' for ExecutionCycleHandle in tracking middleware Replace the explicit try/finally that only existed to call handle.Dispose() with a `using var` declaration — identical semantics (compiler-emitted finally with idempotent dispose), more idiomatic. The regression test HandleReleasedWhenCommitIsElided continues to validate the leak-free property. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime,api): proper 409 conflict shape + shell-scoped pause-persistence key ForceDrain endpoint: the 409 path returned `new ForceDrainResponse()` whose `Outcome` was null at runtime despite the `= null!` annotation, so any strongly-typed client deserializing the conflict body and reading `Outcome.OverallResult` got an NRE. Switch to the existing `ConflictResponse` shape with `Code = "DrainInProgress"` and the current runtime status. Routed via HttpContext.Response.WriteAsJsonAsync because Send.ResponseAsync is constrained to the endpoint's TResponse and cannot send a sibling DTO. QuiescenceSignal persistence key: the DI-registered `IQuiescenceSignal` was constructed with `shellName = null` (DI doesn't inject `string?` defaults), so every shell shared the key `elsa.quiescence.pause.default`. In a CShells multi-shell deployment under PausePersistencePolicy.AcrossReactivations this caused cross-shell contamination — pausing shell A would re-pause shell B on its next activation. Replace the simple AddSingleton<IQuiescenceSignal,...> registration in ShellFeatures/WorkflowRuntimeFeature with a factory that injects `CShells.ShellSettings` and forwards `Settings.Id` as the shell name. The IModule registration is unchanged (no shell platform; null shellName remains correct there). Adds a unit regression test that two QuiescenceSignal instances with different shellNames write to disjoint persistence keys and never to the legacy "default" key. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(quiescence): use TryGetValue for ContainsKey+indexer assertions Combines existence check and value retrieval into a single dictionary lookup, addressing the code-quality bot's repeated suggestion. No behavior change — both PauseWritesKey and PersistenceKeyIncludesShellName still assert the same keys exist with the same content. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Update specs/002-graceful-shutdown/contracts/admin-endpoints.md Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * fix(workflows-runtime): decouple QuiescenceSignal persistence from caller cancellation PersistAsync used to forward the caller's CancellationToken to both _persistenceMutex.WaitAsync and the store I/O. If an HTTP request was cancelled between the in-memory transition (already committed under _sync) and the persistence call, the I/O was silently skipped — leaving AdministrativePause set in memory with no persisted record. The idempotent fast-path on subsequent PauseAsync calls (transitioned == false) meant no retry would happen, so on host restart InitializePersistedStateAsync would find no key and the runtime would come back unpaused, defeating PausePersistencePolicy.AcrossReactivations. Drop the parameter from PersistAsync entirely; use CancellationToken.None for both the semaphore wait and the store I/O. The in-memory transition is already committed by the time PersistAsync runs, so persistence must complete to keep the store consistent with memory. The public PauseAsync/ResumeAsync methods still accept a CancellationToken (interface contract) — it just no longer reaches the persistence layer. Adds a regression test that calls PauseAsync with a pre-cancelled token and asserts both in-memory pause and the persisted key land correctly. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime,api,docs): apply Copilot review feedback batch Code: - DrainOrchestrator.TryForceStopAsync now bounds force-stop with the *remaining* drain budget (deadlineAt - now), not the full overall TimeSpan. A per-source pause that already burned the shutdown window can no longer get another full deadline's worth of force-stop runway. - DrainOrchestrator catch filter narrowed: drop `ex is not InvalidOperationException` exclusion. The "drain already in progress / completed" IOEs are thrown outside the protocol's try block, so they bubble out without entering this handler. Any IOE that lands here is incidental (e.g., from a store inside the drain) and should now be captured into the outcome rather than escaping the whole drain. - ResumeEndpoint 409 path now returns the discriminated ConflictResponse shape (matching ForceDrain) instead of a plain StatusResponse. Routed via HttpContext.Response.WriteAsJsonAsync because Send.ResponseAsync is constrained to TResponse. - Conflict codes aligned to kebab-case across both endpoints to match the contract spec (`runtime-draining` and `drain-in-progress`). Spelling sweep — American English per constitution v1.0.1 III: - DrainOrchestrator.cs: "serialised" → "serialized" - WorkflowInterruptedPayload.cs: "serialised" / "deserialise" → "serialized" / "deserialize" - PassiveIngressSource.cs: "behaviour" → "behavior" - DeadlineBreachEndToEndTests.cs: "serialisable" → "serializable" - specs/002-graceful-shutdown/quickstart.md: "behaviour" → "behavior" - specs/002-graceful-shutdown/checklists/requirements.md: "behaviour" → "behavior" Doc/contract alignment: - quickstart.md: force route corrected from /force to /force-drain. - quiescence-signal.md: removed StateChanged event from contract (interface doesn't define it); corrected persistence section to describe InitializePersistedStateAsync via shell initializer / startup task rather than constructor read; added the per-shell key discriminator. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): correct DI lifetimes for ExecutionCycleTrackingMiddleware and WorkflowRuntimeAdminService Two strict-DI-validation failures surfaced in tests using BuildServiceProvider with validate-on-build: 1. ExecutionCycleTrackingMiddleware was registered as AddSingleton<>, but its constructor takes WorkflowMiddlewareDelegate next — supplied by the workflow execution pipeline builder via UseMiddleware<>(), not from DI. The registration was both unused (no consumer resolves it through the container) and broken (DI fails to construct it because next is unregistered). Removing both registrations. 2. IWorkflowRuntimeAdminService was registered as AddSingleton<> but depends on the scoped INotificationSender (mediator) — captive-dependency violation. All consumers (Pause/Resume/ Status/ForceDrain endpoints) are FastEndpoints, which are scoped per request, so AddScoped is the correct alignment. The other deps (IQuiescenceSignal / IIngressSourceRegistry / IDrainOrchestrator / ISystemClock) are singletons and resolve fine from a scoped consumer. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(workflows-runtime): restore commit-handler-only success disposal + true Cancel idempotency Two issues raised by Copilot's latest review on commit 32a9c0519: 1. ExecutionCycleTrackingMiddleware was disposing the handle at the end of InvokeAsync (via `using var`), but WorkflowRunner runs commit AFTER the pipeline returns (WorkflowRunner.cs:235). That meant the handle was disposed BEFORE the runner's terminal commit, and the drain orchestrator's `await handle.Disposed` would unblock too early — reintroducing the runner-clobber race the original design protected against (see ExecutionCycleAwareCommitStateHandler XML doc). Revert to the original shape: only dispose on exception path. ExecutionCycleAwareCommitStateHandler remains the SOLE success-path disposer, running in its finally block AFTER the inner commit lands. The earlier "leak when commit is elided" concern was a non-issue in production (the standard runner always commits); the buggy `HandleReleasedWhenCommitIsElided` test that specified the wrong contract is removed. The existing `ActiveCountReturnsToZero` test (which uses the real runner end-to-end) already verifies success-path disposal. 2. ExecutionCycleHandle.Cancel() was documented as idempotent but only short-circuited via the _disposed flag. Repeated Cancel() calls before Dispose could trigger the cancel callback multiple times — easy to accidentally fire non-idempotent cancellation side effects more than once. Add an Interlocked _cancelled guard so callback + CTS cancellation run at most once. Existing test that documented the leaky behavior is updated to assert the now-truly-idempotent contract. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Update logging levels and remove unused features in appsettings files * refactor(workflows-runtime): improve graceful shutdown options handling and cleanup solution Refactor the handling of `GracefulShutdownOptions` to ensure options are applied correctly without directly invoking the delegate. Update DI registrations to use appropriate lifetimes and remove redundant wrapper services. Additionally, clean up the solution by removing unused projects and documentation folders. * feat(identity, workflows-runtime): add validation for identity and graceful shutdown options Introduce validation capabilities for `IdentityTokenOptions` and `GracefulShutdownOptions`. Implement extension methods for option validation, enhance service registration, and add unit tests to ensure configurations are validated at startup. Update solution to include new unit test projects. * update(docs): clarify shutdown log message expectations and levels in quickstart.md Optimize explanation of expected log message sequence during graceful shutdown and specify logging levels. * docs: amend constitution to v1.1.0 (SRP, DRY, KISS, conciseness under Principle VII) * refactor(multitenancy): rename and restructure TenantTaskManager to TenantTaskLifecycleCoordinator Rename `TenantTaskManager` to `TenantTaskLifecycleCoordinator` and relocate to a new directory structure, enhancing code organization and test consistency. Retain functional behaviors with no logic alterations. Update unit tests to reflect the naming changes, ensuring consistency with the refactored code structure. * Update logging levels and dependencies - Set default logging level to Debug in appsettings.Development.json - Add missing using directives for Elsa workflows management and runtime features - Update CShells package versions to 0.0.18-preview.104 --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2026-05-02 17:27:08 +00:00
Assert.True(ex.IsFatal());
}
[Fact(DisplayName = "OutOfMemoryException is fatal")]
public void OomFatal() => Assert.True(new OutOfMemoryException().IsFatal());
[Fact(DisplayName = "InsufficientMemoryException (recoverable subclass of OOM) is NOT fatal")]
public void InsufficientMemoryNotFatal() => Assert.False(new InsufficientMemoryException().IsFatal());
[Theory(DisplayName = "Common recoverable exceptions are NOT fatal")]
[InlineData(typeof(InvalidOperationException))]
[InlineData(typeof(ArgumentException))]
[InlineData(typeof(TimeoutException))]
[InlineData(typeof(NullReferenceException))]
[InlineData(typeof(IOException))]
public void RecoverableExceptions(Type exceptionType)
{
var ex = (Exception)Activator.CreateInstance(exceptionType)!;
Assert.False(ex.IsFatal());
}
[Fact(DisplayName = "TypeInitializationException wrapping a fatal cause is itself fatal")]
public void WrappedFatalIsFatal()
{
var inner = new StackOverflowException();
var outer = new TypeInitializationException("X", inner);
Assert.True(outer.IsFatal());
}
[Fact(DisplayName = "TargetInvocationException wrapping a recoverable cause is NOT fatal")]
public void WrappedRecoverableIsNotFatal()
{
var inner = new InvalidOperationException("boom");
var outer = new TargetInvocationException(inner);
Assert.False(outer.IsFatal());
}
}