elsa-core/specs/003-live-server-logs/research.md
Sipke Schoorstra ab3e46bbe2
[codex] Add live server log streaming diagnostics (#7438)
* Add live server logs Spec Kit plan

* Implement live server logs diagnostics module

* Add server log sources and redaction hardening

* Add diagnostics unit tests

* Harden server log hub subscriptions

* Secure server log hub permissions

* Validate server log filter updates

* Add diagnostics logger and source tests

* Add diagnostics integration test project

* Add multi-source diagnostics provider coverage

* Broadcast server log source changes

* Document diagnostics server log streaming

* Add diagnostics sample host wiring

* Record diagnostics validation results

* Address server log PR feedback

* Rename diagnostics module to server logs

* Add server logs shell feature

* Make server logs shell options bindable

* Accept read wildcard for server logs

* Align server logs authorization with API patterns

* Update CShells structure and logging levels, add diagnostics module

* Rename PostgreSql shell feature classes for consistency

* Switch from Sqlite to PostgreSQL for workflow and identity persistence, add QuartzPostgreSql configuration

* Refactor server logs into diagnostics structured logs (#7440)

* Specify diagnostics structured logs refactor

* docs: clarify structured logs spec

* docs: plan diagnostics structured logs

* docs: add diagnostics structured logs tasks

* refactor: rename server logs to diagnostics structured logs

* Refactor PostgreSql persistence features to use centralized entity model handler registration.

* Refactor EFCore persistence features to centralize entity model handler registration for MySql, Sqlite, and Oracle providers.

* Integrate structured logs by renaming server logs, adjusting appsettings, and updating project references.

* Switch from PostgreSQL to Sqlite for workflow and identity persistence, update appsettings configuration.
2026-05-11 00:08:52 +02:00

2.5 KiB

Research: Live Server Log Streaming

R1: Raw console capture vs structured logging

Decision: Capture ILogger events and render them console-style in Studio.

Rationale: Elsa and ASP.NET Core already use ILogger; structured events preserve level, category, scopes, exception data, correlation, and workflow/tenant context. Raw Console.Out cannot reliably support filtering, redaction, or cluster source identity.

Alternatives considered:

  • Redirect Console.Out: simple but fragile, global, hard to secure, and loses structured properties.
  • Require Serilog/Seq/Loki: powerful but would make the feature dependent on a third-party logging stack.

R2: SignalR vs Server-Sent Events

Decision: Use SignalR for live streaming plus REST for backfill.

Rationale: Studio already has SignalR authentication configuration via IHttpConnectionOptionsConfigurator, and Core already has a SignalR precedent for workflow instance updates.

Alternatives considered:

  • SSE: simpler one-way transport but would need separate auth and reconnect conventions.
  • Polling only: easiest but loses the Aspire-like live tail experience.

R3: Cluster support

Decision: Model source identity in the MVP and expose a provider abstraction. Ship in-memory provider first; add shared providers later.

Rationale: Kubernetes support is best achieved through a shared log provider or existing observability backend. The UI contract should not depend on whether events come from memory, Redis, OpenTelemetry, Loki, or another store.

Alternatives considered:

  • Direct Kubernetes API reads: useful for pods but too deployment-specific and unsuitable for non-Kubernetes clusters.
  • Single-node only: simpler but would bake in the wrong event model.

R4: Redaction timing

Decision: Redact before buffering and before publishing.

Rationale: Recent buffers may be exposed later to authorized users. Keeping raw secrets in memory increases accidental disclosure risk.

R5: Source health

Decision: Track source LastSeen and classify health as connected, stale, or disconnected based on provider data and configured timeout.

Rationale: This works for in-memory and cluster providers without requiring a control-plane integration.

R6: Ordering merged streams

Decision: Sort recent queries by event timestamp with sequence/receive order as a deterministic tiebreaker. Live streams deliver provider order.

Rationale: Cluster clocks can skew. A stable tiebreaker is more important than pretending perfect global ordering exists.