Skip to content

Alerts ​

CritterWatch implements a fully event-sourced alert lifecycle. Every alert transition — raised, elevated, reduced, resolved, cleared, plus the operator annotations acknowledged and snoozed — is stored as an immutable event in the console's event store, providing a complete audit trail of system health over time.

⚠️ The console reads that trail back only as a flat feed. Each transition appears as its own Timeline entry; nothing groups them per alert, and the metric value that triggered a transition is not displayed. See What an alert card does not do.

Alert Lifecycle ​

Acknowledge and Snooze are not states. They are annotations recorded against an alert that is still in whichever state it was already in — acknowledging an alert does not resolve or clear it, and neither event touches the active flag. The console nonetheless hides an acknowledged alert immediately, which is the one place its display and the stored record disagree.

States ​

StateDescription
RaisedThreshold first exceeded. Initial alert created.
ElevatedCondition persists beyond the escalation period. Severity increased.
ReducedCondition is improving but not yet resolved.
ResolvedCondition has cleared automatically (system-condition alerts only).
ClearedClosed by an operator. Terminal.

Annotations ​

AnnotationDescription
AcknowledgedAn operator has seen it. The alert stays in its current state; the console stops showing it.
SnoozedSuppressed until a wall-clock time. The alert stays in its current state and resurfaces when the snooze expires.

System-condition alerts (DLQ counts, projection lag, circuit breakers) auto-resolve when the underlying condition clears. Alerts that cannot auto-resolve stay active until an operator clears them — see Operator actions for which types those are.

Alert Types ​

The complete vocabulary is the AlertId.Types enum in Wolverine.CritterWatch: AgentDown, AgentReassignmentStorm, BackPressure, CircuitBreaker, HighWaterAgentRestarted, HighWaterStale, MetricsDlqRate, MetricsExecTime, MetricsFailureRate, MetricsThroughput, NodeFlapping, ProjectionAgentUnassigned, ProjectionDeadLetters, ProjectionLag, ProjectionLagUnmeasurable, ProjectionPaused, ProjectionStale, SelfIngestFailure, StaleListener, TransportDegraded, TransportUnhealthy. The ones an operator meets most often are described below.

Dead Letter Alerts ​

Two distinct alert types, and they measure different things:

MetricsDlqRate is a rate, not a standing count — it fires on how fast dead letters are arriving, so a large old backlog that has stopped growing does not keep shouting:

  • Warning — DlqRateWarningPerHour (default 10 per hour)
  • Critical — DlqRateCriticalPerHour (default 50 per hour)

Both are MetricsAlertDefaults settings, overridable per service and per message type on the Alert Configuration Metrics tab.

ProjectionDeadLetters covers a projection shard's own dead letters, and its two severities are asymmetric on purpose:

  • Critical — a standing backlog at or above DeadLetterCriticalThreshold (default 10), raised immediately and independent of the Warning throttle below
  • Warning — a recency signal, raised when the count increases past DeadLetterWarningThreshold (default 1), and re-raised at most once an hour after a resolve→recur cycle so a flapping shard cannot flood the feed

Projection Stall Alerts ​

Also two types, measuring lag two different ways:

ProjectionLag counts events behind the high-water mark, not seconds:

  • Warning — BehindWarningThreshold (default 1,000 events)
  • Critical — BehindCriticalThreshold (default 10,000 events)

ProjectionStale is the wall-clock interval since the shard last advanced:

  • Warning — StaleWarningThreshold (default 5 minutes)
  • Critical — StaleCriticalThreshold (default 30 minutes)

Both are ProjectionAlertDefaults settings, overridable per service and per shard. When AutoRestartOnStale is enabled, tripping the Critical stale threshold sends RestartProjection to the affected service — once per stale episode. Both auto-resolve when the projection advances again.

Unmeasurable Lag Alerts ​

ProjectionLagUnmeasurable fires when an Async shard reports no high water mark for its scope, so its lag cannot be computed at all.

This is a separate alert type from ProjectionLag on purpose. A gap that could not be computed is not a gap of zero, and the console must not put a number on it — but it must not stay silent either, which is what it used to do.

⚠️ An earlier revision of this page credited the 24 August production incident — a rollout that left 76 of 512 shards with no agent assigned — to this alert type. That holds only if those shards reported no mark. The field report's own wording is that a frozen high water mark subtracted to zero, and a frozen mark is present, so this alert would not have fired on it. The condition that incident describes is covered by ProjectionAgentUnassigned below.

  • Warning — some eligible tenants on the projection cannot be measured. We genuinely do not know whether any of them is behind.
  • Critical — every eligible tenant cannot be measured. That is no longer an unknown: the projection's high water mark is not being reported at all, and the console is blind to it.

Orphaned shards, rebuilding shards, and Inline/Live projections never raise it — none of those has a gap to measure in the first place, which is a different thing from a gap we failed to measure. It resolves as soon as a mark is reported again.

Unassigned Agent Alerts ​

ProjectionAgentUnassigned fires when a projection is registered and nothing is running it — one or more of its shards has no agent assigned at all.

This is the 24 August production incident: a rollout left 76 of 512 shards unassigned, every frozen high water mark subtracted to zero, the screen was entirely green, and no alert fired. Both halves of the console agreed nothing was wrong.

Two separate mechanisms produced that silence, and both are fixed:

  • The per-shard evaluators skip an Action == "Idle" baseline wholesale, on the sound reasoning that a shard which is not running cannot be behind, stale, paused or agent-down. What that reasoning assumes is that not running is fine.

  • AgentDown could not have caught it either. That rule returns early when a shard has no heartbeat at all, so the alert built to detect a dead agent is structurally silent for an agent that was never alive.

  • Warning — some of the projection's shards have no agent.

  • Critical — no agent is running any shard of the projection.

Deliberately quiet in four cases: orphaned shards (a tombstone for an agent that will never run again), rebuilding shards (busy by definition), paused agents (expected state, and ProjectionPaused carries the reason), and Inline/Live projections (no async agent to be missing). It is also suppressed for the first five minutes after a node starts, because a rollout is exactly when shards are legitimately unassigned and an alert on every deploy is how an alert type gets trained away. Standing alerts still resolve during that window; only raising waits.

⚠️ It is scoped to the projection, not the shard: 76 unassigned shards produce one alert naming the tenants, not 76 alerts.

Agent Health Alerts ​

AgentDown is raised when a projection agent's last heartbeat is more than 60 seconds old. It describes an agent that stopped; an agent that never started raises ProjectionAgentUnassigned instead.

  • It is Critical only — there is no Warning tier and no consecutive-check counter.
  • The 60-second window is not configurable.
  • It is deliberately suppressed in two cases where silence is the expected state: an orphaned shard (a projection version bump or removal leaves an agent that will never heartbeat again), and an agent an operator has paused — which raises ProjectionPaused instead, carrying the reason.

Auto-resolves on the next fresh heartbeat.

Service Offline Alerts ​

ServiceOffline is raised when no telemetry from a service has reached the console for five minutes. Heartbeats travel with the rest of a service's telemetry, so this means every node of the service has gone silent. On the Services page the service's pill then reads Offline, not Healthy.

  • It is Critical. The five-minute window matches the "snapshot N min old" badge on the same page, and it is not configurable.
  • It is raised only while the console can show that it is itself ingesting: at least one other service's telemetry is current, and the console's own ingest alarm (SelfIngestFailure) is not active. Otherwise a console that has just restarted, or whose ingest has failed, would declare the whole fleet offline at once. On a console that monitors a single service, it cannot fire, because a silent service leaves nothing to compare against. The per-node heartbeat dots and the stale-snapshot badge still show the silence.
  • A service the console has never received telemetry from raises NoIngestRecord instead. That is a different finding.

Auto-resolves as soon as the service reports again. It is never cleared merely because the console stopped being able to judge. If a service was retired on purpose, evict it.

Shard Ownership Alerts ​

ShardOwnershipConflict is raised when more than one node is running the same projection shard. That is the condition a ProgressionProgressOutOfOrderException reports after the fact, and restarting the shard while both owners are still running cannot fix it.

  • CritterWatch sees it from the node each shard report comes from. Every node reports the shards it runs, so two owners show up as the reporting node alternating: three or more changes within two minutes. An ordinary rebalance changes the reporting node once (twice, if a last report from the old node arrives late), so it never raises this.
  • It is Critical, and it triggers no recovery. Find the node that should not be running the shard on the service's Nodes view, stop or eject it, and only then restart the shard.
  • It resolves once reports have kept arriving for ten minutes without alternating. A shard that has simply stopped reporting keeps the alert.
  • It does not check the reporting node against the node Wolverine assigned. A single owner running on the wrong node is not reported.

Circuit Breaker Alerts ​

Triggered immediately when a circuit breaker trips on any endpoint. Severity is always Critical. Auto-resolves when the circuit breaker resets.

Back Pressure Alerts ​

Triggered when back pressure activates on any endpoint. Auto-resolves when back pressure lifts.

Where alerts appear ​

There is no dedicated Alerts page. Alerts surface in two places an operator can act on, plus a set of counts and indicators that link into them.

Service Details → Health tab — the alert list ​

/service/{serviceName}?tab=health is the only per-alert card list, and it is scoped to one service. The tab is labelled Health (with a badge count); the heading inside the pane is Alerts. Each card carries severity, category, title, message, an optional "raised N days ago" tag, and the action buttons described below.

You arrive here from the per-service alert badge on the Services page, or from a node's alert badge on the Cluster tab. Arriving from a node badge or an endpoint's "Latched" pill sets a read-only focus chip ("endpoint {uri}" / "Node {n}", optionally a severity) that narrows the list; the only control on it is ✕ Clear filter. There is no way to set that focus from the Health tab itself, and no other filter on this surface.

Timeline — the alert history ​

/timeline is a mixed event feed, not an alert list: alert transitions are one of seven categories. Its filter bar is a free-text search, a single-select Severity (Critical / Warning / Info), category chips (Alert, Service, Node, Projection, Listener, Agent, Tenant), and an "Agent & node lifecycle" checkbox that is off by default. The layout's global service selector applies on top.

There is no Status filter and no alert-Type filter on any surface, and no Active / History tab split anywhere. The Health tab shows only active alerts; the Timeline shows every transition including resolutions and clears, but as feed entries rather than as alert rows you can act on.

Counts and indicators ​

SurfaceWhat it showsWhere it goes
Header bella count badge, no dropdown/timeline, with no pre-filter
Dashboard Active Alerts tilecount + critical/warning pills/timeline?severity=…
Dashboard Recent Eventslast 10 timeline entriesthe matching timeline entry
Services page, per servicebadge countthat service's Health tab
Cluster tab, per nodebadge countHealth tab, focused on that node
Endpoints / Projections / Topologya warning icon or coloured dot— (indicator only)

⚠️ Three alert counts legitimately disagree, by design. The bell counts active non-DLQ alerts, coalesced so a per-tenant alert counts once per projection + condition rather than once per tenant. The dashboard's Active Alerts tile includes DLQ alerts. The API returns raw records. So the bell is normally lower than the dashboard tile. Both tooltips explain this in place.

What an alert card does not do ​

Clicking an alert opens nothing. There is no detail panel, drawer, dialog, expandable row, or detail route — only the card's own buttons act. In particular there is no State Timeline on the card: nothing groups an alert's transitions together, and the metric value that triggered a transition is not displayed anywhere.

The transitions themselves are not lost — each one lands on the Timeline as its own entry (raised, elevated, reduced, resolved, cleared), and operator actions additionally land in the Audit Log with an actor. Reconstructing one alert's history means filtering the Timeline to it by hand.

Operator actions ​

Acknowledge / Snooze / Clear / Dismiss ​

ActionWhereEffect
AcknowledgeHealth tab, TimelineRecords AlertAcknowledged and removes the alert from the console immediately.
SnoozeHealth tab, TimelineSuppresses the alert for one hour.
ClearHealth tab onlyRecords AlertCleared and closes the alert.
DismissHealth tab onlyHides the card in this browser only. Nothing is sent to the server.

⚠️ Snooze offers no duration picker. Every call site is hardcoded to one hour and the toast reads "Alert snoozed for 1 hour". The wire message accepts an arbitrary duration, but no UI ever sends anything else.

⚠️ Acknowledge is one-way in the UI. The alert is dropped from the client store the moment it is acknowledged, and no surface can show it again. The record itself stays active — acknowledging does not resolve or close an alert — so the console and the event store disagree about it from that point on.

⚠️ Clear takes no note. ClearAlert carries only the stream id and an actor, and the actor is hardcoded to UI User. There is no note field and no confirmation prompt. The same is true of Acknowledge and Snooze.

Clear and Dismiss appear only on alert types that cannot auto-resolve — MetricsExecTime, MetricsThroughput, MetricsFailureRate, MetricsDlqRate, BackPressure, AgentHealth. Every other type shows "Will auto-resolve when conditions improve" in place of the Clear button.

On a coalesced per-tenant card, Acknowledge and Snooze apply to every member of the group (the toast says "Acknowledged across N tenants"); Clear and Dismiss act only on the representative alert.

Remediation actions ​

Each card offers context buttons for its alert type. The buttons are navigation, with two exceptions — Restart projection and Restart listener, which dispatch a command behind a confirmation prompt and are disabled without an operations license.

Alert typeButtons
MetricsExecTimeView slow message types
MetricsDlqRate, MetricsFailureRateView dead letters
MetricsThroughputView health & metrics
HighWaterStale, HighWaterAgentRestartedView projections / View high-water mark
ProjectionLag, ProjectionStale, AgentDown, ProjectionAgentUnassignedView projection, Restart projection
ProjectionPausedRestart projection
ProjectionDeadLettersView projection
BackPressureView listener
CircuitBreakerView listener, Restart listener
AgentHealthView agents
anything elsenone

⚠️ There is no Replay All, Discard All, Rebuild Projection, Eject Node, or View Service button on an alert. Bulk DLQ replay and discard live on the Dead Letters surface; rebuild lives on the projection pages; eject lives on the Cluster tab. Reaching them from an alert means following its View button and acting there.

The Timeline's alert rows carry the same remediation buttons plus Acknowledge and Snooze, but not Clear, Dismiss, or the View CTA.

How Metrics-Based Alerts Are Determined ​

Some alerts (DLQ count, projection lag, agent health, circuit breaker) are driven by direct counts and trigger as soon as a configured threshold is crossed. Metrics-based alerts — throughput abnormal and execution time abnormal — are different. They compare a current reading against a baseline and only raise when the deviation is large enough and has persisted long enough.

The baseline cascade ​

For each (service, message type) pair the evaluator picks the effective baseline in this order:

  1. Observed history — the average for this service / message type computed from MetricsSample documents in the CritterWatch event store. Only used when it is mature:
    • the oldest sample is at least BaselineMinimumDays old (default 10),
    • there are at least BaselineMinimumSamples buckets (default 100), and
    • for throughput, the observed value is above BaselineMinimumThroughputPerHour (default 1.0/hr) — this stops a single fractional sample from triggering huge multipliers.
  2. Declared baseline — a value supplied by an operator. Comes from one of two sources:
    • configureBaselines on AddCritterWatchMonitoring in the monitored application — see Registration › Declared Baselines.
    • The CritterWatch UI under Settings → Alert Configuration → Baselines.
  3. None → the alert is suppressed during the warmup window. This is the default behaviour for any newly-monitored service that hasn't shipped a declared baseline. Set SuppressThroughputAlertsDuringWarmup / SuppressExecTimeAlertsDuringWarmup to false to opt out (rarely needed — it will alert noisily until the service settles in).

Each baseline change is recorded as a ThroughputBaselineChanged / ExecTimeBaselineChanged event so the audit trail makes the provenance clear: Source = ServiceCapabilities for values advertised by the monitored app at startup, Source = Operator for values typed into the UI.

Threshold multipliers ​

Once an effective baseline exists, the evaluator compares current readings against multipliers / percentages:

SettingDefaultMeaning
ThroughputWarningMultiplier3Warn when current rate is ≥ 3× baseline
ThroughputCriticalMultiplier10Critical when current rate is ≥ 10× baseline
ExecTimeWarningPercent50Warn when avg exec time is ≥ baseline + 50%
ExecTimeCriticalPercent200Critical when avg exec time is ≥ baseline + 200%
FailureRateWarningPercent5Warn when failures > 5% of executions
FailureRateCriticalPercent20Critical when failures > 20% of executions
DlqRateWarningPerHour10Warn when DLQ rate ≥ 10/hr
DlqRateCriticalPerHour50Critical when DLQ rate ≥ 50/hr

Hysteresis (K-of-N) — noise suppression ​

A single noisy evaluation never raises or resolves an alert by itself. Instead the evaluator tracks consecutive passes:

SettingDefaultMeaning
HysteresisRaisePasses2A breach must be observed in this many consecutive passes before the alert is raised (or its severity changed).
HysteresisResolvePasses2A below-threshold reading must be observed in this many consecutive passes before an active alert auto-resolves.

With the default 30-second evaluator cadence this means ~60 seconds of confirmation either way before any state change reaches the user. Increase the values to be more conservative; setting either to 1 disables that side of the hysteresis.

Cascade order ​

Every metrics-based threshold and every baseline-policy knob supports the same three-level cascade — most specific wins:

  1. Per-message-type override (MessageTypeAlertThresholds)
  2. Per-service override (ServiceMetricsAlertOverrides)
  3. Global default (MetricsAlertDefaults)

That includes the multipliers, the failure-rate / DLQ thresholds, the hysteresis pass counts, the warmup-suppression flags, and the declared baselines themselves.

Editing Thresholds ​

There are two ways to change any of the settings above.

From the CritterWatch UI ​

Settings → Alert Configuration exposes three tabs matching the cascade:

  • Global Defaults — applied to every service unless overridden.
  • Service Overrides — pick a service and override any global value. Leave fields blank to inherit from the global default.
  • Message Type Overrides — pick a service + message type combination and override any value. Leave fields blank to inherit from the service or global default.

Edits are stored as documents in the CritterWatch event store and take effect on the next evaluator pass (≤30s). Every edit produces an AlertConfigChanged audit entry, and edits to declared baselines additionally emit ThroughputBaselineChanged / ExecTimeBaselineChanged events with Source = Operator.

Programmatically — configureBaselines ​

For declared baselines specifically, the cleanest way to seed values for a fresh service is at the monitored side:

csharp
opts.AddCritterWatchMonitoring(
    critterWatchUri,
    systemControlUri,
    configureBaselines: baselines => baselines
        .ForService(throughputPerHour: 200, avgExecTimeMs: 25)
        .For<TripBooked>(throughputPerHour: 50, avgExecTimeMs: 40)
);

These flow over on first contact, and an operator can later refine them through the UI (the operator-supplied value will then take precedence in the cascade — configureBaselines only seeds; the UI is authoritative).

Preset profiles ​

Two convenience presets are available from the Settings page:

  • Production Profile — strict thresholds, suitable for production environments
  • Development Profile — relaxed thresholds, reduces noise during development

See Configuration Reference for the full programmatic API.

Free for read-only monitoring. A commercial license is required for administrative actions and the MCP server.