Configuration Reference
Monitored Service Configuration
AddCritterWatchMonitoring()
Call this inside UseWolverine() in each service you want to monitor.
Simple form — the telemetry + control queue URIs (the service identity is the Wolverine ServiceName, not a parameter):
builder.Host.UseWolverine(opts =>
{
// CritterWatch keys this service by its Wolverine ServiceName — there is no
// separate service-name parameter on AddCritterWatchMonitoring.
opts.ServiceName = "my-service";
opts.AddCritterWatchMonitoring(
critterWatchUri: new Uri("rabbitmq://queue/critterwatch"),
systemControlUri: new Uri("rabbitmq://queue/my-service-control"));
});Full options form:
builder.Host.UseWolverine(opts =>
{
opts.AddCritterWatchMonitoring(
// URI of the queue CritterWatch listens on for telemetry
critterWatchUri: new Uri("rabbitmq://queue/critterwatch"),
// URI of the queue this service listens on for incoming commands
systemControlUri: new Uri("rabbitmq://queue/trip-service-control"),
// How this service exports metrics (default: Hybrid)
metricsMode: WolverineMetricsMode.Hybrid
);
});Options Reference
Parameters of AddCritterWatchMonitoring():
| Parameter | Default | Description |
|---|---|---|
critterWatchUri | — | URI of the queue CritterWatch listens on for this service's telemetry |
systemControlUri | (none) | URI of the queue this service listens on for commands from CritterWatch. Null opts out of operator commands. |
metricsMode | Hybrid | How this service exports metrics — see metrics modes |
heartbeatInterval | 30 seconds | Cadence of the WolverineHeartbeat liveness ping each node sends. The dashboard's liveness dot turns amber after one missed beat and red after five. TimeSpan.Zero disables heartbeats, and the service's nodes then render grey/unknown. |
configureBaselines | (none) | Optional callback to declare expected throughput/execution-time baselines for this service. See Alerts › Editing Thresholds. |
configureShardedTopology | (none) | Optional callback declaring sharded telemetry queue slots. Supplying it is what turns clustering on — see Clustering. |
The service's identity is the Wolverine ServiceName, not a parameter here.
Properties on the returned CritterWatchOptions:
| Property | Default | Description |
|---|---|---|
PublishInterval | 2 seconds | How long the observer accumulates telemetry before publishing it as one ServiceUpdates. See Publish cadence below. |
SendDeltaTelemetry | true | Push only the shard states, per-tenant high-water marks and dead-letter counts that changed, with a periodic full pass. See Delta telemetry. |
FullSnapshotInterval | 2 minutes | How often the progression poller emits a complete shard-state pass for reconciliation. |
ScaleCadenceWithFleetSize | true | Scale the publish cadence with this service's observed fan-out. See Cadence scaling. |
MaxPublishInterval | 60 seconds | Ceiling the cadence scaling may back off to. |
MaxItemsPerPush | 5000 | Hard cap on telemetry items folded into a single push, after coalescing. |
MaxPushWireBytes | 196608 (192 KiB) | Largest a single push may be on the wire, compressed. A larger one is published as several. See Push size. |
CaptureConversations | true | Capture per-envelope causality (conversation/correlation/parent-span) for the conversation graph. |
ConversationCaptureBufferSize | 5000 | Bounded client-side buffer of captured hops between flushes; oldest are dropped when full. |
EnableEventStoreExplorer | (on in Development, off elsewhere) | Opt in to every read of this service's event data: the Event Store Explorer (stream history, tag and filter queries), the stream fetch, and the query_events MCP tool. Since 1.1 it covers all of them, not only some (GH-1211). Event payloads are business data — see Event Store Explorer › Opt-in. |
TraceProvider(name) | (none) | Declare the preferred CritterWatch-side trace provider. |
MetricsDataSource(name) | (none) | Declare the preferred CritterWatch-side metrics data source. |
EnableSqlQuery | (follows EnableEventStoreExplorer) | Opt in to the read-only SQL view and the query_sql MCP tool (GH-1235). Set it explicitly to refuse SQL while leaving the rest of the explorer on. See SQL query. |
EnableDocumentExplorer | (follows EnableEventStoreExplorer) | Opt in to the Document Explorer's reads of stored documents (type list, pages, one by id) and the query_documents / get_document MCP tools (GH-1307). Its own switch because a service can have document stores and no event store. See Document explorer opt-in. |
CompactStreams(...) | (none) | Declare a stream compaction policy — "compact any stream of aggregate type X whose un-compacted growth exceeds N, on this cron schedule". See Stream compaction. |
UseMessagePackWireFormat | false | Send steady-state telemetry batches as MessagePack instead of JSON. See Wire format — mind the upgrade ordering. |
AllowDurableEndpoints | false | Escape hatch for the startup guard that refuses to run CritterWatch's own endpoints in Durable mode. See Durable endpoint guard — the right fix is nearly always to scope the offending policy instead. |
Agent health polling is a process-wide setting rather than a per-service option: CritterWatchObserver.HealthCheckInterval (default 60 seconds) and CritterWatchObserver.StateSnapshotInterval (default 60 seconds).
Publish cadence and console write volume
Every publish is one write of this service's ServiceSummary document in the console's store, and that document grows with the number of endpoints, handlers and message types the service reports — hundreds of kilobytes is normal for a large service. Because the projection is Inline, each write is a full document rewrite, so PostgreSQL write volume, TOAST churn and autovacuum work all scale linearly with the publish rate. A single-service deployment publishing every second was measured holding a one-row mt_doc_servicesummary table at 370 MB with 19,266 autovacuums.
Widening PublishInterval never drops telemetry — updates merge into the same batch, so the only effect is up to that much extra latency before the console sees them. The default is far inside the console's own cadences (30 s alert evaluation, 60 s health-check and state-snapshot intervals), so raising it is cheap:
builder.Host.UseWolverine(opts =>
{
var critterWatch = opts.AddCritterWatchMonitoring(
critterWatchUri: new Uri("rabbitmq://queue/critterwatch"),
systemControlUri: new Uri("rabbitmq://queue/trip-service-control"));
// Each publish is a full rewrite of this service's ServiceSummary document in the
// console's store, so write volume scales linearly with this cadence. Nothing is
// dropped by widening it — updates merge into the same batch.
critterWatch.PublishInterval = TimeSpan.FromSeconds(10);
});Large deployments — many endpoints, many projection agents, or a console store shared with the monitored application's database — should raise it. Values below 250 ms are clamped.
Delta telemetry and reconciliation
The leader node polls every event database's progression rows every 15 seconds and reports a baseline for every registered shard, a high-water mark per tenant and per database, and a dead-letter count per shard. On a large fleet the overwhelming majority of that re-states marks that did not move — a service with 518 message databases and 2,173 tenants makes this the second-largest traffic term after metrics, and the console discards nearly all of it as a no-op write.
With SendDeltaTelemetry on (the default) the poller sends only what changed, and emits a complete pass every FullSnapshotInterval. Treat that interval as a correctness parameter rather than a tuning knob: the periodic full snapshot is truth and the deltas are commentary, so it is the upper bound on how long a console that dropped a delta can disagree with the service. Every full pass restates everything, which is what heals it.
Two things are deliberately never filtered:
- The real-time projection observer. Its snapshots carry the
LastAdvanced/LastHeartbeattimestamps the console computes the ProjectionStale and AgentDown alerts from, so withholding one would let a stored timestamp age past a threshold reality had not crossed. - A shard-identity reconciliation marker's contents. The marker is the console's only authority for removing a shard, so it is sent whole or withheld whole — never partially. An unchanged marker is withheld because the console has already applied exactly that pruning, and a full pass always resends it.
Each push declares its own completeness on the wire, so a console never has to infer it. A console that predates the declaration, or one reading a push that carries none, keeps its existing behaviour: shard state is merged, and absence from a push has always meant "unchanged", never "gone".
Cadence scaling by fleet size
A monitored service's telemetry volume is roughly message types × destinations × tenants, and a flat PublishInterval is blind to all three. With ScaleCadenceWithFleetSize on (the default), the client derives its effective publish cadence from the largest fan-out dimension it can observe about itself — tenants, message databases, or endpoints — using a fixed, published table:
| Observed fan-out | Multiplier | Effective interval at the 2 s default |
|---|---|---|
| under 250 | 1x | 2 s (unchanged) |
| 250 – 999 | 2x | 4 s |
| 1,000 – 4,999 | 5x | 10 s |
| 5,000 and up | 10x | 20 s |
The result is clamped to MaxPublishInterval. A small service is completely unaffected; a 2,000-tenant service lands on a 10-second cadence without an operator having to discover that number themselves. Batches held back by the scaling are coalesced — a superseded gauge reading replaces the earlier one rather than queueing behind it — and capped at MaxItemsPerPush so no single push can grow without bound. Additive metrics samples are never coalesced away.
The governor logs once when it changes cadence, naming the observed fan-out and the option that turns it off, so a console updating every 20 seconds is always explainable from the service's own log. Set ScaleCadenceWithFleetSize = false to pin PublishInterval exactly.
builder.Host.UseWolverine(opts =>
{
var critterWatch = opts.AddCritterWatchMonitoring(
critterWatchUri: new Uri("rabbitmq://queue/critterwatch"),
systemControlUri: new Uri("rabbitmq://queue/trip-service-control"));
// All four of these are ON by default at the values shown; set them only to
// override the defaults.
// Send only the shard states, per-tenant high-water marks and dead-letter
// counts that actually changed since the last push...
critterWatch.SendDeltaTelemetry = true;
// ...with a COMPLETE pass this often, which is the upper bound on how long a
// console that dropped a delta can disagree with this service.
critterWatch.FullSnapshotInterval = TimeSpan.FromMinutes(2);
// Pace publishes by this service's own observed fan-out (tenants, message
// databases, endpoints) instead of a flat PublishInterval regardless of size.
critterWatch.ScaleCadenceWithFleetSize = true;
// The ceiling that scaling may back off to.
critterWatch.MaxPublishInterval = TimeSpan.FromSeconds(60);
});Metrics are a separate lever
Cadence scaling and delta telemetry shape the progression and health telemetry. On a large multi-tenant fleet the dominant term is usually still per-tenant metrics — for those, raise Wolverine's Metrics.SamplingPeriod, or switch to metricsMode: WolverineMetricsMode.SystemDiagnosticsMeter and let an existing Prometheus stack carry them. See metrics modes.
Service Name Must Be Unique
The ServiceName is used as the Marten event stream key. Two services with the same name will overwrite each other's state. Use a name that uniquely identifies the service across your entire deployment.
Durable endpoint guard
CritterWatch's own telemetry and control endpoints refuse to start in Wolverine's EndpointMode.Durable, and the refusal is deliberate (GH-954). A durable CritterWatch endpoint writes every telemetry envelope through the monitored application's own message store, multiplying its database traffic — which turns the monitoring tool into the outage it exists to report.
The usual cause is not a CritterWatch setting at all: an application-wide policy such as UseDurableInboxOnAllListeners() reaches into system endpoints it was never meant to cover. The right fix is to scope that policy:
// Prefer this — leave CritterWatch's endpoints alone.
opts.Policies.UseDurableInboxOnAllListeners();
opts.ListenToRabbitQueue("critterwatch").BufferedInMemory();AllowDurableEndpoints = true suppresses the guard if you genuinely need it. Setting it means accepting the write amplification above, so treat it as a last resort rather than a way to get past a failed startup.
Wire format
Steady-state telemetry batches are JSON by default. UseMessagePackWireFormat = true switches them to MessagePack. It does not make them smaller — both formats ride the same Brotli compression, and on the wire they come out within a few percent of each other. What it saves is the console's ingest cost: measured on the shipped format, the console decodes a MessagePack batch about 3x faster with roughly half the allocation, which matters on a high-volume fleet where decoding telemetry shows up in the console's CPU and heap. If the problem is message size against a broker limit, see Push size instead.
Upgrade ordering
Leave this off until the console is on a version that reads the MessagePack frames (1.0.0-rc.8+). An older console receiving them cannot deserialize them. The console always accepts both formats, so a mixed fleet — some services on JSON, some on MessagePack — is fine, and the flag can be flipped one service at a time.
First-activation pushes (the ones carrying Capabilities) stay JSON regardless of this flag; only the steady-state batches change format.
Push size
A broker caps the size of one message — Amazon SQS at 256 KiB, Azure Service Bus standard tier at 256 KB — and on a large fleet a single telemetry push can reach it. MaxItemsPerPush does not prevent that, because it counts items rather than bytes: one agent-health report is one item however many agents it covers. Since 1.1 the push is also bounded in the broker's own units (GH-1267):
- Pushes are compressed harder. Telemetry was Brotli-compressed at the fastest setting; it now uses quality 6, which measured 55–62% smaller on a modelled 2,173-tenant, 514-database push for a few extra milliseconds per push on the monitored service. Consoles and services on either side of the change read each other's frames.
- A push that would still exceed
MaxPushWireBytesis published as several. Shard states, agent health, persisted counts and metrics are divided across the pieces; the console merges them entry by entry, so nothing is lost or reordered. The reconciliation markers that remove stale shards always travel whole, and the handshake, the node roster and change events ride the first piece. The service logs once, at Information, when it starts dividing pushes.
The 192 KiB default leaves room under SQS's cap for the envelope's own headers. Raise it only on a transport with a larger limit, or set it to 0 to never divide. A single section that cannot be divided and is still too large — in practice only a very large capabilities announcement — is published as it is and logged at Error, as before.
Document explorer opt-in
EnableDocumentExplorer gates the Document Explorer's reads of stored documents: the type list, a page of a type, one document by id, and the query_documents / get_document MCP tools. It follows EnableEventStoreExplorer unless you set it, which is how it behaved before it existed. It is a switch of its own for two reasons. A service can have document stores and no event store at all. And whole business documents are worth refusing separately from event reads: set EnableDocumentExplorer = false to keep the event explorer and refuse documents. Refused reads are answered by name, never with an empty page.
On a console with permissions wired, a page or a by-id read also needs the documents.read capability (mcp.documents.read for the tools), and every read writes an audit entry.
SQL query opt-in
EnableSqlQuery gates the Event Explorer's SQL view and the query_sql MCP tool. It follows EnableEventStoreExplorer unless you set it, so a service that has opted into the explorer has also opted into SQL. Set it to false to keep the rest of the explorer while refusing SQL:
critterWatch.EnableEventStoreExplorer = true;
critterWatch.EnableSqlQuery = false; // explorer yes, ad-hoc SQL noA SELECT reaches every projection table in the store, so each run writes an audit entry — "who read what" is the audit log's question. The statement itself is guarded: a single statement, table references allow-listed to the objects the store declares, a read-only transaction that is always rolled back, a statement timeout, and a row cap reported as truncated rather than presented as a total.
CritterWatch Server Configuration
AddCritterWatch()
var builder = WebApplication.CreateBuilder(args);
builder.AddCritterWatch(
builder.Configuration.GetConnectionString("critterwatch")!,
opts =>
{
opts.UseRabbitMq(new Uri("amqp://localhost")).AutoProvision();
opts.ListenToRabbitQueue("critterwatch").ProcessInParallelWithNativeAcks().UseCritterWatchSerializer();
});
// Single-node is the default — nothing else to configure. For a multi-node
// cluster, supply configureClusterShardedTopology (see Deployment › Clustering).
var app = builder.Build();
app.UseCritterWatch();
app.Run();UseCritterWatch()
Maps all CritterWatch middleware into the ASP.NET Core pipeline:
app.UseCritterWatch();
// With a custom SignalR route:
app.UseCritterWatch(signalRRoute: "/my-hub");UseCritterWatch() registers:
- Wolverine HTTP endpoints under
/api/critterwatch/* - SignalR hub at
/api/messages(configurable) - Static file serving for the embedded Vue SPA
- Client-side routing fallback for the SPA
Storage schema
All CritterWatch data — its documents (service summaries, alerts, metrics rollups) and its event store — is isolated to a dedicated database schema, so it never collides with the host application's own tables. The default schema is critterwatch, and the name is configurable via the schemaName parameter:
builder.AddCritterWatch(
builder.Configuration.GetConnectionString("critterwatch")!,
schemaName: "monitoring"); // default: "critterwatch"The same schemaName parameter is available on the lower-level opts.AddCritterWatchServices(...) registration (both the Marten/PostgreSQL and Polecat/SQL Server flavors). CritterWatch only ever creates and migrates tables inside that one schema.
Why a dedicated schema matters
This is what lets CritterWatch share a database with the application it monitors without stepping on it. Point schemaName at any schema you like; CritterWatch keeps all of its storage there.
Docker Compose
A complete docker-compose.yml for local development:
services:
postgres:
image: postgres:16
environment:
POSTGRES_USER: postgres
POSTGRES_PASSWORD: postgres
POSTGRES_DB: critterwatch
ports:
- "5432:5432"
volumes:
- postgres_data:/var/lib/postgresql/data
rabbitmq:
image: rabbitmq:3-management
ports:
- "5672:5672" # AMQP
- "15672:15672" # Management UI
environment:
RABBITMQ_DEFAULT_USER: guest
RABBITMQ_DEFAULT_PASS: guest
volumes:
postgres_data:Connection String Formats
PostgreSQL
Host=localhost;Port=5432;Database=critterwatch;Username=postgres;Password=postgresFor cloud providers:
Host=my-postgres.postgres.database.azure.com;Database=critterwatch;Username=app@my-postgres;Password=secret;SSL Mode=RequireConnection pool ceiling
Npgsql defaults Max Pool Size to 100 per process. The console runs a number of independent readers (projection progression, rebuild batches and cells, table-size estimates, envelope counts, metrics) plus Wolverine's durability agent; a single console pod monitoring one service was measured holding ~25 connections, most of them idle.
CritterWatch therefore applies its own default ceiling of 50 when your connection string doesn't specify one. An explicit Max Pool Size in the connection string always wins — including a larger one. To change the default without touching the connection string:
{
"CritterWatch": {
"Postgres": {
"MaxPoolSize": 25
}
}
}Consoles sharing a database with the monitored application
If the console's store lives on the same instance as the application it monitors, size this deliberately. Exhausting max_connections there takes down the monitored application, not just the console — the failure mode is FATAL: remaining connection slots are reserved for roles with privileges of the "pg_use_reserved_connections" role in the application's logs.
RabbitMQ
amqp://guest:guest@localhost:5672/
amqps://user:pass@my-rabbit.cloud:5671/ # TLSAmazon SQS
AddCritterWatchMonitoring works identically against the SQS transport. Telemetry pushes are kept under SQS's 256 KiB message limit by MaxPushWireBytes. Two queues are involved per monitored service:
| Queue | Direction | Purpose |
|---|---|---|
critterWatchUri | service → CritterWatch | Metrics, heartbeats, capability snapshots |
systemControlUri | CritterWatch → service | Pause / restart listeners, rebuild projections, DLQ ops |
opts.UseAmazonSqsTransport();
opts.AddCritterWatchMonitoring(
critterWatchUri: SqsEndpointUri.Queue("critterwatch"),
systemControlUri: SqsEndpointUri.Queue("critterwatch-control-trip-service"));Dead letter queues with AutoProvision() off
By default the WolverineFx.AmazonSqs transport attaches every listener to its DefaultDeadLetterQueueName (wolverine-dead-letter-queue). With AutoProvision() enabled the broker creates that queue on startup, so the control listener AddCritterWatchMonitoring installs is wired up cleanly.
In production environments with AutoProvision() off, that default DLQ typically isn't pre-provisioned. The broker startup then fails:
Wolverine.AmazonSqs.WolverineSqsTransportException: Error while trying to
initialize Amazon SQS queue 'wolverine-dead-letter-queue'
---> Amazon.SQS.Model.QueueDoesNotExistExceptionThree ways to handle it, pick whichever fits your infra automation:
1. Provision the default DLQ alongside your application queues (CDK / Terraform / etc.). Nothing to change in code.
2. Point the control listener at an existing DLQ you already provision. Re-open the same endpoint after AddCritterWatchMonitoring and chain DeadLetterQueueName:
opts.AddCritterWatchMonitoring(critterWatchUri, systemControlUri);
opts.ListenToSqsQueue("critterwatch-control-trip-service", q =>
{
q.DeadLetterQueueName = "trip-service-dlq";
});3. Disable native DLQs on this transport entirely (rely on Wolverine's durability instead):
opts.UseAmazonSqsTransport(t => t.DisableAllNativeDeadLetterQueues());
opts.AddCritterWatchMonitoring(critterWatchUri, systemControlUri);Same shape applies to other native-DLQ transports
Azure Service Bus and the other transports with a default DLQ behave the same way. Whichever DLQ strategy you pick for the rest of your application's listeners, the queues AddCritterWatchMonitoring installs inherit the same contract — there's no special CritterWatch DLQ to provision.
Native DLQs and CritterWatch management
CritterWatch's Dead Letter Queue explorer manages the durable (database) DLQ only — the failed messages Wolverine persists to its message store. It has no visibility into broker-native dead-letter queues (Amazon SQS DLQ, Azure Service Bus $DeadLetterQueue, RabbitMQ DLX). A message that dead-letters natively stays at the broker and won't appear in CritterWatch until it's forwarded into the Wolverine database.
So for CritterWatch to manage a service's dead letters, those failures need to land in — or be propagated into — the durable store. Two approaches:
A. Send failures straight to the durable store (no native DLQ). Simplest for CritterWatch; it changes where your dead letters live, so weigh it against any existing native-DLQ tooling or alarms you rely on.
// Amazon SQS — disable native DLQs transport-wide; failures go to the
// Wolverine durability database instead.
opts.UseAmazonSqsTransport(t => t.DisableAllNativeDeadLetterQueues());On RabbitMQ the per-endpoint equivalent is the WolverineStorage dead-letter mode (failures bypass the native DLX and go to the durable store).
B. Keep the native DLQ and forward it into the durable store. Preserves your existing native dead-lettering and surfaces those messages in CritterWatch — the better fit when retrofitting onto a running system.
RabbitMQ has this built in — Wolverine listens on the native DLQ, reconstructs each envelope, and writes it to the durability database where CritterWatch can replay or discard it:
csharpopts.UseRabbitMq(/* ... */).EnableDeadLetterQueueRecovery();Amazon SQS / Azure Service Bus don't yet have a built-in equivalent (tracked upstream in wolverine#3103). Until it lands, either use approach A, or stand up a listener on the native DLQ that forwards each message to the store via
IMessageInbox.MoveToDeadLetterStorageAsync(...).
Retrofitting without disrupting the host
Adding CritterWatch shouldn't change queues your application already owns. Keep AutoProvision() scoped (or pre-provision CritterWatch's critterwatch and control queues through your infrastructure automation), and use per-endpoint DLQ settings so CritterWatch never re-declares or alters the host's existing dead-letter infrastructure.
appsettings.json
The recommended approach for managing connection strings:
{
"ConnectionStrings": {
"critterwatch": "Host=localhost;Database=critterwatch;Username=postgres;Password=postgres",
"rabbitmq": "amqp://localhost"
}
}builder.AddCritterWatch(
builder.Configuration.GetConnectionString("critterwatch")!,
opts =>
{
var rabbitUri = new Uri(
builder.Configuration.GetConnectionString("rabbitmq")!);
opts.UseRabbitMq(rabbitUri).AutoProvision();
opts.ListenToRabbitQueue("critterwatch").ProcessInParallelWithNativeAcks().UseCritterWatchSerializer();
});Timeline retention
The Health Timeline (and the dashboard's Recent Events widget) is backed by an append-only TimelineEntry document per lifecycle fact. CritterWatch sweeps that table on a background timer — hourly by default, on the cluster leader only — so it stays bounded without an operator ever having to run a cleanup script.
The sweep makes three passes, each independently configurable under CritterWatch:Timeline:
{
"CritterWatch": {
"Timeline": {
"RetentionPeriod": "30.00:00:00",
"PruneInterval": "01:00:00",
"CompactRedundantEntries": true,
"CompactableEventTypes": [ "AgentStarted" ],
"CompactionGracePeriod": "00:15:00",
"MaxEntriesPerService": 0,
"DeleteBatchSize": 500,
"MaxDeletesPerSweep": 50000
}
}
}| Key | Default | Meaning |
|---|---|---|
RetentionPeriod | 30.00:00:00 (30 days) | Entries older than this are deleted. 00:00:00 keeps entries forever. |
PruneInterval | 01:00:00 (1 hour) | How often the sweep runs. Every pass is idempotent, so the cadence isn't load-bearing. |
CompactRedundantEntries | true | Collapse a run of consecutive identical entries — same service, event type, subject, title, severity, description — down to the first entry of the run. |
CompactableEventTypes | [ "AgentStarted" ] | Which event types compaction applies to. Configuring this replaces the default list. |
CompactionGracePeriod | 00:15:00 | Entries newer than this are never compacted, so the sweep can't race the live feed. |
MaxEntriesPerService | 0 (off) | Opt-in hard cap: keep only the newest N entries per service, regardless of age. |
DeleteBatchSize | 500 | Documents deleted per transaction. Pruning a large backlog stays chunked instead of taking one long table-wide lock. |
MaxDeletesPerSweep | 50000 | Ceiling on deletions per sweep. A big accumulated backlog drains over successive sweeps rather than in one transaction storm. 0 = no ceiling. |
Why compaction matters more than the age window. Before 1.0.0-beta.4, every keep-alive from a monitored service re-reported its running agents, and each re-report was materialized as a fresh "Agent started" entry — thousands per hour for a modest fleet, burying the events an operator actually cares about. That emit path is fixed (an AgentStarted is now written only when the agent → node assignment really changed), but existing deployments still carry the accumulated rows, and those rows are all inside any sane retention window — no age policy would ever remove them. Compaction is what reclaims them: it keeps the entry that recorded the transition and deletes the re-reports behind it. A genuine change (an agent moving to a different node) breaks the run and is always preserved.
Keeping long history
Set RetentionPeriod to 00:00:00 to keep timeline entries indefinitely. Compaction still runs, so the redundant re-reports go away while every real transition is kept.
Metrics storage and sizing
Read this before running the console against a system under load
The built-in metrics collection is the largest thing the console writes, by roughly two orders of magnitude. A production deployment measured mt_doc_metricssample at 2,590 MB / 2.65 million rows in two and a half hours — about 1 GB/hour — against 20 MB for the next-largest document table. On that fleet it accounted for ~99% of the console's database traffic.
The default is the right choice for getting started: zero extra infrastructure, works out of the box. It is not the right choice for a system under real load. See Use an external metrics stack under load below, and size the knobs in this section deliberately if you stay on the built-in mode.
What gets stored
The console persists one MetricsSample row per (service, message type, destination, tenant, bucket). Multiply those out before estimating — the tenant axis is what surprises people:
- Each 1-minute bucket writes a row per (message type × destination) per service.
- On a multi-tenant service, the console also writes one row per real tenant alongside the store-global aggregate, so per-tenant evaluation cannot be averaged away by a busy neighbour. Since 1.0 the tenant rows are capped at
PerTenantSampleCeiling(default 25) per series-bucket — the busiest tenants keep their own rows, the remainder folds into one"*OTHER*"row — so a service with 850 active tenants writes bounded per-tenant rows instead of 850 per series. - Only active combinations materialize rows, so the table grows with tenant activity, not with tenant count — quiet tenants becoming active push it up without any fleet change.
Rows fold to hourly granularity past SampleHotWindow, and whole monthly partitions drop past SampleRetentionPeriod — but both act on data you have already written and paid to store, and the hourly fold preserves every dimension of the key, tenant included. The hourly tail is smaller only by the 60:1 bucket fold; it still carries the full tenant multiplier.
Steady-state row count, using per-bucket figures you can measure on your own deployment (select count(*) from mt_doc_metricssample group by bucket_start on a recent minute):
rows ≈ active (message type × destination × (capped tenants + 1)) combos
× (hot-window minutes + retained-tail hours)A field data point for scale: one monitored service with 512 shard databases and ~530 active tenants measured ~4,300 rows per minute-bucket (~912 bytes each) — ~12.5M resident rows for a 2-day hot window plus a 43-day hourly tail of similar order, before the per-tenant cap shipped.
Retention and sampling knobs
Metric samples are bucketed, rolled up and pruned on background timers. The knobs bind from CritterWatch:Metrics:
{
"CritterWatch": {
"Metrics": {
"SampleRetentionPeriod": "45.00:00:00",
"SampleHotWindow": "2.00:00:00",
"SampleBucketWidth": "00:01:00",
"BaselineLookback": "28.00:00:00"
}
}
}| Key | Default | Description |
|---|---|---|
SampleRetentionPeriod | 45.00:00:00 | How long MetricsSample rows are kept. The table is range-partitioned by month, so retention drops whole partitions rather than deleting rows. 00:00:00 disables pruning. |
SampleHotWindow | 2.00:00:00 | How long samples stay at 1-minute granularity. Past this, the leader folds each hour into one hourly row. Baselines read across both granularities, so totals are unaffected — you lose only within-hour shape for older data. |
SampleBucketWidth | 00:01:00 | Quantization width for incoming samples — one row per window instead of ~60 one-second inserts. Widening it thins the throughput evaluator's 15-minute read window (2–3 closed buckets at 00:05:00 instead of ~15) and can under-report throughput enough to trip low-throughput alerts falsely — prefer the hot window and tenant ceiling as sizing levers. |
PerTenantSampleCeiling | 25 | Cap on distinct per-tenant rows per (service × message type × destination × bucket) series. The busiest tenants keep their own rows; the remainder folds into one "*OTHER*" row, so totals stay exact while the tenant multiplier is bounded. Per-tenant metrics alerting only evaluates tenants with their own rows. 0 removes the cap. |
SampleFlushInterval | 00:00:15 | How often the accumulator flushes closed in-memory buckets to the store. |
PersistOpenBuckets | false | Write still-accumulating buckets on every flush. Off by default: at the defaults it wrote each row four times per minute, three of them an intermediate total superseded before anything read it. Turn it on only if you need the open bucket visible to readers within the bucket width. |
SamplePruneInterval | 01:00:00 | How often the retention prune runs (leader only). |
SampleRollupInterval | 01:00:00 | How often aged minute rows are folded to hourly (leader only). |
MaxRollupHoursPerPass | 24 | Cap on hours folded per rollup pass, so a large backlog drains over successive passes instead of in one long tick. |
BaselineLookback | 28.00:00:00 | How far back the hour-of-day / day-of-week baseline lookup reads. Four weeks gives every (hour, weekday) cell four observations. 00:00:00 reads all retained history. |
Shrinking SampleHotWindow is the highest-leverage knob
Minute rows outnumber their hourly fold 60:1, and the hot window decides how many of them are resident. Dropping it from 7 days to 2 removes roughly 70% of the resident minute rows at no cost to any shipped view — GetThroughputSeriesAsync reads both granularities and slots by bucket end, so totals are identical either way.
Sizing BaselineLookback
Keep it comfortably above the 10-day baseline minimum and below SampleRetentionPeriod — reading past retention can only find rows that are about to be pruned. Widening it directly widens the amount of data the baseline query scans on every alert evaluation, and on a busy multi-month table that read is expensive even with the (ServiceName, HourOfDay, DayOfWeek) index in place.
Use an external metrics stack under load
If you are tuning SampleRetentionPeriod because the table is too big, that is the signal to change mode rather than the knob. A purpose-built time-series store does this job better and cheaper, and keeps the load off the database your monitored application is using.
Storage is only half the cost. Native metrics publishing is also the dominant term in the console's telemetry ingest volume: every export interval, each monitored service folds one metrics sample per active (message type × destination × tenant) combination into its ServiceUpdates batches, so the message volume arriving at the console multiplies with tenant count exactly like the row count above does. On a high-tenant fleet that multiplication is what saturates the console first — a field deployment with ~2,200 active tenants measured metrics-dominated telemetry arriving at ~7,900 messages/minute, far beyond what the console could ingest, with the listener permanently backlogged — and, as the storage warning above notes, metrics account for ~99% of the console's database traffic.
Switching the service to WolverineMetricsMode.SystemDiagnosticsMeter severs all of this at the source: Wolverine never starts the metrics accumulator, so the samples — per-tenant breakdowns included — are never produced, never published, never ingested, and never written. It removes the transport volume and the console's ingest and write load, not just the stored rows.
Recommended posture for heavy load and high tenant counts
For any deployment with meaningful throughput or a large tenant population, prefer a third-party APM backend — Prometheus, VictoriaMetrics, Datadog, or Application Insights — via CritterWatch's external metrics support, instead of native metrics publishing. The metrics views and alert evaluators keep working (they read through the bound source), while the console's largest ingest and database cost disappears entirely.
On the monitored service, stop persisting samples in the console's store:
// SystemDiagnosticsMeter: the monitored service persists NOTHING in the
// console's store — metrics flow only through .NET's System.Diagnostics.Metrics
// for an external scraper (Prometheus, VictoriaMetrics, Datadog agent, ...).
// Bind the service to a metrics data source on the console and every metrics
// view + alert evaluator reads through the external store instead.
var builder = WebApplication.CreateBuilder(args);
builder.Host.UseWolverine(opts =>
{
opts.UseRabbitMq(new Uri("amqp://localhost")).AutoProvision();
opts.AddCritterWatchMonitoring(
critterWatchUri: new Uri("rabbitmq://queue/critterwatch"),
systemControlUri: new Uri("rabbitmq://queue/trip-service-control"),
metricsMode: WolverineMetricsMode.SystemDiagnosticsMeter
);
});
builder.Build().Run();Then register the external source on the console host and bind the service to it. For Prometheus (or VictoriaMetrics, which speaks the same query API):
services.AddCritterWatchMetricsDataSource<PrometheusMetricsDataSource, PrometheusMetricsDataSourceOptions>(
"prometheus",
opts => opts.BaseUrl = "http://prometheus:9090");
services.SetDefaultCritterWatchMetricsDataSource("prometheus");For Azure, register the Application Insights source instead (AddCritterWatchMetricsDataSource<AppInsightsMetricsDataSource, AppInsightsMetricsDataSourceOptions> — the required WorkspaceId and the ingest-lag allowance are covered in Application Insights metrics); Datadog is covered in Datadog metrics. Per-service bindings — including the service-declared .MetricsDataSource("...") push and the operator override flow — are managed via Settings → Metrics Data Sources. The console polls the external store instead of accumulating its own samples.
What you keep: every metrics view, and alerting — the evaluators read through the bound data source. What stops applying: the retention/sampling knobs above, because the console is no longer the one storing the data. Its retention becomes your time-series store's retention.
A service in this posture — SystemDiagnosticsMeter mode and bound to an external source — defaults to live-only persistence: the console writes no MetricsSample / MessagingMetricsBucket rows for it at all. An explicit External metrics persistence override on the service (Settings, or the alert-overrides API) wins in either direction; Persist same as internal restores the old always-persist behaviour. Before this defaulting (#938), live-only was a buried per-service opt-in and the console silently paid the ~1 GB/hour storage cost for data nothing read.
Update check
CritterWatch asks nuget.org once a day whether a newer version of the package it is running has been published, and shows a small badge on the About widget if so. It is on by default.
What is sent
A single unauthenticated GET for a public document:
GET https://api.nuget.org/v3-flatcontainer/{package-id}/index.jsonNo API key, no query string, no headers identifying the install, and no request body. The package id is the one this console ships as — critterwatch, critterwatch.sqlserver or critterwatch.sqlite — and nothing else about the deployment leaves the process. The response is a list of published version strings.
The probe runs server-side, from the console host, not from the operator's browser. That means it inherits the host's proxy, TLS and handler configuration, and no operator's browser is put in touch with nuget.org.
In a cluster every node checks independently — one CDN request per node per day. There is no leader election and no shared state, because coordinating it would cost more than the request it saves.
Turning it off
{
"CritterWatch": {
"UpdateCheck": {
"Enabled": false
}
}
}| Knob | Default | Notes |
|---|---|---|
Enabled | true | The single off switch. |
Interval | 24:00:00 | The answer changes at most daily. |
PackageSourceIndexUrl | https://api.nuget.org/v3-flatcontainer/ | Repoint at a private mirror, or blank it to disable the request while leaving the feature registered. |
IncludePrerelease | (unset) | Unset mirrors the running version: a stable install hears only about stable releases, a release-candidate install hears about anything newer. Set true/false to force it. |
Failure is silent, by design
No network, blocked DNS, a proxy that intercepts TLS, a 404, a 500, malformed JSON — every one of them leaves the reported status as unavailable, logs at Debug, and shows nothing in the UI. An air-gapped install must not emit a warning a day forever, and a failed update check is never raised as an alert or a health-check finding: the alert system describes fleet conditions, and this is not one.
Not visible in development
The check is skipped entirely when IHostEnvironment.IsDevelopment(). A build from source always reports the placeholder version from Directory.Build.props, so without this every developer would be told about a release they are already ahead of. To see the feature by hand, run the console in a non-Development environment.
Alert thresholds
Most alert thresholds are tuned in the UI rather than in configuration files — see Alert Configuration for the live preview, history tab, and three-level cascade (global → per-service → per-message-type).
Defaults that ship with the console:
| Threshold | Default |
|---|---|
| DLQ count Warning / Critical | 10 / 100 |
| Projection lag Warning / Critical | 30s / 300s |
| Agent unhealthy Warning / Critical | 2 / 5 consecutive checks |
| DLQ rate / hour Warning / Critical | 10 / 50 |
| Failure rate Warning / Critical | 5% / 20% |
| Throughput multiplier Warning / Critical | 3× / 10× of baseline |
| Exec time Warning / Critical | +50% / +200% over baseline |
For services that need different defaults baked in (rather than tuned post-deploy), declare baselines from AddCritterWatchMonitoring — see Registration → Declared Baselines.
