Upgrade Notes
Version-to-version changes that affect a deployed console's database schema or stored data — plus compatibility notes for monitored hosts — what CritterWatch does about them automatically, and the manual path when it can't.
This page is deliberately narrow
It covers only what an operator must do to move a running console between versions. For what is new and what was fixed in each release, see the Release Notes. A change appears here only when it has an upgrade consequence; everything else lives there.
1.1.0-beta.5 → 1.1.0-beta.6: evict once more any service an earlier version evicted
Who is affected: a console on which you evicted a service (the Evict action on its detail page) while running any version before 1.1.0-beta.6, including 1.0.x, and where that service is still running or will run again.
What happened: eviction archived the service's event streams. Every store refuses to append to an archived stream, so each time that service reported again its update failed, and it never reappeared on the console, however often it restarted. Upgrading does not change streams that are already archived, so such a service stays missing after the upgrade too.
What to do: after upgrading, evict the service once more. Eviction now deletes the service's history rather than archiving it, including the streams an earlier eviction archived: the service's own stream, its listener history, and the streams of the alerts it had raised. So the service registers again the next time it reports, and those alerts can fire again. Without this step the service stays missing, and even once it is back, any alert it had before the old eviction would fail to raise. Nothing else is needed, and a service you never evicted is unaffected.
1.1.0-beta.6 → 1.1.0-beta.7: a console over its licensed service cap parks the surplus services
What changed. A license's Max monitored services limit is now enforced against the whole fleet. Before, it was checked only when a brand-new service first connected, so a fleet that was already over the limit (for example, one that ran unlicensed and then took a smaller key) stayed monitored in full.
Who is affected. Only a console whose license sets a service limit and that has more services registered than the limit allows. Everyone else sees no change.
What happens on upgrade. Within 30 seconds of starting, the console keeps monitoring the services that registered earliest, up to the limit, and parks the rest. A parked service keeps all of its data and stays listed, marked Not monitored — license limit, but its new telemetry is not applied and commands to it are refused. Nothing is deleted or evicted.
What to do. If the earliest-registered services are not the ones you want monitored, choose them with Keep monitored on each service's page, evict services you no longer need, or move to a license with a higher limit. The license is read once at startup, so a new key takes effect on the next restart. See Licensing.
1.1.0-beta.6 → 1.1.0-beta.7: the document explorer hides soft-deleted documents and asks for every tenant explicitly (GH-1304)
What changed. The three stores now agree on what a document read returns (JasperFx 2.78's document diagnostics contract), and CritterWatch follows it:
- Soft-deleted documents are hidden unless you turn on Include deleted, and then each one is marked. Before, Marten and Polecat returned them as though they were live, and Fisher hid them.
- No tenant means the default tenant. The explorer's All tenants choice now sends an explicit "every tenant" request. Before, Marten and Polecat read every tenant for no tenant, and Fisher read the default tenant, which on a store whose documents are all tenanted is nothing.
- Naming a sub-class returns only that sub-class, and a base type's page says which rows are sub-classes.
- Reading one document by id takes its tenant directly, and returns a soft-deleted document marked rather than missing.
The query_documents MCP tool has matching allTenants and includeDeleted parameters and returns each document's tenant, type and deleted state in rows. Omitting tenantId reads the default tenant.
Who is affected. Anyone who reads documents on a Marten or Polecat store with soft deletes, or on a conjoined multi-tenant store. An agent or script that called query_documents without a tenant to see every tenant now sees the default tenant only: pass allTenants: true.
What to do. Nothing on the console for the behaviour change. But see the next note: the monitored services must move their store with the client.
1.1.0-beta.6 → 1.1.0-beta.7: a monitored service must run Marten 9.43, Polecat 5.34 or Fisher 1.15 or later
What changed. The 1.1 client package (Wolverine.CritterWatch) depends on JasperFx 2.78. JasperFx 2.77 added members to the document-store diagnostics interface that every store implements, so a store built against an older JasperFx cannot load beside it.
Who is affected. A monitored service that upgrades the client and keeps an older store. Restore and build succeed, because NuGet lifts JasperFx under the old store without complaint, and then the service fails at startup:
System.TypeLoadException: Method 'LoadDocumentAsync' in type 'Marten.DocumentStore' from assembly
'Marten, Version=9.42.0.0' does not have an implementation.(Measured with Marten 9.42.0 and JasperFx 2.78.0; Polecat below 5.34 and Fisher below 1.15 are built the same way. Tracked upstream as jasperfx#933.)
What to do. Upgrade the store in the same change as the client: Marten 9.43.0, Polecat 5.34.0 or Fisher 1.15.0, or later.
1.1.0-beta.6 → 1.1.0-beta.7: reading documents needs documents.read when permissions are wired (GH-1307)
What changed. Reading stored documents in the Document Explorer (a page of a type, or one document by id) now requires the new documents.read capability, scoped to the store the request names or else the service, and every such read writes an audit entry. The query_documents and get_document MCP tools require mcp.documents.read the same way. Type lists and mapping DDL are schema and stay ungated. On the monitored service, a new EnableDocumentExplorer switch governs document reads. It follows EnableEventStoreExplorer when unset, so nothing changes there unless you set it.
Who is affected. Only consoles with permissions wired, using an authorizer that denies what it is not told to allow, which is the documented recipe. On those consoles, document reads stop working after the upgrade until you grant documents.read (and mcp.documents.read for agents) to the roles that should have it. The refusal says which capability is missing. Consoles without permissions wired, the default, are unaffected. This lands inside the 1.1 beta window on purpose: after 1.1.0 it would be a break inside a stable release.
What to do. Add documents.read wherever your authorizer grants event-store-explorer.view to people who should also see stored documents. documents.write and documents.delete are reserved for the explorer's future write surface and are not checked yet.
1.0.x → 1.1: if you wired real permissions, check your capability assignments
Only affects hosts that register an ICritterWatchAuthorizer. If you have not, nothing changes: the default allows everything, and it always did.
The console now enforces capabilities on every operator command it relays or handles. Several were previously checked for an AI agent over MCP and not for a person clicking the same button:
| command | capability |
|---|---|
| Dead-letter replay / edit-and-replay | dlq.replay |
| Dead-letter discard | dlq.discard |
| Pause / restart / drain one listener | listener.pause / .restart / .drain |
⚠️ An operator who could run one of these before may now be denied. The refusal is explicit — the console answers with a rejection naming the missing capability rather than failing silently — but it will be a surprise if the capability was never assigned.
What to do: review the capabilities your authorizer grants against the table above before upgrading, and grant what your operators actually need. Read-only monitoring is unaffected; queries were never gated and still are not.
1.0.x → 1.1: the SQL view refuses a bare table name that resolves outside the store's schema
Only affects the Event Explorer's SQL view and the query_sql MCP tool. The allow-list used to match bare table names and ignore the schema, so select … from mt_events resolved through PostgreSQL's search_path. A bare reference is now qualified against the store's own schema and matched qualified-only.
⚠️ A statement that relied on search_path resolution now fails. On the fleet that found this, select count(*) from mt_events on one service had been answering from an unrelated public.mt_events table; it now reads the store's own table, and a reference that would resolve anywhere else is refused with the list of what the store declares. Qualified references (myschema.mt_events) are unchanged.
1.0.x → 1.1: a shard stopped by a progression conflict reports Conflicted, which is a new wire field
Affects the projection-health verdicts the console shows; no action is required to upgrade. A shard stopped by ProgressionProgressOutOfOrderException — two processes both believe they own it — used to render as Paused. It now renders as Conflicted, ranked above Failing, because the remedy is to find the other owner rather than to press Restart.
The classification rides a new field on the shard-state message. A monitored service on an older Wolverine.CritterWatch sends nothing there and the console falls back to the exception type, which reaches the same verdict — so a mixed fleet is fine. Bump the monitored service's package when convenient to move it onto the classified path.
1.0.x → 1.1: two capability families now name a different resource (#1258)
Only affects hosts that register an ICritterWatchAuthorizer and decide using the resource argument. An authorizer that grants on the capability alone and ignores the resource is unaffected.
CritterWatch asks an authorizer two things: which capability, and which resource it applies to. Two gates run for an operator command — the MCP tool, for an AI agent, and the relay, for every command whoever sent it — and for two command families those gates named different resources for the same action. An authorizer written against one was asked a question by the other that it had no basis to answer, and answered no.
They agree now, so the resource you are asked about has changed:
| capability family | the MCP tool used to ask about | the relay used to ask about | both ask about now |
|---|---|---|---|
tenant.add · .enable · .disable · .remove · .hard-delete | the service name | the tenant id | the tenant id |
projection.pause · .restart · .rebuild | {service} or {service}:{tenant} | the agent URI | the agent URI, plus @{tenant} when the command names one |
The direction was chosen by the rule the relay's table already stated and enforced with a test: a capability names the most specific identifier available, so a grant can say "may disable THIS tenant" or "may rebuild THIS projection". For projections neither side could express what the other could — the relay distinguished projections but not tenants, the tool tenants but not projections — so the resource now carries both rather than one winning.
projection.eject, the listener family, and every other capability are unchanged.
⚠️ The agent URI's host is lower-cased — event-subscription://tripservice/Trip:All — because Uri.ToString() normalizes it. This is long-standing relay behaviour rather than a 1.1 change, but it is now visible to an MCP-originated grant too. An authorizer that compares a service name case-sensitively must account for it.
What to do: if your authorizer matches on resource, update the patterns for these two families before upgrading. A grant that no longer matches fails closed — the operator is refused with the capability and resource named in the refusal, rather than silently succeeding — so the failure is visible, but it is still a refusal.
1.0.x → 1.1: the legacy critterwatch queue's leader pin is yours to write (GH-1024)
Only affects multi-node consoles that listen on the unsharded critterwatch queue.
That queue has no per-service hashing, so its only protection against competing consumers is .ListenOnlyAtLeader() — and that call lives in your own host's Wolverine configuration, beside your ListenToRabbitQueue("critterwatch"). CritterWatch does not own your listener configuration and cannot add it for you.
opts.ListenToRabbitQueue("critterwatch").ListenOnlyAtLeader();Without it, every node appends to the same ServiceSummary streams — which corrupts them rather than merely contending, and surfaces as intermittent stream-version collisions long before anyone connects it to the listener configuration.
What to do: nothing, if the pin is already there — it has been in the recommended configuration for some time. This release adds a runtime check that logs at Critical and names the offending endpoints when it finds more than one node and an unprotected listener, so if you are missing it you will now be told rather than left to diagnose the symptom.
1.0.x → 1.1: an embedded console has its own wolverine_* tables (GH-1139)
Only affects the new embedded mode. Nothing about a standalone console changed.
⚠️ This corrects guidance that was wrong before 1.1 shipped. The embedded documentation previously said the console's messages ride your application's existing envelope tables and that "you do not get a second set of wolverine_* tables to reason about or clean up." That is not what happens. Sharing occurs only when the host has set opts.Durability.MessageStorageSchemaName — which is not the same property as the MessageStorageSchemaName you set inside IntegrateWithWolverine(o => ...), and setting the latter looks like it worked because it does take effect for your own tables.
What to do: if you run any maintenance that resets or rebuilds the message store, iterate every store rather than the main one:
foreach (var store in await host.GetRuntime().Stores.FindAllAsync())
{
await store.Admin.RebuildAsync();
}A second set of envelope tables that nothing sweeps is how a previous defect went unnoticed for weeks — stale envelopes in an unswept schema replaying a previous run's telemetry on every boot.
1.0.1 / 1.0.2 → 1.1: extended progression tracking is turned on for every store (GH-1188, GH-1183)
What changes. The Wolverine.CritterWatch client package now switches on extended progression tracking for every event store registered in a monitored service's container, through each store's own supported setting, rather than inferring it from the DI descriptors. The console does the same for its own stores. On first startup after the upgrade, each of those stores' progression tables gains the extra columns wherever schema management is left to the store.
Why you have to know. The upstream default for this flag is off, deliberately and permanently — defaulting it on is breaking for stores that never asked to be monitored, and upstream now has a guard test saying not to flip it again outside a major version (marten#5318 reversed an earlier restoration). CritterWatch therefore sets it rather than inheriting it, and that is load-bearing rather than belt-and-braces. Before this release, an ancillary store that kept a bare progression table made the console error on every progression read and write against it.
What to do.
- Schema management on (the default): nothing. The columns are added on first startup.
AutoCreate.Noneor externally managed schemas: apply each store's schema migration before starting the upgraded service, as you would for any other schema change — whatever route you already use to emit and review store DDL.- Read-only database users: the account CritterWatch connects with needs
ALTER TABLEon the progression table for the one startup that applies it, or the migration must be applied out of band.
TIP
The added columns are small and always-populated; there is no backfill and no rebuild. If you see the console log that a store still reports extended progression as off, that is the switch failing to take rather than a read failure, and it will name the store.
1.0.0 / 1.0.1 → 1.1: the console no longer forces its transport schema (GH-1126)
(Carried from the unreleased 1.0.2, which is folded into this release.)
The console's queue tables move from critterwatch_wolverine to whatever the transport is configured with — Wolverine's wolverine_queues default unless you pass something.
- If your fleet worked by setting
critterwatch_wolverineon every service, it keeps working: that value is still honoured, it is simply no longer forced. - Otherwise drop the argument on both halves.
- Leftover
critterwatch_wolverine.wolverine_queue_*tables hold at most a few seconds of already-expired telemetry and can be dropped.
See Transport channels — the transport schema is a fleet-wide rendezvous, and every service and the console must agree on it. This is the change that, while the schema was forced, silently stranded all telemetry on database-queue fleets: services wrote, the console drained a different same-named table, and neither side logged anything.
1.0.x → 1.1: scheduling a rebuild needs a current client package and a durable queue (GH-19)
Scheduled rebuilds are a new command, not a new field on an existing one. That is deliberate: an optional property on the existing rebuild contract would be silently ignored by a service running an older Wolverine.CritterWatch, which would then rebuild immediately — an operator schedules for 03:00 and gets a rebuild in business hours, with nothing reporting a problem. A new command instead gets a no-handler report, which the console surfaces.
So: a monitored service must be on the 1.1 client package for "Rebuild later…" to work against it. Services on older packages are unaffected in every other respect — this is additive.
The service also needs somewhere durable to keep the schedule. The scheduled message lands in the monitored service's own message store, which is what makes it visible and cancellable on the Scheduled Messages page. On a default in-memory local queue it lives only in process, so a redeploy before 03:00 drops it with no error and it never appears on that page either. Enable AlwaysMakeScheduledMessagesDurable() or give the queue a durable inbox. A service with no message store at all cannot honour a schedule under any configuration, and CritterWatch refuses the schedule rather than accepting one it will drop.
1.0.x → 1.1: the new event-store query filters need a current client package (GH-1198)
The Event Store Explorer's by-metadata search and the query_events tool gained multi-type unions, timestamp windows and sequence windows this release. A monitored service must be on the 1.1 client package to honour them.
A service on an older Wolverine.CritterWatch does not merely lack the feature — it drops an unknown filter while deserializing the request, so the filter never reaches the event store to be refused. The store's own guard against silently-ignored filters is one layer too low to see this, and unfiltered results would come back looking exactly like filtered ones.
So the service echoes back the filters it actually applied, and the console and the tool compare that against what was sent. Against an older service you get the dropped filters named and no rows at all, rather than a page that reads as an answer. This is additive: existing unfiltered and single-type queries behave exactly as before against any version, and never raise the warning.
1.0.x → 1.1: the Event Store Explorer opt-in now covers every event read (GH-1211)
What changed. CritterWatchOptions.EnableEventStoreExplorer — set on each monitored service, on by default in Development and off everywhere else — used to gate only some of the reads that return event data: the stream, tag, document and projection-stepper views. The by-metadata event query, which returns full event bodies and backs both the Explorer's filter view and the query_events MCP tool, never checked it. As of 1.1 every read of event data respects it.
Who is affected. A monitored service that is not running in Development and has not opted in. On 1.0.x the Explorer's By Metadata view and query_events returned its events anyway; on 1.1 they answer "Event store explorer is disabled on this service" (ExplorerDisabled from the MCP tool) — the same answer the other Explorer views already gave that service.
What to do. Nothing, if you did not mean those reads to work. To keep them, opt the service in:
builder.Host.UseWolverine(opts =>
{
opts.ServiceName = "my-service";
var critterWatch = opts.AddCritterWatchMonitoring(
critterWatchUri: new Uri("rabbitmq://queue/critterwatch"),
systemControlUri: new Uri("rabbitmq://queue/my-service-control"));
// On by default in Development, OFF everywhere else. Event payloads are business data:
// turning this on lets anyone with console access (and MCP callers holding
// mcp.events.query) read this service's events.
critterWatch.EnableEventStoreExplorer = true;
});Opting in exposes that service's event payloads to anyone with console access, and to MCP callers holding mcp.events.query. That trade-off is the reason the default is off outside Development, and the reason one setting now governs all of it rather than some of it.
Needs both halves on 1.1. The check runs on the monitored service, so a service still on a 1.0.x Wolverine.CritterWatch keeps its old behaviour until it upgrades.
1.0.x → 1.1: SQL over the event store is a new read, and it follows the explorer opt-in (GH-1235)
What changed. The Event Store Explorer's new SQL view and the query_sql MCP tool run a read-only SELECT on the monitored service, against the tables its event store declares. It is gated by a new switch on the monitored service, CritterWatchOptions.EnableSqlQuery, which when left unset follows EnableEventStoreExplorer — so a service that has opted into the explorer has, on upgrading its Wolverine.CritterWatch package, also opted into SQL.
Who is affected. A service that opted into the explorer outside Development and would rather not expose SQL. SQL reads by the table — projection documents included — which is more than the structured views return. Set EnableSqlQuery = false to keep the explorer's structured views and refuse SQL; the refusal is answered by name (Disabled), never silently.
builder.Host.UseWolverine(opts =>
{
opts.ServiceName = "my-service";
var critterWatch = opts.AddCritterWatchMonitoring(
critterWatchUri: new Uri("rabbitmq://queue/critterwatch"),
systemControlUri: new Uri("rabbitmq://queue/my-service-control"));
// The structured explorer views stay on; a SELECT over this service's tables is refused
// by name. Left unset, EnableSqlQuery follows EnableEventStoreExplorer.
critterWatch.EnableEventStoreExplorer = true;
critterWatch.EnableSqlQuery = false;
});Permissions and audit. On a console with permissions wired, the SQL view requires the new sql.query capability scoped to the store, and the MCP tool mcp.sql.query scoped to the service — neither is implied by the event-query grants. Every run is written to the audit log with the statement.
Needs both halves on 1.1. The handler lives in the monitored service, so a service on a 1.0.x Wolverine.CritterWatch has no SQL view: the console reports that it did not answer, and the MCP tool NoResponse.
1.0.x → 1.1: some stores answer to a new identity, so their shard rows are keyed anew (marten#5409, fisher#279, polecat 5.26)
The console keys every persisted shard-progression row on the store's subject URI and the database identifier the store reports. Three of those identities changed upstream in the 1.1 pin, each because the old one was unusable as a key:
| Store | Was | Now | Who is affected |
|---|---|---|---|
| Marten, a named primary store | marten://main regardless of StoreName | marten://{storename} | only hosts that set StoreOptions.StoreName on the primary store; an unnamed one is still marten://main, and ancillary stores were already per-store |
| Fisher | the database uri, sqlite://localhost/{path} | fisher://{storename} | every Fisher store |
| Polecat, multi-database | sqlserver://host/db | sqlserver://host/db/dbo (a schema segment) | Polecat reports the change renames rows keyed on that URI; treat it as affected until you have checked yours |
Once a monitored service upgrades its client package it reports under the new identity, and the console creates fresh rows for it. Rows under the old identity are not reconciled away — the reconciliation marker a service sends is scoped to the store it now names, so the old store's rows are never claimed by anyone. They sit as stale cards until you evict the service from the console (the Evict action on the service's detail page) after its upgrade, which drops everything the console holds for it and lets the next heartbeat rebuild it clean. Nothing in the console's own schema changes.
1.0.x → 1.1: no action, but worth knowing
Per-tenant metric roll-ups are no longer scraped and broadcast by default. The per-tenant answer is requested on demand instead. If you built anything on the 30-second broadcast, it now needs to ask.
The Operations screen is shard-grained. Saved links or bookmarks into the old projection-grained inventory still resolve, but the screen opens on exceptions rather than on the full list.
A recurring schedule's live state needs a current client package (GH-1233). A monitored service on this release of
Wolverine.CritterWatchreports its schedules' runtime state — paused, next occurrence, last handled — on its own telemetry cadence. A service on an older package reports none, and its Schedule Explorer rows readannounced <age>rather thanas of <age>: the values are its capability announcement's copy, frozen when it registered its schedules. That is rendered as a distinct claim rather than as a stale-looking live reading, so nothing is silently wrong — but the age on those rows says nothing about the schedule until the service is upgraded.A leftover database-less high-water row on a single-database store (1.1.0-beta.5). Earlier 1.1 betas seeded a single-database store's high-water mark under a row with no database, beside the live mark the store reports under its database's name. The seed never moved again. From beta.5 the seed lands on the live row itself, but a console upgraded from an earlier beta keeps the old row until its progression data is reset. Nothing reads it as the store's mark (the higher of two marks wins), so it is safe to leave; delete it only if you query the table yourself and select the mark by the row's shape.
rc.9 → rc.10: a third console package, CritterWatch.Sqlite, and a new shared core id (GH-1014, GH-1015)
Nothing breaks, and existing consumers do not have to move. CritterWatch is still the Marten / PostgreSQL console package and still exposes AddCritterWatch(). Two things are new:
CritterWatch.SqlitejoinsCritterWatchandCritterWatch.SqlServeras a third store flavour, backed by Fisher / SQLite. It needs no database service at all — it writes a file — which makes it the cheapest way to stand a console up. All three expose the sameAddCritterWatch(WebApplicationBuilder)/UseCritterWatch(WebApplication)pair with an identical signature, so moving between stores is a package swap and no code change.CritterWatch.Servicesis now a published package id of its own: the store-agnostic core all three flavours share. You never install it directly — each flavour depends on it — but it will appear in your dependency graph where it previously did not.
Who is affected: nobody upgrading in place. The only thing to know is that CritterWatch.Services is the core, not the Postgres console — if you had scripts or docs pointing dotnet pack or a package reference at src/CritterWatch.Services expecting the console, that is now src/CritterWatch.
rc.8 → rc.9: the standalone MessageHandlingMetrics publish route is gone (GH-955)
AddCritterWatchMonitoring no longer declares a publish route for Wolverine's MessageHandlingMetrics type. The route was dead wiring: the CritterWatch observer folds every exported metrics batch into ServiceUpdates (and deliberately never publishes the batches individually), so the route carried zero traffic on every deployment — but it declared a contract that repeatedly misled analysis about what the console's leader-pinned legacy queue carries.
Who is affected: nobody, behaviorally. Metrics keep arriving inside ServiceUpdates exactly as before, in every metrics mode. If you had custom Wolverine routing that keyed off the MessageHandlingMetrics message type on the monitored side, that type no longer has a CritterWatch-declared route — but nothing publishes it standalone, so such routing was already inert.
rc.6 → rc.7: external-source services default to live-only metrics persistence (GH-938)
A service that reports WolverineMetricsMode.SystemDiagnosticsMeter and is bound to a non-Postgres metrics data source (Prometheus, VictoriaMetrics, Datadog, Application Insights) now defaults to ExternalMetricsPersistence.LiveOnly: the console stops writing MetricsSample / MessagingMetricsBucket rows for it. Previously live-only was a per-service opt-in buried in the alert overrides, so this posture silently kept paying the console-side storage cost (~1 GB/hour measured in the field) for data nothing read — the metrics views and alert evaluators for such a service read the external store.
Who is affected: only deployments that already run services in SystemDiagnosticsMeter mode with an external source binding and did NOT set the live-only override. After upgrading, those services' new samples stop accumulating; existing rows age out through normal retention.
To keep the old behaviour for a specific service, set its External metrics persistence override to PersistSameAsInternal (Settings → the service's alert overrides, or the alert-config API) — an explicit override always wins over the default, in either direction.
rc.6 → rc.7: SampleRetentionPeriod now actually drives the metrics partition window (GH-935)
Before this fix, both store registrations built the MetricsSample rolling partition window from the hardcoded 45-day default, ignoring a configured CritterWatch:Metrics:SampleRetentionPeriod. Consequences on earlier versions:
- Shortening retention below 45 days reclaimed space by row deletes only (dead tuples + vacuum churn) — the O(1) monthly partition drops still waited for the 45-day floor.
- Lengthening retention past ~56 days was destructive: the drop pass retired month-partitions that still held rows inside the configured window.
After upgrading, the window follows the configured value on both PostgreSQL and SQL Server, and the startup partition pass is additive only — provisioning happens at boot, while partition drops run solely from the hourly prune, which skips them when retention is 00:00:00 (keep forever) or when any per-service retention override outlives the global window. Existing partitions are adopted in place; no data migration runs. If you overrode SampleRetentionPeriod on an earlier version expecting partition-level reclaim, the first prune after upgrade will retire every partition past your configured window — verify that window is what you intend before deploying the upgrade.
Monitored hosts using AddOpenApi(): upgrade Wolverine to 6.17.2+ (GH-689)
The problem
On WolverineFx / WolverineFx.Http 6.17.1 and earlier, any read of ASP.NET Core's ApiExplorer that lands before the web server starts permanently freezes every OpenAPI document as empty ("paths": {}) for the lifetime of the host — ASP.NET's version-keyed description cache serves the first (empty) answer forever (wolverine#3371).
CritterWatch's monitoring satellite is an in-box early reader: the capability snapshot (ServiceCapabilities.ReadFrom) consults ApiExplorer from the observer's boot-time batching loop, which can win the race against web-server startup — reported at roughly 40% incidence on 2-core CI runners, rarer but possible anywhere. The result is that a plain Wolverine + CritterWatch + AddOpenApi() host can break its own OpenAPI documents just by being monitored, with no user code involved. HTTP routing keeps working; only the OpenAPI documents (Microsoft.AspNetCore.OpenApi, Swashbuckle, versioned documents) are affected, and only until the process restarts and wins the race.
Why the read happens at all
CritterWatch describes the monitored host's non-Wolverine endpoints — minimal API, MVC, Razor Pages, SignalR — by matching each RouteEndpoint against ASP.NET's ApiExplorer (IApiDescriptionGroupCollectionProvider), which is what populates the ASP.NET Endpoints page. That lookup is what touches ApiExplorer, and the capability snapshot runs it during the observer's boot handshake. This matters for what the fix below does and does not cover.
The fix
Upgrade the monitored service to Wolverine 6.17.2 or later. Wolverine's ApiDescription provider now enumerates its own HttpGraph — complete as soon as MapWolverineEndpoints() returns — instead of the DI EndpointDataSource that ASP.NET only populates at server start, so a pre-start read returns the same complete answer as a post-start one (wolverine#3373).
No CritterWatch configuration change is needed, and no CritterWatch version change is involved — the fix is entirely on the monitored host's Wolverine version.
Hybrid hosts: 6.17.2 does not close this completely
If your OpenAPI document also contains minimal-API, MVC, or Razor Pages endpoints, upgrading is not the whole answer. 6.17.2 makes only Wolverine's descriptions start-independent. ASP.NET's own providers still compose their descriptions at server start, and the ApiExplorer read is still happening pre-start, and ASP.NET still caches the whole collection for the lifetime of the host.
So on a hybrid host the pre-start read now freezes a document that has the Wolverine endpoints but none of the minimal-API / MVC ones — where before 6.17.2 it froze fully empty. That is a better failure, and a more deceptive one: the document is populated, so it looks fine until someone notices their minimal-API routes are missing. This is called out in the fix's own behavior notes.
A pure Wolverine.Http host (every described endpoint comes from MapWolverineEndpoints()) is fully covered by 6.17.2 and needs nothing further. A hybrid host should defer its boot-time ApiExplorer reads as well — see below.
How to tell you've been bitten
Hit the OpenAPI document on a running host (/openapi/v1.json, or your configured route):
"paths": {}— the fully-empty freeze. Monitored host on Wolverine ≤ 6.17.1.- Paths present, but every minimal-API / MVC route missing while Wolverine routes are there — the hybrid-host freeze described above.
In both cases HTTP routing itself keeps working — only the documents are wrong — and both clear on a process restart that wins the race, which is exactly what makes this easy to dismiss as a flake. It reproduced at roughly 40% on 2-core CI runners and far more rarely on developer machines, so a document that is fine locally and empty in CI is the signature.
Interim mitigation
For a host pinned to Wolverine ≤ 6.17.1, or a hybrid host on any version that wants to be certain:
- Preferred — make sure nothing reads ApiExplorer before
ApplicationStarted. Wrap the registeredIApiDescriptionGroupCollectionProviderso that pre-start reads are served from a throwaway provider instance (a fresh instance has a fresh cache, so ASP.NET's lifetime cache is never poisoned) and post-start reads delegate to the real one. - Simplest — don't attach CritterWatch monitoring to an OpenAPI-exposing host until it is on 6.17.2+ (and, if hybrid, until its boot-time ApiExplorer reads are deferred).
Either way this is a property of the monitored host, not of the console: no CritterWatch setting changes it, because the read is intrinsic to describing non-Wolverine endpoints at snapshot time.
beta.1 → beta.2+: mt_doc_metricssample becomes a partitioned table (GH-685)
What changed
beta.1 created the metrics history table (<schema>.mt_doc_metricssample, PostgreSQL/Marten consoles) as a plain heap table. beta.2 redesigned it for retention at scale: the table is now a range-partitioned parent (monthly child partitions plus a DEFAULT catch-all) with duplicated, indexed service_name / bucket_end columns that the metrics reads filter on.
PostgreSQL cannot ALTER a plain table into a partitioned one, so a console database first provisioned on beta.1 and then upgraded is stuck in the old shape. On beta.2 the symptoms were:
GET /api/critterwatch/throughput-seriesreturns HTTP 500 for every parameter combination, so the dashboard's System Throughput chart is permanently empty (the generated SQL references thebucket_endcolumn, which the old shape doesn't have).- Retention leaks: partition-based pruning has no partitions to drop and the row-level prune fails on the same missing column, so
mt_doc_metricssamplegrows without bound (a single low-volume service accumulated 1.9M rows in the original report). - Both failures were silent: writes kept working, the host booted cleanly, and the only trace was a startup warning.
What happens automatically now
On startup (and again on every retention prune cycle), the metrics partitioner inspects the table shape:
Current partitioned shape — nothing to do.
beta.1 heap shape (an ordinary table with Marten's
id+datacolumns) — it is auto-migrated, with row counts logged at each step:- The legacy table (and its constraints/indexes) is renamed aside to
mt_doc_metricssample_pre_partition. - The partitioned parent — and the matching upsert functions and indexes — are rebuilt from Marten's configured DDL. This works under any
AutoCreatepolicy, includingNone. - The monthly +
DEFAULTchild partitions are created. - Rows are backfilled from the renamed table, bounded by the retention horizon (
CritterWatch:Metrics:SampleRetentionPeriod, default 45 days). Rows already past retention — including the leaked backlog the drift itself caused — are deliberately left behind: they would be pruned immediately anyway, and the bound keeps the migration a one-time bounded cost. - The renamed table is dropped.
Startup waits only for the schema phase (steps 1–3 — fast catalog work). The backfill and drop (steps 4–5) run in the background after the console is up and serving, with the
metrics-storecomponent onGET /api/critterwatch/system-healthreporting degraded until they complete. This matters on a large legacy heap: earlier builds ran the whole copy inside host startup, so on a multi-GB table orchestrator liveness probes killed the "hung" console before the copy finished and every restart began it again (#928). Don't restart the console whilemetrics-storereports degraded unless you have to — the migration resumes where it left off, but each interruption re-scans.The migration is idempotent and resumable: each step keys off catalog state and the backfill uses
ON CONFLICT DO NOTHING, so an interruption (crash, deploy, connection loss) between any two steps resumes cleanly on the next startup.- The legacy table (and its constraints/indexes) is renamed aside to
Anything else — a shape that is neither the beta.1 heap nor the current partitioned parent — CritterWatch refuses to touch the table and fails loudly instead of warn-and-continue: an error log, plus a degraded
metrics-storecomponent onGET /api/critterwatch/system-healthand on thecritterwatch-systemASP.NET Core health check. The rest of the console keeps working.
During the (one-time) migration window, an in-flight metrics flush can log a transient write error between the rename and the function rebuild; it retries on the next flush interval and succeeds.
Manual migration path
If you'd rather migrate by hand — or your table is in a shape the auto-migration refuses — run the following against the console database (adjust critterwatch to your configured schema name). Metrics history is statistical/ephemeral data, so the simplest safe path is drop-and-recreate:
-- 1. Move the old table aside (or DROP TABLE it if you don't care about recent history)
ALTER TABLE critterwatch.mt_doc_metricssample RENAME TO mt_doc_metricssample_old;
ALTER TABLE critterwatch.mt_doc_metricssample_old
RENAME CONSTRAINT pkey_mt_doc_metricssample TO pkey_mt_doc_metricssample_old;Then restart the console. On boot it recreates mt_doc_metricssample in the correct partitioned shape (this is the same "resume from the renamed table" path the auto-migration uses when the leftover table is named mt_doc_metricssample_pre_partition; with any other name the fresh table is simply created empty). Optionally backfill recent history:
-- 2. (Optional) copy recent rows into the new partitioned table
INSERT INTO critterwatch.mt_doc_metricssample
(id, data, mt_last_modified, mt_version, mt_dotnet_type, service_name, bucket_end)
SELECT id, data, mt_last_modified, mt_version, mt_dotnet_type,
data ->> 'serviceName',
(data ->> 'bucketEnd')::timestamptz
FROM critterwatch.mt_doc_metricssample_old
WHERE (data ->> 'bucketEnd')::timestamptz >= now() - interval '45 days'
ON CONFLICT DO NOTHING;
-- 3. Clean up
DROP TABLE critterwatch.mt_doc_metricssample_old;How to tell it worked
GET /api/critterwatch/throughput-series?hours=24&bucketMinutes=60returns200with a bucket array, and the dashboard's System Throughput chart renders.GET /api/critterwatch/system-healthreports"status": "healthy".SELECT relkind FROM pg_class WHERE relname = 'mt_doc_metricssample'returnsp(partitioned), and\d+ critterwatch.mt_doc_metricssampleshows monthly partitions plusmt_doc_metricssample_default.- The leaked backlog is gone after the migration (out-of-retention rows are not carried over), and from then on normal retention keeps the table bounded by dropping aged monthly partitions.
