Cross-application MCP server
CritterWatch ships a Model Context Protocol server (CritterWatch.Mcp) that lets AI agents query and operate the distributed system through the same data plane the SPA reads from. One MCP endpoint covers every monitored Wolverine / Marten / Polecat service — the aggregator is CritterWatch itself.
This page covers the cross-application server. Per-application MCP servers (Marten.Mcp, Polecat.Mcp, WolverineFx.Mcp) ship in their own NuGet packages and expose store / runtime details inside a single service; this page is about the CritterWatch-level surface that spans all of them.
Commercial license required
The entire MCP surface is a paid-tier feature. Every MCP tool is license-gated — including the read/diagnostic tools — so there is no free MCP tier. Without a valid commercial license, every tool returns a "license missing" envelope. (Read-only monitoring through the SPA, by contrast, is always free.)
Looking for "how do I connect"?
If you're on the consumer side — pointing Claude Desktop, MCP Inspector, or a custom MCP agent at a running BFF — see MCP Quick Start for the connection examples, the discovery flow, the deny envelope, and tenant-scoped invocations. This page is the server-side reference.
What's mounted
The BFF mounts the MCP server at /api/mcp (the CritterWatchMcpExtensions.DefaultRoute constant). Clients connect over streamable HTTP; no SSE fallback is required.
POST http://your-critterwatch-host/api/mcpFor multi-package composition (CritterWatch tools alongside per-store tools), see Composition below.
Tool catalog
Read tools
Always license-gated. RBAC-free except the dead-letter reads, the projection tools, and the event query, which are gated on mcp.dlq.read / mcp.projection.step / mcp.events.query — see below.
| Family | Tools | Capability |
|---|---|---|
| Alerts | list_active_alerts, get_alert, summarize_active_alerts | — |
| Health | summarize_cluster_health, get_service_health, list_degraded_surfaces | — |
| Performance | get_backlog_state, list_backlog_hotspots, get_projection_lag, list_shard_health | — |
| OpenTelemetry traces | query_recent_traces, get_trace, check_trace_provider_health | — |
| Message routing | get_message_routing, list_message_routing, explain_message_routing | — |
| Document explorer | list_document_types, query_documents, get_document | mcp.documents.read (the last two; every call audited) |
| Lifecycle | describe_lifecycle | — |
| Dead letters | summarize_dead_letters, query_dead_letters | mcp.dlq.read |
| Projections | diagnose_projection, run_projection_stepper, get_projection_apply_source | mcp.projection.step |
| Event query | query_events | mcp.events.query |
| Scheduled messages | query_scheduled_messages | — |
| Recurring schedules | list_recurring_schedules | — |
| Tenants | list_tenants | — |
| Stream compaction | list_compaction_policies, get_compaction_activity | — |
Dead-letter reads
summarize_dead_letters gives grouped counts by message type and exception type — the triage view. query_dead_letters returns individual envelopes including their ids, which are exactly what replay_dead_letters / discard_dead_letters take. Together they close the investigate → explain → replay loop that the action tools alone could not: before these existed an agent had no way to discover an envelope id short of a human pasting it out of the console.
They carry a capability where the other read tools do not, because dead letters contain message bodies and exception detail — business data. "May look at them" is worth granting separately from "may act on them".
Always check databasesAnswered against databasesAnnounced
Both tools fan out across every physical message database a service owns, and both report how many were asked versus how many replied, plus a partial flag and a warning string.
A store that fails to answer is otherwise indistinguishable from a store with no rows. That is exactly how the console came to assert "No dead letter queue entries found." over a queue holding 42 dead letters on an eight-database deployment (#915). An empty result with partial: true means some stores did not report — not that the queue is empty. Do not conclude a service is clean on the strength of a partial read.
Projection troubleshooting
The three tools are one runbook, in order, and each hands the next its inputs:
diagnose_projection answers "is this projection running?" from the console's own data — no round trip to your service, so it works even when the service is unreachable. It returns the projection's shards worst first, each with the shared health verdict and the reason behind it, the latched apply error's type and message, the shard's position and the high-water mark it was measured against, dead-letter counts, and any alerts already firing about it.
run_projection_stepper answers "is it producing the right answer?", which nothing in a progression table can. It replays the monitored service's own production projection code over a slice of its real events and returns the projected state before and after every event, per aggregate instance. Nothing is written — the replay is in-memory on the service and does not move any shard's position. Source modes are stream, streamslice (fromVersion/toVersion, much cheaper once you know where to look) and tagquery (DCB tags).
get_projection_apply_source fetches the exact Apply(<EventType>) overload that ran on a step the stepper flagged. Source comes from the running assembly's PDB, so a service published without symbols answers found: false with an explanation rather than code.
A null gap is not a caught-up shard
gap and highWaterMark are null when nothing could be measured — never 0. mark − sequence is zero both for a level shard and for one with no denominator at all, and collapsing the two once put a green pill on twenty projections whose lag the tile two inches away called unmeasurable (#1113 / #1138). "gap": 0 means caught up; "gap": null with "status": "Unmeasurable" means you do not know.
An empty shards array has three causes that are indistinguishable from the array alone — an unknown service (serviceKnown: false), a misspelled projection name, and a service whose shards CritterWatch is not tracking at all. None is a clean bill of health. Read serviceKnown and warning.
Stepper responses truncate by default
Every step row carries the projected state twice and a stream can run to thousands of events, so maxRowsPerBucket defaults to 25 (max 200) and maxBuckets to 10 (max 50). Compare rowsReturned against rowsTotal and bucketsReturned against bucketsTotal; when either is short, truncated is true and warning says so in prose. Pass includeState: false for a compact outline when the states are large.
Per-step apply failures never populate error — they ride the individual rows and the timeline continues past them, because the state after a failed apply is part of the diagnosis. A response with ran: true and totalFailedSteps: 3 ran fine and found three broken applies.
run_projection_stepper and get_projection_apply_source additionally require CritterWatchOptions.EnableEventStoreExplorer = true on the monitored service; diagnose_projection has no such precondition.
Event query
query_events (#1186) queries the raw events in a monitored service's event store — identically whether that service runs Marten (PostgreSQL), Polecat (SQL Server) or Fisher (SQLite), because the query executes on the monitored service through the JasperFx.Events abstractions. The agent needs no database credentials and no dialect; the console's message channel is the access path. It is the "what actually happened?" tool: did the command emit the events you think it did, in the order you think, with the payloads you think — the write side, where run_projection_stepper is the read side.
Every filter combines with every other in one query, and they all AND together: tags (a JSON object, DCB semantics — every tag must match), streamId, eventTypeName / eventTypeNames, the timestamp and sequence windows, and correlationId / causationId / userName. Every answer is a page plus totalCount — of the combined match — on Marten, Polecat and Fisher alike (GH-1211). All of it takes tenantId, storeUri (an ancillary store's Subject URI, exactly as CritterWatch reports it), paging, and includeData: false for an envelope-only survey.
A tags-only query against a monitored service too old to honour tag filters on the combined path is detected — its applied-filter echo lacks TagValues — and re-asked on the older standalone tag path, which that service does speak. Every other dropped filter stays a loud FilterMismatch.
There is deliberately no aggregation: totalCount per filtered call is the counting primitive (count-by-type is one cheap call per event type), and the reasoning belongs to the connected model.
Read the honesty fields before concluding anything
NoResponse means the service did not answer — not that it has no events. A tenant filter against a store without multi-tenancy is an error (EventQueryFailed naming the reason), never an empty page. hasMore: true / a totalCount larger than the page means the page is a floor. Every query_events call requires CritterWatchOptions.EnableEventStoreExplorer = true on the monitored service — on by default only in Development — and answers ExplorerDisabled otherwise, which is a configuration answer, never "no events". Before 1.1 only the tag path checked it (GH-1211).
Event payloads are business data, so the tool carries its own capability (mcp.events.query) — the same reasoning as mcp.projection.step, whose projected states are derived from these very events. The CLI twin is the events-query command in JasperFx.Events' command line (beside projection-run), which runs on the application host against its own store — the no-console path, and the right tool inside a test loop.
SQL query
query_sql (#1235) runs one read-only SELECT against a monitored service's event-store database, on the service's own connection, and returns the rows. It is the escape hatch beside query_events: use it for the question the structured filters cannot express — a join to a projection document table, a GROUP BY over a payload field, a count per stream — at the price of writing the target store's own dialect (PostgreSQL for Marten, T-SQL for Polecat, SQLite for Fisher; the answer's engine says which). Prefer query_events for anything its filters can express.
What the monitored service enforces: a single SELECT (or WITH … SELECT) — a second statement and INTO are refused; table references allow-listed to the objects the store declares, with a refusal that lists them (allowedTables) so the next attempt can be right; a read-only transaction that is always rolled back; a statement timeout (timeoutSeconds, default 30, max 120) and a row cap (maxRows, default 100, max 1,000) reported as truncated rather than presented as a total. There is no paging past the cap — it is the contract, not a missing parameter — so a statement that needs rows beyond it windows them itself with ORDER BY plus LIMIT/OFFSET or OFFSET … FETCH. parameters is a JSON object referenced as @name in the statement, typed by inference. Cells come back typed where the column type allows — integers, decimals and booleans as numbers and booleans, JSON columns parsed into JSON — and as text otherwise.
Addressing: storeUri names one event store and is required on a service with more than one (StoreAmbiguous lists them — SQL cannot fan out across stores the way query_events does); on a database-per-tenant store give tenantId or databaseIdentifier (DatabaseAmbiguous lists the databases; the service never picks one).
Honesty fields
NoResponse means the service did not answer — not zero rows. Disabled means the service has not opted in: CritterWatchOptions.EnableSqlQuery follows EnableEventStoreExplorer unless set, so it is on by default only in Development. Timeout means the statement was cancelled, not that it returned nothing. truncated: true means the rows are a floor in the statement's own order — put an ORDER BY and a LIMIT / TOP in the statement.
This reads business data by the table, so it carries its own capability, mcp.sql.query, scoped to the target service — separate from mcp.events.query, because an operator trusted to read events by filter is not thereby trusted to SELECT * from every projection. Every attempt that reaches the wire is written to the audit log with the statement and the caller. There is no CLI twin inside the stack: psql / sqlcmd / sqlite3 against your own store are that twin.
Scheduled messages
query_scheduled_messages (#19 tail) lists the pending scheduled messages in a monitored service's own message store — envelope metadata only (id, message type, scheduled time, destination, attempts), never message bodies, which is why it needs no dedicated mcp.* capability. Its ids are what the cancel/reschedule action tools take, and it is where a rebuild deferred with schedule_projection_rebuild becomes visible: the service schedules the deferred command into its own store precisely so this surface — and the console's Scheduled Messages page — can see, cancel, and move it.
Honesty fields
NoResponse means the service did not answer — not an empty schedule. A multi-database service answers once per database and the unscoped read returns the first answer, labelled as such — pass databaseUri to read one database exactly. And scheduling has no duplicate guard (the store exposes no message bodies to compare): calling schedule_projection_rebuild twice schedules two rebuilds, both visible and cancellable here.
Trace tools route through the per-service ITraceProvider binding cascade — operators bind Jaeger, Datadog, etc. to specific services and the tools surface whichever provider is configured for the service in the query.
Action tools (RBAC-gated)
Each takes a target serviceName (or resource id) and runs the RBAC enforcement pipeline before publishing the existing Wolverine command. On allow, returns an Accepted JSON envelope; on deny, returns a Forbidden / LicenseMissing envelope.
| Family | Tools | Capability |
|---|---|---|
| DLQ | replay_dead_letters, discard_dead_letters | dlq.replay, dlq.discard |
| Projection | pause_projection, restart_projection, rebuild_projection, eject_projection | projection.pause, projection.restart, projection.rebuild, projection.eject |
| Recurring schedules | pause_recurring_schedule, resume_recurring_schedule, trigger_recurring_schedule | recurring-schedule.pause, recurring-schedule.resume, recurring-schedule.trigger |
| Stream compaction | run_compaction_policy | compaction-policy.run |
| Stream compaction | pause_compaction_policy, resume_compaction_policy | recurring-schedule.pause, recurring-schedule.resume |
| Scheduled messages | schedule_projection_rebuild, cancel_scheduled_message, reschedule_scheduled_message | projection.rebuild (scheduling a rebuild IS a rebuild action), scheduled-message.cancel, scheduled-message.reschedule |
| Tenant | add_tenant, enable_tenant, disable_tenant, remove_tenant, hard_delete_tenant | tenant.add, tenant.enable, tenant.disable, tenant.remove, tenant.hard-delete |
| Alert | acknowledge_alert, snooze_alert, clear_alert | alert.acknowledge, alert.snooze, alert.clear |
| ChaosMonkey | enable_chaos_monkey, disable_chaos_monkey, set_chaos_monkey_failure_rate, set_chaos_monkey_slow_handler, set_chaos_monkey_projection_failure_rate, set_chaos_monkey_projection_poison, clear_chaos_monkey_projection_poison, seed_dead_letters | chaos-monkey.toggle (on/off), chaos-monkey.configure (rate / delay / poison / seeding) |
| Listener | pause_listener, restart_listener, drain_listener | listener.pause, listener.restart, listener.drain |
| Metrics | delete_metrics_samples | metrics.delete |
| Service | evict_service | service.evict |
ChaosMonkey and Tenant deliberately split their capabilities so an operator trusted to stop chaos isn't automatically trusted to crank it higher, and a tenant cleanup grant doesn't extend to dropping the tenant's database. See the RBAC page for the full rationale.
Deterministic chaos: prefer it over the rates
Three of the ChaosMonkey tools produce a known failure rather than a probable one. Reach for them whenever you need to reason about a specific failure — which is most of the time.
| Tool | Use instead of | Why |
|---|---|---|
set_chaos_monkey_projection_poison | set_chaos_monkey_projection_failure_rate | The rate is a dice roll per apply, so the alert, the dead-letter drill-in and the projection stepper each land on a different random event. A poison names one event — by type, or by stream and version — so all three agree, and the stepper reproduces it on the exact step instead of failing randomly under itself. It fires independently of the rate, so you can reproduce with the rate at 0. |
seed_dead_letters | waiting on set_chaos_monkey_failure_rate | The failure rate has produced Handled=905, DeadLetter=0 after minutes on a service whose handlers mostly succeed. Seeding writes a known number immediately, with varied message types, exception types, attempt counts and ages so the DLQ summary groups meaningfully. |
clear_chaos_monkey_projection_poison | — | Disarms the poison and leaves the probabilistic rate alone. Use disable_chaos_monkey to stop everything. |
Seeded dead letters are replayable: a configurable fraction is seeded recoverable and succeeds on replay, while the rest throw again and return to the queue. That makes "replay the recoverable ones, discard the rest" a real decision rather than a scripted gesture where everything replays cleanly.
Both write to real storage on a real service. They are chaos-monkey.configure-gated and license-gated, like the rest of the family — use them against staging and dev services.
Per-tenant scoping on projection action tools
pause_projection, restart_projection, and rebuild_projection accept an optional tenantId argument. When supplied, the action runs against only that tenant's projection shard (using the same per-tenant daemon path the SPA's per-tenant Rebuild button uses). When omitted on a multi-tenant service, the action targets the store-global shard — typically what you want on a single-tenant service and almost never what you want when tenants are partitioned.
For authorizers, the RBAC resource scope shifts from serviceName to {serviceName}:{tenantId} when the agent supplies a tenant id. That lets you write per-tenant grants without overloading service-level ones:
public Task<bool> IsAllowedAsync(
ClaimsPrincipal principal, string capability, string? resource, CancellationToken ct)
{
if (capability == Capabilities.ProjectionRebuild
&& resource is { } scope && scope.Contains(':'))
{
var (service, tenantId) = SplitScope(scope);
return Task.FromResult(principal.HasGrant(service, tenantId, capability));
}
// … other paths
return Task.FromResult(false);
}This is the same scope shape used by the SignalR-routed commands and the per-tenant HTTP API calls — one resource convention across the three surfaces.
RBAC enforcement
Each action tool calls a single helper before publishing:
var gate = await McpAuthorizationContext.EnforceAsync(
httpContextAccessor, authorizer,
Capabilities.DlqReplay, resource: serviceName, ct);
if (gate.IsDenied) return gate.DenyPayload!;The helper runs the license guard + RBAC check (RbacGuard.IsAllowedAsync against HttpContext.User). Deny produces a stable-shape JSON envelope:
| Field | Meaning |
|---|---|
error | "LicenseMissing" or "Forbidden" |
message | Human-readable explanation |
capability | The capability string the caller is missing (only on Forbidden) |
resource | The resource scope the deny was evaluated against (only on Forbidden, only when supplied) |
Off-mode hosts (no custom authorizer registered) see the DefaultAllowAuthorizer and every grant succeeds — same shape as the HTTP enforcement layer. See RBAC for the operator- facing detail.
Why stateless transport
The MCP server runs with HttpServerTransportOptions.Stateless = true. This is required for RBAC enforcement: in the default (stateful) mode the HttpContext reachable via IHttpContextAccessor is the one that initialised the MCP session, not the one of the current tool invocation. HttpContext.User would be stale (or empty) for every action after init.
Stateless mode also removes the need for session affinity on multi-node deployments — a useful side benefit for clustered CritterWatch installations.
Licensing
All MCP tools are license-gated. Read tools and action tools both check McpLicenseGuard.IsAllowed() before doing any work; the first check caches the result for the process lifetime. The license is the same JASPERFX-signed license CritterWatch's core uses (see Licensing).
In tests, pre-seed the cache via CritterWatch.Mcp.Licensing.McpLicenseGuard.SetForTesting(true) at assembly load — see src/McpTests/LicenseSetup.cs for the module- initialiser pattern.
Composition
Standalone (default)
builder.Services.AddCritterWatchMcp();
// …
app.MapCritterWatchMcp();This:
- Registers an MCP server with the streamable-HTTP transport configured stateless.
- Adds every tool in the read + action families. The registered catalog is test-enforced by
McpToolRegistrationTests, which asserts registered == documented == defined — so the livetools/listresponse is the count, and this page is the description. - Wires
AddHttpContextAccessor()so action tools can resolve the caller's principal.
Chained alongside per-store servers
If your host already composes an MCP server (e.g. with Marten.Mcp + WolverineFx.Mcp), chain CritterWatch's tools onto the existing builder:
builder.Services
.AddMcpServer()
.WithHttpTransport(o => o.Stateless = true)
.AddCritterWatchMcp()
.AddMartenMcp()
.AddWolverineMcp();When composing, the host is responsible for setting Stateless = true on the transport and for registering AddHttpContextAccessor() itself. The IMcpServerBuilder overload of AddCritterWatchMcp() only chains the tools; the standalone IServiceCollection overload sets up both for you.
Connecting an MCP client
Any client that speaks streamable HTTP MCP can connect — for example the @modelcontextprotocol/inspector:
npx @modelcontextprotocol/inspector
# URL: http://localhost:5173/api/mcp (or your BFF's address)For Claude Desktop or similar, configure the MCP server in the client's config to point at /api/mcp on your CritterWatch host. The transport is streamable HTTP, not stdio — pick the matching client setting.
If your host has RBAC enforcement on, the MCP client needs to present an authenticated principal that the host's authentication layer recognises (OIDC bearer token, signed header from a reverse proxy, etc.). Anonymous calls hit the fail-closed-on-no-principal branch and get a Forbidden envelope.
Testing
McpTests in the CritterWatch repo covers every tool with the same matrix: happy-path publish, RBAC deny, fail-closed-on-no-principal, license-missing, plus per-tool bad-request cases. Tools are invoked directly with substituted IHttpContextAccessor / authorizer / IMessageBus — no MCP server stand-up needed for unit-style tests. See src/McpTests/Mcp/DlqActionToolsTests.cs for the reference shape.
