ASP.NET Core Health Checks
This page is for operators wiring a Wolverine service into Kubernetes, load-balancer drain logic, or any other system that consumes the standard ASP.NET Core /healthz endpoint.
What this is for
The Wolverine.HealthChecks package plugs Wolverine's runtime state into the ASP.NET Core health-check pipeline. When you enable it, anything that asks your service "are you healthy?" via the standard health endpoint gets a yes/no answer derived from Wolverine's actual state — not just "the web server is up."
The two big consumers of that answer:
- Kubernetes liveness probes. If the probe fails, the pod gets restarted. Use this for "the runtime is in a state we can't recover from, please cycle me."
- Kubernetes readiness probes (and any HTTP load balancer with a similar concept). If the probe fails, the pod stays running but is pulled out of the rotation. Use this for "I'm here but don't send me traffic right now" — startup, drain, or a transient broker outage.
Other common consumers: blue-green deploy gates, canary analysis, and the "is it safe to recycle this node" check that operators run before manual maintenance.
How this differs from CritterWatch's own health view
CritterWatch and the ASP.NET Core health endpoint answer related but distinct questions.
| Question | Where to look |
|---|---|
| Is this pod safe to keep in the rotation right now? | /healthz (Wolverine.HealthChecks) — one bit per node, consumable by k8s and load balancers |
| Why is this service slow? Which listener is jammed? Which projection is lagging? | CritterWatch — per-agent, per-endpoint, per-broker detail for human operators |
The ASP.NET Core health endpoint is for machines making routing decisions. CritterWatch is for humans diagnosing problems.
When the CritterWatch Service Details Overview tab shows the green Registered badge, your service is reporting its Wolverine state to anything that asks via /healthz — that's how Kubernetes decides whether to route traffic to a pod. When it shows the amber Not registered badge, k8s has no signal from Wolverine; the service is either invisible to the orchestrator or relying on a generic "the web server is up" check that won't catch most failure modes.
Wiring it up
Three lines in Program.cs. Add the package reference:
dotnet add package Wolverine.HealthChecksRegister the check and map the endpoint:
builder.Services.AddHealthChecks()
.AddWolverine();
// ...
app.MapHealthChecks("/healthz");Refresh your service's CritterWatch page after the next deploy. The Health checks card on the Overview tab should flip to green Registered, listing the check name (wolverine by default) and any tags you added.
Splitting liveness and readiness
The conventional Kubernetes pattern is two endpoints driven by tags:
builder.Services.AddHealthChecks()
.AddWolverine(tags: new[] { "live", "ready" });
app.MapHealthChecks("/healthz/live", new HealthCheckOptions
{
Predicate = check => check.Tags.Contains("live"),
});
app.MapHealthChecks("/healthz/ready", new HealthCheckOptions
{
Predicate = check => check.Tags.Contains("ready"),
});Then in your Kubernetes manifest:
livenessProbe:
httpGet:
path: /healthz/live
port: 8080
readinessProbe:
httpGet:
path: /healthz/ready
port: 8080The CritterWatch Health checks card shows your registered tags as chips so you can confirm the split is in place at a glance.
Listener-level health
The AddWolverineListeners() extension is strictly stronger than AddWolverine(). The bus check tells you "the runtime started"; the listener check tells you "every active listener is still accepting messages." Use it when one stuck listener should pull the pod out of rotation, even though the runtime itself is still alive.
builder.Services.AddHealthChecks()
.AddWolverine(tags: new[] { "live", "ready" })
.AddWolverineListeners(tags: new[] { "ready" });The listener check reports:
Healthy— every listener isAccepting.Degraded— at least one listener isTooBusyorGloballyLatched.Unhealthy— every listener has stopped, or no listeners exist when one is expected.
Tag the listener check with ready only (not live). A jammed listener should bleed traffic away, but it's not necessarily a reason to restart the pod — let the readiness probe recover when the back-pressure clears.
Probing the CritterWatch console itself
Everything above is about the services CritterWatch watches. The console is also a Wolverine app on a pod, and it needs a probe of its own.
Do not point a liveness probe at /. That route serves the SPA shell and returns 200 for as long as Kestrel is alive — which says nothing about whether the console can process a message. A production console once ran for the better part of an hour at READY 1/1, RESTARTS 0, probes green, failing every inbound message with OutOfMemoryException while its metrics queue climbed to 866 000. .NET fails the managed allocation before the kernel can OOM-kill the container, so nothing restarted the pod and nothing looked wrong from outside.
The console maps two routes for this:
| Route | Reports |
|---|---|
/health/live | The critterwatch-liveness check only — 503 when a restart is the right response |
/health/ready | Every registered CritterWatch check, including critterwatch-system |
UseCritterWatch() maps both. If you wire the console's endpoints by hand:
// UseCritterWatch() already maps these. Call MapCritterWatchHealthChecks()
// yourself only when you wire the console's endpoints by hand instead of
// going through UseCritterWatch().
//
// /health/live — liveness. Store reachable + within the memory limit.
// /health/ready — readiness. Every registered CritterWatch check.
app.MapCritterWatchHealthChecks();livenessProbe:
httpGet:
path: /health/live
port: 8080
initialDelaySeconds: 30
periodSeconds: 15
readinessProbe:
httpGet:
path: /health/ready
port: 8080What makes /health/live return 503
Two signals, both deliberately traffic-independent — a console with no monitored services registered processes nothing, forever, and is perfectly healthy. The probe generates its own work rather than watching message traffic:
- Store round-trip failure. A bounded query through the console's own document store, on a timer. Four consecutive failures (about a minute at the defaults) reports
Unhealthy. - Sustained memory pressure. Occupancy above 95% of the process's memory limit — the container limit, read from the runtime, not the node's physical RAM — for eight consecutive probes (about two minutes). This is the signal that catches the failure above, where a small query may well still succeed while real work cannot allocate.
Both need several consecutive bad probes. A liveness probe that restarts a pod on one blip is its own outage.
A degraded console subsystem (critterwatch-system, e.g. an unmigratable MetricsSample schema) reports Degraded, and /health/live still returns 200. That state is worth surfacing — check /health/ready or GET /api/critterwatch/system-health — but restarting the process does not fix it.
Tuning
Bind the CritterWatch:Liveness configuration section:
{
"CritterWatch": {
"Liveness": {
"ProbeInterval": "00:00:15",
"ProbeTimeout": "00:00:05",
"ConsecutiveProbeFailuresBeforeUnhealthy": 4,
"MemoryPressureThreshold": 0.95,
"ConsecutiveMemoryPressureBeforeUnhealthy": 8,
"Enabled": true
}
}
}If the memory signal trips on a healthy console, the console is undersized for your topology rather than broken — the ServiceSummary a large fleet assembles is the usual driver. Raise the pod's memory limit first; lowering MemoryPressureThreshold only makes the alarm quieter, not the console larger.
Operator action when CritterWatch shows "Not registered"
The amber Not registered badge means CritterWatch found a service that's wired up to it for monitoring but isn't reporting Wolverine state to ASP.NET Core. Two paths forward:
- You don't need k8s probes for this service — for example, a background worker without an HTTP surface. Acknowledge the warning; nothing is broken.
- You do need probes — follow the wiring snippet above. The card refreshes automatically after the next deploy.
If you've registered the package but still see "Not registered," check that the Wolverine.HealthChecks package version matches your Wolverine runtime version. CritterWatch detects the registrations by full type name, so a stale package reference (typo'd namespace, forked build) won't show up as registered even though /healthz works.
