Skip to main content

Recovery

An alert tells your team that a resource became unhealthy. Recovery closes the loop: once the resource is healthy again, Mission Control updates the same Slack thread or email conversation that carried the original alert.

Recovery is opt-in for each notification through onResolved.

How it works​

  1. A health event fires and Mission Control delivers the notification as usual, applying waitFor, filter, silences, inhibitions and repeatInterval.
  2. Mission Control remembers where that alert landed — the Slack message, or the email recipients and subject.
  3. The resource turns healthy. Mission Control waits for onResolved.waitFor of continuous health, so a resource that flaps straight back to unhealthy produces no recovery update.
  4. Mission Control updates the original destination: a threaded Slack reply, a Slack reaction, or an email that replies to the original alert.

Recovery works on the delivery, not on the trigger. A healthy event does not need to appear in spec.events, and the original filter, silences, inhibitions and repeat limits are not applied a second time.

What counts as healthy

Only a real transition to healthy resolves an outage. Unknown health, deletion of the resource, and a resource that stops matching the notification's filter all leave the outage open.

A change between warning and unhealthy stays inside the same outage. A full healthy-to-unhealthy transition starts a new outage, which gets a conversation of its own and resets repeatInterval.

Supported destinations​

Recovery updates a destination that already received the original alert, so the transport must support a follow-up message:

RecipientSupported
to.connection pointing at a Slack ConnectionThreaded reply, emoji reaction, or both
to.connection pointing at an SMTP ConnectionReply to the original email
to.email or to.person over system SMTPReply to the original email
to.team, where the team rule selects one of the transports aboveSame as the selected transport
fallback, when the fallback delivery is the one that succeededSame as the selected transport

Grouped notifications, to.playbook, to.webhook, generic slack:// URLs and every other Shoutrrr transport are rejected when the notification is validated.

Supported events

onResolved accepts config.healthy, config.warning, config.unhealthy, config.degraded, the matching component.* events, check.passed and check.failed. groupBy and onResolved cannot be combined.

Examples​

Reply in the Slack thread​

This notification alerts on unhealthy Kubernetes deployments through a Slack Connection, and marks the recovery on the same message.

recovery-slack.yaml
# Set the Slack channel and create recovery-slack-token with key token first.
apiVersion: mission-control.flanksource.com/v1
kind: Connection
metadata:
name: recovery-slack
namespace: mission-control
spec:
slack:
channel: C0123456789
token:
valueFrom:
secretKeyRef:
name: recovery-slack-token
key: token
---
apiVersion: mission-control.flanksource.com/v1
kind: Notification
metadata:
name: deployment-health-recovery-slack
namespace: mission-control
spec:
events: [config.unhealthy, config.warning]
filter: config.type == 'Kubernetes::Deployment'
waitFor: 1m
to:
connection: connection://mission-control/recovery-slack
onResolved:
enabled: true
waitFor: 30s
slack:
reply: true
reaction: white_check_mark

This example:

  1. Uses waitFor: 1m to hold the initial alert until the deployment has been unhealthy for a minute.
  2. Uses onResolved.waitFor: 30s to require 30 seconds of continuous health before the recovery update.
  3. Uses slack.reply to post the update in the original thread and slack.reaction to add a :white_check_mark: to the original message.

Give reaction the emoji name on its own — white_check_mark, not :white_check_mark:. Set reply, reaction, or both; a slack block that disables reply without a reaction fails validation.

Reply to the original email​

This notification emails an address over the shared system SMTP settings, then replies to that email once the deployment recovers.

recovery-system-smtp.yaml
# Configure the shared system SMTP Connection (or SMTP env/CLI settings) first.
# Replace the recipient before applying. No per-notification system-connection grant is needed.
apiVersion: mission-control.flanksource.com/v1
kind: Notification
metadata:
name: deployment-health-recovery-email
namespace: mission-control
spec:
events: [config.unhealthy, config.warning]
filter: config.type == 'Kubernetes::Deployment'
waitFor: 1m
to:
email: operations@example.com
onResolved:
enabled: true
waitFor: 30s

This example:

  1. Uses to.email so delivery goes through the system SMTP Connection or the SMTP environment settings.
  2. Uses onResolved.enabled on its own, because the slack block does not apply to email delivery.
  3. Achieves a recovery mail that carries the original Re: subject and the In-Reply-To and References headers, so mail clients thread it with the alert.

Mission Control records the recipients, sender and subject at the time of the original send. Reserve the Message-ID, In-Reply-To, References, From, To and Subject headers for Mission Control — overriding them through to.properties breaks threading.

Write your own recovery message​

This notification delivers through a named SMTP Connection and replaces the default recovery body with a template.

recovery-named-smtp.yaml
# Replace the host/addresses and create recovery-smtp-credentials first.
# Grant this notification read permission only on this named Connection.
apiVersion: mission-control.flanksource.com/v1
kind: Connection
metadata:
name: recovery-smtp
namespace: mission-control
spec:
smtp:
host: smtp.example.com
port: 587
encryption: ExplicitTLS
auth: plain
username:
valueFrom:
secretKeyRef:
name: recovery-smtp-credentials
key: username
password:
valueFrom:
secretKeyRef:
name: recovery-smtp-credentials
key: password
fromAddress: alerts@example.com
fromName: Mission Control
toAddresses: [operations@example.com]
---
apiVersion: mission-control.flanksource.com/v1
kind: Notification
metadata:
name: deployment-health-recovery-named-email
namespace: mission-control
spec:
events: [config.unhealthy, config.warning]
filter: config.type == 'Kubernetes::Deployment'
waitFor: 1m
to:
connection: connection://mission-control/recovery-smtp
onResolved:
enabled: true
waitFor: 30s
template: 'Recovered {{.resource.id}} at {{.recoveredAt}} (outage {{.outageDuration}}).'

This example:

  1. Uses to.connection to send through the recovery-smtp Connection instead of the system SMTP settings.
  2. Uses onResolved.template to replace the default body.
  3. Achieves a one-line recovery mail naming the resource, the recovery time and the length of the outage.

Without template, the body names the resource and reports the recovery time and the outage duration.

Template variables​

A recovery template reads the same variables as the notification template, plus:

FieldDescriptionScheme
outageDuration

How long the resource stayed unhealthy

Duration

recoveredAt

When the resource became healthy

time.Time

resolvedAt

When Mission Control sent the recovery update

time.Time

resource

The recovered resource, with id, type and health

map[string]any

Permissions​

A notification that delivers through a named Connection needs read permission on that Connection during recovery, not only during the original send. Credentials are read again for every recovery attempt, so revoking the permission or deleting the Connection stops the recovery update. A missing named Connection never falls back to system SMTP.

Delivery over system SMTP — to.email, to.person, or the system SMTP Connection — uses the shared application mail settings and needs no per-notification grant.

Retries​

A failed recovery update retries with exponential backoff — 2, 4, 8, 16, 32, 64 and 128 seconds. Set notification.recovery.max-retries to change the budget; the default is 7 retries after the first attempt. Set it to 0 to disable retries.

Once the budget runs out, Mission Control stops retrying that update, and raising notification.recovery.max-retries afterwards does not restart it. The cause is usually a revoked permission on the Connection, a deleted Connection, or a deleted notification.

Retention​

Mission Control deletes recovery state once it is no longer needed. Every property below takes the prefix notification.recovery. and accepts Go duration syntax along with the d, w and y units, so 720h, 30d and 4w2d are the same value.

PropertyDefaultDescription
retention.resolved720hAge at which a delivered recovery update is deleted
retention.exhausted2160hAge at which a recovery update that ran out of retries is deleted
retention.episodes720hAge at which a finished outage is deleted
retention.deleted-states720hGrace period before the health record of a deleted resource is deleted. Values below 720h are raised to 720h

Set a duration to 0 to keep that category forever. Mission Control only deletes a recovery update after it is delivered or has run out of retries.

Fields​

FieldDescriptionScheme
enabled

Set to true to send a recovery update once the resource becomes healthy again. (Default: false)

boolean

slack.reaction

Emoji name to add to the original message, for example white_check_mark. Give the name on its own, without the surrounding colons.

string

slack.reply

Post the recovery update as a threaded reply to the original message. (Default: true)

boolean

template

Replaces the default recovery body. In addition to the usual template variables, resource, recoveredAt, resolvedAt and outageDuration are available.

Go Template

waitFor

How long the resource must stay continuously healthy before Mission Control sends the recovery update. (Default: 0)

Duration