Recovery
An alert tells your team that a resource became unhealthy. Recovery closes the loop: once the resource is healthy again, Mission Control updates the same Slack thread or email conversation that carried the original alert.
Recovery is opt-in for each notification through onResolved.
How it works
- A health event fires and Mission Control delivers the notification as usual, applying
waitFor,filter, silences, inhibitions andrepeatInterval. - Mission Control remembers where that alert landed — the Slack message, or the email recipients and subject.
- The resource turns healthy. Mission Control waits for
onResolved.waitForof continuous health, so a resource that flaps straight back to unhealthy produces no recovery update. - Mission Control updates the original destination: a threaded Slack reply, a Slack reaction, or an email that replies to the original alert.
Recovery works on the delivery, not on the trigger. A healthy event does not need to appear in spec.events, and the
original filter, silences, inhibitions and repeat limits are not applied a second time.
Only a real transition to healthy resolves an outage. Unknown health, deletion of the resource, and a resource that
stops matching the notification's filter all leave the outage open.
A change between warning and unhealthy stays inside the same outage. A full healthy-to-unhealthy transition starts
a new outage, which gets a conversation of its own and resets repeatInterval.
Supported destinations
Recovery updates a destination that already received the original alert, so the transport must support a follow-up message:
| Recipient | Supported |
|---|---|
to.connection pointing at a Slack Connection | Threaded reply, emoji reaction, or both |
to.connection pointing at an SMTP Connection | Reply to the original email |
to.email or to.person over system SMTP | Reply to the original email |
to.team, where the team rule selects one of the transports above | Same as the selected transport |
fallback, when the fallback delivery is the one that succeeded | Same as the selected transport |
Grouped notifications, to.playbook, to.webhook, generic slack:// URLs and every other Shoutrrr transport are
rejected when the notification is validated.
onResolved accepts config.healthy, config.warning, config.unhealthy, config.degraded, the matching
component.* events, check.passed and check.failed. groupBy and onResolved cannot be combined.
Examples
Reply in the Slack thread
This notification alerts on unhealthy Kubernetes deployments through a Slack Connection, and marks the recovery on the same message.
recovery-slack.yaml# Set the Slack channel and create recovery-slack-token with key token first.
apiVersion: mission-control.flanksource.com/v1
kind: Connection
metadata:
name: recovery-slack
namespace: mission-control
spec:
slack:
channel: C0123456789
token:
valueFrom:
secretKeyRef:
name: recovery-slack-token
key: token
---
apiVersion: mission-control.flanksource.com/v1
kind: Notification
metadata:
name: deployment-health-recovery-slack
namespace: mission-control
spec:
events: [config.unhealthy, config.warning]
filter: config.type == 'Kubernetes::Deployment'
waitFor: 1m
to:
connection: connection://mission-control/recovery-slack
onResolved:
enabled: true
waitFor: 30s
slack:
reply: true
reaction: white_check_mark
This example:
- Uses
waitFor: 1mto hold the initial alert until the deployment has been unhealthy for a minute. - Uses
onResolved.waitFor: 30sto require 30 seconds of continuous health before the recovery update. - Uses
slack.replyto post the update in the original thread andslack.reactionto add a:white_check_mark:to the original message.
Give reaction the emoji name on its own — white_check_mark, not :white_check_mark:. Set reply, reaction, or
both; a slack block that disables reply without a reaction fails validation.
Reply to the original email
This notification emails an address over the shared system SMTP settings, then replies to that email once the deployment recovers.
recovery-system-smtp.yaml# Configure the shared system SMTP Connection (or SMTP env/CLI settings) first.
# Replace the recipient before applying. No per-notification system-connection grant is needed.
apiVersion: mission-control.flanksource.com/v1
kind: Notification
metadata:
name: deployment-health-recovery-email
namespace: mission-control
spec:
events: [config.unhealthy, config.warning]
filter: config.type == 'Kubernetes::Deployment'
waitFor: 1m
to:
email: operations@example.com
onResolved:
enabled: true
waitFor: 30s
This example:
- Uses
to.emailso delivery goes through the system SMTP Connection or the SMTP environment settings. - Uses
onResolved.enabledon its own, because theslackblock does not apply to email delivery. - Achieves a recovery mail that carries the original
Re:subject and theIn-Reply-ToandReferencesheaders, so mail clients thread it with the alert.
Mission Control records the recipients, sender and subject at the time of the original send. Reserve the Message-ID,
In-Reply-To, References, From, To and Subject headers for Mission Control — overriding them through
to.properties breaks threading.
Write your own recovery message
This notification delivers through a named SMTP Connection and replaces the default recovery body with a template.
recovery-named-smtp.yaml# Replace the host/addresses and create recovery-smtp-credentials first.
# Grant this notification read permission only on this named Connection.
apiVersion: mission-control.flanksource.com/v1
kind: Connection
metadata:
name: recovery-smtp
namespace: mission-control
spec:
smtp:
host: smtp.example.com
port: 587
encryption: ExplicitTLS
auth: plain
username:
valueFrom:
secretKeyRef:
name: recovery-smtp-credentials
key: username
password:
valueFrom:
secretKeyRef:
name: recovery-smtp-credentials
key: password
fromAddress: alerts@example.com
fromName: Mission Control
toAddresses: [operations@example.com]
---
apiVersion: mission-control.flanksource.com/v1
kind: Notification
metadata:
name: deployment-health-recovery-named-email
namespace: mission-control
spec:
events: [config.unhealthy, config.warning]
filter: config.type == 'Kubernetes::Deployment'
waitFor: 1m
to:
connection: connection://mission-control/recovery-smtp
onResolved:
enabled: true
waitFor: 30s
template: 'Recovered {{.resource.id}} at {{.recoveredAt}} (outage {{.outageDuration}}).'
This example:
- Uses
to.connectionto send through therecovery-smtpConnection instead of the system SMTP settings. - Uses
onResolved.templateto replace the default body. - Achieves a one-line recovery mail naming the resource, the recovery time and the length of the outage.
Without template, the body names the resource and reports the recovery time and the outage duration.
Template variables
A recovery template reads the same variables as the notification template, plus:
| Field | Description | Scheme |
|---|---|---|
outageDuration | How long the resource stayed unhealthy | |
recoveredAt | When the resource became healthy |
|
resolvedAt | When Mission Control sent the recovery update |
|
resource | The recovered resource, with |
|
Permissions
A notification that delivers through a named Connection needs read permission on that Connection during recovery, not only during the original send. Credentials are read again for every recovery attempt, so revoking the permission or deleting the Connection stops the recovery update. A missing named Connection never falls back to system SMTP.
Delivery over system SMTP — to.email, to.person, or the system SMTP Connection — uses the shared application
mail settings and needs no per-notification grant.
Retries
A failed recovery update retries with exponential backoff — 2, 4, 8, 16, 32, 64 and 128 seconds. Set
notification.recovery.max-retries to change the budget; the default is 7 retries after the first attempt. Set it to
0 to disable retries.
Once the budget runs out, Mission Control stops retrying that update, and raising
notification.recovery.max-retries afterwards does not restart it. The cause is usually a revoked permission on the
Connection, a deleted Connection, or a deleted notification.
Retention
Mission Control deletes recovery state once it is no longer needed. Every property below takes the prefix
notification.recovery. and accepts Go duration syntax along with the d, w and y units, so 720h, 30d and
4w2d are the same value.
| Property | Default | Description |
|---|---|---|
retention.resolved | 720h | Age at which a delivered recovery update is deleted |
retention.exhausted | 2160h | Age at which a recovery update that ran out of retries is deleted |
retention.episodes | 720h | Age at which a finished outage is deleted |
retention.deleted-states | 720h | Grace period before the health record of a deleted resource is deleted. Values below 720h are raised to 720h |
Set a duration to 0 to keep that category forever. Mission Control only deletes a recovery update after it is
delivered or has run out of retries.
Fields
| Field | Description | Scheme |
|---|---|---|
enabled | Set to |
|
slack.reaction | Emoji name to add to the original message, for example |
|
slack.reply | Post the recovery update as a threaded reply to the original message. (Default: |
|
template | Replaces the default recovery body. In addition to the usual template variables, | |
waitFor | How long the resource must stay continuously healthy before Mission Control sends the recovery update. (Default: |