← All playbooks
Tier 4: Total Clusterfuck

What should you say when your deploy takes down production?

The short answer

Post a status message immediately with impact, current action, and the next update time — then keep that promise even when there is nothing new. Blameless facts during, timeline after. Never speculate about cause while the site is still down.

What's actually at stake

  • 01Leadership panics about silence, not about outages. Regular updates buy you enormous patience.
  • 02Cause claims made in minute five are wrong in minute forty and quoted forever.
  • 03The person who writes the timeline controls the narrative of the retro.

Do

  • · Update every 15 minutes, even with 'no change yet'.
  • · State user impact in human terms, not error codes.
  • · Volunteer to write the postmortem before anyone assigns it.

Don't

  • · Do not name a root cause you have not confirmed.
  • · Do not say 'should be fixed' — say what is verified.
  • · Do not disappear into the fix and stop posting.

A draft you can send now

Incident update, 14:20: checkout is failing for roughly 30% of users since 13:41. Cause not confirmed yet; we've rolled back the 13:38 deploy and are watching error rates. I own the comms here. Next update 14:35 regardless of status.
Tailor this to my exact disaster →

Questions people actually ask

Should I say it was my deploy?
During the incident, only the facts of what changed. Ownership belongs in the postmortem, and yes, put it there.
Who should send updates?
One named person, not the person fixing it. Split those roles the moment the incident starts.

Take it further

Related incidents

All CYABot tools →