When the Alarms Go Off: Incident Response Lessons Written in Blood (and Breach Notifications)
Photo: cybersecurity incident response team computer monitors emergency, via julielohre.com
Opinion: Nobody wants to learn incident response from their own breach. The tuition is too high and the exam is open-book in the worst possible way — the attackers already read all your playbooks.
But here's the thing: the last two years have generated an almost obscene amount of hard-won knowledge, courtesy of organizations that got hit hard, went through the process, and — if we're lucky — published post-incident analyses honest enough to be useful. As someone who has watched these incidents unfold in real time, consulted on a handful of response engagements, and read every public post-mortem I could get my hands on, I want to share what I actually believe separates the organizations that contained damage from those that became case studies in what not to do.
This isn't a checklist post. Those exist, and most of them are fine. This is about the decisions — the ones made at 2 a.m. with incomplete information and executives demanding answers nobody has yet.
The Preparation Gap Is Still the Whole Game
I'll say something that will sound obvious and isn't: most organizations that got destroyed in 2024 breaches were not destroyed because the attackers were particularly clever. They were destroyed because the defenders had never actually practiced responding to a real incident.
Having an incident response plan in a PDF on a SharePoint site is not incident response preparedness. I cannot stress this enough. When the Change Healthcare ransomware incident unfolded earlier this year, the scale of disruption — affecting pharmacy operations across the entire country — reflected not just the attackers' capability but the profound unpreparedness of systems that had never been stress-tested for failure.
The organizations that fared best in recent incidents — and there are some, even if they don't make headlines — share one consistent trait: they had run tabletop exercises that were uncomfortable. Not the kind where the CISO plays along politely and everyone agrees the response would have gone great. The kind where a red team facilitator spends three hours poking holes in every assumption the IR team makes.
Actionable takeaway: Run a tabletop exercise this quarter that specifically tests your communication chain under the assumption that your primary communication tools (email, Slack, Teams) are compromised or unavailable. Because in a real ransomware incident, they often are.
Detection Timing Is Your Destiny
The 2024 Snowflake credential-stuffing campaign that hit Ticketmaster, AT&T, and dozens of other organizations illustrates a pattern that keeps appearing in post-breach analyses: the difference between a manageable incident and a catastrophic one is often measured in hours, not days.
In the Snowflake cases, attackers were operating in environments for extended periods before detection. The access vectors — compromised credentials without MFA — were not sophisticated. The persistence wasn't clever. What made it damaging was dwell time.
Mean time to detect (MTTD) is a metric every security team tracks, but what I see less attention paid to is what happens in the first 30 minutes after detection. That window is where incidents get contained or spiral. Your SIEM fires an alert. Who sees it? What do they do with it? Do they have the authority to isolate systems immediately, or do they have to navigate three approval layers while the attacker is pivoting laterally?
The lesson from real breaches: Empower your tier-1 analysts to take containment actions — specifically network isolation of suspected compromised endpoints — without waiting for management approval. Build the policy, train to it, and accept that you will occasionally isolate a clean machine. That's a far better outcome than the alternative.
Communication Failures Are Breach Multipliers
This one is harder to talk about because it's organizational and political rather than technical, but it's responsible for more secondary damage than almost any technical failure.
In a ransomware incident affecting a mid-sized healthcare provider I'm aware of — details obscured for obvious reasons — the security team identified the initial compromise and began containment within a reasonable timeframe. The technical response was actually solid. What turned it into a prolonged disaster was a 14-hour delay in notifying the legal and communications teams because nobody could agree on who had authority to make that call.
During those 14 hours, the attacker completed their data exfiltration. The technical response had contained the ransomware deployment, but the data was already gone.
What the survivors do differently: They pre-define a notification tree that activates automatically when certain incident severity thresholds are crossed — not when the security team feels confident enough to brief leadership. Legal, PR, and executive leadership get looped in at the start of a potential major incident, not after the security team has built a complete picture. You will never have a complete picture at the start. Make peace with that.
Ransomware Negotiation: The Part Nobody Puts in the Playbook
Let's talk about the elephant in the server room. A significant percentage of organizations hit with ransomware in 2024 paid. The FBI will tell you not to. Your cyber insurance carrier will have opinions. Your lawyers will have opinions. Your board will want to know why operations are still down.
I'm not going to tell you whether to pay — that's a decision with legal, ethical, and practical dimensions that vary enormously by organization. What I will tell you is that the organizations that navigated this most effectively had made the decision in advance, in the abstract, before they were sitting across from a countdown timer and a Bitcoin wallet address.
Decide now: Under what circumstances would your organization consider payment? What's the threshold? Who makes the call? If you haven't answered those questions in a calm boardroom setting, you will answer them in the worst possible way under the worst possible conditions.
The Recovery Phase Gets No Respect
Everybody focuses on detection, containment, and eradication. Recovery — the process of actually rebuilding systems and restoring operations — is where I see organizations fall apart in slow motion after the acute crisis has passed.
The MGM Resorts incident showed what extended recovery looks like at scale. The initial compromise and encryption event was dramatic, but the weeks of degraded operations that followed were the real cost. Slot machines down. Reservations broken. Customer data exposed. The extended recovery period reflected inadequate backup architecture and restoration procedures that had never been tested at full scale.
Test your backups. Restore from them. Time it. Not theoretically — actually pull your production backup and restore a critical system from scratch and measure how long it takes. If you've never done this, I promise you the number is going to be alarming.
The CISO's Real Job in an Incident
Here's my honest take after watching a lot of incident responses: the CISO's job during an active incident is not technical. Your team handles the technical response. Your job is to manage upward, outward, and across — keeping the board informed without panicking them, coordinating with legal and comms, making resource allocation decisions, and ensuring your team doesn't burn out during a response that might last weeks.
The CISOs who struggled most in high-profile 2024 incidents were those who tried to be the technical lead and the organizational interface simultaneously. You cannot do both well under sustained pressure.
Build your team so you can step back from the technical details and do the organizational work. That's not weakness. That's the job.
Final Word
None of this is secret knowledge. The principles of good incident response have been documented for decades. What separates organizations that execute well from those that don't is almost never information — it's the unglamorous work of preparation, practice, and pre-made decisions.
The attackers are ready. The question is whether you are.