← All field notes
incident responseleadershipidentityfor executives

Containment under fire: leading an identity breach in the first hour

When identity itself is compromised, the first hour is about command, scope, and containment, not restoration. Here is how leadership should run it: the decisions only an executive can own, the containment tradeoff, and the record the board will ask for.

When the compromised layer is identity itself, the incident is different in kind from a single breached server, and the difference is a leadership problem before it is a technical one. A breached server is contained by isolating it. A compromised identity plane means the attacker can become anyone - and the containment moves available to you, isolating domain controllers or pausing hybrid synchronization, carry real business impact and have to be taken while the attacker is still active. Someone has to authorize hurting the business to stop the bleeding, and that someone is not the incident-response lead acting alone.

The mental model for the first hour is that you are running a command decision under uncertainty, not a repair job. The instinct of every capable technical team is to fix and restore, and in an identity breach that instinct, followed too early, rebuilds the house on top of the intruder. Leadership’s job is to hold the line between “stop and understand” and “restore,” long enough to make the few decisions only an executive can make.

Why is the first hour not about restoration?

Because you cannot safely restore what you have not yet scoped, and in an identity breach the attacker’s foothold lives where a restore does not reach. Bring systems back before you understand the compromise and you are very likely bringing them back with the attacker still inside - the forged tickets, the extra credentials, the federation or trust changes that a reboot and a password reset leave untouched. Restoration is not the first hour’s job; it is what you earn once you know what you are restoring.

So spend the first hour on three things. Stand up a command structure and an incident bridge, so decisions have an owner and information has one place to land instead of scattering across channels. Get a preliminary, fact-checked scope before any external communication, because the first numbers are almost always wrong and a wrong number said early is remembered long after it is corrected. And authorize containment with a clear owner, the resources to execute it, and a rollback plan, so the person told to isolate the domain controllers has the authority and the safety net to actually do it. None of those three is a technical task; all three are leadership.

What is the decision that is actually yours?

Containment that hurts the business versus dwell time that expands the damage - and which way to lean, right now, with incomplete information. This is the executive decision at the center of an identity breach, and it cannot be delegated to the technical lead as if it were a configuration choice, because it is a business risk decision with consequences the business will feel either way.

The math usually favors containment. Every hour an identity attacker remains active, the blast radius grows: more accounts, more systems, more persistence planted, more that has to be rebuilt later. The cost of dwell time compounds, while the cost of containment - the outage, the paused integration, the disrupted work - is usually bounded and recoverable. But “usually” is not “always,” and the point is not that containment always wins; it is that the tradeoff has to be named explicitly and owned by someone with the authority to accept the business impact. An IR lead who isolates a domain controller without that authority is exposed, and an executive who leaves the call to them has abdicated the one decision that was theirs.

What should leadership not do?

Three failures recur, and each is a way of avoiding the decision rather than making it. Do not brief the board or the public on guesses. Early scope is provisional, and an inaccurate figure or a premature “we believe it is contained” becomes the thing everyone remembers and holds you to. Brief on what is fact-checked, and say plainly what is still unknown.

Do not delegate the whole call to the technical lead without giving them authority and resources. Handing someone the responsibility for a business-impacting containment action while withholding the authority to take it, or the resources to execute and roll back, is not delegation - it is leaving the decision unmade and the risk on the wrong person. And do not freeze the response while you over-analyze. Waiting for certainty that will not arrive in the first hour is itself a decision, and it is the one the attacker is counting on, because every minute of analysis paralysis is a minute of dwell time they use to spread.

How do you make the containment call with incomplete information?

By treating it as a risk decision with a bias toward action, framed and owned, not as a technical certainty to wait for. You will not have full scope in the first hour, and you do not need it to decide. What you need is a clear statement of the tradeoff - what containment costs the business versus what an additional hour of attacker dwell likely costs - an owner empowered to accept that cost, and a rollback plan so the action is reversible if the scope turns out smaller than feared.

The identity-specific containment levers are drastic by nature: isolating or shutting domain controllers, pausing hybrid sync so a cloud compromise cannot ride into on-prem or vice versa, mass-revoking sessions and tokens, resetting the krbtgt account to invalidate forged tickets. Each hurts, and each is sometimes exactly right. The leadership skill is not knowing in advance which lever to pull; it is creating the conditions - command, scope, authority, rollback - under which the right lever can be pulled quickly and reversed if wrong, rather than debated until the window closes.

Why is early scope so hard to trust, and how do you brief anyway?

Because in the first hour you are looking at partial telemetry through the fog of an active intrusion, and identity breaches are especially deceptive: the attacker is using legitimate credentials, so the “normal” activity and the malicious activity look alike until you have done real work to separate them. Your first scope is a hypothesis, not a finding. The pressure - from the board, from customers, from regulators, from your own desire to sound in control - is to convert that hypothesis into a confident number and say it out loud. That is the trap.

The discipline is to brief on confidence levels, not guesses. Say what you know and have verified, say what you suspect and are still confirming, and say plainly what you do not yet know - and give a time by which you will know more, rather than a number you will have to retract. A stakeholder can act on “we have confirmed X, we are investigating whether Y, next update in two hours.” A stakeholder cannot un-hear “we believe it is contained” when it was not.

This is why external communication comes after a fact-checked preliminary scope, not before, and why the incident commander holds that line even under pressure to say something reassuring. The cost of an accurate “we do not know yet” is a moment of discomfort. The cost of a confident wrong answer is your credibility for the rest of the incident and the inquiry that follows it. Scope humbly, brief precisely, and update on a cadence.

Why does the record matter as much as the decision?

Because after the incident the board asks two questions - when did you know, and what did you do - and the quality of your answer is set during the first hour, not reconstructed afterward. A timestamped log of your containment decisions and the tradeoffs you weighed is two things at once: better governance in the moment, because it forces the decision to be explicit and owned, and your defensible record later, because it shows a deliberate response rather than a scramble.

The record is also what separates a well-led incident from a lucky one in the eyes of regulators, insurers, and your own board. Two organizations can take the same containment action; the one that can show it decided deliberately, weighed the business impact, owned the call at the right level, and kept the receipts is the one that is judged to have managed the crisis rather than been managed by it. The through-line of the first hour is that leadership under fire is not about having the answers - it is about running the decision: command, scope, a named owner for the hard tradeoff, and a record of how you made it.

Frequently asked questions

Why is an identity breach different from a normal server compromise?

Because the compromised layer is the thing everything else trusts. A single breached server is contained by isolating it; a compromised identity plane means the attacker can authenticate as anyone, and containment can require isolating domain controllers or pausing hybrid sync - actions with real business impact taken while the attacker is still active. That makes it a leadership decision, not only a technical one.

What should leadership do in the first hour of an identity breach?

Resist the pull to restore. Stand up a command structure and an incident bridge, get a preliminary and fact-checked scope before any external communication, and authorize containment with a named owner, the resources to execute, and a rollback plan. Restoration comes after you understand what you are restoring.

How do you weigh containment cost against letting the attacker dwell?

Every hour an active identity attacker stays in widens the blast radius, so the cost of delay usually exceeds the cost of containment - but containment that hurts the business is a risk decision with consequences. Frame it explicitly as that tradeoff, assign an executive to own it, and document the reasoning, rather than freezing while you over-analyze or delegating it without authority.

Why does restoring first make an identity breach worse?

Because restoring before you have scoped the compromise rebuilds on top of the attacker's persistence. In an identity breach that persistence lives in things a rebuild or password reset does not touch - forged tickets, extra credentials, federation changes - so you can bring systems back only to have the attacker still inside. Understand and contain first, restore second.

What will the board ask afterward, and how do you prepare?

The two questions are always when did you know and what did you do. Prepare by keeping a timestamped log of your containment decisions and the tradeoffs you weighed as the incident unfolds. Do not brief on early guesses, which are usually wrong and are remembered; brief on fact-checked scope, and let the decision log be your defensible record.