The uncomfortable arithmetic
There is a single fact that determines most of what is useful in denial-of-service defence, and it is routinely learned during an incident rather than before one.
If an attack delivers more traffic than your circuit can carry, the excess is discarded by a router upstream of you. It never reaches your equipment. No firewall rule, rate limit, autoscaling policy or application change can recover packets that were dropped before arrival. At that point your infrastructure is not under load; it is unreachable, which is a different problem with a different owner.
This is why volumetric defence is largely a procurement and architecture question decided in advance. The runbook matters, but much of it describes who to call and what to tell them. That is not a failure of engineering — it is an accurate reflection of where the control sits.
What money actually buys
Transit capacity and headroom. Provisioning close to peak leaves nothing to absorb a surge. Headroom is not waste; it is the margin that keeps a moderate attack from becoming an outage. How much is a business decision about tolerable downtime, not a technical constant.
Scrubbing. Traffic is diverted to infrastructure with the capacity to filter it and the clean remainder returned. The technical part is well understood. The parts that decide whether it helps are contractual: how quickly it activates, whether it is always-on or triggered, who is authorised to trigger it, and what happens at three in the morning on a public holiday.
Anycast footprint. Announcing the same address from many locations divides an attack across all of them instead of concentrating it. This is the most structurally elegant answer to volumetric attack, and it is an architectural commitment made long before it is needed.
Provider diversity. A single upstream is a single point of failure whose capacity ceiling becomes yours.
The clauses that matter
Contracts are where mitigation quietly succeeds or fails, and the questions worth asking are unglamorous.
Always-on or reactive? Reactive diversion is cheaper and adequate against sustained floods. It is structurally unable to handle short bursts: detection, diversion and filtering take minutes, and an attack built as repeated sixty-second bursts is over before mitigation engages, then repeats once it stands down. If burst attacks are in your threat model, reactive scrubbing is not a cheaper version of the same protection — it is a different and much weaker one.
How long does activation take, measured how? "Rapid" is not a number. Ask for the figure, ask whether it is measured from detection or from your phone call, and ask what it has actually been.
Who can trigger it? If activation requires approval from someone who may be asleep, that delay is part of your response time. Pre-authorise it.
What is the capacity ceiling, and what happens above it? Every service has a limit. Knowing whether you are shed or best-effort above it is better learned in a meeting than in an incident.
Does anything cost extra during an attack? Discovering that mitigation is metered while deciding whether to enable it is an unpleasant position to be in.
Is there a tested activation path? A contact who has never been called is a hypothesis, not a plan.
What still belongs to engineering
None of this means the technical work is unimportant — it means it addresses a different portion of the problem.
State-exhaustion attacks consume connection tables rather than bandwidth, and frequently succeed at volumes that never threaten a circuit. Those are yours to solve: handshake proxying at the edge, connection state created as late as possible, stateful devices kept out of paths that do not need them.
Application-layer attacks arrive as valid requests in volumes low enough that no network control will distinguish them. Only you know which endpoints are expensive, which are cacheable, and what a legitimate request pattern looks like.
And visibility is entirely yours. During an incident, the protocol mix, packet sizes, source distribution and timing pattern are what let an upstream provider filter precisely rather than bluntly. Without that detail they can only apply a blunt instrument, and blunt instruments discard legitimate users alongside the attack.
A short readiness test
- Do you know your circuit capacity, and what fraction normal peak consumes?
- Would you detect an attack from your own telemetry, or would a customer tell you?
- Is scrubbing contracted? Is it always-on? Have you activated it deliberately as a test?
- Who is authorised to trigger it, and are they reachable outside business hours?
- Can you produce protocol and source detail during an incident, at sub-minute resolution?
- Does your mitigation design require routing changes under pressure? Have they been rehearsed?
- Has the runbook been tested against something other than a single steady flood?
Anything answered with "probably" is worth converting to a documented answer before it is tested by an attacker.
The summary
Denial-of-service defence is unusual among security problems in that the most important decisions are commercial and architectural rather than configurational, and they must be made before the event. Capacity, diversity, anycast and a scrubbing arrangement with tested activation determine the outcome of a volumetric attack far more than anything configured on the day.
The engineering half is real and necessary — state exhaustion and application-layer abuse are genuine problems with genuine technical answers. But no amount of it rescues a saturated circuit, and time spent tuning the far end of one is time not spent on the arrangements that would have kept it clear.