loader
Newsroom
Cybersecurity

Your Cloud Provider Publishes Its IP Ranges

The major cloud and CDN operators publish exactly which addresses are theirs, in machine-readable form, updated as they grow. A surprising amount of security tooling still guesses instead.

If you want to know whether an address belongs to Amazon, Amazon will tell you. So will Google, Microsoft, Oracle, Cloudflare, Fastly, DigitalOcean and most of the rest. They publish machine-readable documents listing precisely which ranges are theirs, they maintain them as they acquire more address space, and they do it for a practical reason: their customers need to write firewall rules.

These documents have existed for years. They are free, unauthenticated, and linked from the operators' own documentation.

A good deal of security tooling ignores them and guesses instead.

What guessing looks like

The guess usually takes the form of a large, round CIDR block written by hand. A /9 or a /10, covering millions of addresses, recorded as belonging to a single provider because a lot of it demonstrably does.

The reasoning is not stupid. Someone looked at a sample of addresses, saw a consistent owner, and wrote down the enclosing block. It is fast, it needs no integration, and it is right most of the time — which is exactly what makes it durable, because the cases where it is wrong do not announce themselves.

But address space is not allocated in tidy blocks, and it moves. Ranges are transferred between organisations. A /9 written down two years ago contains allocations that have since changed hands, and every address in those is now confidently mislabelled. A single line in a configuration file can put several million addresses under the wrong name, and nothing in the system will ever flag it, because a hand-written range carries no indication that it was ever a guess.

Why the shortcut survives

Partly inertia: the ranges were written early, they worked, and nothing forced a revisit.

Mostly, though, it is that the failure is invisible from the inside. A guess and a published fact look identical once they are rows in the same table. Both produce a match. Neither carries its provenance. Unless the system deliberately records where a claim came from, there is no way to notice that some of your data was published by the operator and some was typed by a person, and no way to prefer the first when they disagree.

That is the part worth fixing, and it is not really about cloud ranges. It is that a dataset which cannot distinguish its strongest evidence from its weakest cannot improve — because every addition dilutes it, and nobody can tell which rows to trust when two of them conflict.

What to ask

If you rely on datacenter or cloud detection, two questions are worth putting to whoever supplies it.

Where did this range come from? If the answer is a general list of methods rather than the provenance of this specific range, the dataset is not tracking it.

How quickly does it follow the provider? Cloud operators add and return address space continuously. A dataset refreshed from published documents tracks that within days. A hand-maintained one tracks it whenever somebody remembers.

Neither question is about coverage, and both predict how often the answer will be confidently wrong.