loader
Research
Vulnerability Research

A Monitoring Agent Is Not an API Client

Continuous telemetry and request-driven API calls arrive over the same transport and look identical at the door. Metering them the same way builds a system where a customer's own monitoring throttles their production integration — a denial of service nobody had to attack you to cause.

Two things arrive at a platform over HTTPS with a key attached: a request some code made because it wanted an answer, and a heartbeat an agent made because a timer fired. At the door they are indistinguishable. Metering them as though they are the same thing is a design error with consequences well beyond the invoice.

This is a note about a class of mistake, not an incident report. It is easy to make, and it is much easier to design out than to unwind.

The two shapes of traffic

Request-driven traffic is bounded by intent. An integration makes a call because something happened — a user signed up, a page loaded, a job ran. Volume tracks the customer's real activity. Metering per call works precisely because a call means something happened.

Continuous telemetry is bounded by a clock. An agent reports on a fixed cadence forever, whether or not anything is happening. Its volume is a function of how often it ticks and how many agents there are, and it is completely disconnected from business activity. An idle monitored host produces exactly as much traffic as a busy one — that is the entire point of monitoring.

Multiply any small per-tick cost by "forever" and it becomes a large number. This is the arithmetic that catches people out: a cadence chosen for graph smoothness, measured in seconds, produces five-figure daily event counts per agent. A rate that is obviously trivial per event is obviously not trivial per day. Nobody does that multiplication at design time, because the cadence is chosen by whoever is building the graphs and the pricing is chosen by someone else.

The failure that is not about money

Overcharging is the visible symptom, and it is the less serious one. The real problem appears when both kinds of traffic are counted against the same allowance.

Consider what that means in practice. A customer installs monitoring across their infrastructure. The agents run continuously and steadily consume the shared allowance. When it is exhausted the platform starts refusing requests — correctly, by its own rules. But the requests being refused are the customer's integration, the thing they are actually paying for, and the thing that consumed the allowance is the observability they added in order to be responsible.

The result is an outage the customer inflicted on themselves by adopting a second product, and it gets worse the more diligently they monitor. Every additional host they instrument brings it closer. Whatever error the platform returns will be technically accurate and completely useless: it will say the quota was exceeded, and the customer will look at their own modest API usage and conclude the platform is broken.

There is a related trap in the other direction. If telemetry consumes a shared budget, then an agent's cadence becomes a lever — and telemetry that can spend a budget is telemetry that can be used to exhaust one.

Classify by kind, not by transport

The fix is not a bigger allowance. A bigger allowance moves the collision further away without removing it, and it helps least the customers who instrument the most — which is to say, the ones using the product properly.

Telemetry and request-driven work should be separate categories with separate accounting, decided by what the traffic is rather than which door it arrived through. A few implications are worth stating plainly.

Cadence should not be a billing decision. If sampling interval changes what a customer owes, they will tune it for the invoice rather than for the resolution they actually need, and the product gets worse at its job. Telemetry ingestion is better governed by concurrency, retention and resolution — things that genuinely cost more to provide.

Fair-use limits still belong there. "Not metered per event" is not the same as "unbounded". But those limits should be enforced against the telemetry category, where hitting one degrades monitoring, rather than against a shared pool where hitting one takes production down as collateral.

Enforcement and accounting must agree. If traffic is admitted under one category and recorded under another, usage reports will disagree with what the customer experiences, and support will spend its time reconciling two numbers that were never measuring the same thing.

The general rule

Ask what governs the volume. If it is the customer's activity, meter it per event. If it is a clock, do not — because a clock never stops, never gets bored, and does not care what else was relying on the budget it is spending.