The uncomfortable premise
A commercial captcha-solving service does not break the captcha. It employs people to solve it, and sells the answers by the thousand at prices that clear a fraction of a cent per solve.
This is worth stating plainly because it invalidates the intuitive response. If the challenge is being answered by an actual person, then no amount of making it harder to answer addresses the failure. It raises the cost slightly, frustrates real users considerably, and leaves the attack fully intact. Any defence premised on "design a puzzle a machine cannot solve" has already conceded, because there is no machine on the other side at the moment of solving.
The useful question is different: the challenge was answered by a human, but was it answered by this visitor?
Two shapes of relay
Relay attacks split into two families, and conflating them produces defences that miss.
In the first, the operator's software hands the entire challenge to a worker, who loads and solves it on their own machine. Whatever proof results is then carried back and redeemed by the operator's infrastructure. The defining property is that the context that solved the challenge and the context that redeems it are different — different network, different machine, different session.
In the second, the operator's software drives everything and outsources only the perception. It fetches the challenge itself, ships an image of it to a worker, receives an answer like "the two on the left," and submits from its own infrastructure. Here the network is consistent, the session is consistent, and the headers are whatever the operator chooses to send. Nothing about the request contradicts itself.
The first family is comparatively tractable: a proof that is only valid in the context that earned it does not survive being moved. Origin controls, binding to the requesting network, single use, and a short validity window all attack it directly, and together they make the straightforward version of this attack uneconomic.
The second family defeats every one of those controls, because none of them are being violated.
What is actually left
Once network, origin and session all check out, the remaining evidence is behavioural — the difference between a page that was used and a page that was merely driven.
We are deliberately not going to enumerate which signals Sentinel weighs or how it combines them. Publishing that list converts a detection problem into a checklist, and a checklist is something an adversary can work through. What is worth saying is the principle behind the selection, because the principle is durable even when specific signals are not:
Prefer signals a relay must actively reconstruct over signals it merely has to avoid. A signal that is absent when automation is present can be supplied. A signal that requires a coherent story to be maintained across the entire life of an interaction is far more expensive to fake, because the fabrication must remain internally consistent with everything else the client has already claimed. Contradictions between what a client asserted at one moment and what it asserts at another are unusually informative, because a genuine session has no reason to produce them.
Two further rules follow from operating this in production.
First, nothing behavioural should be decisive alone. Real traffic is strange. Corporate middleboxes rewrite things. Privacy tooling alters things. Assistive technology produces interaction patterns that look nothing like the median. Any single signal treated as proof will eventually lock out a legitimate user who did nothing wrong, and that failure is invisible to you and infuriating to them. Signals should accumulate into a judgement, and the judgement should have somewhere to go other than "denied."
Second, degrade rather than block. The correct response to a suspicious-but-not-conclusive profile is more friction, not a wall.
The economic layer
The durable defence is not detection. It is price.
A solving service is priced per solve because its cost is per solve — a person spends seconds of attention, and those seconds are the product. The margin on that business is thin and the volume assumption is aggressive. Anything that increases the seconds required per successful attempt attacks the business model directly, and it does so without needing to detect anything at all.
This reframes what a good challenge is. The goal is not maximum difficulty; difficulty punishes your real users at least as much as your attackers, and usually more, because your real users are not being paid to be there. The goal is maximum asymmetry: work that a genuine visitor absorbs in a couple of seconds as part of a task they already wanted to complete, and that an attacker must purchase, at scale, per attempt.
Several properties push in that direction. A challenge whose answer is a set rather than a single target costs more attention per solve and cannot be brute-forced by selecting everything. A challenge that varies what it asks prevents a worker from settling into a rhythm. Chaining rounds multiplies per-attempt cost linearly while adding only seconds for someone who is already halfway through signing up.
That last one is why a raised floor during an incident is a genuine mitigation rather than theatre. It does not make the attack impossible. It makes it cost several times more per success at exactly the moment the attacker has committed resources — and an operation running on fractions of a cent has very little room to absorb that.
What this means for anyone building or buying
Three conclusions we would hold to regardless of vendor.
Treat "unsolvable by machines" as a category error. Anyone selling that is either misunderstanding the threat or hoping you do. The question is what a solve costs and whether the proof survives being moved.
Assume any client-supplied claim can be forged. Headers that identify origin or context are useful signals and terrible authorisations. Anything a non-browser client can set, a non-browser client will set. Build so that forging them is necessary but not sufficient.
Budget for false positives before you deploy. Every real detection system has them. The question is not whether yours will, but whether a legitimate user who trips it has a path forward that does not require contacting support. If the answer is no, the system is not finished.