loader
Research
Privacy & Identity

Why Credential Stuffing Survives MFA

Multi-factor authentication stops password reuse from becoming account takeover — until the attacker stops trying to defeat the factor and starts working around it.

The attack has changed shape, not stopped

Credential stuffing is mechanically simple. An attacker takes username and password pairs exposed in one breach and replays them against unrelated services, relying on reuse. It requires no vulnerability. The credentials are valid; the only error was made by the user, years earlier, on a different site.

Multi-factor authentication is the standard answer and a genuinely effective one. A correct password alone no longer produces a session. What follows is not that attackers defeated the factor, but that they stopped attacking it directly — and the paths around it are more numerous than most deployments account for.

Working around the factor

Endpoints that never asked for it. MFA is typically enforced on the interactive login form. Legacy protocols, mobile API endpoints, token exchanges and integration paths frequently authenticate with a password alone, because they were built before the requirement or exempted to avoid breaking a client. An attacker enumerates authentication surfaces and uses whichever one asks for least. This is the single most common bypass and it is a coverage failure, not a cryptographic one.

Fatigue and coercion. Push-based approval asks a human to make a security decision with almost no context. Repeated prompts, particularly at inconvenient hours, produce approvals — not through carelessness so much as through a reasonable desire to make the noise stop. Prompts arriving alongside a convincing phone call from "IT" succeed at a much higher rate. The factor works exactly as designed; the human is the interface being attacked.

Real-time relay. A phishing page that proxies to the genuine service in real time collects the password, presents the authentic MFA challenge, relays the user's response, and captures the resulting session. Every factor is satisfied legitimately. What the attacker steals is not the credential but the session that follows it, which is why this defeats time-based codes and push approval equally — neither is bound to the site the user is actually talking to.

Recovery flows. Account recovery exists to restore access when a factor is lost, which means it is a documented path to authentication without that factor. If recovery relies on email possession, phone control or knowledge answers, its strength is the strength of those — and phone control is negotiable through carrier social engineering.

Session token theft. Once a session exists, the authentication is finished. Tokens stolen from a device, a compromised browser extension or an exposed log grant access without touching the login flow at all.

Why detection is the second half

If the authentication decision is the only control, every one of those paths ends in a valid session and nothing more is examined. The practical countermeasure is to treat authentication as evidence rather than proof, and keep evaluating after login.

Credential stuffing has a distinctive shape at the population level even when each attempt looks ordinary in isolation:

  • Failure distribution. Normal traffic produces failures concentrated on a few accounts — people mistyping their own passwords repeatedly. Stuffing produces a single failure across an enormous number of distinct accounts, which is the inverse and is unmistakable in aggregate.
  • Address behaviour. Attempts spread across residential proxy pools show many addresses each making very few attempts. Per-address rate limits never trigger; the aggregate is obvious.
  • Client uniformity. Large-scale automation tends to be uniform in ways real browser populations are not — TLS fingerprint, header ordering, capability support clustering far more tightly than genuine traffic.
  • Timing. Human login is irregular. Automated attempts arrive with machine regularity, and even deliberate jitter usually retains a recognisable distribution.
  • Success anomalies. A successful login from a new address, region and device simultaneously, on an account with no history of any of them, deserves scrutiny regardless of which factors were satisfied.

Controls in order of effect

Close the coverage gap first. Enumerate every path that produces a session and confirm each enforces the same requirements. This costs nothing beyond the audit and removes the most-used bypass. Legacy endpoints kept alive for one client are the usual finding.

Prefer origin-bound factors. Passkeys and hardware security keys are cryptographically bound to the site requesting authentication. A relay proxy is a different origin, so the credential does not produce a valid assertion. This is the one control that structurally defeats real-time phishing rather than making it harder, and it is the highest-leverage change available.

Add context to approvals. Number matching, and showing location, application and requesting address, converts a reflexive tap into a decision with information attached. It does not eliminate fatigue but it substantially raises the cost.

Rate-limit on the right dimension. Per-address limits are ineffective against distributed attempts. Limit per account, per credential pair, and on aggregate failure rate across the population.

Score at population level. Individual attempts are unremarkable; the campaign is not. Reputation and behavioural signals evaluated across all traffic catch what per-request logic cannot.

Treat recovery as a primary path. Apply the same scrutiny to recovery as to login, because an attacker certainly will. Recovery that is weaker than the front door defines your actual security level.

Watch sessions, not just logins. Bind tokens to client characteristics where practical, re-authenticate before consequential actions, and treat abrupt changes in address or device mid-session as a signal.

What this means for deployment

Enabling MFA and considering the problem closed is the failure mode worth naming. It is a large improvement that shifts attacker effort rather than ending it, and the paths it shifts effort toward — uncovered endpoints, recovery flows, relay phishing, stolen sessions — are exactly the ones that receive least attention precisely because MFA is assumed to have handled the risk.

The systems that hold up in practice pair strong, origin-bound authentication with continuous evaluation of what happens afterwards. Neither half is sufficient. Authentication decides whether to issue a session; detection decides whether to keep trusting it.