Case study · 2026
Flipping the auth provider without logging everyone out
A move to a new identity provider across several environments, with real users who must not notice. The sequence that made a one-way door two-way: validate first, keep the old provider live, verify each environment separately.
Context
A move to a new identity provider across multiple environments, with real users who must not notice. The tempting version — change the runtime setting and redeploy — can lock every user out in one shot.
The problem
A provider switch is a one-way door if you get the ordering wrong. The client can start issuing tokens the backend can’t yet validate. The shortcut of “just reset everyone’s credentials” introduces an email-verification race on first sign-in, and the assumption that dev, staging and production share a tenant or an audience value is usually wrong.
What I did
- Verified, by reading the code, that dependent services would register and validate tokens from the new provider unconditionally, before touching any client-facing flag.
- Ran the full cross-system flow in a staging spike first, confirming issuer and audience acceptance and, crucially, the negative case: unauthorized tokens get rejected.
- Kept the old provider active in the list through the cutover, so authentication could never be left fully broken mid-switch.
- Read issuer, audience and enabled methods directly from each live environment instead of trusting recalled config, and confirmed production separately from dev and staging.
- Where an account had to be bootstrapped manually, marked its email verified in the same step to sidestep the verification-versus-sign-in race.
The trade-off
Preparation steps that are additive and reversible go in first; the one behavior-changing switch is gated behind explicit confirmation with a fallback still standing. Running two providers side by side costs some configuration surface and a retirement task later. The alternative is an outage nobody can roll forward from.
Outcome
A cutover with a live fallback and a tested rollback, verified per environment rather than assumed uniform.
What I’d do differently
Build the negative-path test — unauthorized token rejected — into the spike checklist as a first-class artifact, not an afterthought. And script the per-environment config read so it’s repeatable instead of a set of manual lookups.