Flipping the auth provider without logging everyone out
There’s a category of change where the naive version and the safe version look almost identical in the diff, but one of them can lock every user out of your product. Switching identity providers is the cleanest example I know.
The naive version: change the runtime setting that says “use provider B,” redeploy, done. The problem is ordering. If the frontend starts getting tokens from provider B before the backend knows how to validate provider B’s tokens, every authenticated request fails at once. And unlike a normal bug, you can’t just roll forward — your users are already holding tokens the system rejects.
Here’s the sequence I use instead.
1. Prove the backend can validate the new tokens — by reading the code, not by hoping
Before touching anything client-facing, I confirm that the services which verify tokens will accept provider B unconditionally. Not “there’s a config option for it.” Actually trace that the verification path registers the new issuer and audience. If the backend can’t validate B’s tokens yet, flipping the client is a guaranteed outage.
2. Run the whole flow in a staging spike first — including the negative case
It’s not enough to prove a valid token from B gets in. You have to prove an invalid token gets rejected. A cutover that accepts everything is worse than the problem you started with. The negative test is the one people skip and the one that matters.
3. Keep the old provider live during the switch
This is the big one. Don’t remove provider A when you add provider B. Run both, so that at no point in the cutover is authentication fully broken. A user mid-session on an A token keeps working; new sessions get B. You retire A later, deliberately, once you’ve watched the metrics.
4. Read the config from each live environment — don’t trust your memory
I’ve been burned assuming dev, staging, and prod share a tenant or an audience value. They often don’t. Read the issuer, audience, and enabled methods directly from each environment and confirm production separately. Parity is an assumption, and assumptions about auth are expensive.
5. Beware the credential-reset race
A tempting shortcut when migrating users is “just have everyone reset their credentials.” But if your platform requires email verification, you’ve now got a race between the verification email and the first sign-in attempt, and some fraction of users land in a broken state on their first try. If a pre-verified path exists — a federated sign-in, an invite-acceptance flow — use it. When I do have to bootstrap an account manually, I mark its email verified in the same step, so there’s no window for the race to happen.
Separate the reversible from the irreversible
Most of the work above is additive and safe: adding B’s validation, adding B to the provider list, reading config. All of that can go in ahead of time and be verified at leisure. There is exactly one behavior-changing step — the flip that makes B the default — and that one gets an explicit confirmation and a fallback still standing behind it.
The meta-lesson: for one-way doors, spend your effort making them two-way. A cutover with a live fallback and a tested rollback isn’t a cutover anymore. It’s just a config change you happen to be watching closely.