← Work

Case study · 2026

Flipping the auth provider without logging everyone out

A move to a new identity provider across several environments, with real users who must not notice. The sequence that made a one-way door two-way: validate first, keep the old provider live, verify each environment separately.

Period
2026
Stack
  • Node.js
  • OIDC/JWT
  • AWS Secrets Manager
  • GitHub Actions
Note
Flipping the auth provider without logging everyone out

Context

A move to a new identity provider across multiple environments, with real users who must not notice. The tempting version — change the runtime setting and redeploy — can lock every user out in one shot.

The problem

A provider switch is a one-way door if you get the ordering wrong. The client can start issuing tokens the backend can’t yet validate. The shortcut of “just reset everyone’s credentials” introduces an email-verification race on first sign-in, and the assumption that dev, staging and production share a tenant or an audience value is usually wrong.

What I did

  • Verified, by reading the code, that dependent services would register and validate tokens from the new provider unconditionally, before touching any client-facing flag.
  • Ran the full cross-system flow in a staging spike first, confirming issuer and audience acceptance and, crucially, the negative case: unauthorized tokens get rejected.
  • Kept the old provider active in the list through the cutover, so authentication could never be left fully broken mid-switch.
  • Read issuer, audience and enabled methods directly from each live environment instead of trusting recalled config, and confirmed production separately from dev and staging.
  • Where an account had to be bootstrapped manually, marked its email verified in the same step to sidestep the verification-versus-sign-in race.

The trade-off

Preparation steps that are additive and reversible go in first; the one behavior-changing switch is gated behind explicit confirmation with a fallback still standing. Running two providers side by side costs some configuration surface and a retirement task later. The alternative is an outage nobody can roll forward from.

Outcome

A cutover with a live fallback and a tested rollback, verified per environment rather than assumed uniform.

What I’d do differently

Build the negative-path test — unauthorized token rejected — into the spike checklist as a first-class artifact, not an afterthought. And script the per-environment config read so it’s repeatable instead of a set of manual lookups.