Case study · 2025 – 2026
The subscription that went deaf
A Salesforce platform-event subscription stopped delivering after days of uptime with no errors and a green health check. The bug was in the CometD transport; the fix was mitigations now and a protocol change next.
Context
A NestJS backend on ECS Fargate subscribes to Salesforce platform events that drive downstream work. After several days of uptime, events stopped arriving. No error callbacks, no reconnect logs, health check green. A restart fixed it — until the next time.
The problem
The first theory was token expiry. Migrating from JWT Bearer to Client Credentials (which we wanted anyway) didn’t change the behavior. So it wasn’t auth. Something below the application layer was dropping the subscription without telling anyone.
What I did
- Traced it to the CometD transport in jsforce: on a
403::Unknown clientre-handshake, the underlying Faye library reconnects but never re-subscribes. At least five open issues describe this, the oldest from 2015. - Confirmed that switching to Change Data Capture wouldn’t help — same transport, same bug.
- Shipped mitigations: a watchdog on last-event age, a scheduled resubscribe every four hours, replay IDs persisted in Redis so recovery resumes where it stopped, and a process restart as last resort.
- Specified the real fix: the Pub/Sub API over gRPC, with gRPC keepalive tuned under the AWS NAT Gateway’s 350-second idle timeout and the token endpoint pointed at the org’s My Domain URL, not the generic test host.
- Wrote the runbook for creating the integration app across sandboxes and production, including the post-refresh checklist.
The trade-off
Patching the transport is cheap and immediate; changing the protocol is the only thing that actually fixes it. Doing both, in that order, was the right sequence.
Outcome
The silent gaps are detectable and self-healing while the Pub/Sub migration lands. Each mitigation emits a metric, so the team can now see how often the transport goes deaf instead of finding out from a stalled queue.
What I’d do differently
Emit a “seconds since last event” metric from day one. “No errors” looked like “healthy” for far longer than it should have.