← Work

Case study · 2025 – 2026

The subscription that went deaf

A Salesforce platform-event subscription stopped delivering after days of uptime with no errors and a green health check. The bug was in the CometD transport; the fix was mitigations now and a protocol change next.

Period
2025 – 2026
Stack
  • NestJS
  • @jsforce/jsforce-node v3
  • Salesforce Platform Events
  • Pub/Sub API (gRPC)
  • Redis
  • ECS Fargate
  • CloudWatch
Note
The subscription that went deaf

Context

A NestJS backend on ECS Fargate subscribes to Salesforce platform events that drive downstream work. After several days of uptime, events stopped arriving. No error callbacks, no reconnect logs, health check green. A restart fixed it — until the next time.

The problem

The first theory was token expiry. Migrating from JWT Bearer to Client Credentials (which we wanted anyway) didn’t change the behavior. So it wasn’t auth. Something below the application layer was dropping the subscription without telling anyone.

What I did

  • Traced it to the CometD transport in jsforce: on a 403::Unknown client re-handshake, the underlying Faye library reconnects but never re-subscribes. At least five open issues describe this, the oldest from 2015.
  • Confirmed that switching to Change Data Capture wouldn’t help — same transport, same bug.
  • Shipped mitigations: a watchdog on last-event age, a scheduled resubscribe every four hours, replay IDs persisted in Redis so recovery resumes where it stopped, and a process restart as last resort.
  • Specified the real fix: the Pub/Sub API over gRPC, with gRPC keepalive tuned under the AWS NAT Gateway’s 350-second idle timeout and the token endpoint pointed at the org’s My Domain URL, not the generic test host.
  • Wrote the runbook for creating the integration app across sandboxes and production, including the post-refresh checklist.

The trade-off

Patching the transport is cheap and immediate; changing the protocol is the only thing that actually fixes it. Doing both, in that order, was the right sequence.

Outcome

The silent gaps are detectable and self-healing while the Pub/Sub migration lands. Each mitigation emits a metric, so the team can now see how often the transport goes deaf instead of finding out from a stalled queue.

What I’d do differently

Emit a “seconds since last event” metric from day one. “No errors” looked like “healthy” for far longer than it should have.