← Notes

The subscription that went deaf

For months we had a bug that never showed up as a bug.

Our NestJS backend subscribes to Salesforce platform events. Records change in the CRM, an event fires, our side picks it up and does work. After several days of uptime, the subscription would stop delivering. No error callback. No reconnect log. CPU flat, memory flat, health check green. On the Salesforce side, events kept publishing. On ours, nothing arrived until someone restarted the task, at which point everything came back and the clock started again.

The first wrong theory

When something works for days and then quietly doesn’t, the first suspect is token lifecycle. We were using the JWT Bearer flow, so we migrated to Client Credentials. We wanted that anyway: JWT Bearer breaks after every sandbox refresh, because the pre-authorization records get wiped and nobody remembers to recreate them until the integration is dead on a Monday. Client Credentials has its own quirk (the token endpoint has to be your org’s My Domain URL, not the generic test host), but once it worked, it kept working.

The subscription still went deaf. That ruled out auth and pointed at the transport.

What the library actually does

jsforce’s streaming client speaks CometD over Faye: HTTP long-polling on the Bayeux protocol. When Salesforce decides a client is stale, it answers 403::Unknown client and forces a re-handshake. Faye handles the handshake correctly. What it does not do is re-subscribe to the channels afterward. You end up with a perfectly live connection subscribed to nothing, and nothing in the callback surface tells you.

There are at least five open jsforce issues describing this exact pattern. The oldest is from 2015. This isn’t a configuration problem or a version problem; it’s the behavior of the transport, and it’s been the behavior for a decade.

The tempting next move was Change Data Capture. Different channel, richer payloads, a changedFields bitmap, echo suppression. All true, and all irrelevant, because CDC over the same CometD transport inherits the same re-handshake bug. The channel was never the problem.

What we shipped first

You can’t wait for a protocol migration when events are being lost in production. So we shipped mitigations, and every one of them is ugly:

  • A watchdog on last-event age. If nothing has arrived in longer than the expected gap for that channel, force a resubscribe.
  • A scheduled resubscribe every four hours regardless of what the watchdog thinks.
  • Replay IDs persisted in Redis, so a resubscribe or a restart resumes from the last event we saw. Salesforce retains events for 72 hours, which is plenty if you actually use it.
  • A process restart as the last resort, because a Fargate task that dies gets replaced.

None of this fixes the bug. All of it makes the bug survivable, and — this is the part that mattered — measurable. Each mitigation emits a metric. We can now see how often the transport goes deaf, which we could not before.

The real fix

The Pub/Sub API is Salesforce’s replacement for the streaming API: gRPC bidirectional streaming, Avro-encoded payloads, explicit replay semantics, and gap and overflow events when something is missed. It doesn’t go silent; it tells you. Philippe Ozil’s salesforce-pubsub-api-client does most of the heavy lifting in Node, and it supports the Client Credentials flow natively.

Two details worth knowing before you migrate. First, this is a protocol change, not a channel-name swap: HTTP/1.1 long-polling to gRPC streaming, with everything that implies for your network path. Second, if your service sits behind an AWS NAT Gateway, idle connections are dropped at 350 seconds; set gRPC keepalive comfortably under that or you’ll build a fresh mystery.

The lesson

“No errors” is not the same as “healthy.” A subscriber’s health is the age of its last message, not the state of its socket. The bug was invisible for as long as it was because everything we monitored was about the process, and the failure was about the absence of work.

If a system’s failure mode is silence, you have to instrument the silence. That’s a seconds_since_last_event metric with an alarm on it, from day one, before you know anything is wrong. It costs almost nothing, and it would have turned months of “huh, restart it” into a ticket with a root cause in a week.