← Notes

Fail soft by default: integrating an LLM with systems that can go down

The first time I wired a language model up to real system state, I did the obvious thing: fetch the state, put it in the prompt, send the turn. It worked beautifully in the demo and then, a week later, the service I was fetching from had a bad thirty seconds and every single chat turn returned an error. Not “the AI gave a worse answer.” An error. The user typed a message and got a spinner that never resolved.

That taught me the rule I now apply to every AI integration: the model is the least of your reliability problems. The moment you connect a probabilistic system to deterministic ones — a database, a Lambda, an internal API — you’ve built a distributed system, and distributed systems fail. The only question is whether they fail loud or soft.

What “fail soft” actually means here

Concretely: if the thing you’re fetching to enrich the prompt is unavailable, malformed, or missing a precondition, the turn should proceed without the enrichment. The user gets a slightly less informed answer instead of no answer. That’s almost always the right trade. A chat assistant that occasionally forgets to proactively offer you something is fine. A chat assistant that returns 500 because a downstream cache blinked is not.

In practice this is a try/catch that returns null and a caller that treats null as “inject nothing.” It sounds trivial. It is trivial. But you have to decide it’s the contract, write the test that proves the null path continues the conversation, and resist the urge to “surface the error to the user so they know.” They don’t need to know. The system knew, logged it, and kept going.

Distinguish “failed” from “not applicable”

Here’s a subtlety that bit me. There are two different reasons the enrichment might be absent:

  1. The dependency failed — timeout, parse error, throttle. This is the fail-soft path.
  2. The enrichment doesn’t apply — there’s genuinely no state for this user yet, so there’s nothing to fetch.

These look the same at the call site (no data) but they’re not the same event. The first is an error you want in your logs and metrics because a spike means something’s wrong. The second is a normal precondition you log at debug and forget. Collapse them into one path and you either alarm on nothing or go blind to real failures. I now model them separately: a precondition-skip logs quietly, a genuine failure logs as an error with the exception attached.

Keep the wire format out of your domain

The other thing I always do now is put a strict schema at the boundary. The external service speaks its shape (snake_case, whatever fields it likes); my code speaks mine. A strict parser at the edge means an unexpected field or a changed type is caught right there, as a clean parse failure that routes into the fail-soft path, instead of propagating three layers deep as an undefined that blows up somewhere unrelated at 2am.

The bonus: when the upstream team evolves their response, the schema change is a single small diff in one file. Every consumer keeps speaking the domain shape and doesn’t care.

Ship it dark

Finally, none of this goes out hot. It goes behind a feature flag that defaults to off, rolled out per environment — staging, then a small cohort, then everyone. If anything’s wrong, the fix is flipping a flag, not reverting code and waiting for a deploy. The single most calming sentence in an incident is “we can turn it off in ten seconds,” and you only get to say it if you built it that way from the start.

The demo-to-production gap for AI features isn’t about the model. It’s about everything around the model. Build the “around” to fail soft and you can actually sleep.