Don't let the model decide the checkout
Chat-led interfaces are seductive because the model handles the hard part: understanding what the person means and responding like a human would. The temptation that follows is to let it handle everything else too. Need the user to book an appointment? Give the model a scheduleAppointment tool and let it decide when to offer it. Need a document uploaded? Another tool. Let the conversation flow.
We didn’t do that, and I want to explain why, because the alternative is less glamorous and much more reliable.
Two kinds of actions
In the customer app I work on, there’s a category of action I’d call business-critical: things the person must do for their case to move forward, in a particular order, at a particular time. Scheduling a call. Uploading a specific document. Signing something. If the model forgets to offer one, or offers it at the wrong moment, or offers it twice, that’s not a rough edge in the conversation. It’s a person’s claim stalling.
Then there’s everything else: questions, explanations, reassurance, “what happens next.” That’s where the model is excellent and where I want it to have as much room as possible.
The design question was whether those two categories should go through the same mechanism. We considered two options.
Option A — deterministic resolution. A service on the backend looks at the authoritative case state and computes the list of tasks that apply right now. That list is handed to the frontend as structured data. The model narrates around it; it doesn’t produce it.
Option B — model-driven. The model has tools for each action and decides, from the conversation and whatever context it’s given, when to call them.
We chose A for anything business-critical, and I’d choose it again.
Why the model shouldn’t be the workflow engine
Language models are probabilistic. Most of the time that’s fine — “most of the time” is what “conversational” means. But a workflow engine has to be right every time, and “the model usually surfaces the scheduling step” is a bug report waiting to be written. Every guard you add to the prompt to make it more reliable (“ALWAYS offer scheduling when the case is qualified”) is business logic living in the least testable place in your stack.
There’s also the question of what the model actually knows. Case state lives in a system of record owned by another layer. Making the model responsible for reading that state correctly, interpreting it, and acting on it adds a chain of inference to something that should be a query.
How the deterministic version works
A task resolver on the backend reads the case context and returns a list of task descriptors. Each has a type, a priority tier (critical, high, standard) and an integer order within the tier, so the UI can sort without ambiguity. The composite key matters: a single priority number becomes a mess the first time someone needs “this one, but before that one.”
The tasks travel to the frontend as message annotations on the AI SDK’s data stream — structured metadata attached to the message rather than text the model generated. The React side renders them as real CTAs with real handlers. The model can talk about them, and its system context tells it what’s pending so its narration is coherent, but it can’t invent one, drop one, or reorder them.
The opening message follows the same principle. Instead of sending a sentinel “init” message and letting the model respond to it (a pattern that leaks a fake user turn into your transcript), the server generates the greeting once with generateText and hands it to useChat as an initial message. Deterministic where it matters, generated where it doesn’t.
One placement decision worth stating plainly: the streaming route handlers live in Next.js, where the AI SDK expects them. The NestJS layer is the gateway — audit logging, request enrichment, rate limits, credential isolation — not the place where chat logic goes. Getting that boundary wrong the first time cost us a design round.
Where Option B is still right
I’m not against tool calls. For exploratory, low-stakes actions — “show me my recent uploads,” “explain this letter” — letting the model decide is exactly right, and the AI SDK’s human-in-the-loop approval flag covers the ones that need a confirmation step. The rule I landed on is simple: if a missed or duplicated action would be a support ticket, it doesn’t go through the model.
The lesson
The model is very good at language and very bad at being a state machine. Keep the “must happen” logic in code you can unit test, hand the results to the model as facts, and let it do the part it’s actually good at. Your prompt gets shorter, your transcripts get cleaner, and the day someone asks “why didn’t the user get the scheduling button?” you’ll answer with a query instead of a guess.