← Work

Case study · 2026

A chat-led app next to the dashboard

A conversational mode added to a customer app without a rewrite. Next.js Multi-Zone, a mandatory NestJS gateway, token streaming from Lambda Function URLs, and business-critical actions resolved in code rather than by the model.

Period
2026
Stack
  • Next.js App Router
  • Vercel AI SDK
  • NestJS
  • AWS Lambda (Function URLs)
  • EventBridge
  • SQS
  • Nx
  • Prisma
Note
Don't let the model decide the checkout

Context

The customer app needed a conversational mode as a primary way to interact, alongside the existing dashboard. The AI logic already lived in Python Lambdas behind EventBridge and SQS; the frontends deploy to Vercel; every AI interaction has to be auditable.

The problem

Two UIs in one Nx monorepo with shared sessions, token streaming from Lambda to the browser, no credentials in the frontend, no bypass of the audit trail — and business-critical actions surfaced in chat at the right moment, every time.

What I did

  • Chose Next.js Multi-Zone over Module Federation: native to Vercel, independent deploys, no runtime module resolution to debug.
  • Made a NestJS gateway mandatory between the browser and any model or Lambda — credential isolation, audit logging, request enrichment, rate limits.
  • Streamed from Lambda Function URLs in RESPONSE_STREAM mode instead of API Gateway WebSocket: no connection-state table, less overhead.
  • Kept the AI SDK route handlers in Next.js, where they belong, and used the SDK’s useChat and message annotations for structured payloads.
  • Resolved critical actions deterministically from the authoritative case state (priority tier plus order), not through LLM tool calls, and pre-generated the opening message server-side instead of a sentinel message.

The trade-off

Function URLs give up bidirectional messaging, which this flow doesn’t need, for roughly 8–10 ms of overhead versus 15–30 ms through API Gateway. Deterministic tasks give up some conversational flexibility for guaranteed behavior where it matters.

Outcome

Accepted as an ADR with explicit targets: p95 first token under 500 ms, chat error rate under 1%, session continuity across zones above 99%. The implementation follows the ADR’s phased plan, gateway first.

What I’d do differently

Write down which layer owns case state before anyone drafts code. One design round went in the wrong direction on exactly that assumption.